Search bioRxiv⌕ Search

Biology subjects

Bharj, G.

Publications and source records attributed to Bharj, G..

3 recordsLinked to original sources

Rapid and Interpretable AMR Diagnostics via Genomics and Cell Painting using Differential Geometry-based Directed-Simplicial Neural Networks on Multimodal Data

Antimicrobial resistance (AMR) remains a critical global health challenge, particularly in high-prevalence regions such as India, where rapid and interpretable diagnostic tools are urgently needed. To address this challenge, we present a computational framework for AMR prediction that integrates genomic and cellular phenotypic data using an in-house developed differential geometry-based Directed Simplicial Neural Network (Dg-Dir-SNNs) applied to multimodal datasets. Using this framework, we analyzed 384 clinically relevant AMR isolates, including Escherichia coli and Klebsiella pneumoniae, integrating 256 genomic k-mer features with 503 cellular morphology descriptors derived from high-content Cell Painting assays. The Dg-Dir-SNNs model constructs an inferred-causal network of top-ranked biomarker-driving features, predicting potential directional dependencies among genomic motifs and phenotypic features. Network analysis identified kmer_TATG as the top-ranked driver associated with predicted resistance, with a local neighborhood including other genomic motifs (kmer_TTTT, kmer_CGTG, kmer_TCAC, kmer_CGTA, kmer_GAAA, kmer_TAAA, kmer_TACA, kmer_TGTG, kmer_TGAG, kmer_AAAA) and a key morphological feature (Cells_correlation_ER_Brightfield). These relationships suggest potential mechanistic associations in which specific genomic motifs may influence cellular phenotypes linked to antimicrobial resistance. Although not yet clinically deployed, this approach demonstrates the potential of multimodal AI-driven modeling for rapid in silico AMR prediction. By providing interpretable, biologically grounded insights, the framework may support future diagnostic development, targeted surveillance strategies, and experimental validation in high-resistance healthcare settings.

microbiology↗

MAP-PRS: Multi-Ancestry Portfolio-Based Polygenic Risk Scores

Polygenic Risk Scores (PRS) are emerging tools for predicting an individuals genetic risk for complex diseases. However, their usefulness in clinical practice remains limited because most existing models are based on data from people of European ancestry, leading to reduced accuracy and stability in other populations. This imbalance restricts the equitable use of PRS in precision medicine. To overcome these limitations, we introduce Multi-Ancestry Portfolio-Based Polygenic Risk Scores (MAP-PRS)--a new framework that combines mathematical modeling and data science principles to improve both fairness and reliability in genetic risk prediction across populations. MAP-PRS treats each ancestry-specific PRS as part of a "portfolio," similar to how investments are managed in finance, balancing two key aspects: predictive return (how well the score predicts disease) and risk complexity (how uncertain or ancestry-specific the prediction is). By jointly optimizing these factors, MAP-PRS identifies the best combination of ancestry-informed PRS models that maximize predictive accuracy while minimizing bias and instability. This approach also uses advanced computational tools--such as Bayesian modeling, machine learning, and generative neural networks--to refine risk estimates, incorporate environmental and lifestyle factors, and increase representation from under-studied populations. In doing so, MAP-PRS supports more inclusive, equitable, and interpretable precision medicine. As an initial demonstration, MAP-PRS has been applied to predict Type 2 Diabetes (T2D) risk in European ancestry populations, establishing a foundation for broader, multi-ancestry implementation. Future extensions will include additional diseases, such as cervical cancer and HPV susceptibility, endometrioid ovarian cancer, and Alzheimers disease--bringing us closer to clinically actionable and globally equitable genetic risk prediction.

genetics↗

Interpretable Machine Learning and Comparative Genomics Reveal Microbial Plastic-Degrading (Microbeyt) Potential

Plastic pollution poses a critical environmental threat, and microbial enzymes represent a sustainable strategy for polymer degradation. We present a computational pipeline that integrates orthogroup-based genomic analysis with machine learning and interpretable feature importance to identify microbial strains with high plastic-degrading potential. Using presence or absence matrices and SHAP-derived feature contributions to the MTP visualization, the workflow highlights conserved gene modules driving predictive classification. Application to a single genus revealed strains harboring versatile enzymatic repertoires capable of targeting diverse polymers, including polyethylene, polyethylene terephthalate, polyurethane, and polyhydroxyalkanoates. These findings provide a rational framework for prioritizing candidate strains for experimental validation and bioremediation strategies. Overall, this study demonstrates how integrating comparative genomics with interpretable machine learning can guide the systematic discovery of microbial solutions to plastic pollution.

genomics↗