Search bioRxiv⌕ Search

Biology subjects

Parsa, M. S.

Publications and source records attributed to Parsa, M. S..

4 recordsLinked to original sources

Structure-Aware Mapping of Disease-Relevant Missense Variation in Nuclear Pore complex Genes

Missense variation in the nuclear pore complex (NPC) remains difficult to interpret because sequence change, structural context, and sparse clinical labels all interact in nontrivial ways. We study three functionally distinct nucleoporins GLE1, NUP214, and NUP62 and build a reproducible pipeline that binds variants to canonical UniProt coordinates, overlays AlphaFold2 per-residue confidence, and assigns domain/feature labels from UniProtKB/Pfam. Primary inferences rely strictly on curated Clin-Var assertions, while a separate high-confidence pseudo-labeled cohort is created for sensitivity analyses using a guarded weak-supervision scheme: a centroid-cosine scorer over handcrafted sequence-structural features is ensembled with a positive-unlabeled classifier, and only variants passing conservative probability gates are promoted. Across genes, curated data reveal coherent structure-function signals: pathogenic substitutions concentrate in specific domains and structurally ordered regions, while the pseudo-labeled cohort preserves these trends under expanded sample size without entering into hypothesis tests. The result is a transparent workflow that cleanly separates ground truth from weak supervision, avoids leakage, and produces interpretable, domainlevel effect estimates. We argue that this combination of principled labeling, structural context, and simple, auditable models offers a practical path for variant interpretation in nucleoporins and, more broadly, in proteins rich in intrinsically disordered and repeat-containing regions.

genomics↗

KODA: Agentic Framework for Microbiome Drug Target Discovery

The gut microbiome plays a crucial role in human health and disease, influencing diverse biological processes such as immune regulation and nutrient metabolism. However, the complexity of micro-bial interactions and their metabolic cross-feeding dynamics remains poorly understood. This study proposes KODA, an agentic framework that integrates large language models (LLMs) and knowledge graphs (KGs) to facilitate the discovery of targets in antimicrobial drugs in the gut microbiome. Our approach employs a multi-agent system to interpret natural language queries and translate them into precise graph database queries, enabling intuitive interactions with complex microbiome data. Focusing on KEGG orthologies related to essential microbial genes, KODA identifies potential antimicrobial drug targets by analyzing microbial metabolic pathways. The system employs a Neo4j-based microbiome KG, which integrates microbial interaction data, metabolic models, and KEGG annotations. A dedicated evaluation framework, which incorporates LLM-based reviewers, assesses the quality of generated queries and analytical reports. Our results demonstrate the efficacy of KODA in providing actionable insights for antimicrobial research, particularly in identifying conserved essential genes as potential drug targets. This framework holds the potential to democratize microbiome research by lowering technical barriers and accelerating hypothesis generation in drug discovery.

bioinformatics↗

SIMBA-GNN: Simulation-augmented Microbiome Abundance Graph Neural Network

Understanding gut microbiome dynamics gut requires deciphering complex, metabolically driven interactions beyond taxonomic profiles. We present SIMBA, a novel framework that integrates mechanistic metabolic simulations with a graph neural network (GNN) to predict microbial abundances and uncover cross-feeding relationships. By simulating pairwise interactions among gut microbes using metabolic networks, we generate biologically grounded graphs that capture metabolite cross-feeding and functional relationships. Our custom GNN, enhanced with edge-aware attention, is trained through a multi-stage pipeline combining self-supervised learning, simulation-based pretraining, and fine-tuning on real microbial abundance data. SIMBA achieves state-of-the-art performance (Spearman {rho} = 0.85) and enables interpretable insights into keystone taxa and metabolic bottlenecks. This work demonstrates the power of combining metabolic networks with deep learning for precision microbiome analysis.

systems biology↗

A GENERALIZED PROTEIN DESIGN ML MODEL ENABLESGENERATION OF FUNCTIONAL DE NOVO PROTEINS

Traditional protein design is fundamentally constrained by known sequences and folds. To break free from these limitations, we introduce a new alternative: designing proteins directly from plain-language specifications. To achieve this, we trained MP4, a transformer-based model that maps natural language prompts to protein sequences, on a dataset of 3.2 billion points and 138k tokens. In a benchmark of 96 prompts representing a wide array of functions and contexts, MP4 excelled by simultaneously improving on three key metrics: sequence realism, predicted fold quality, and alignment to the requested function. This high performance is particularly significant as it was achieved using only text as input which is a major departure from other models. Experimental validation confirmed our computational predictions: two de novo designs were experimentally shown to be both expressible and thermostable, with high-resolution crystallography (1.30 [A] and 1.77 [A]) ultimately revealing one to possess a paradigm-shifting novel fold. Functionally, the designs were also active, demonstrating both ATP binding and hydrolysis in vitro. This work demonstrates the realization of natural-language intent as functional proteins that express, crystallize, and catalyze. Although the underlying approach is still in early development with incomplete coverage and controllability, MP4 delivers a profound impact: it lowers the barrier to protein design and vastly expands the space for creative exploration in molecular programming.

biochemistry↗