Search bioRxiv⌕ Search

Biology subjects

Dehal, P. S.

Publications and source records attributed to Dehal, P. S..

2 recordsLinked to original sources

Data coherence over data volume drives generalisable genome-based prediction of microbial carbon utilisation

Microbial carbon utilisation is a foundational ecological phenotype that remains difficult to predict from genomes despite well-characterised pathways. Machine-learning models generalise poorly across datasets, a failure usually attributed to training-set size and taxonomic bias. To test this, we integrated binary growth phenotypes for 819 strains across 240 carbon sources from four datasets. Balanced accuracy fell from 0.86 within datasets to 0.62 across them, and testing on close relatives recovered only 0.03 of that drop, so mechanistically inconsistent genotype-phenotype relationships drove models to dataset-correlated shortcuts. Restricting training to concordant samples (measured growth matched their annotated pathway) doubled the carbon sources recovering known pathway genes across datasets (6 to 12 of 15), whereas matched random subsets did not. Adding over 8000 literature-curated BacDive genomes to the training set did not improve cross-dataset performance more than the smaller concordant set, suggesting coherence matters more than volume. Because such filtering requires a mechanistic predictor, we tested a mechanism-free alternative combining phylogenetic agreement and experimental labels, which recovered part of the gain but not the recall advantage. Concordance-trained models were bounded specialists, rescuing mechanistic false negatives twice as often as false positives (44\% versus 19\%), mostly metabolic generalists. To locate those bounds, model confidence defined an applicability domain, and prioritising low-confidence genomes for training improved cross-dataset accuracy more than random or diversity-based sampling, especially for the weaker phenotypes. This recasts generalisation in biological machine learning as a problem of label-mechanism agreement and applicability-domain definition, alongside data volume and algorithm choice.

bioinformatics↗

KBase Research Agent: Automated Multi-Agent Workflow Construction for Reproducible Genome Analysis

Constructing multi-step bioinformatics workflows, from read quality control through genome assembly to functional annotation, requires expertise in both biology and computational tool selection, creating a bottleneck for scalable and reproducible analysis. We present the KBase Research Agent, a multi-agent system for automating such workflows within the DOE Systems Biology Knowledgebase (KBase). Given a set of sequencing reads and a research objective, the agent constructs an analysis plan grounded in KBase documentation and a Knowledge Graph (KG) of the KBase application catalog, then selects, parameterizes, validates and executes appropriate KBase applications to carry out the workflow. The resulting analysis is preserved as a reproducible KBase Narrative. We evaluate the systems planning and execution quality against ground truth constructed from reference workflows derived from peer-reviewed Microbiology Resource Announcements. We further apply the agent to 100 previously unanalyzed bacterial isolate genomes from the JGI IMG/M database, where it autonomously performed read quality control, genome assembly, taxonomic classification with GTDB-Tk, and downstream analysis producing annotated genomes, reproducible Narratives, and draft manuscripts without human intervention. Across these experiments, the KBase Research Agent demonstrates the feasibility of domain-grounded, end-to-end scientific workflow automation in a production bioinformatics platform.

bioinformatics↗