Search bioRxiv⌕ Search

Biology subjects

Kulkarni, S. V.

Publications and source records attributed to Kulkarni, S. V..

2 recordsLinked to original sources

A simple and accurate method for inferring missing ploidy information from sequence data

Polyploidy can be a critical factor for explaining plant trait variation, niche diversification, or speciation. However, inferring ploidy from silica-dried or historical samples using chromosome counts or flow cytometry is not possible, and scaling up ploidy estimation to population-level fresh contemporary samples can be challenging as well. Thus, we present a new method for estimating ploidy levels directly from sequencing data using machine learning; the Polyploid Population Genomics Tool Kit (PPGTK). The machine-learning approach is advantageous as it relaxes the assumptions of previous probabilistic methods and provides per-sample probabilities, allowing investigators to evaluate uncertainty in their system of interest.. We demonstrate performance and accuracy of the method on simulated and empirical data. Simulations showed above 99% accuracy, even for low coverage data, as long reads were mappable to the reference genome. For empirical analyses, we used target enrichment data from blueberry wild relatives (Vaccinium sect. Cyanococcus) and whole-genome data from sweetpotato wild relatives (Ipomoea ser. Batatas). Ploidy was recovered with 99% accuracy across 70 Vaccinium individuals and 97% across 82 Ipomoea individuals. Analysis of many individuals is fast and requires only a multisample VCF, which is presumably generated for the research anyway, and some samples of known ploidy for training the classifier. The approach implemented in PPGTK is promising for collections-based research as well, enabling ploidy classification of historical specimens based on present-day observations. The method is implemented in a new Python package as a single command that can run on a conventional laptop.

evolutionary biology↗

BiomarkerKB: FAIR and Integrated Biomarker Knowledge Connecting Biomolecular and Clinical Data Types

Biomarkers are essential tools for disease detection, risk assessment, therapeutic monitoring, and precision medicine. However, biomarker data are dispersed across heterogeneous resources, inconsistently reported in the literature, and rarely standardized for computational use. This fragmentation limits reproducibility, cross-study integration, and the discovery of novel biomarker and disease relationships. We developed BiomarkerKB, a knowledgebase designed to harmonize and integrate biomarker information under a standardized data model. The model follows the FDA-NIH BEST biomarker definition and captures both core fields (biomarker entity, disease/condition, exposure agent) and contextual metadata (specimen, biomarker role, evidence, provenance). Biomarker data and related annotations were either curated from publications or collected from public resources (e.g., OpenTargets, GWAS Catalog, ClinVar, CIViC, OncoMX) and were also contributed by the Common Fund Data Coordinating Centers and the Early Detection Research Network (EDRN). Standardization was achieved using ontologies and reference resources such as Disease Ontology, UBERON, UniProtKB, and HUGO Gene Nomenclature Committee (HGNC) gene symbols. BiomarkerKB data were ingested into a Neo4j-based knowledge graph and integrated with the Common Fund Data Ecosystem (CFDE) Knowledge Graph. The initial release of BiomarkerKB contains over 200,000 biomarker-disease associations spanning genes, proteins, metabolites, glycans, and chemical elements. The knowledge graph comprises more than 300,000 nodes and 1.2 million edges, enabling structured exploration of biomarker relationships within CFDE data as demonstrated through the knowledge graph query-based use cases presented in this study. A publicly accessible web portal (https://biomarkerkb.org) provides keyword search, filtering, data downloads, and access to graph visualization to support both researchers and computational analyses. BiomarkerKB addresses a critical gap in biomarker informatics by providing an integrated, FAIR (Findable, Accessible, Interoperable, and Reusable), and unified framework for biomarker knowledge exploration and discovery.

bioinformatics↗