Search bioRxiv⌕ Search

Biology subjects

Boldirev, G.

Publications and source records attributed to Boldirev, G..

2 recordsLinked to original sources

A comprehensive benchmark of discrepancies across microbial genome reference databases

Metagenomic analysis of microbial communities relies significantly on the quality and completeness of reference genomes, which allow researchers to compare sequencing reads against reference genome collections to reveal essential community characteristics. However, the reliability of these analyses is often compromised by substantial discrepancies across existing reference resources, including differences in genome content, assembly fragmentation, taxonomic representation, and metadata completeness. While these inconsistencies are known to introduce bias, the extent of divergence between major databases remains largely unknown. Here, we present a comprehensive benchmark of discrepancies across multiple widely used microbial genome reference resources. We developed the Cross-DB Genomic Comparator (CDGC), which utilizes reference genome alignments to systematically capture discrepancies in genome assemblies across reference databases. Applying this framework, we found that 99% of viral genomes were identical across databases, indicating strong consistency in viral reference resources. In contrast, fungal genomes showed substantially greater variability: although 82% of assemblies exhibited at least 90% similarity, only 7% were identical across databases. More concerning, we identified a subset of 461 assemblies with less than 50% similarity, suggesting the presence of technical artifacts, incomplete assemblies, or damaged genome files that require closer examination. Collectively, these results demonstrate that systematic cross-database benchmarking provides a critical mechanism for refining the accuracy of individual reference databases and advancing efforts towards more unified and reliable universal reference genomes.

bioinformatics↗

Leveraging a hybrid cross-disciplinary training model to accelerate global bioinformatics capacity

Disparities in formal bioinformatics training exacerbate the global skills gap, impeding the democratized application of advanced genomic technologies. To bridge this divide, we introduce a scalable, hybrid training framework designed to rapidly accelerate regional bioinformatics capacity. We exemplify this approach through the Eastern European Bioinformatics and Genomics (EEBG) workshop series -- a cross-disciplinary initiative that pairs international faculty with local institutions to deliver modular, hands-on curricula. Functioning as a structured knowledge-transfer pipeline, the series has catalyzed a sustainable educational ecosystem, evidenced by the establishment of multiple independent summer schools across the region. The assessment of the 2025 EEBG workshop in Krakow, Poland, validates the models viability; participant metrics confirm high efficacy in skill acquisition (mean satisfaction: 4.4/5.0) and community building. Crucially, the hybrid delivery mode dismantled geographic barriers, serving as a vital mechanism for maintaining scientific continuity for researchers facing displacement and crisis. Synthesizing these outcomes, we define the core features of a replicable blueprint for scientific readiness in resource-constrained environments. We conclude by presenting a strategic roadmap -- organized around infrastructure standardization, governance sustainability, and geographical expansion -- for adapting this regional proof-of-concept into a global export-ready model, offering a critical path toward ensuring universal access to genomic innovation.

scientific communication and education↗