Search bioRxiv⌕ Search

Biology subjects

Boev, N.

Publications and source records attributed to Boev, N..

3 recordsLinked to original sources

Diversity and Genomic Organization of Non-B DNA Motifs in Haplotype-Resolved Human Genome Assemblies

Long-read sequencing and telomere-to-telomere genome assemblies now enable the exploration of previously inaccessible repetitive and structurally complex regions of the human genome. Using 130 haplotype-resolved genome assemblies from 65 individuals across diverse populations, we systematically analyzed six major classes of non-B DNA motifs. By evaluating their biophysical stability, we distinguished structurally stable motifs from those forming unstable secondary structures and mapped their distribution across individuals and genomic contexts. Our work revealed significant variation at the population level in both motif abundance and predicted structural stability, uncovering previously unrecognized diversity in non-B DNA landscapes. Non-B DNA motifs exhibit notable, structure-specific enrichment in highly repetitive and evolutionarily dynamic regions that remain largely unresolved in short-read-based genomes, including centromeres, segmental duplications, structural variant breakpoints, and mobile element insertions. Our findings provide a refined view of the potential secondary-structure organization within repetitive regions of the human genome and highlight structural stability as a key factor shaping the distribution of non-B DNA motifs in regions linked to genome instability, evolution, and human variation.

genomics↗

DNA shape and epigenomics distinguish the mechanistic origin of human genomic structural variations

The recent advent of long-read whole genome sequencing has enabled us to create an accurate telomere-to-telomere reference genome, construct pangenome graphs, and compile precise catalogs of genomic structural variations (SVs). These comprehensive SV repositories provide an excellent opportunity to explore the role of SVs in genotype-phenotype associations and examine the mechanisms by which SVs are introduced through double-strand break (DSB) repair. Here, we employed comprehensive SV catalogs identified through various short- and long-read whole genome sequencing efforts to infer the underlying mechanisms of SV introduction based on their genomic and epigenomic profiles. Our findings indicate that high local DNA methylation and DNA shape-related features, such as low variations in propeller twist, support the origins of homology-driven SVs. Subsequently, we utilized an active-learning-based unsupervised clustering approach, revealing that the homology-dependent SVs show greater evidence of retaining ancestral recombination patterns compared to their homology-independent counterparts. Finally, our comparison of inherited and de novo SVs from healthy populations and rare disease cohorts showed distinct upstream H3K27me3 levels in de novo SVs from individuals with ultra-rare disorders. These findings highlight genome-wide characteristics that may influence the choice of repair mechanisms linked to heritable SV origins.

genomics↗

MicroNucML: A machine learning approach for micronuclei segmentation and the refinement of nuclei-micronuclei relationships

Micronuclei (MN) are structures containing small fragments of DNA, arising from mitotic errors or failed DNA repair attempts. Therefore, MN serve as markers of genomic instability and are typically quantified either manually or through threshold-based methods, which can be tedious and inaccurate, leading to varying degrees of success and throughput. By employing a two-phase labeling approach that utilizes polygon and brush segmentation, along with refinement using SAM2, we developed a high-quality MN segmentation tool. Subsequent data augmentation, which captured heterogeneity in image quality and color diversity, enabled us to train a generalizable Mask-RCNN model optimized for small object detection, achieving state-of-the-art performance in MN detection. Finally, we applied our model to immunofluorescence data obtained from cell lines exposed to DNA damage conditions to gain biological insights into MN dynamics and their role in inducing genome instability. In summary, this work establishes an accessible resource for systematically studying genome instability with significantly greater fidelity and sensitivity, enabling insights into damage biology that were previously unresolved.

bioinformatics↗