Search bioRxiv⌕ Search

Biology subjects

Smeds, L.

Publications and source records attributed to Smeds, L..

5 recordsLinked to original sources

Variation and selection at predicted G-quadruplexes across the human pangenome

G-quadruplexes (G4s), non-canonical DNA structures whose sequence motifs occupy approximately 1% of the human genome, are important for myriad cellular functions, including regulating transcription and replication. Yet they also contribute to genomic instability by increasing mutations and structural variation. Despite their significance, G4 motifs have not been studied in detail across multiple human genomes. Here, we conducted a comprehensive analysis of presence/absence and sequence variation, measured selection strength, and evaluated gene expression regulation potential for predicted G4s (pG4s) across population groups in the second release of the Human Pangenome Reference Consortium dataset, comprising high-quality, near-telomere-to-telomere diploid genomes from 231 individuals worldwide, along with three reference assemblies. Across the human pangenome, we identified over 353 million pG4s, including 1.15 million pG4s absent from reference assemblies but shared across other haplotypes. Our analysis revealed that pG4 sharing patterns recapitulate human population structure: African individuals displayed lower levels of pG4 sharing than non-Africans, whereas East Asian individuals exhibited higher levels of sharing. By analyzing the site frequency spectrum across various genomic annotations, we computed and compared selection coefficients (Sd) at pG4 vs. non-pG4 sites. As expected, the strongest purifying selection (Sd [&ge;] 10) was detected at protein-coding exons, where pG4 sites had similar or lower selection coefficients compared with those for pG4 sites. Strikingly, this pattern reversed at regulatory regions: although purifying selection was weaker overall at promoters, introns, enhancers, and replication origins (1 [&le;] Sd < 10), pG4 sites at these regions experienced stronger selection than non-pG4 sites--suggesting that pG4s play functional roles outside coding sequences. Additionally, by integrating pG4 data with long-read transcriptome data profiles from this large cohort, we found that pG4s located at promoters and at (or near) exon-intron junctions may influence variation in gene expression levels and transcript isoforms, respectively, across the human pangenome individuals. Leveraging extensive population-scale data, our research illuminates the fundamental importance and functional relevance of G4s across human genomes.

genomics↗

Comparative analysis of single-stranded and non-canonical DNA formation in human and other ape cells with telomere-to-telomere genomes

Non-canonical (non-B) DNA secondary structures, e.g., G-quadruplexes and triplex DNA, are mutation hotspots and genome regulators contributing to disease and evolution. Yet they remain uncharacterized in complete genomes in vivo. Here we exploited the fact that many non-B DNA structures form single-stranded DNA (ssDNA). Using permanganate/S1 footprinting across 14 cell lines, we generated ssDNA profiles for human and six non-human ape telomere-to-telomere (T2T) genomes. Newly resolved satellite arrays--e.g., at ribosomal DNA and centromeres--displayed high ssDNA levels, implicating non-B DNA in satellite expansion and function. Hidden Markov Models applied to our ssDNA data revealed active genomic domains with specific functions--e.g., replication, transcription, or recombination--each enriched in particular non-B DNA types. Human-specific ssDNA domains correlated with nervous system genes, whereas cancer and embryonic cells showed increased ssDNA in transposable elements. Our ssDNA analysis across ape T2T genomes uncovered conserved and species-specific DNA structural dynamics central to genome regulation.

genomics↗

Non-canonical DNA in bird telomere-to-telomere genomes

Non-canonical (non-B) DNA motifs are sequences that can fold into structures (e.g., G-quadruplexes and Z-DNA) distinct from the canonical right-handed helix. In mammals, these structures regulate gene expression, act as mutation hotspots, and are associated with cancer, yet they remain undercharacterized in other species. Because non-B DNA motifs are difficult to sequence, many are absent from incomplete genome assemblies, limiting functional analyses. Here, we present the first comprehensive analysis of non-B DNA motifs in birds, using the telomere-to-telomere genome of zebra finch, the near-complete chicken genome, and high-quality genomes of six additional bird species. We show that, first, unlike in mammals, the non-B DNA landscape in birds differs markedly among chromosome groups: gene-rich and extremely small dot chromosomes show the highest coverage (15.1-30.1% in zebra finch), microchromosomes--intermediate coverage (6.4-18.1%), and macrochromosomes--the lowest (5.9-6.9%). Non-B DNA motif coverage on dot chromosomes negatively correlates with PacBio sequencing depth, potentially explaining their assembly challenges. Second, similar to mammals, in zebra finch, G-quadruplexes are enriched at promoters and 5'UTRs, implying regulatory roles. We experimentally validated four common G-quadruplexes and predicted others using long-read methylation data. Overall, non-B DNA distribution reflects distinct features of avian genome architecture, suggests a role in gene regulation, and informs strategies for complete bird genome sequencing.

genomics↗

The complete genome of a songbird

Bird genomes are the smallest among amniotes, but remain challenging to assemble due to their structural complexity. This study presents the first fully phased, diploid, telomere-to-telomere (T2T) reference genome for the zebra finch (Taeniopygia guttata), a model organism for neuroscience and evolutionary genomics. Combining multiple sequencing strategies resulted in closing nearly all gaps, adding [~]90 Mbp of previously missing sequence (7.8%). This includes T2T assemblies for all microchromosomes, including dot chromosomes, and the previously almost entirely missing chr16. The T2T genome is comprehensively annotated for genes, repeats, structural variants, and long-read methylation calls. Complete centromeric structures were assembled and annotated along with kinetochore binding sites. Relative to the previous high-quality reference of the Vertebrate Genomes Project, 2,778 (8.51%) previously unassembled or unannotated genes were identified, of which 9% overlap with segmental duplications. This first complete genome of a songbird, now the new public reference, illuminates avian genome architecture and function.

genomics↗

Non-canonical DNA in human and other ape telomere-to-telomere genomes

Non-canonical (non-B) DNA structures--e.g., bent DNA, hairpins, G-quadruplexes (G4s), Z-DNA, etc.--which form at certain sequence motifs (e.g., A-phased repeats, inverted repeats, etc.), have emerged as important regulators of cellular processes and drivers of genome evolution. Yet, they have been understudied due to their repetitive nature and potentially inaccurate sequences generated with short-read technologies. Here we comprehensively characterize such motifs in the long-read telomere-to-telomere (T2T) genomes of human, bonobo, chimpanzee, gorilla, Bornean orangutan, Sumatran orangutan, and siamang. Non-B DNA motifs are enriched at the genomic regions added to T2T assemblies, and occupy 9-15%, 9-11%, and 12-38% of autosomes, and chromosomes X and Y, respectively. G4s and Z-DNA are enriched at promoters and enhancers, as well as at origins of replication. Repetitive sequences harbor more non-B DNA motifs than non-repetitive sequences, especially in the short arms of acrocentric chromosomes. Most centromeres and/or their flanking regions are enriched in at least one non-B DNA motif type, consistent with a potential role of non-B structures in determining centromeres. Our results highlight the uneven distribution of predicted non-B DNA structures across ape genomes and suggest their novel functions in previously inaccessible genomic regions. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=92 SRC="FIGDIR/small/610891v3_ufig1.gif" ALT="Figure 1"> View larger version (18K): org.highwire.dtl.DTLVardef@1a359forg.highwire.dtl.DTLVardef@b66b36org.highwire.dtl.DTLVardef@38dae8org.highwire.dtl.DTLVardef@abdf84_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗