Search bioRxiv⌕ Search

Biology subjects

Vasquez, K. M.

Publications and source records attributed to Vasquez, K. M..

5 recordsLinked to original sources

Landscape and mutational dynamics of G-quadruplexes in the complete human genome and in haplotypes of diverse ancestry

G-quadruplexes (G4s) are alternative DNA structures with diverse biological roles, but their examination in highly repetitive parts of the human genome has been hindered by the lack of reliable sequencing technologies. Recent long-read based genome assemblies have enabled their characterization in previously inaccessible parts of the human genome. Here, we examine the topography and genomic instability of potential G4-forming sequences in the gap-less, reference human genome assembly and in 88 haplotypes of diverse ancestry. We report that G4s are highly enriched in specific repetitive regions, including in certain centromeric and pericentromeric repeat types, and in ribosomal DNA arrays, and experimentally validate the most prevalent G4s detected. G4s tend to have lower methylation than expected throughout the human genome and are genomically unstable, showing an excess of all mutation types, including substitutions, insertions and deletions and most prominently structural variants. Finally, we show that G4s are consistently enriched at PRDM9 binding sites, a protein involved in meiotic recombination. Together, our findings establish G4s as dynamic and functionally significant elements of the human genome and highlight new avenues for investigating their contributions to human disease and evolution.

genomics↗

ZSeeker: An optimized algorithm for Z-DNA detection in genomic sequences

Z-DNA is an alternative left-handed helical form of DNA with a zigzag-shaped backbone that differs from the right-handed canonical B-DNA helix. Z-DNA has been implicated in various biological processes, including transcription, replication, and DNA repair, and can induce genetic instability. Repetitive sequences of alternating purines and pyrimidines have the potential to adopt Z-DNA structures. ZSeeker is a novel computational tool developed for the accurate detection of potential Z-DNA-forming sequences in genomes, addressing limitations of prior methods. By introducing a novel methodology informed and validated by experimental data, ZSeeker enables the refined detection of potential Z-DNA-forming sequences. Built both as a standalone Python package and as an accessible web interface, ZSeeker allows users to input genomic sequences, adjust detection parameters, and view potential Z-DNA sequence distributions and Z-scores via downloadable visualizations. Our Web Platform provides a no-code solution for Z-DNA identification, with a focus on accessibility, user-friendliness, speed and customizability. By providing efficient, high-throughput analysis and enhanced detection accuracy, ZSeeker has the potential to support significant advancements in understanding the roles of Z-DNA in normal cellular functions, genetic instability, and its implications in human diseases. AvailabilityZSeeker is released as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/ZSeeker. A web-interface of ZSeeker is publicly available at https://zseeker.netlify.app/.

bioinformatics↗

Characterization of hairpin loops and cruciforms across 118,065 genomes spanning the tree of life

Inverted repeats (IRs) can form alternative DNA secondary structures called hairpins and cruciforms, which have a multitude of functional roles and have been associated with genomic instability. However, their prevalence across diverse organismal genomes remains only partially understood. Here, we examine the prevalence of IRs across 118,065 complete organismal genomes. Our comprehensive analysis across taxonomic subdivisions reveals significant differences in the distribution, frequency, and biophysical properties of perfect IRs among these genomes. We identify a total of 29,589,132 perfect IRs and show a highly variable density across different organisms, with strikingly distinct patterns observed in Viruses, Bacteria, Archaea, and Eukaryota. We report IRs with perfect arms of extreme lengths, which can extend to hundreds of thousands of base pairs. Our findings demonstrate a strong correlation between IR density and genome size, revealing that Viruses and Bacteria possess the highest density, whereas Eukaryota and Archaea exhibit the lowest relative to their genome size. Additionally, the study reveals the enrichment of IRs at transcription start and termination end sites in prokaryotes and Viruses and underscores their potential roles in gene regulation and genome organization. Through a comprehensive overview of the distribution and characteristics of IRs in a wide array of organisms, this largest-scale analysis to date sheds light on the functional significance of inverted repeats, their contribution to genomic instability, and their evolutionary impact across the tree of life.

genomics↗

Quadrupia: Derivation of G-quadruplexes for organismal genomes across the tree of life

G-quadruplex DNA structures exhibit a profound influence on essential biological processes, including transcription, replication, telomere maintenance, and genomic stability. These structures have demonstrably shaped organismal evolution. However, a comprehensive, organism-wide G-quadruplex map encompassing the diversity of life has remained elusive. Here, we introduce Quadrupia, the most extensive and well-characterized G-quadruplex database to date, facilitating the exploration of G-quadruplex structures across the evolutionary spectrum. Quadrupia has identified G-quadruplex sequences in 108,449 reference genomes, with a total of 140,181,277 G-quadruplexes. The database also hosts a collection of 319,784 G-quadruplex clusters of 20 or more members, annotated by taxonomic distributions, multiple sequence alignments, profile Hidden Markov Models and cross-references to G-quadruplex 3D structures. Examination of G-quadruplexes across functional genomic elements in different taxa indicates preferential orientation and positioning, with significant differences between individual taxonomic groups. For example, we find that G-quadruplexes in bacteria with a single replication origin display profound preference for the leading orientation. Finally, we experimentally validate the most frequently observed G-quadruplexes using CD-spectroscopy, UV melting, and fluorescent-based approaches. Quadrupia is publicly available through https://www.pavlopoulos-lab.org/quadrupia.

genomics↗

MoCoLo: a testing framework for motif co-localization

Sequence-level data offers insights into biological processes through the interaction of two or more genomic features from the same or different molecular data types. Within motifs, this interaction is often explored via the co-occurrence of feature genomic tracks using fixed-segments or analytical tests that respectively require window size determination and risk of false positives from over-simplified models. Moreover, methods for robustly examining the co-localization of genomic features, and thereby understanding their spatial interaction, have been elusive. We present a new analytical method for examining feature interaction by introducing the notion of reciprocal co-occurrence, define statistics to estimate it, and hypotheses to test for it. Our approach leverages conditional motif co-occurrence events between features to infer their co-localization. Using reverse conditional probabilities and introducing a novel simulation approach that retains motif properties (e.g., length, guanine-content), our method further accounts for potential confounders in testing. As a proof-of-concept, MoCoLo confirmed the co-occurrence of histone markers in a breast cancer cell line. As a novel analysis, MoCoLo identified significant co-localization of oxidative DNA damage within non-B DNA forming regions that significantly differed between non-B DNA structures. Altogether, these findings demonstrate the potential utility of MoCoLo for testing spatial interactions between genomic features via their co-localization.

bioinformatics↗