Search bioRxiv⌕ Search

Biology subjects

Georgakopoulos Soares, I.

Publications and source records attributed to Georgakopoulos Soares, I..

6 recordsLinked to original sources

Simultaneous epigenomic profiling and regulatory activity measurement using e2MPRA

Cis-regulatory elements (CREs) have a major effect on phenotypes including disease. They are identified in a genome-wide manner by analyzing the binding of transcription factors (TFs), various co-factors and histone modifications in DNA using assays such as ChIP-seq, Cut&Tag and ATAC-seq. However, these assays are descriptive and require high-throughput technologies, such as massively parallel reporter assays (MPRAs), to test the functional activity and variant effect on these sequences. Currently, technologies that can simultaneously analyze both the regulatory function of a specific sequence and the TFs, cofactors and epigenomic modifications that determine it do not exist. Here, we developed enrichment followed by epigenomic profiling MPRA (e2MPRA), a novel technology that utilizes lentivirus-based MPRA to enrich for the integration of specific CREs into the genome followed by Cut&Tag or ATAC-seq targeted specifically for these sequences. This method allows to simultaneously analyze in a high-throughput manner regulatory activity, protein binding and epigenetic modification of thousands of candidate CREs and their variants. We demonstrate that e2MPRA can be used to dissect the epigenetic functions of TF motifs arranged in synthetic enhancers, as well as to analyze the effect of enhancer sequence variants on epigenetic modifications. In summary, this technology will increase our understanding of the regulatory code, its effect on the epigenome and how its alteration can lead to a variety of phenotypes including human disease.

genomics↗

Landscape and mutational dynamics of G-quadruplexes in the complete human genome and in haplotypes of diverse ancestry

G-quadruplexes (G4s) are alternative DNA structures with diverse biological roles, but their examination in highly repetitive parts of the human genome has been hindered by the lack of reliable sequencing technologies. Recent long-read based genome assemblies have enabled their characterization in previously inaccessible parts of the human genome. Here, we examine the topography and genomic instability of potential G4-forming sequences in the gap-less, reference human genome assembly and in 88 haplotypes of diverse ancestry. We report that G4s are highly enriched in specific repetitive regions, including in certain centromeric and pericentromeric repeat types, and in ribosomal DNA arrays, and experimentally validate the most prevalent G4s detected. G4s tend to have lower methylation than expected throughout the human genome and are genomically unstable, showing an excess of all mutation types, including substitutions, insertions and deletions and most prominently structural variants. Finally, we show that G4s are consistently enriched at PRDM9 binding sites, a protein involved in meiotic recombination. Together, our findings establish G4s as dynamic and functionally significant elements of the human genome and highlight new avenues for investigating their contributions to human disease and evolution.

genomics↗

metagRoot: A comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71,091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, Hidden Markov Models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org

bioinformatics↗

Unraveling diversity by isolating peptide sequences specific to distinct taxonomic groups

The identification of succinct, universal fingerprints that enable the characterization of individual taxonomies can reveal insights into trait development and can have widespread applications in pathogen diagnostics, human healthcare, ecology and the characterization of biomes. Here, we investigated the existence of peptide k-mer sequences that are exclusively present in a specific taxonomy and absent in every other taxonomic level, termed taxonomic quasi-primes. By analyzing proteomes across 24,073 species, we identified quasi-prime peptides specific to superkingdoms, kingdoms, and phyla, uncovering their taxonomic distributions and functional relevance. These peptides exhibit remarkable sequence uniqueness at six- and seven-amino- acid lengths, offering insights into evolutionary divergence and lineage-specific adaptations. Moreover, we show that human quasi-prime loci are more prone to harboring pathogenic variants, underscoring their functional significance. This study introduces taxonomic quasi-primes and offers insights into their contributions to proteomic diversity, evolutionary pathways, and functional adaptations across the tree of life, while emphasizing their potential impact on human health and disease.

bioinformatics↗

Ribosomal DNA arrays are the most H-DNA rich element in the human genome

Repetitive DNA sequences can form non-canonical structures such as H-DNA which is an intramolecular triplex DNA structure. The new Telomere-to-Telomere (T2T) genome assembly for the human genome has eliminated gaps, enabling the examination of highly repetitive regions including centromeric and pericentromeric repeats and ribosomal DNA arrays. This gapless assembly allows for the examination of the distribution of H-DNA sequences in parts of the human genome that were not previously annotated. We find that H-DNA appears once every 30,000 bps in the human genome. Its distribution is highly inhomogeneous with H-DNA motif hotspots being detectable in acrocentric chromosomes. Ribosomal DNA arrays in acrocentric chromosomes are the genomic element with the highest H-DNA enrichment, with 13.22% of total H-DNA motifs being found in ribosomal DNA arrays, representing a 42.65-fold enrichment over what would be expected by chance. Across the acrocentric chromosomes we report that 55.87% of all H-DNA motifs found in these chromosomes are in rDNA array loci. The H-DNA motifs are primarily found in the intergenic spacer regions of the ribosomal DNA arrays, generating repeated clusters. We also discover that binding sites for PRDM9, a protein that regulates the formation of double-strand breaks and determines the meiotic recombination hotspots in humans and most mammals, are over 5-fold enriched for H-DNA motifs. Finally, we provide evidence that our findings are consistent in other non-human great ape genomes. We conclude that ribosomal DNA arrays are the most enriched genomic loci for H-DNA sequences in human and other great ape genomes.

genomics↗

Alternative splicing modulation by G-quadruplexes

Alternative splicing is central to metazoan gene regulation but the regulatory mechanisms involved are only partially understood. Here, we show that G-quadruplex (G4) motifs are enriched ~3-fold both upstream and downstream of splice junctions. Analysis of in vitro G4-seq data corroborates their formation potential. G4s display the highest enrichment at weaker splice sites, which are frequently involved in alternative splicing events. The importance of G4s in RNA as supposed to DNA is emphasized by a higher enrichment for the non-template strand. To explore if G4s are involved in dynamic alternative splicing responses, we analyzed RNA-seq data from mouse and human neuronal cells treated with potassium chloride. We find that G4s are enriched at exons which were skipped following potassium ion treatment. We validate the formation of stable G4s for three candidate splice sites by circular dichroism spectroscopy, UV-melting and fluorescence measurements. Finally, we explore G4 motifs across eleven representative species, and we observe that strong enrichment at splice sites is restricted to mammals and birds.

genomics↗