Search bioRxiv⌕ Search

Biology subjects

Potapova, T. A.

Publications and source records attributed to Potapova, T. A..

8 recordsLinked to original sources

KaryoScope: rapid, alignment-free sequence annotation for the pangenome era

The pangenome era is producing long-read sequencing data and complete genome assemblies (1-3) at a pace that current annotation methods cannot match. Existing tools were each built for a single feature class (repeats, centromeric satellites, or genes) and falter precisely where the genome is most variable and harbours clinically important variation: the centromeres, subtelomeres, and acrocentric short arms. Here we present KaryoScope, an alignment-free method to annotate an assembly at base-pair resolution across any desired feature classes in a single pass, completing in minutes on a standard workstation. Applied to the Human Pangenome Reference Consortium Release 2 assemblies (3), KaryoScope identifies the SST1 macrosatellite as the recurrent sequence at Robertsonian translocation fusion points (4, 5), delivers the first pangenome-wide census of D4Z4 macrosatellite structural diversity at the 4q and 10q subtelomeres relevant to facioscapulohumeral muscular dystrophy (6), and reveals previously uncharacterised centromere structural polymorphism, including chromosome-specific satellite loss and megabase-scale rearrangement validated by fluorescence in situ hybridization. A pre-built KaryoScope database for the human genome is distributed alongside the tool, and additional databases can be built for any reference genome or annotation source. Together, these capabilities bring the most variable regions of the genome within reach for comparative, clinical, and pangenome-scale analysis. KaryoScope is available at https://github.com/barthel-lab/KaryoScope.

bioinformatics↗

Haplotype-resolved centromeric chromatin organization from a complete diploid human genome

Centromeres ensure proper chromosome segregation during cell division, yet the organization and regulation of centromeric chromatin within satellite DNA arrays remain incompletely understood. Here, we leverage the complete diploid human genome benchmark (T2T-HG002) to provide a detailed study of centromeric sequence and chromatin architecture on individual haplotypes. Using adaptive-sampling-enriched, ultra-long-read DiMeLo-seq, we achieve single-molecule chromatin profiling across all centromeres, revealing that along single chromatin fibers, CENP-A, the histone variant specifying centromere identity, forms multiple discrete subdomains within hypomethylated centromere dip regions (CDRs) that are flanked by H3K9me3-enriched heterochromatin. Despite underlying sequence variation, CDRs localize to sequence-homogeneous domains and maintain relatively balanced CENP-A dosage and aggregate length across all chromosomes and between haplotypes. Further, we show that bidirectional changes to centromeric and pericentromeric DNA methylation are accompanied by changes to centromeric chromatin architecture. In passaged cells with centromeric hypomethylation, subdomain boundaries are eroded, and adjacent CENP-A domains tend to merge and expand. Conversely, in pluripotent stem cells with centromeric hypermethylation, CDRs are fundamentally reorganized, such that discrete hypomethylated domains are frequently consolidated into broader contiguous tracts. These methylation-associated CDR restructuring events suggest that DNA methylation acts as a principal regulator of human centromere organization, with implications for understanding centromere plasticity, epigenetic inheritance, and chromosomal instability in development and disease.

genomics↗

A Complete Genome for the Common Marmoset

The common marmoset is a New World monkey (NWM) commonly used as a model organism to investigate questions in primate evolution and human disease, including Alzheimers and other neurodegenerative diseases, as well as neuropsychiatric disorders. Here we present the first telomere-to-telomere (T2T) reference genome for the common marmoset, adding over 88 Mb of sequence and resolving challenging genomic regions. An additional near-T2T assembly from a second unrelated individual yields a total of four high-quality haplotypes for analysis. The improved contiguity and accuracy of these assemblies enable unprecedented insights into complex and rapidly evolving genomic regions such as centromeres, sex chromosomes, ribosomal DNA (rDNA) structure, and the major histocompatibility complex (MHC). We fully resolved all marmoset centromeres, uncovering dimeric alpha satellites with chromosomal specificity and stratified inactive layers documenting ancestral centromere turnover. We assembled six acrocentric autosomes with gene-poor, satellite-rich short arms and provide evidence that most of them can harbor rDNA and all of them share large pseudo-homologous regions (PHRs). The Y chromosome, but not the X chromosome, carries active rDNA and PHRs, and the rDNA copy number is sexually dimorphic. Chromosomes that share PHRs also share closely related centromeric satellite DNA, supporting a model of ongoing recombinational exchange between heterologous chromosomes facilitated by rDNA. We discovered multiple novel, marmoset-specific MHC genes that are predicted to protect against pathogens encountered in its environment. Leveraging this complete reference, we further identified over 500 transcribed genes with transcript models or expansions specific to the marmoset lineage. Together with additional long-read marmoset assemblies, these genomes were used to construct a marmoset pangenome, providing a robust reference framework for short-read mapping across diverse individuals. This resource will improve the utility of the common marmoset as a biomedical model organism and fill key gaps in our understanding of primate evolution.

genomics↗

Origin and evolution of acrocentric chromosomes in human and great apes

The short arms of human acrocentric chromosomes are characterized by nucleolar organizer regions essential for ribosome biogenesis, but their highly repetitive nature has hindered genomic analysis. Leveraging the recently completed genomes of all major ape lineages, we identified recurrent features of their acrocentrics, including enriched repeat classes, centromere repositioning by whole-arm inversion, interchromosomal sequence exchange, and birth-and-death evolution of multiple gene families. Together, these processes have enabled the repeated amplification and diversification of the FRG1 gene family over 25 million years of ape evolution, and, in gorilla, the formation and amplification of a novel IGSF3-GGT fusion gene under positive selection. Similar evolutionary events also explain the distribution of segmental duplications and heterochromatin in the modern human genome, predisposing it to karyotypic abnormalities such as Robertsonian translocations. Our findings highlight acrocentric chromosomes as key drivers of evolution in the great apes, with implications for speciation, adaptation, and clinical genomics.

genomics↗

Complete genomes of a multi-generational pedigree to expand studies of genetic and epigenetic inheritance

Pedigree analysis remains the gold standard for rare disease diagnostics, yet whole genome sequencing studies typically omit critical regions like centromeres, telomeres, and acrocentric chromosome p-arms. Here, we present telomere-to-telomere (T2T) reference genomes for four self-identified African American individuals of admixed ancestry spanning three generations. Our parent-of-origin assigned, chromosome-level assemblies revealed precise meiotic recombination breakpoints in previously inaccessible regions, including recombination events across acrocentric and subtelomeric sequences. Centromeric regions were highly stable, with multi-megabase arrays inherited intact across three generations, while the position of kinetochore assembly sites remained consistent and predominantly associated with the p-arm proximal region. The relative lengths of telomeres on individual chromosomes were maintained across generations. Using a targeted rDNA assembly approach, we reconstructed a complete megabase-scale ribosomal DNA (rDNA) array corresponding to the paternal chromosome 14. This openly available pedigree provides a benchmark dataset for studying recombination and genetic and epigenetic variation across the complete genome.

genomics↗

The formation and propagation of human Robertsonian chromosomes

Robertsonian chromosomes are a type of variant chromosome found commonly in nature. Present in one in 800 humans, these chromosomes can underlie infertility, trisomies, and increased cancer incidence. Recognized cytogenetically for more than a century, their origins have remained mysterious. Recent advances in genomics allowed us to assemble three human Robertsonian chromosomes completely. We identify a common breakpoint and epigenetic changes in centromeres that provide insight into the formation and propagation of common Robertsonian translocations. Further investigation of the assembled genomes of chimpanzee and bonobo highlights the structural features of the human genome that uniquely enable the specific crossover event that creates these chromosomes. Resolving the structure and epigenetic features of human Robertsonian chromosomes at a molecular level paves the way to understanding how chromosomal structural variation occurs more generally, and how chromosomes evolve.

genomics↗

Epigenetic control and inheritance of rDNA arrays

Ribosomal RNA (rRNA) genes exist in multiple copies arranged in tandem arrays known as ribosomal DNA (rDNA). The total number of gene copies is variable, and the mechanisms buffering this copy number variation remain unresolved. We surveyed the number, distribution, and activity of rDNA arrays at the level of individual chromosomes across multiple human and primate genomes. Each individual possessed a unique fingerprint of copy number distribution and activity of rDNA arrays. In some cases, entire rDNA arrays were transcriptionally silent. Silent rDNA arrays showed reduced association with the nucleolus and decreased interchromosomal interactions, indicating that the nucleolar organizer function of rDNA depends on transcriptional activity. Methyl-sequencing of flow-sorted chromosomes, combined with long read sequencing, showed epigenetic modification of rDNA promoter and coding region by DNA methylation. Silent arrays were in a closed chromatin state, as indicated by the accessibility profiles derived from Fiber-seq. Removing DNA methylation restored the transcriptional activity of silent arrays. Array activity status remained stable through the iPS cell re-programming. Family trio analysis demonstrated that the inactive rDNA haplotype can be traced to one of the parental genomes, suggesting that the epigenetic state of rDNA arrays may be heritable. We propose that the dosage of rRNA genes is epigenetically regulated by DNA methylation, and these methylation patterns specify nucleolar organizer function and can propagate transgenerationally.

genomics↗

Distinct states of nucleolar stress induced by anti-cancer drugs

Ribosome biogenesis is a vital and energy-consuming cellular function occurring primarily in the nucleolus. Cancer cells have an especially high demand for ribosomes to sustain continuous proliferation. This study evaluated the impact of existing anticancer drugs on the nucleolus by screening a library of anticancer compounds for drugs that induce nucleolar stress. For a readout, a novel parameter termed "nucleolar normality score" was developed that measures the ratio of the fibrillar center and granular component proteins in the nucleolus and nucleoplasm. Multiple classes of drugs were found to induce nucleolar stress, including DNA intercalators, inhibitors of mTOR/PI3K, heat shock proteins, proteasome, and cyclin-dependent kinases (CDKs). Each class of drugs induced morphologically and molecularly distinct states of nucleolar stress accompanied by changes in nucleolar biophysical properties. In-depth characterization focused on the nucleolar stress induced by inhibition of transcriptional CDKs, particularly CDK9, the main CDK that regulates RNA Pol II. Multiple CDK substrates were identified in the nucleolus, including RNA Pol I - recruiting protein Treacle, which was phosphorylated by CDK9 in vitro. These results revealed a concerted regulation of RNA Pol I and Pol II by transcriptional CDKs. Our findings exposed many classes of chemotherapy compounds that are capable of inducing nucleolar stress, and we recommend considering this in anticancer drug development. Types of nucleolar stresses identified in this study O_FIG O_LINKSMALLFIG WIDTH=193 HEIGHT=200 SRC="FIGDIR/small/517150v3_ufig1.gif" ALT="Figure 1"> View larger version (80K): org.highwire.dtl.DTLVardef@18724b6org.highwire.dtl.DTLVardef@17b60e7org.highwire.dtl.DTLVardef@117004dorg.highwire.dtl.DTLVardef@114ecec_HPS_FORMAT_FIGEXP M_FIG (1) DNA intercalators and RNA Pol inhibitors induced canonical nucleolar stress manifested by partial dispersion of granular component (GC) and segregation of rDNA and fibrillar center (FC) components UBF, Treacle, and POLR1A within nucleolar stress caps. (2) Inhibition of mTOR and PI3K growth pathways induced a metabolic suppression of function accompanied by the decrease in nucleolar normality score, size, and rRNA production, without dramatic re-organization of nucleolar anatomy. (3) Inhibitors targeting HSP90 and proteasome induced proteotoxicity, resulting in the disruption of protein homeostasis and the accumulation of misfolded and/or undegraded proteins. These effects were accompanied by a decrease in nucleolar normality score, rRNA output, and in some cases formation of protein aggregates (aggresomes) inside the nucleolus. (4) Inhibition of transcriptional CDK activity led to the disruption of interactions between rDNA, RNA Pol I, and GC proteins. This resulted in almost complete nucleolar dissolution, leaving behind an extended bare rDNA scaffold with only a few associated FC proteins remaining. UBF and PolI-recruiting protein Treacle remained associated with the rDNA, while POLR1A and GC dispersed in the nucleoplasm. rRNA production ceased and the nucleolar normality score was greatly reduced. C_FIG

cell biology↗