Search bioRxivSearch

bioRxiv · 10.1101/2020.08.28.272245

Centromeric repeats of the Western European house mouse I: high sequence diversity among monomers at local and global spatial scales

Abstract

Previous work found that the centromeric repeats of the Western European house mouse (Mus musculus domesticus) are composed predominantly of a 120 bp monomer that is shared by the X and autosomes. Polymorphism in length and sequence was also reported. Here I quantified the length and sequence polymorphism of the centromeric repeats found on the X and autosomes. The levels of local and global sequence variation were also compared. I found three length variants: a 64mer, 112mer and 120mer with relative frequencies of 2.4%, 8.6%, and 89%, respectively. There was substantial sequence variation within all three length variants with a rank-order of: 64mer < 120mer < 112mer. The 64mer was never found alone on long Sanger traces, and was arranged predominantly as a 176 bp higher-order repeat composed of a 64/112mer dimer. Reanalysis of archived ChIP-seq reads found that all three length variants were enriched with the foundational centromere protein CENP-A, but the enrichment was far higher for the 120mer. This pattern indicates that only the 120mer contributes substantially to the functional centromeres, i.e., to the kinetochore-binding, centric cores of the centromeric repeat arrays. Despite only moderate sequence divergence among random pairs of 120mers (averaging 5.9%), other measures of sequence diversity were exceptionally high: i) variant richness (numerical diversity) -on average, one new sequence variant was observed every 4th additional monomer randomly sampled (in N = 7.2 x 103 monomers), and ii) variant evenness -all of the nearly 2 x 103 observed sequence variants were at low frequency, with the most common variant having a frequency of only 5.7%. I next used long Sanger trace data from the Mouse Genome Project to assess the pattern of monomer diversity among neighboring 120mers. Unexpectedly, side-by-side monomers were rarely identical in sequence, and sequence divergence between these neighbors was nearly as high as that between random pairs taken from the genome-wide pool of all 120mers. I also used long Sanger traces to determine sequence variation among neighborhoods of 5 contiguous 120 bp monomers. Sequence diversity within these small regions typically spanned most of the entire range of that found genome-wide. Despite high sequence variation within these neighborhoods, the density of monomers with functional binding motifs for CENP-B (i.e., b-boxes with sequence NTTCGNNNNANNCGGGN) was strongly conserved at about 50%. The overarching pattern of monomer structure at the centromeric repeats of this subspecies is: i) high homogeneity in the density CENP-B binding sites, and ii) high heterogeneity in monomer sequence at both local and global levels.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rice, W. R.. 2020-08-29. Centromeric repeats of the Western European house mouse I: high sequence diversity among monomers at local and global spatial scales. https://doi.org/10.1101/2020.08.28.272245

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Geometry of antigenic evolution improves influenza vaccine selection

Anticipating antigenic evolution is essential for selecting effective seasonal influenza A/H3N2 vaccine strains. To this end, we integrated hemagglutination-inhibition and neutralization titers spanning 2002 to 2025 into a unified Bayesian antigenic map. The map resolves twelve antigenic clusters advancing in discrete steps, with several clusters co-circulating in most seasons. In 15 of 21 seasons, the WHO-recommended vaccine belonged to an earlier cluster than the dominant circulating cluster. The direction of each vaccine update relative to recent viral drift predicted vaccine effectiveness one season ahead in out-of-sample forecasts. Antigenic distance, the conventional measure of vaccine-virus match, was weakly associated with effectiveness until update direction was accounted for. Retrospectively ranking candidate strains by predicted effectiveness would have selected a strain predicted to outperform the WHO recommendation in every season, raising mean predicted effectiveness by 10 percentage points.

evolutionary biology

Evolutionary replay of duplicate-gene retention across independent whole-genome duplications

Whole-genome duplications repeatedly expose ancestral gene lineages to the same broad evolutionary outcome-retention or loss of duplicated copies-but it remains unclear whether this history replays similarly across evolutionary scales. We placed duplicate retention in shared hierarchical orthologous-group coordinates and compared percentile ranks defined within each event-wide mapped universe. Three independent angiosperm whole-genome duplications showed reproducible replay (global rank effect T-replay = 0.210, bootstrap 95% confidence interval 0.172-0.248; permutation P = 1/100,001). A plant reference-panel score specified before target outcomes were examined predicted retention after the Apple/Pear duplication ({rho} = 0.169, n = 373). Deep transfer was heterogeneous: the teleost-genome-duplication estimate was positive but unresolved ({rho} = 0.107, n = 151, 95% confidence interval -0.050 to 0.260), whereas transfer to the ancient budding-yeast whole-genome duplication (yeast WGD) was supported ({rho} = 0.280, n = 186). Independently reconstructed animal outcomes also replayed between teleost and Stylommatophora duplications (r = 0.226, n = 146, P = 0.00326), although the effect remained below a prespecified strong-effect threshold. A strict plant-animal comparison was limited to 25 deeply one-to-one lineages and was unresolved (r = 0.033, 95% confidence interval -0.303 to 0.340). Thus, ancestral gene-lineage identity contributes reproducibly to duplicate retention after independent whole-genome duplications, but replay is structured by evolutionary lineage and modified by event-specific history rather than governed by one universal gene-fate ranking.

evolutionary biology

A Hymenoptera-restricted gene mediating ant castes co-opts deeply conserved machinery to control organ size

Lineage-specific genes are widespread and have been implicated as phenotypic innovation inducers, but how they acquire complex developmental functions remains poorly understood. Ant queens and workers develop dramatically different organ sizes from identical genomes under juvenile hormone (JH) control, yet the molecular effectors translating JH signalling into caste-specific organ growth remain unknown. Here we identify torch, a Hymenoptera-restricted gene, as the most consistently gyne-biased and JH-responsive gene across 68 ant species. Knockdown of torch in virgin queens of Monomorium pharaonis produces a worker-like, multi-organ growth-restricted phenotype. Mechanistically, torch harbours an E-box-like motif activated by the JH receptor Gce-Tai and acts as a GA-repeat-binding transcription factor that regulates Hippo signalling, the deeply conserved organ-size control pathway in animals. Expressing torch heterologously in mice and a growth-restricted Drosophila background shows that the gene retained its general growth-promoting activity across more than 700 million years of animal evolution in lineages that lack the gene, establishing that its function is mediated through conserved rather than ant-specific machinery. A lineage-specific gene can therefore acquire complex morphogenetic function by co-opting ancient organ-size circuitry, providing a general route by which novel genes can drive phenotypic innovation.

evolutionary biology