Search bioRxivSearch

EXPLORE THE ARCHIVE

Genomics

Find preprints about genomes, sequencing and genetic variation.

11 recordsLinked to original sources

Evolutionary origins of protein novelty across an entire yeast subphylum

Novel protein-coding sequences fuel molecular and cellular evolutionary innovations and frequently contribute to species-defining characteristics. They can originate either de novo from previously noncoding sequences or through extreme divergence of already coding ones. How frequently each mechanism occurs and how they shape the structural and functional potential of the resulting proteins remains unclear. Here, we conducted a broad computational investigation of genetic and protein novelty throughout the entire subphylum of Saccharomycotina yeasts. We detected more than 5,000 robust de novo genes across 332 species and compared them to more than 6,000 novel genes resulting from extreme sequence divergence, revealing two quantitatively similar but qualitatively distinct modes of evolution of novelty. A remarkable 40% of de novo proteins are predicted to localize to mitochondria compared to only 20% of divergent, with the latter also being substantially longer and more disordered. A detailed analysis of conservatively predicted tertiary structures of novel proteins shows that ''invention'' of new folds occurs more frequently through de novo emergence. We also illustrate cases of evolutionary ''re-invention'' of existing protein folds from noncoding sequences. Our work deepens our understanding of the origins and importance of novel proteins, opening new directions for further structural and functional characterization.

genomics

Microsecond molecular dynamics of SOD1 variants suggest a structural basis for divergent ALS clinical outcomes

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease characterised by progressive motor neuron degeneration. Mutations in the SOD1 gene represent the second most common genetic cause of ALS (ALS), and distinct SOD1 missense variants present with markedly different clinical profiles. A4V leads to an aggressive form of the disease (median survival [~]1y), H46R confers a mild, slowly progressive course and I113T exhibits an intermediate phenotype. The molecular basis by which these mutations produce divergent clinical outcomes remains poorly understood. We performed extensive classical molecular dynamics simulations of wild-type SOD1 and the three ALS-associated variants in the apo monomeric state to attempt to investigate the mechanisms behind such phenotypic differences. Structural stability, global compactness, and conformational flexibility, as well as analysis of collective motions between residues and estimation of free energy, were assessed. The H46R, A4V, and I113T variants exhibited distinct dynamic behaviours, highlighting differences in structural stability, local flexibility, and intramolecular interactions. These findings suggest that specific structural regions may contribute differently to protein dysfunction and could represent key elements for understanding the relationship between molecular dynamic properties and the differing clinical severity associated with these variants. Most strikingly, H46R exhibited exceptional structural stability across every analytical level, the lowest global deviation, most attenuated local flexibility, strongest internal dynamic coordination, and the deepest, most confined free energy basins of any system examined. This convergent multi-layered evidence of structural restraint provides a compelling mechanistic basis for the mild and slowly progressive clinical course of H46R ALS, suggesting that enhanced conformational rigidity, rather than bulk destabilisation, is the defining biophysical feature of this variant, and that its pathogenic mechanism operates through a route fundamentally decoupled from the aggregation-driven toxicity that characterises the more aggressive SOD1-ALS mutations.

genomics

PhageTAILor leverages machine learning for phage tail-like elements detection and classification in plant-associated bacteria

Phage tail-like elements (PTEs) -- tailocins, bacterial type VI secretion systems (T6SS), and extracellular contractile injection systems (eCIS) -- are contractile nanomachines that bacteria use to kill their neighbors and compete within their micro-ecosystems. PTEs help shape microbial community composition. Most PTE detection tools only detect a single PTE class. Moreover, most tailocin detection methods are largely restricted to Pseudomonas, leaving a key part of tailocin diversity uncharacterized. In this work, we present PhageTAILor (https://github.com/hjcho-bio/PhageTAILor), an integrative and fully automated pipeline that detects and classifies prophages and 3 PTE classes from bacterial genomes. PhageTAILor combines a 6-detector homology-based candidate search (geNomad, tail-gene, PHROGs-tail, SecReT6, eCIStem, and a divergence-tolerant tail-HMM detector) with a LightGBM classifier comprising 1 multiclass and 3 binary heads, trained on 6,501 bacterial genomes carrying 13,082 prophages and PTEs. A phylogeny-free feature matrix used in our model keeps predictions reproducible between model construction and user inference. PhageTAILor performs strongly at the genome level and generalizes beyond its Pseudomonas-rich training set. On a 76-strain cross-clade benchmark, PhageTAILor detected tailocins at F1 = 0.955. Furthermore, it identified 12 of 13 experimentally validated tailocins spanning five genera versus 2 of 13 for a Pseudomonas-restricted tool TattleTail. PhageTAILor also demonstrated sensitivity equivalent to viral detection tool geNomad while avoiding its higher false-positive rate. Applied to 7,925 plant- and soil-associated bacterial isolates, PhageTAILor showed that prophages in the phyllosphere and tailocins in plant-associated bacteria, whereas eCIS are enriched in soil. PhageTAILor is distributed as an open-source, modular pipeline with a command-line interface.

microbiology

AmPair: automating housekeeping-gene primer design for species-level metataxonomics

Amplicon sequencing of the 16S rRNA gene is the most widely used approach for profiling bacterial communities, but its taxonomic resolution is typically limited to the genus level. Many species carry multiple divergent 16S rRNA alleles that overlap across species boundaries, an ambiguity that even full-length, long-read sequencing cannot fully resolve. Shotgun metagenomics achieves species-level resolution but remains costly, particularly when only a single genus is of interest. Amplicon sequencing of rapidly evolving, protein-coding housekeeping genes offers a cost-effective alternative, yet no tool exists to identify suitable primer sets for a given target taxon. Here we present AmPair, a Snakemake pipeline that, given a target genus and one or more candidate housekeeping genes, designs and ranks primer pairs binding conserved regions while flanking a variable region capable of species-level discrimination, and validates them in silico across all available genomes. Using the genus Bacillus and the housekeeping gene tuf as a case study, the primer set recommended by AmPair amplified 99% of 2,392 genomes; only 0.04% carried multiple alleles and none showed inter-species allele overlap, compared with 91.41% and 69.49%, respectively, for the standard 16S rRNA V1-V9 region. Applied to a Bacillus community profiled by Nanopore sequencing, the same primers resolved closely related species. AmPair thus offers a generalizable and accessible route to species-level community profiling.

bioinformatics

A replicated patient-specific component of tumour telomere length across two pan-cancer cohorts

Bulk telomere length measured from tumour sequencing is routinely interpreted as a property of the cancer cells. However, a tumour specimen is a mixture, and the patient who supplies it has a telomere length of their own. Here I re-analyse published pan-cancer telomere estimates and ask how much of a tumour's telomere length is patient-specific. A calibration step comes first. Whole-genome and low-pass estimates recover the known cross-sectional attrition of leukocyte telomeres with age, at 26.6 bp per year in blood normals, whereas whole-exome estimates do not. After adjustment for cancer type, sequencing centre and sex, the exome slope is minus 0.6 bp per year. In 684 blood-normal aliquots sequenced by both assays, the whole-genome estimate declines at 38.9 bp per year, whereas the exome estimate from the same DNA shows no detectable decline. The difference between assays is 41.5 bp per year, with P = 3 x 10^-10. Because exome data constitute 78.6% of the original resource, downstream analyses use only whole-genome and low-pass libraries. Within those data, tumour telomere length tracks the patient's matched-normal telomere length. The Spearman correlation is 0.395 in TCGA, with positive associations in 22 of 23 cancer types. This finding replicates in PCAWG using a different telomere estimator, with a correlation of 0.472 and positive associations in all 24 histologies examined. Adjustment for cancer type, sequencing centre and library type leaves a regression coefficient of 0.385. The association is also stable after adjustment for age, sex, tumour purity, leukocyte fraction, ploidy, sequencing coverage and continental ancestry, with coefficients ranging from 0.406 to 0.429. Pure normal-cell admixture is rejected as the sole explanation. Under a two-compartment mixture model, the coefficient for host telomere length is expected to equal 1 and the host-by-purity interaction to equal minus 1. These restrictions are jointly rejected with P = 0.001. Tumour purity, leukocyte fraction and age each explain only about 1 to 3% of within-cohort variance and do not alter the cross-cancer ranking. By contrast, the between-cohort coefficient is not directly interpretable. Its apparent near one-to-one relationship with tissue-associated telomere length depends strongly on which tissue supplies the matched-normal reference and on the statistical spread of that predictor, falling to 0.44 when organ-matched solid tissue is used. Bulk tumour telomere length is therefore a composite phenotype containing a replicated patient-specific component. Telomere biomarker studies should include matched-normal telomere length as a covariate rather than treating tumour telomere length as exclusively tumour-intrinsic.

cancer biology

OMICON: a community resource for studying gene coexpression networks in normal and neoplastic human brain samples

Genome-wide coexpression analysis of intact tissue samples is a powerful approach for identifying reproducible signatures of cell types and states, since it can survey vast numbers of individuals, cells, and transcripts. However, it can be difficult to optimize gene coexpression network construction and compare results from independent analyses. To address these challenges, we developed OMICON (theomicon.ucsf.edu) for research on human brain gene coexpression networks. OMICON contains gene expression data from >17K normal and neoplastic human brain samples with standardized metadata. Systematic analysis of independent datasets identified >250K gene coexpression modules, which were characterized and compared via enrichment analysis with >40K gene sets. All modules are discoverable via an advanced search engine that can filter by genes, metadata, and enrichment results. Analyses can also be browsed with an interactive workflow visualization tool, and users can communicate within OMICON using @mention functionality to support communal research on human brain gene coexpression networks.

neuroscience

Distinct functions of Nup93 paralogs in tumor growth and Polycomb-mediated repression of JAK/STAT signaling

Nuclear pore complexes (NPCs) are nuclear envelope (NE)-embedded protein assemblies that mediate nucleocytoplasmic exchange and interact with the genome, including binding of an NPC component Nup93 to Polycomb chromatin domains. Here, we investigated the in vivo relevance of this relationship in Drosophila, which unusually contains two distinct paralogs of Nup93. Interestingly, we identified a Nup93-2-specific tumorigenic phenotype in larval wings, where depletion of Nup93-2, but not Nup93-1, led to tumor-like overgrowth, reminiscent of Polycomb mutations. Consistently, our transcriptomic analysis revealed a wide-spread loss of gene silencing in Nup93-2-depleted wings, particularly in a Nup93-bound Polycomb domain spanning genes for activators of JAK/STAT signaling. Nup93 paralogs were not found to differ in their effect on NPC biogenesis but strikingly, showed differences in subnuclear localization patterns. While Nup93-1 co-localized exclusively with fully assembled NPCs, Nup93-2 exhibited only partial co-localization and was found at additional NE locations in a tissue-specific manner. Together, our results identify an in vivo silencing role of a Nup93 paralog and suggest that Nup93-2 may form a unique NE-associated complex that targets a subset of Polycomb domains containing growth-promoting genes.

developmental biology

Cross-Kingdom Control: Yeast Prion Protein Modulates Host Physiology in Drosophila

Prions, once mainly studied for their pathogenic roles, are now gaining recognition as adaptive elements in microbial physiology. Over one-third of wild yeast isolates harbor prion proteins, yet their impact on host-microbe interactions remains poorly characterized. Given the ecological dominance of yeasts in the Drosophila mycobiome, we leveraged the Drosophila melanogaster-Saccharomyces cerevisiae system to investigate how the mycobiome-derived prion, [MRPL10+], modulates host physiology. We show that flies exposed to [MRPL10+] yeast exhibit significantly enhanced cold tolerance and increased locomotor activity. This effect persists with heat-killed yeast and diluted culture, suggesting a stable, potent bioactive factor. Using the genetically diverse Drosophila Global Diversity Lines (GDL), we identified natural variation in responsiveness to [MRPL10+] yeast. Genome-wide association and functional RNAi screening revealed a gut-brain signaling axis involving genes critical for digestion, intercellular communication, transcription regulation, and neural transmission. Notably, serotonin and octopamine pathways were essential for [MRPL10+]-induced changes in cold tolerance and locomotion, implicating neuromodulatory circuits in prion-mediated microbial signaling. Our findings establish a mechanistic link between a fungal prion and host metabolic and neural adaptation. This work provides the first genetic dissection of a prion-mediated host-microbe interaction, laying the groundwork for investigating beneficial prions in complex microbial communities and highlighting a new dimension of the mycobiomes influence on animal physiology.

evolutionary biology

The interaction between NC(p7)1-55 and p6 may regulate interactions with nucleic acids during assembly through modulation of Gag folding.

We present the solution structures of HIV-1 proteins NC(p7)1-55 corresponding to the full-length NC(p7) and mature p6. The studies were carried in water and, to mimic the membrane, in micellar DPC (Dodecylphosphocholine) conditions. Our results unravel for the first time the structure adopted by the N-terminal amino acids of the free NC(p7)1-55, with the formation of a small helix spanning residues F6 to R10. Our NMR and Fluorescence Anisotropy data disclose an interaction between NC(p7)1-55 and p6 both in water and DPC, with respective Kd of 2.5mM and 370 mM at 23{degrees}C. The interaction is thus strengthened in lipidic conditions. Protein p6 stabilizes the N-terminus of NC(p7)1-55 while increasing at the same time the dynamic of the first zinc finger. Although the entire p6 sequence is involved in the interaction, we show that its C-terminal region is particularly sensitive to the presence of NC(p7)1-55, with a propensity of forming a a helix ranging from amino acids S111 to F116. This study brings experimental evidence of a direct protein-protein interaction between p6 and the N-terminal region of NC(p7)1-55. We further show that such interaction is readily accommodated within the NC(p15) framework and hypothesize that it may facilitate the selective assembly of assembly of the viral genomic RNA (gRNA) in the cell.

biophysics

An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

bioinformatics

Comparative Transcriptional Responses of Human Blood to Neutron and Photon Irradiation

Despite the well-known health risks of neutron exposures, key gaps remain in understanding neutron-induced molecular responses and identifying reliable biodosimetric markers that distinguish neutrons from photon exposure. We provide the first genome-wide analysis of the human blood transcriptional response to an accelerator-derived fission-like spectrum of neutrons versus photons, evaluating transcriptomic relative biological effectiveness (RBE) and radiation quality-discriminating gene signatures. Whole blood from healthy donors was irradiated ex vivo with X-rays (140 kV, 0-4 Gy, n = 3) or neutrons (0.1-8 MeV, 0-1 Gy, n = 2), incubated for 6 h or 24 h, and processed for RNA sequencing from peripheral blood mononuclear cells (PBMCs). Neutrons were markedly more potent than X-rays at inducing differentially expressed genes (DEGs) at equal doses, showing a peak response 6 h post-irradiation followed by a decline. In contrast, X-rays caused a continuous increase in DEGs up to 24 h (neutrons vs. X-rays at 1 Gy: 1,449 vs. 121 DEGs at 6 h; 996 vs. 621 DEGs at 24 h). A universal p53-centered 34-gene signature, including FDXR, EDA2R, GADD45A, and ZMAT3, showed highly monotonic dose responses (Spearman correlation coefficient {approx} 1) across donors, radiation qualities, and timepoints. Additionally, difference-in-differences analysis identified radiation quality-discriminating genes only at 6 h, with transcriptional convergence observed by 24 h, suggesting a very narrow time window for biodosimetric differentiation. We identified a neutron-specific gene signature driven by cGAS-STING-NF-{kappa}B signaling (RELB, NFKB1, C3, MALAT1) and suppression of B-cell and myeloid identity genes (IGHD, TCL1A, CLEC7A, TLR2), defining a biologically coherent neutron quality index with distinct immunomodulatory effects. For the first time, we assessed neutron RBEs at the gene, pathway, and global transcriptomic levels in a human blood model, reporting a global transcriptomic neutron RBE of 1.30 (95% CI: 1.14-1.49) at 6 h and 1.21 (95% CI: 1.14-1.28) at 24 h, providing a valuable basis for biodosimetry in mixed-field exposure scenarios. Our findings advance the mechanistic understanding of neutron radiation responses and support the development of biodosimetric approaches for mixed-field exposure scenarios.

biophysics