Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.07.24.666385

Disease-associated genetic variants can cause mutations in tissue-specific protein isoforms

Abstract

Genetic variants can cause protein-coding mutations that result in disease. Variants are typically interpreted using the reference transcript for a gene. However, most human multi-exon genes encode alternative isoforms. Here, we show that coding exons in alternative isoforms harbour more population variants than exons of reference isoforms, consistent with their reduced evolutionary constraint, and that these variants are more likely to cause nonsynonymous coding mutations. Common and rare disease-associated variants mapping to alternative transcripts can lead to amino acid substitutions predicted to be structurally damaging in the corresponding protein isoform. The alternative transcripts to which disease-associated variants map demonstrate high tissue-specific expression, with many unannotated in reference human genomes, revealed only by long-read RNA-sequencing. As an example, we report an unannotated alternative transcript of the inflammasome regulator DPP9 that is lung epithelium-specific and which harbours a common genetic variant associated with severe COVID-19 and lung fibrosis. The variant causes a p.Leu8Pro missense mutation in an alternative first exon, predicted to disrupt the encoded alpha helix. These findings highlight the importance of considering alternative isoforms, their tissue-specific expression, and full-length transcripts in variant interpretation, with implications for uncovering underappreciated mechanisms of both common and rare disease.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Weykopf, G., Badonyi, M., Friman, E. T., Nguyen, J. M. H., Ioannou, A., Livesey, B. J., Coutts, A., Hird, E. F., Wham, M., Stanton, C., Vitart, V., Su, J., Murphy, L., Baillie, J. K., Gorrell, M. D., Marsh, J. A., Bickmore, W. A., Biddie, S. C.. 2025-07-26. Disease-associated genetic variants can cause mutations in tissue-specific protein isoforms. https://doi.org/10.1101/2025.07.24.666385

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

ssJSD: A fusion of sparsity and spatial information for HiC single-cell clustering

Single-cell high-throughput chromatin conformation capture (scHiC) enables profiling three dimensional genome architecture at cellular resolution, providing insights into cell-to-cell variability and cellular functions. Recent frameworks utilize spatial interaction patterns to derive dissimilarity measures for downstream tasks such as cell clustering. However, the inherent sparsity and ultra-high dimensionality of scHiC contact matrices pose significant challenges. A central hurdle is that existing measures typically treat all zeros without distinction, failing to differentiate biologically meaningful structural zeros (SZs) from technical dropouts. Here, we introduce ssJSD (spatial and sparsity informed Jensen-Shannon Divergence), a computational framework designed to explicitly account for scHiC-specific sparsity patterns. By integrating band-wise contact frequency profiles with SZ-induced sparsity matrices, ssJSD leverages both spatial interaction patterns and biological absence of contacts. We adopted two complementary integration strategies: an early fusion approach that concatenates information into a single representation, and a late fusion approach that integrates JSD-based dissimilarities through diverse averaging methods. Through simulations and applications to human cell lines and prefrontal cortex data, we demonstrate that ssJSD improves clustering accuracy and effectively distinguishes cell types. Our study indicates that integrating SZ patterns is important for accurately quantifying cell-to-cell variability in 3D genomics.

genomics↗

Interrogation of noncoding schizophrenia risk variants using CRISPR-based functional genomics

Schizophrenia (SCZ) is a highly heritable complex disorder influenced by coding and noncoding genetic variation. Its genetic causes, particularly those involving noncoding variation, are largely unknown. High-throughput CRISPR screens enable dissection of disease-associated loci and identification of noncoding regulatory elements and variants that modulate gene expression. We screened SCZ GWAS loci linked to genes that are also associated in whole-exome sequencing studies to identify regulatory elements and variants impacting expression of disease-relevant genes. We used CRISPRi paired with HCR-FlowFISH to epigenetically silence 333 putative regulatory elements and measure the downstream effects on gene expression of causal SCZ genes, FAM120A, SV2A, and STAG1, in iPSCs and iPSC-derived neurons (iNeurons). We identified 78 regulatory elements that significantly alter expression of a SCZ gene, including noncoding enhancers/silencers as well as promoters of genes and lncRNAs. Pooled prime editing screens interrogated noncoding variant influence on gene expression for SCZ-associated variants and uncharacterized common variants from diverse population studies. We find that a common variant in the promoter of SV2A, rs112851681:A>G (MAF = 3.56%, 1000 Genomes) enhances transcriptional activity in iPSCs and iNeurons. These findings show distinct noncoding mechanisms that map within GWAS signals, and provide a path forward for interrogating noncoding regulatory elements and variants in disease loci.

genomics↗

The automated eukaryotic pangenome pipeline EukPan reveals accessory genome differentiation beyond core-gene phylogeny in Aspergillus oryzae

Pangenome analysis reveals recurrent gene-content variation beyond a single reference genome, but its application to eukaryotes is constrained by inconsistent gene annotation. ANNEVO predicts gene models from genome FASTA assemblies without RNA-seq data. We developed EukPan, an automated post-annotation pipeline that standardizes GFF/GTF files, selects representative isoforms, constructs proteomes, infers orthogroups, builds a concatenated single-copy core-protein alignment, and summarizes shared accessory orthogroups while excluding orthogroups detected in only one genome. Applied with ANNEVO to 123 Aspergillus oryzae genomes, EukPan identified 11,245 core and 4,407 shared accessory orthogroups. The core-protein phylogeny broadly recovered the reported A-H classification, whereas accessory-genome analyses clearly separated the 33 group-A strains from the other 90 strains. Directional analysis identified 62 group-A-associated and 158 group-A-depleted orthogroups, with major facilitator superfamily (MFS) transporter and fungal Zn2Cys6 transcription-factor domains prominent in the depleted set. Among 93 orthogroups present in all non-A strains and absent from all group-A strains, 59 mapped to 10 segments of RIB40, the standard A. oryzae reference genome and a non-A (group-F) strain. EukPan therefore enables reproducible, coordinated core- and accessory-pangenome analysis from eukaryotic genome assemblies.

genomics↗