Search bioRxiv⌕ Search

Biology subjects

Pozo, F.

Publications and source records attributed to Pozo, F..

2 recordsLinked to original sources

APPRIS principal isoforms and MANE Select transcripts in clinical variant interpretation

Most coding genes are able to generate multiple alternatively spliced transcripts. Determining which of these transcript variants produces the main protein isoform, and which of a genes multiple splice variants are functionally important, is crucial in comparative genomics and essential for clinical variant interpretation. Here we show that the principal isoforms chosen by APPRIS and the MANE Select variants provide the best approximations of the main cellular protein isoforms. Principal isoforms are predicted from conservation and from protein features, and MANE transcripts are chosen from the consensus between teams of expert manual curators. APPRIS principal isoforms coincide in over 94% of coding genes with MANE Select transcripts and the two methods are particularly discriminating when they agree on the main splice variant. Where the two methods agree, the splice variants coincide with the main isoform detected in proteomics experiments in 98.2% of genes with multiple protein isoforms. We also find that almost all ClinVar pathogenic mutations map to MANE Select or APPRIS principal isoforms. Where APPRIS and MANE agree on the main isoform, 99.93% of validated pathogenic variants map to principal rather than alternative exons. MANE Plus Clinical transcripts cover most validated pathogenic mutations in alternative coding exons. TRIFID functional importance scores are particularly useful for distinguishing clinically important alternative isoforms: the highest scoring TRIFID isoforms are more than 300 times more likely to have validated pathogenic mutations. We find that APPRIS, MANE and TRIFID are important for determining the biological relevance of splice isoforms and should be an essential part of clinical variant interpretation.

genomics↗

Phylodynamics of SARS-CoV-2 transmission in Spain

ObjectivesSARS-CoV-2 whole-genome analysis has identified three large clades spreading worldwide, designated G, V and S. This study aims to analyze the diffusion of SARS-CoV-2 in Spain/Europe. MethodsMaximum likelihood phylogenetic and Bayesian phylodynamic analyses have been performed to estimate the most probable temporal and geographic origin of different phylogenetic clusters and the diffusion pathways of SARS-CoV-2. ResultsPhylogenetic analyses of the first 28 SARS-CoV-2 whole genome sequences obtained from patients in Spain revealed that most of them are distributed in G and S clades (13 sequences in each) with the remaining two sequences branching in the V clade. Eleven of the Spanish viruses of the S clade and six of the G clade grouped in two different monophyletic clusters (S-Spain and G-Spain, respectively), with the S-Spain cluster also comprising 8 sequences from 6 other countries from Europe and the Americas. The most recent common ancestor (MRCA) of the SARS-CoV-2 pandemic was estimated in the city of Wuhan, China, around November 24, 2019, with a 95% highest posterior density (HPD) interval from October 30-December 17, 2019. The origin of S-Spain and G-Spain clusters were estimated in Spain around February 14 and 18, 2020, respectively, with a possible ancestry of S-Spain in Shanghai. ConclusionsMultiple SARS-CoV-2 introductions have been detected in Spain and at least two resulted in the emergence of locally transmitted clusters, with further dissemination of one of them to at least 6 other countries. These results highlight the extraordinary potential of SARS-CoV-2 for rapid and widespread geographic dissemination.

microbiology↗