Search bioRxiv⌕ Search

Biology subjects

Zehr, J. D.

Publications and source records attributed to Zehr, J. D..

6 recordsLinked to original sources

Natural selection differences detected in key protein domains between non-pathogenic and pathogenic Feline Coronavirus phenotypes

Feline Coronaviruses (FCoVs) commonly cause mild enteric infections in felines worldwide (termed Feline Enteric Coronavirus [FECV]), with around 12% developing into deadly Feline Infectious Peritonitis (FIP; Feline Infectious Peritonitis Virus [FIPV]). Genomic differences between FECV and FIPV have been reported, yet the putative genotypic basis of the highly pathogenic phenotype remains unclear. Here, we used state-of-the-art molecular evolutionary genetic statistical techniques to identify and compare differences in natural selection pressure between FECV and FIPV sequences, as well as to identify FIPV and FECV specific signals of positive selection. We analyzed full length FCoV protein coding genes thought to contain mutations associated with FIPV (Spike, ORF3abc, and ORF7ab). We identified two sites exhibiting differences in natural selection pressure between FECV and FIPV: one within the S1/S2 furin cleavage site, and the other within the fusion domain of Spike. We also found 15 sites subject to positive selection associated with FIPV within Spike, 11 of which have not previously been suggested as possibly relevant to FIP development. These sites fall within Spike protein subdomains that participate in host cell receptor interaction, immune evasion, tropism shifts, host cellular entry, and viral escape. There were 14 sites (12 novel) within Spike under positive selection associated with the FECV phenotype, almost exclusively within the S1/S2 furin cleavage site and adjacent C domain, along with a signal of relaxed selection in FIPV relative to FECV, suggesting that furin cleavage functionality may not be needed for FIPV. Positive selection inferred in ORF7b was associated with the FECV phenotype, and included 24 positively selected sites, while ORF7b had signals of relaxed selection in FIPV. We found evidence of positive selection in ORF3c in FCoV wide analyses, but no specific association with the FIPV or FECV phenotype. We hypothesize that some combination of mutations in FECV may contribute to FIP development, and that is unlikely to be one singular "switch" mutational event. This work expands our understanding of the complexities of FIP development and provides insights into how evolutionary forces may alter pathogenesis in coronavirus genomes.

evolutionary biology↗

Evolutionary shortcuts via multi-nucleotide substitutions and their impact on natural selection analyses.

Inference and interpretation of evolutionary processes, in particular of the types and targets of natural selection affecting coding sequences, are critically influenced by the assumptions built into statistical models and tests. If certain aspects of the substitution process (even when they are not of direct interest) are presumed absent or are modeled with too crude of a simplification, estimates of key model parameters can become biased, often systematically, and lead to poor statistical performance. Previous work established that failing to accommodate multi-nucleotide (or multi-hit, MH) substitutions strongly biases dN/dS-based inference towards false positive inferences of diversifying episodic selection, as does failing to model variation in the rate of synonymous substitution (SRV) among sites. Here we develop an integrated analytical framework and software tools to simultaneously incorporate these sources of evolutionary complexity into selection analyses. We found that both MH and SRV are ubiquitous in empirical alignments, and incorporating them has a strong effect on whether or not positive selection is detected, (1.4-fold reduction) and on the distributions of inferred evolutionary rates. With simulation studies, we show that this effect is not attributable to reduced statistical power caused by using a more complex model. After a detailed examination of 21 benchmark alignments and a new high-resolution analysis showing which parts of the alignment provide support for positive selection, we show that MH substitutions occurring along shorter branches in the tree explain a significant fraction of discrepant results in selection detection. Our results add to the growing body of literature which examines decadesold modeling assumptions (including MH) and finds them to be problematic for comparative genomic data analysis. Because multi-nucleotide substitutions have a significant impact on natural selection detection even at the level of an entire gene, we recommend that selection analyses of this type consider their inclusion as a matter of routine. To facilitate this procedure, we developed, implemented, and benchmarked a simple and well-performing model testing selection detection framework able to screen an alignment for positive selection with two biologically important confounding processes: site-to-site synonymous rate variation, and multi-nucleotide instantaneous substitutions.

bioinformatics↗

Human HspB1, HspB3, HspB5 and HspB8: Shaping these Disease Factors during Vertebrate Evolution

Small heat shock proteins (sHSPs) emerged early in evolution and occur in all domains of life and nearly in all species, including humans. Mutations in four sHSPs (HspB1, HspB3, HspB5, HspB8) are associated with neuromuscular disorders. The aim of this study is to investigate the evolutionary forces shaping these sHSPs during vertebrate evolution. We performed comparative evolutionary analyses on a set of orthologous sHSP sequences, based on the ratio of non-synonymous: synonymous substitution rates for each codon. We found that these sHSPs had been historically exposed to different degrees of purifying selection, decreasing in this order: HspB8 > HspB1, HspB5 > HspB3. Within each sHSP, regions with different degrees of purifying selection can be discerned, resulting in characteristic selective pressure profiles. The conserved -crystallin domains were exposed to the most stringent purifying selection compared to the flanking regions, supporting a dimorphic pattern of evolution. Thus, during vertebrate evolution the different sequence partitions were exposed to different and measurable degrees of selective pressures. Among the disease-associated mutations, most are missense mutations primarily in HspB1 and to a minor extent in the other sHSPs. Our data provide an explanation for this disparate incidence. Contrary to the expectation, most missense mutations cause dominant disease phenotypes. Theoretical considerations support a connection between the historic exposure of these sHSP genes to a high degree of purifying selection and the unusual prevalence of genetic dominance of the associated disease phenotypes. Our study puts the genetics of inheritable sHSP-borne diseases into the context of vertebrate evolution.

evolutionary biology↗

RASCL: Rapid Assessment Of SARS-CoV-2 Clades Through Molecular Sequence Analysis

An important component of efforts to manage the ongoing COVID19 pandemic is the Rapid Assessment of how natural selection contributes to the emergence and proliferation of potentially dangerous SARS-CoV-2 lineages and CLades (RASCL). The RASCL pipeline enables continuous comparative phylogenetics-based selection analyses of rapidly growing clade-focused genome surveillance datasets, such as those produced following the initial detection of potentially dangerous variants. From such datasets RASCL automatically generates down-sampled codon alignments of individual genes/ORFs containing contextualizing background reference sequences, analyzes these with a battery of selection tests, and outputs results as both machine readable JSON files, and interactive notebook-based visualizations. AvailabilityRASCL is available from a dedicated repository at https://github.com/veg/RASCL and as a Galaxy workflow https://usegalaxy.eu/u/hyphy/w/rascl. Existing clade/variant analysis results are available here: https://observablehq.com/@aglucaci/rascl. ContactDr. Sergei L Kosakovsky Pond (spond@temple.edu). Supplementary informationN/A

bioinformatics↗

Conserved recombination patterns across coronavirus subgenera

Recombination contributes to the genetic diversity found in coronaviruses and is known to be a prominent mechanism whereby they evolve. It is apparent, both from controlled experiments and in genome sequences sampled from nature, that patterns of recombination in coronaviruses are non-random and that this is likely attributable to a combination of sequence features that favour the occurrence of recombination breakpoints at specific genomic sites, and selection disfavouring the survival of recombinants within which favourable intra-genome interactions have been disrupted. Here we leverage available whole-genome sequence data for six coronavirus subgenera to identify specific patterns of recombination that are conserved between multiple subgenera and then identify the likely factors that underlie these conserved patterns. Specifically, we confirm the non-randomness of recombination breakpoints across all six tested coronavirus subgenera, locate conserved recombination hot- and cold-spots, and determine that the locations of transcriptional regulatory sequences are likely major determinants of conserved recombination breakpoint hot-spot locations. We find that while the locations of recombination breakpoints are not uniformly associated with degrees of nucleotide sequence conservation, they display significant tendencies in multiple coronavirus subgenera to occur in low guanine-cytosine content genome regions, in non-coding regions, at the edges of genes, and at sites within the Spike gene that are predicted to be minimally disruptive of Spike protein folding. While it is apparent that sequence features such as transcriptional regulatory sequences are likely major determinants of where the template-switching events that yield recombination breakpoints most commonly occur, it is evident that selection against misfolded recombinant proteins also strongly impacts observable recombination breakpoint distributions in coronavirus genomes sampled from nature.

bioinformatics↗

Recent zoonotic spillover and tropism shift of a Canine Coronavirus is associated with relaxed selection and putative loss of function in NTD subdomain of spike protein.

A recent study reported the occurrence of Canine Coronavirus (CCoV) in nasopharyngeal swabs from a small number of patients hospitalized with pneumonia during a 2017-18 period in Sarawak, Malaysia. Because the genome sequence for one of these isolates is available, we conducted comparative evolutionary analyses of the spike gene of this strain (CCoV-HuPn-2018), with other available Alphacoronavirus 1 spike sequences. The most N-terminus subdomain (0-domain) of the CCoV-HuPn-2018 spike protein has sequence similarity to Transmissible Gastroenteritis Virus (TGEV) and CCoV2b strains, but not to other members of the type II Alphacoronaviruses (i.e., CCoV2a and Feline CoV2-FCoV2). This 0-domain in CCoV-HuPn-2018 has evidence for relaxed selection pressure, an increased rate of molecular evolution, and a number of unique amino acid substitutions relative to CCoV2b and TGEV sequences. A region of the 0-domain determined to be key to sialic acid binding and pathogenesis in TGEV had clear differences in amino acid sequences in CCoV-HuPn-2018 relative to both CCoV2b (enteric) and TGEV (enteric and respiratory). The 0-domain of CCoV-HuPn-2018 also had several sites inferred to be under positive diversifying selection, including sites within the signal peptide. Downstream of the 0-domain, FCoV2 shared sequence similarity to the CCoV2b and TGEV sequences, with analyses of this larger alignment identifying positively selected sites in the putative Receptor Binding Domain (RBD) and Connector Domain (CD). Recombination analyses strongly implicated a particular FCoV2 strain in the recombinant history of CCoV-HuPn-2018 with molecular divergence times estimated at around 60 years ago. We hypothesize that CCoV-HuPn-2018 had an enteric origin, but that it has lost that particular tropism, because of mutations in the sialic acid binding region of the spike 0-domain. As selection pressure on this region was reduced, the virus evolved a respiratory tropism, analogous to other Alphacoronavirus 1, such as Porcine Respiratory Coronavirus (PRCV), that have lost this region entirely. We also suggest that signals of positive selection in the signal peptide as well as other changes in the 0-domain of CCoV-HuPn-2018 could represent an adaptive role in this new host and that this could be in part due to the different spatial distribution of the N-linked glycan repertoire for this strain.

evolutionary biology↗