Search bioRxiv⌕ Search

Biology subjects

DeMontigny, W. C.

Publications and source records attributed to DeMontigny, W. C..

2 recordsLinked to original sources

OmegaSwitch: Bayesian Markov-Modulated Codon Models for Estimating dN/dS

Selective pressures can vary across both sites and evolutionary lineages; however, most codon models accommodate heterogeneity along only one of these dimensions and require the number of selective regimes to be specified in advance. Here, we introduce OmegaSwitch, a Bayesian phylogenetic software framework for inferring changes in the nonsynonymous-to-synonymous substitution-rate ratio (dN/dS) across sites and through evolutionary time. We implement a Markov-modulated codon model in which lineages transition among discrete dN/dS regimes and use reversible-jump Markov chain Monte Carlo to infer the number of regimes simultaneously. We further develop a Dirichlet-process mixture extension that allows the parameters governing these time-heterogeneous processes to vary among sites. Ancestral sampling produces joint posterior distributions of dN/dS across sites and nodes of the phylogeny, enabling lineage- and site-specific summaries with quantified uncertainty. Simulation analyses showed that both the posterior intervals for dN/dS and the number of evolutionary regimes were well calibrated under both models. We demonstrate OmegaSwitch using vertebrate alpha- and beta-globins. OmegaSwitch therefore provides a flexible Bayesian framework for investigating how selective pressures vary across protein-coding sequences and phylogenetic history.

evolutionary biology↗

Inferring Gene Presence in Incomplete Data via Phylogenetic Occupancy Modeling

Increasing access to genomic data has revolutionized our understanding of biology. Organisms that were previously unculturable or otherwise difficult to study have been investigated using metagenomic sequencing and bioinformatic assemblies, illuminating biological diversity that was previously invisible. However, as the availability of genomic data has grown, so has the challenge posed by incomplete genomes. Many genomes obtained from metagenomic assemblies or mixed cultures are of poor quality and establishing fully complete genomes requires substantial effort. Incomplete genomes pose difficulties for several common analyses, particularly gene-inventory and core-genome analyses. When genomes are incomplete, distinguishing true gene absence from non-detection becomes difficult. For relatively complete genomes, gene absences are inferred to be true absences, but for highly incomplete genomes, researchers often exclude such data entirely. Probabilistic models have attempted to address this issue in core genome analyses. In particular, mO-TUpan utilizes an iterative algorithm to infer genome completeness and categorize genes as core or accessory. In this work, we substantially improve upon this approach by integrating a well-established class of ecological models, called occupancy models, with evolutionary modeling. Our "phylogenetic occupancy model" defines a probability distribution over gene presence that accounts for shared information across related genomes. This framework simultaneously estimates genome completeness and the probability that a gene is present but unobserved. This model substantially outperforms competing methods for core genome inference and enables inference of single-gene presence/absence and ancestral-state reconstruction. Alongside this paper, we provide our model as a Python package.

evolutionary biology↗