Search bioRxivSearch

Biology subjects

Ingraham, J. B.

Publications and source records attributed to Ingraham, J. B..

2 recordsLinked to original sources

The EVcouplings Python framework for coevolutionary sequence analysis

SummaryCoevolutionary sequence analysis has become a commonly used technique for de novo prediction of the structure and function of proteins, RNA, and protein complexes. This approach requires extensive computational pipelines that integrate multiple tools, databases, and data processing steps. We present the EVcouplings framework, a fully integrated open-source application and Python package for coevolutionary analysis. The framework enables generation of sequence alignments, calculation and evaluation of evolutionary couplings (ECs), and de novo prediction of structure and mutation effects. The application has an easy to use command line interface to run workflows with user control over all analysis parameters, while the underlying modular Python package allows interactive data analysis and rapid development of new workflows. Through this multi-layered approach, the EVcouplings framework makes the full power of coevolutionary analyses available to entry-level and advanced users.\n\nAvailabilityhttps://github.com/debbiemarkslab/evcouplings\n\nContactsander.research@gmail.com, debbie@hms.harvard.edu

bioinformatics

Deep generative models of genetic variation capture mutation effects

The functions of proteins and RNAs are determined by a myriad of interactions between their constituent residues, but most quantitative models of how molecular phenotype depends on genotype must approximate this by simple additive effects. While recent models have relaxed this constraint to also account for pairwise interactions, these approaches do not provide a tractable path towards modeling higher-order dependencies. Here, we show how latent variable models with nonlinear dependencies can be applied to capture beyond-pairwise constraints in biomolecules. We present a new probabilistic model for sequence families, DeepSequence, that can predict the effects of mutations across a variety of deep mutational scanning experiments significantly better than site independent or pairwise models that are based on the same evolutionary data. The model, learned in an unsupervised manner solely from sequence information, is grounded with biologically motivated priors, reveals latent organization of sequence families, and can be used to extrapolate to new parts of sequence space.

genetics