Search bioRxivSearch

Biology subjects

Beauchamp, K. A.

Publications and source records attributed to Beauchamp, K. A..

8 recordsLinked to original sources

The Dynamic Conformational Landscapes of the Protein Methyltransferase SETD8

Elucidating conformational heterogeneity of proteins is essential for understanding protein functions and developing exogenous ligands for chemical perturbation. While structural biology methods can provide atomic details of static protein structures, these approaches cannot in general resolve less populated, functionally relevant conformations and uncover conformational kinetics. Here we demonstrate a new paradigm for illuminating dynamic conformational landscapes of target proteins. SETD8 (Pr-SET7/SET8/KMT5A) is a biologically relevant protein lysine methyltransferase for in vivo monomethylation of histone H4 lysine 20 and nonhistone targets. Utilizing covalent chemical inhibitors and depleting native ligands to trap hidden high-energy conformational states, we obtained diverse novel X-ray structures of SETD8. These structures were used to seed massively distributed molecular simulations that generated six milliseconds of trajectory data of SETD8 in the presence or absence of its cofactor. We used an automated machine learning approach to reveal slow conformational motions and thus distinct conformational states of SETD8, and validated the resulting dynamic conformational landscapes with multiple biophysical methods. The resulting models provide unprecedented mechanistic insight into how protein dynamics plays a role in SAM binding and thus catalysis, and how this function can be modulated by diverse cancer-associated mutants. These findings set up the foundation for revealing enzymatic mechanisms and developing inhibitors in the context of conformational landscapes of target proteins.

biophysics

Clinical Impact and Cost-Effectiveness of a 176-Condition Expanded Carrier Screen

PurposeCarrier screening identifies couples at high risk for conceiving offspring affected with serious heritable conditions. Minimal screening guidelines mandate testing for cystic fibrosis and spinal muscular atrophy, but expanded carrier screening (ECS) assesses reproductive risk for hundreds of conditions simultaneously. Although medical societies consider ECS an acceptable practice, the health economics of ECS remain incompletely characterized.\n\nMethodsThe clinical impact and cost-effectiveness of a 176-condition ECS panel were investigated using a decision-tree model comparing minimal screening and ECS in a preconception setting. Carrier rates from >50,000 patients informed disease-incidence estimates, while cost and life-years-lost data were aggregated from the literature and a cost-of-care database. Model robustness was evaluated using one-way and probabilistic sensitivity analyses.\n\nResultsFor every 100,000 pregnancies, 300 are predicted to be affected by ECS-panel conditions, which, on average, individually incur $1,300,000 in lifetime costs and increase mortality by 26 undiscounted life-years on average. Relative to minimal screening, ECS reduces the affected-birth rate and is cost-effective (i.e., <$50,000 incremental cost per life-year), findings robust to reasonable model-parameter perturbation.\n\nConclusionECS is predicted to reduce the population burden of Mendelian disease in a cost-effective manner compared to many other common medical interventions.

genetics

Open Force Field Consortium: Escaping atom types using direct chemical perception with SMIRNOFF v0.1

Here, we focus on testing and improving force fields for molecular modeling, which see widespread use in diverse areas of computational chemistry and biomolecular simulation. A key issue affecting the accuracy and transferrability of these force fields is the use of atom typing. Traditional approaches to defining molecular mechanics force fields must encode, within a discrete set of atom types, all information which will ever be needed about the chemical environment; parameters are then assigned by looking up combinations of these atom types in tables. This atom typing approach leads to a wide variety of problems such as inextensible atom-typing machinery, enormous difficulty in expanding parameters encoded by atom types, and unnecessarily proliferation of encoded parameters. Here, we describe a new approach to assigning parameters for molecular mechanics force fields based on the industry standard SMARTS chemical perception language (with extensions to identify specific atoms available in SMIRKS). In this approach, each force field term (bonds, angles, and torsions, and nonbonded interactions) features separate definitions assigned in a hierarchical manner without using atom types. We accomplish this using direct chemical perception, where parameters are assigned directly based on substructure queries operating on the molecule(s) being parameterized, thereby avoiding the intermediate step of assigning atom types -- a step which can be considered indirect chemical perception. Direct chemical perception allows for substantial simplification of force fields, as well as additional generality in the substructure queries. This approach is applicable to a wide variety of (bio)molecular systems, and can greatly reduce the number of parameters needed to create a complete force field. Further flexibility can also be gained by allowing force field terms to be interpolated based on the assignment of fractional bond orders via the same procedure used to assign partial charges. As an example of the utility of this approach, we provide a minimalist small molecule force field derived from Mercks parm@Frosst (an Amber parm99 descendant), in which a parameter definition file only {approx} 300 lines long can parameterize a large and diverse spectrum of pharmaceutically relevant small molecule chemical space. We benchmark this minimalist force field on the FreeSolv small molecule hydration free energy set and calculations of densities and dielectric constants from the ThermoML Archive, demonstrating that it achieves comparable accuracy to the Generalized Amber Force Field (GAFF) that consists of many thousands of parameters.

biophysics

Quantifying configuration-sampling error in Langevin simulations of complex molecular systems

While Langevin integrators are popular in the study of equilibrium properties of complex systems, it is challenging to estimate the timestep-induced discretization error: the degree to which the sampled phase-space or configuration-space probability density departs from the desired target density due to the use of a finite integration timestep. In [1], Sivak et al. introduced a convenient approach to approximating a natural measure of error between the sampled density and the target equilibrium density, the KL divergence, in phase space, but did not specifically address the issue of configuration-space properties, which are much more commonly of interest in molecular simulations. Here, we introduce a variant of this near-equilibrium estimator capable of measuring the error in the configuration-space marginal density, validating it against a complex but exact nested Monte Carlo estimator to show that it reproduces the KL divergence with high fidelity. To illustrate its utility, we employ this new near-equilibrium estimator to assess a claim that a recently proposed Langevin integrator introduces extremely small configuration-space density errors up to the stability limit at no extra computational expense. Finally, we show how this approach to quantifying sampling bias can be applied to a wide variety of stochastic integrators by following a straightforward procedure to compute the appropriate shadow work, and describe how it can be extended to quantify the error in arbitrary marginal or conditional distributions of interest.

biophysics

Development and validation of an expanded carrier screen that optimizes sensitivity via full-exon sequencing and panel-wide copy-number-variant identification

PurposeBy identifying pathogenic variants across hundreds of genes, expanded carrier screening (ECS) enables prospective parents to assess risk of transmitting an autosomal recessive or X-linked condition. Detection of at-risk couples depends on the number of conditions tested, the diseases respective prevalences, and the screens sensitivity for identifying disease-causing variants. Here we present an analytical validation of a 235-gene sequencing-based ECS with full coverage across coding regions, targeted assessment of pathogenic noncoding variants, panel-wide copy-number-variant (CNV) calling, and customized assays for technically challenging genes.\n\nMethodsNext-generation sequencing, a customized bioinformatics pipeline, and expert manual call review were used to identify single-nucleotide variants, short insertions and deletions, and CNVs for all genes except FMR1 and those whose low disease incidence or high technical complexity precludes novel variant identification or interpretation. Variant calls were compared to reference and orthogonal data.\n\nResultsValidation of our ECS data demonstrated >99% analytical sensitivity and >99% specificity. A preliminary assessment of 15,177 patient samples reveals the substantial impact on fetal disease-risk detection attributable to novel CNV calling (13.9% of risk) and technically challenging conditions (15.5% of risk), such as congenital adrenal hyperplasia.\n\nConclusionValidated, high-fidelity identification of different variant types--especially in diseases with complicated molecular genetics--maximizes at-risk couple detection.

genetics

OpenMM 7: Rapid Development of High Performance Algorithms for Molecular Dynamics

OpenMM is a molecular dynamics simulation toolkit with a unique focus on extensibility. It allows users to easily add new features, including forces with novel functional forms, new integration algorithms, and new simulation protocols. Those features automatically work on all supported hardware types (including both CPUs and GPUs) and perform well on all of them. In many cases they require minimal coding, just a mathematical description of the desired function. They also require no modification to OpenMM itself and can be distributed independently of OpenMM. This makes it an ideal tool for researchers developing new simulation methods, and also allows those new methods to be immediately available to the larger community.

biophysics

MSMBuilder: Statistical Models for Biomolecular Dynamics

MSMBuilder is a software package for building statistical models of high-dimensional time-series data. It is designed with a particular focus on the analysis of atomistic simulations of biomolecular dynamics such as protein folding and conformational change. MSMBuilder is named for its ability to construct Markov State Models (MSMs), a class of models that has gained favor among computational biophysicists. In addition to both well-established and newer MSM methods, the package includes complementary algorithms for understanding time-series data such as hidden Markov models (HMMs) and time-structure based independent component analysis (tICA). MSMBuilder boasts an easy to use command-line interface, as well as clear and consistent abstractions through its Python API (application programming interface). MSMBuilder is developed with careful consideration for compatibility with the broader machine-learning community by following the design of scikit-learn. The package is used primarily by practitioners of molecular dynamics but is just as applicable to other computational or experimental time-series measurements. http://msmbuilder.org

biophysics

Systematic Design and Comparison of Expanded Carrier Screening Panels

Purpose: The recent growth in pan-ethnic expanded carrier screening (ECS) has raised questions about how such panels might be designed and evaluated systematically. Design principles for ECS panels might improve clinical detection of at-risk couples and facilitate objective discussions of panel choice.\n\nMethods: Guided by medical-society statements, we propose a method for the design of ECS panels that aims to maximize the aggregate and per-disease sensitivity and specificity across a range of Mendelian disorders considered serious by a systematic classification scheme. We evaluated this method retrospectively using results from 474,644 de-identified carrier screens. We then constructed several idealized panels to highlight strengths and limitations of different ECS methodologies.\n\nResults: Based on modeled fetal risks for \"severe\" and \"profound\" diseases, a commercially available ECS panel (Counsyl) is expected to detect 183 affected conceptuses per 100,000 US births. A screens sensitivity is greatly impacted by two factors: (1) the methodology used (e.g., full-exon sequencing finds up to 46 more affected fetuses per 100,000 than targeted genotyping with an optimal 50 variant panel), and (2) the detection rate of the screen for diseases with high prevalence and complex molecular genetics (e.g., fragile X syndrome, spinal muscular atrophy, 21-hydroxylase deficiency, and alpha-thalassemia account for 54 affected fetuses per 100,000).\n\nConclusion: The described approaches allow principled, quantitative evaluation of which diseases and methodologies are appropriate for pan-ethnic expanded carrier screening.

genetics