Search bioRxivSearch

Biology subjects

Mukherjee, S.

Publications and source records attributed to Mukherjee, S..

22 records · Page 2Linked to original sources

High-dimensional regression over disease subgroups

We consider high-dimensional regression over subgroups of observations. Our work is motivated by biomedical problems, where disease subtypes, for example, may differ with respect to underlying regression models, but sample sizes at the subgroup-level may be limited. We focus on the case in which subgroup-specific models may be expected to be similar but not necessarily identical. Our approach is to treat subgroups as related problem instances and jointly estimate subgroup-specific regression coefficients. This is done in a penalized framework, combining an{ell} 1 term with an additional term that penalizes differences between subgroup-specific coefficients. This gives solutions that are globally sparse but that allow information-sharing between the subgroups. We present algorithms for estimation and empirical results on simulated data and using Alzheimers disease, amyotrophic lateral sclerosis and cancer datasets. These examples demonstrate the gains our approach can offer in terms of prediction and the ability to estimate subgroup-specific sparsity patterns.

bioinformatics

Bridging microscopic and macroscopic mechanisms of p53-MDM2 binding using molecular simulations and kinetic network models

Under normal cellular conditions, the tumor suppressor protein p53 is kept at a low levels in part due to ubiquitination by MDM2, a process initiated by binding of MDM2 to the intrinsically disordered transactivation domain (TAD) of p53. Although many experimental and simulation studies suggest that disordered domains such as p53 TAD bind their targets nonspecifically before folding to a tightly-associated conformation, the molecular details are unclear. Toward a detailed prediction of binding mechanism, pathways and rates, we have performed large-scale unbiased all-atom simulations of p53-MDM2 binding. Markov State Models (MSMs) constructed from the trajectory data predict p53 TAD peptide binding pathways and on-rates in good agreement with experiment. The MSM reveals that two key bound intermediates, each with a non-native arrangement of hydrophobic residues in the MDM2 binding cleft, control the overall on-rate. Using microscopic rate information from the MSM, we parameterize a simple four-state kinetic model to (1) determine that induced-fit pathways dominate the binding flux over a large range of concentrations, and (2) predict how modulation of residual p53 helicity affects binding, in good agreement with experiment. These results suggest new ways in which microscopic models of bound-state ensembles can be used to understand biological function on a macroscopic scale.\n\nAUTHOR SUMMARYMany cell signaling pathways involve protein-protein interactions in which an intrinsically disordered peptide folds upon binding its target. Determining the molecular mechanisms that control these binding rates is important for understanding how such systems are regulated. In this paper, we show how extensive all-atom simulations combined with kinetic network models provide a detailed mechanistic understanding of how tumor suppressor protein p53 binds to MDM2, an important target of new cancer therapeutics. A simple four-state model parameterized from the simulations shows a binding-then-folding mechanism, and recapitulates experiments in which residual helicity boosts binding. This work goes beyond previous simulations of small-molecule binding, to achieve pathways and binding rates for a large peptide, in good agreement with experiment.

biophysics

Development and assessment of fully automated and globally transitive geometric morphometric methods, with application to a biological comparative dataset with high interspecific variation

Automated geometric morphometric methods are promising tools for shape analysis in comparative biology: they improve researchers abilities to quantify biological variation extensively (by permitting more specimens to be analyzed) and intensively (by characterizing shapes with greater fidelity). Although use of these methods has increased, automated methods have some notable limitations: pairwise correspondences are frequently inaccurate or lack transitivity (i.e., they are not defined with reference to the full sample). In this study, we reassess the accuracy of two previously published automated methods, cPDist [1] and auto3Dgm [2], and evaluate several modifications to these methods. We show that a substantial fraction of alignments and pairwise maps between specimens of highly dissimilar geometries were inaccurate in the study of Boyer et al. [1], despite a taxonomically sensitive variance structure of continuous Procrustes distances. We also show these inaccuracies can be remedied by utilizing a globally informed methodology within a collection of shapes, instead of only comparing shapes in a pairwise manner (c.f. [2]). Unfortunately, while global information generally enhances maps between dissimilar objects, it can degrade the quality of correspondences between similar objects due to the accumulation of numerical error. We explore a number of approaches to mitigate this degradation, quantify the performance of these approaches, and compare the generated pairwise maps (as well as the shape space characterized by these maps) to a \"ground truth\" obtained from landmarks manually collected by geometric morphometricians. Novel methods both improve the quality of the pairwise correspondences relative to cPDist, and achieve a taxonomic distinctiveness comparable to auto3Dgm.

bioinformatics

HOMINID: A framework for identifying associations between host genetic variation and microbiome composition

Recent studies have uncovered a strong effect of host genetic variation on the composition of host-associated microbiota. Here, we present HOMINID, a computational approach based on Lasso linear regression, that given host genetic variation and microbiome composition data, identifies host SNPs that are correlated with microbial taxa abundances. Using simulated data we show that HOMINID has accuracy in identifying associated SNPs, and performs better compared to existing methods. We also show that HOMINID can accurately identify the microbial taxa that are correlated with associated SNPs. Lastly, by using HOMINID on real data of human genetic variation and microbiome composition, we identified 13 human SNPs in which genetic variation is correlated with microbiome taxonomic composition across body sites. In conclusion, HOMINID is a powerful method to detect host genetic variants linked to microbiome composition, and can facilitate discovery of mechanisms controlling host-microbiome interactions.\n\nAvailability and implementationSoftware, code, tutorial, installation and setup details, and synthetic data are available in the project homepage: https://github.com/blekhmanlab/hominid.\n\nReal dataset used here is from Blekhman et al. (Blekhman et al. 2015); 16S rRNA gene sequence data and OTU tables are available on the HMP DACC website (www.hmpdacc.org), and host genetic data are deposited in dbGaP under project number phs000228.

genomics