Search bioRxivSearch

Biology subjects

Shi, H.

Publications and source records attributed to Shi, H..

9 recordsLinked to original sources

AIDE: annotation-assisted isoform discovery and abundanceestimation from RNA-seq data

Genome-wide accurate identification and quantification of full-length mRNA isoforms is crucial for investigating transcriptional and post-transcriptional regulatory mechanisms of biological phenomena. Despite continuing efforts in developing effective computational tools to identify or assemble full-length mRNA isoforms from second-generation RNA-seq data, it remains a challenge to accurately identify mRNA isoforms from short sequence reads due to the substantial information loss in RNA-seq experiments. Here we introduce a novel statistical method, AIDE (Annotation-assisted Isoform DiscovEry), the first approach that directly controls false isoform discoveries by implementing the testing-based model selection principle. Solving the isoform discovery problem in a stepwise and conservative manner, AIDE prioritizes the annotated isoforms and precisely identifies novel isoforms whose addition significantly improves the explanation of observed RNA-seq reads. We evaluate the performance of AIDE based on multiple simulated and real RNA-seq datasets followed by a PCR-Sanger sequencing validation. Our results show that AIDE effectively leverages the annotation information to compensate the information loss due to short read lengths. AIDE achieves the highest precision in isoform discovery and the lowest error rates in isoform abundance estimation, compared with three state-of-the-art methods Cufflinks, SLIDE, and StringTie. As a robust bioinformatics tool for transcriptome analysis, AIDE will enable researchers to discover novel transcripts with high confidence.

bioinformatics

The C. elegans SMOC-1 protein acts cell non-autonomously to promote bone morphogenetic protein signaling

Bone morphogenetic protein (BMP) signaling regulates many different developmental and homeostatic processes in metazoans. The BMP pathway is conserved in Caenorhabditis elegans, and is known to regulate body size and mesoderm development. We have identified the C. elegans smoc-1 (Secreted MOdular Calcium binding protein-1) gene as a new player in the BMP pathway. smoc-1(0) null mutants have a small body size, while overexpression of smoc-1 led to a long body size and increased expression of the RAD-SMAD BMP reporter, suggesting that SMOC-1 acts as a positive modulator of BMP signaling. Using double mutant analysis, we showed that SMOC-1 antagonizes the function of the glypican LON-2 and acts through the BMP ligand DBL-1 to regulate BMP signaling. Moreover, SMOC-1 appears to specifically regulate BMP signaling without significant involvement in a TGF{beta}-like pathway that regulates dauer development. We found that smoc-1 is expressed in multiple tissues, including cells of the pharynx, intestine, and posterior hypodermis, and that the expression of smoc-1 in the intestine is positively regulated by BMP signaling. We further established that SMOC-1 functions cell non-autonomously to regulate body size. Human SMOC1 and SMOC2 can each partially rescue the smoc-1(0) mutant phenotype, suggesting that SMOC-1s function in modulating BMP signaling is evolutionarily conserved. Together, our findings highlight a conserved role of SMOC proteins in modulating BMP signaling in metazoans.\n\nARTICLE SUMMARYBMP signaling is critical for development and homeostasis in metazoans, and is under tight regulation. We report the identification and characterization of a Secreted MOdular Calcium binding protein SMOC-1 as a positive modulator of BMP signaling in C. elegans. We established that SMOC-1 antagonizes the function of LON-2/glypican and acts through the DBL-1/BMP ligand to promote BMP signaling. We identified smoc-1-expressing cells, and demonstrated that SMOC-1 acts cell non-autonomously and in a positive feedback loop to regulate BMP signaling. We also provide evidence suggesting that the function of SMOC proteins in the BMP pathway is conserved from worms to humans.

genetics

Phenotype-specific enrichment of Mendelian disorder genes near GWAS regions across 62 complex traits

Although recent studies provide evidence for a common genetic basis between complex traits and Mendelian disorders, a thorough quantification of their overlap in a phenotype-specific manner remains elusive. Here, we quantify the overlap of genes identified through large-scale genome-wide association studies (GWAS) for 62 complex traits and diseases with genes known to cause 20 broad categories of Mendelian disorders. We identify a significant enrichment of phenotypically-matched Mendelian disorder genes in GWAS gene sets. Further, we observe elevated GWAS effect sizes near phenotypically-matched Mendelian disorder genes. Finally, we report examples of GWAS variants localized at the transcription start site or physically interacting with the promoters of phenotypically-matched Mendelian disorder genes. Our results are consistent with the hypothesis that genes that are disrupted in Mendelian disorders are dysregulated by noncoding variants in complex traits, and demonstrate how leveraging findings from related Mendelian disorders and functional genomic datasets can prioritize genes that are putatively dysregulated by local and distal non-coding GWAS variants.

genetics

A unifying framework for joint trait analysis under a non-infinitesimal model

MotivationA large proportion of risk regions identified by genome-wide association studies (GWAS) are shared across multiple diseases and traits. Understanding whether this clustering is due to sharing of causal variants or chance colocalization can provide insights into shared etiology of complex traits and diseases.\n\nResultsIn this work, we propose a flexible, unifying framework to quantify the overlap between a pair of traits called UNITY (Unifying Non-Infinitesimal Trait analYsis). We formulate a Bayesian generative model that relates the overlap between pairs of traits to GWAS summary statistic data under a non-infinitesimal genetic architecture underlying each trait. We propose a Metropolis-Hastings sampler to compute the posterior density of the genetic overlap parameters in this model. We validate our method through comprehensive simulations and analyze summary statistics from height and BMI GWAS to show that it produces estimates consistent with the known genetic makeup of both traits.\n\nAvailabilityThe UNITY software is made freely available to the research community at: https://github.com/bogdanlab/UNITY\n\nContactruthjohnson@ucla.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Conservation of conformational dynamics across prokaryotic actins

The actin family of cytoskeletal proteins is essential to the physiology of virtually all archaea, bacteria, and eukaryotes. While X-ray crystallography and electron microscopy have revealed structural homologies among actin-family proteins, these techniques cannot probe molecular-scale conformational dynamics. Here, we use all-atom molecular dynamic simulations to reveal conserved dynamical behaviors in four prokaryotic actin homologs: MreB, FtsA, ParM, and crenactin. We demonstrate that the majority of the conformational dynamics of prokaryotic actins can be explained by treating the four subdomains as rigid bodies. MreB, ParM, and FtsA monomers exhibited nucleotide-dependent dihedral and opening angles, while crenactin monomer dynamics were nucleotide-independent. We further determine that the opening angle of ParM is sensitive to a specific interaction between subdomains. Steered molecular dynamics simulations of MreB, FtsA, and crenactin dimers revealed that changes in subunit dihedral angle lead to intersubunit bending or twist, suggesting a conserved mechanism for regulating filament structure. Taken together, our results provide molecular-scale insights into the nucleotide and polymerization dependencies of the structure of prokaryotic actins, suggesting mechanisms for how these structural features are linked to their diverse functions.\n\nSignificance StatementSimulations are a critical tool for uncovering the molecular mechanisms underlying biological form and function. Here, we use molecular-dynamics simulations to identify common and specific dynamical behaviors in four prokaryotic homologs of actin, a cytoskeletal protein that plays important roles in cellular structure and division in eukaryotes. Dihedral angles and opening angles in monomers of bacterial MreB, FtsA, and ParM were all sensitive to whether the subunit was bound to ATP or ADP, unlike in the archaeal homolog crenactin. In simulations of MreB, FtsA, and crenactin dimers, changes in subunit dihedral angle led to bending or twisting in filaments of these proteins, suggesting a mechanism for regulating the properties of large filaments. Taken together, our simulations set the stage for understanding and exploiting structure- function relationships of bacterial cytoskeletons.

biophysics

Probabilistic fine-mapping of transcriptome-wide association studies

Transcriptome-wide association studies (TWAS) using predicted expression have identified thousands of genes whose locally-regulated expression is associated to complex traits and diseases. In this work, we show that linkage disequilibrium (LD) among SNPs induce significant gene-trait associations at non-causal genes as a function of the overlap between eQTL weights used in expression prediction. We introduce a probabilistic framework that models the induced correlation among TWAS signals to assign a probability for every gene in the risk region to explain the observed association signal while controlling for pleiotropic SNP effects and unmeasured causal expression. Importantly, our approach remains accurate when expression data for causal genes are not available in the causal tissue by leveraging expression prediction from other tissues. Our approach yields credible-sets of genes containing the causal gene at a nominal confidence level (e.g., 90%) that can be used to prioritize and select genes for functional assays. We illustrate our approach using an integrative analysis of lipids traits where our approach prioritizes genes with strong evidence for causality.

genetics

Cell size regulation through tunable geometric localization of the bacterial actin cytoskeleton

In the rod-shaped bacterium Escherichia coli, the actin-like protein MreB localizes in a curvature-dependent manner and spatially coordinates cell-wall insertion to maintain cell shape across changing environments, although the molecular mechanism by which cell width is regulated remains unknown. Here, we demonstrate that the bitopic membrane protein RodZ regulates the biophysical properties of MreB and alters the spatial organization of E. coli cell-wall growth. The relative expression levels of MreB and RodZ changed in a manner commensurate with variations in growth rate and cell width. We carried out single-cell analyses to determine that RodZ systematically alters the curvature-based localization of MreB and cell width in a manner dependent on the concentration of RodZ. Finally, we identified MreB mutants that we predict using molecular dynamics simulations to alter the bending properties of MreB filaments at the molecular scale similar to RodZ binding, and showed that these mutants rescued rod-like shape in the absence of RodZ alone or in combination with wild-type MreB. Together, our results show that E. coli controls its shape and dimensions by differentially regulating RodZ and MreB to alter the patterning of cell-wall insertion, highlighting the rich regulatory landscape of cytoskeletal molecular biophysics.

biophysics

A Bayesian Framework for Multiple Trait Colocalization from Summary Association Statistics

MotivationMost genetic variants implicated in complex diseases by genome-wide association studies (GWAS) are non-coding, making it challenging to understand the causative genes involved in disease. Integrating external information such as quantitative trait locus (QTL) mapping of molecular traits (e.g., expression, methylation) is a powerful approach to identify the subset of GWAS signals explained by regulatory effects. In particular, expression QTLs (eQTLs) help pinpoint the responsible gene among the GWAS regions that harbor many genes, while methylation QTLs (mQTLs) help identify the epigenetic mechanisms that impact gene expression which in turn affect disease risk. In this work we propose multiple-trait-coloc (moloc), a Bayesian statistical framework that integrates GWAS summary data with multiple molecular QTL data to identify regulatory effects at GWAS risk loci.\n\nResultsWe applied moloc to schizophrenia (SCZ) and eQTL/mQTL data derived from human brain tissue and identified 52 candidate genes that influence SCZ through methylation. Our method can be applied to any GWAS and relevant functional data to help prioritize disease associated genes.\n\nAvailabilitymoloc is available for download as an R package (https://github.com/clagiamba/moloc). We also developed a web site to visualize the biological findings (icahn.mssm.edu/moloc). The browser allows searches by gene, methylation probe, and scenario of interest.\n\nContactclaudia.giambartolomei@gmail.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

genomics

Local genetic correlation gives insights into the shared genetic architecture of complex traits

Although genetic correlations between complex traits provide valuable insights into epidemiological and etiological studies, a precise quantification of which genomic regions contribute to the genome-wide genetic correlation is currently lacking. Here, we introduce{rho} -HESS, a technique to quantify the correlation between pairs of traits due to genetic variation at a small region in the genome. Our approach only requires GWAS summary data and makes no distributional assumption on the causal variant effects sizes while accounting for linkage disequilibrium (LD) and overlapping GWAS samples. We analyzed large-scale GWAS summary data across 35 complex traits, and identified 27 genomic regions that contribute significantly to the genetic correlation among these traits. Notably, we find 7 genomic regions that contribute to the genetic correlation of 12 pairs of traits that show negligible genome-wide correlation, further showcasing the power of local genetic correlation analyses. Finally, we leverage the distribution of local genetic correlations across the genome to assign putative direction of causality for 15 pairs of traits.

genetics