Search bioRxivSearch

Biology subjects

Wen, X.

Publications and source records attributed to Wen, X..

10 recordsLinked to original sources

D-GPM: a deep learning method for gene promoter methylation inference

BackgroundGene promoter methylation plays a critical role in a wide range of biological processes, such as transcriptional expression, gene imprinting, X chromosome inactivation, etc. Whole-genome bisulfite sequencing generates a comprehensive profiling of the gene methylation levels but is limited by a high cost. Recent studies have partitioned the genes into landmark genes and target genes and suggested that the landmark gene expression levels capture adequate information to reconstruct the target gene expression levels. Moreover, the methylation level of the promoter is usually negatively correlated with its corresponding gene expression. This result inspired us to propose that the methylation level of the promoters might be adequate to reconstruct the promoter methylation level of target genes, which would eventually reduce the cost of promoter methylation profiling.\n\nResultsHere, we developed a deep learning model (D-GPM) to predict the whole-genome promoter methylation level based on the methylation profile of the landmark genes. We benchmarked D-GPM against three machine learning methods, namely, linear regression (LR), regression tree (RT) and support vector machine (SVM), based on two criteria: the mean absolute deviation (MAE) and the Pearson correlation coefficient (PCC). After profiling the methylation beta value (MBV) dataset from the TCGA, with respect to MAE and PCC, we found that D-GPM outperforms LR by 9.59% and 4.34%, RT by 27.58% and 22.96% and SVM by 6.14% and 3.07% on average, respectively. For the number of better-predicted genes, D-GPM outperforms LR in 92.65% and 91.00%, RT in 95.66% and 98.25% and SVM in 85.49% and 81.56% of the target genes.\n\nConclusionsD-GPM acquires the least overall MAE and the highest overall PCC on MBV-te compared to LR, RT, and SVM. For a genewise comparative analysis, D-GPM outperforms LR, RT, and SVM in an overwhelming majority of the target genes, with respect to the MAE and PCC. Most importantly, D-GPM predominates among the other models in predicting a majority of the target genes according to the model distribution of the least MAE and the highest PCC for the target genes.

bioinformatics

Genetic and behavioural requirements for structural brain plasticity

1Human MRI studies show that experience can lead to changes in the volume of task-specific brain regions; however, the behavioural and molecular processes driving these changes remain poorly understood. Here, we used in-vivo mouse MRI and RNA sequencing to investigate the neuroanatomical and transcriptional changes induced by environmental enrichment, exercise, and social interaction. Additionally, we asked whether the volume changes require CREB, a transcription factor critical for memory formation and neuronal plasticity. Enrichment rapidly increased cortical and hippocampal volume, and these effects were not attributable to exercise or social interaction. Instead, they likely arise from learning and sensorimotor experience. Nevertheless, the volume changes were not attenuated in mice with memory impairments caused by loss of CREB, indicating that these effects are driven by processes distinct from this canonical learning and memory pathway. Finally, within brain regions that underwent volume changes, enrichment increased the expression of genes associated with axonogenesis, dendritic spine development, synapse structural plasticity, and neurogenesis, suggesting these processes underlie the volume changes detected with MRI.

neuroscience

Identification, Genotyping, and Pathogenicity of Trichosporon spp. Isolated from Giant Pandas

Trichosporon is the dominant genus of epidermal fungi in giant pandas and causes local and deep infections. To provide the information needed for the diagnosis and treatment of trichosporosis in giant pandas, the sequence of ITS, D1/D2, and IGS1 loci in 29 isolates of Trichosporon spp. which isolated from the body surface of giant pandas were combination to investigate interspecies identification and genotype. Morphological development was examined via slide culture. Additionally, mice were infected by skin inunction, intraperitoneal injection, and subcutaneous injection for evaluation of pathogenicity. The twenty-nine isolates of Trichosporon spp. were identified as belonging to 11 species, and Trichosporon jirovecii and T. asteroides were the commonest species. Four strains of T. laibachii and one strain of T. moniliiforme were found to be of novel genotypes, and T. jirovecii was identified to be genotype 1. T. asteroides had the same genotype which involved in disseminated trichosporosis. The morphological development processes of the Trichosporon spp. were clearly different, especially in the processes of single-spore development. Pathogenicity studies showed that 7 species damaged the liver and skin in mice, and their pathogenicity was stronger than other 4 species. T. asteroides had the strongest pathogenicity and might provoke invasive infection. The pathological characteristics of liver and skin infections caused by different Trichosporon spp. were similar. So it is necessary to identify the species of Trichosporon on the surface of giant panda. Combination of ITS, D1/D2, and IGS1 loci analysis, and morphological development process can effectively identify the genotype of Trichosporon spp.

microbiology

Efficient inference of single cell expression profiles with overlapping pooling and compressed sensing

Plate-based single cell RNA-Seq (scRNA-seq) methods can detect a comprehensive profile for gene expression but suffers from high library cost of each single cell. Although cost can be reduced significantly by massively parallel scRNA-seq techniques, these approaches lose sensitivity for gene detection. Inspired by group testing and compressed sensing, here, we designed a computational framework to close the gap between sensitivity and library cost. In our framework, single cells were overlapped assigned into plenty of pools. Expression profile of each pool was then obtained by using plate-based sequence approach. The expression profile of all single cells was recovered based on the pool expression and the overlapped pooling design. The inferred expression profile showed highly consistency with the original data in both accuracy and cell types identification. A parallel computing scheme was designed to boost speed when processing the enormous single cells, and elastic net regression was combined with compressed sensing to auto-adapt for both sparsely and densely expressed genes.

bioinformatics

Bayesian Multi-SNP Genetic Association Analysis: Control of FDR and Use of Summary Statistics

Multi-SNP genetic association analysis has become increasingly important in analyzing data from genome-wide association studies (GWASs) and molecular quantitative trait loci (QTL) mapping studies. In this paper, we propose novel computational approaches to address two outstanding issues in Bayesian multi-SNP genetic association analysis: namely, the control of false positive discoveries of identified association signals and the maximization of the efficiency of statistical inference by utilizing summary statistics. Quantifying the strength and uncertainty of genetic association signals has been a long-standing theme in statistical genetics. However, there is a lack of formal statistical procedures that can rigorously control type I errors in multi-SNP analysis. We propose an intuitive hierarchical representation of genetic association signals based on Bayesian posterior probabilities, which subsequently enables rigorous control of false discovery rate (FDR) and construction of Bayesian credible sets. From the perspective of statistical data reduction, we examine the computational approaches of multi-SNP analysis using z-statistics from single-SNP association testing and conclude that they likely yield conservative results comparing to using individual-level data. Built on this result, we propose a set of sufficient summary statistics that can lead to identical results as individual-level data without sacrificing power. Our novel computational approaches are implemented in the software package, DAP-G (https://github.com/xqwen/dap), which applies to both GWASs and genome-wide molecular QTL mapping studies. It is highly computationally efficient and approximately 20 times faster than the state-of-the-art implementation of Bayesian multi-SNP analysis software. We demonstrate the proposed computational approaches using carefully constructed simulation studies and illustrate a complete workflow for multi-SNP analysis of cis expression quantitative trait loci using the whole blood data from the GTEx project.

genetics

An eQTL landscape of kidney tissue in human nephrotic syndrome

Expression quantitative trait loci (eQTL) studies illuminate the genetics of gene expression and, in disease research, can be particularly illuminating when using the tissues directly impacted by the condition. In nephrology, there is a paucity of eQTLs studies of human kidney. Here, we used whole genome sequencing (WGS) and microdissected glomerular (GLOM) & tubulointerstitial (TI) transcriptomes from 187 patients with nephrotic syndrome (NS) to describe the eQTL landscape in these functionally distinct kidney structures.\n\nUsing MatrixEQTL, we performed cis-eQTL analysis on GLOM (n=136) and TI (n=166). We used the Bayesian \"Deterministic Approximation of Posteriors\" (DAP) to fine-map these signals, eQtlBma to discover GLOM-or TI-specific eQTLs, and single cell RNA-Seq data of control kidney tissue to identify cell-type specificity of significant eQTLs. We integrated eQTL data with an IgA Nephropathy (IGAN) GWAS to perform a transcriptome-wide association study (TWAS).\n\nWe discovered 894 GLOM eQTLs and 1767 TI eQTLs at FDR <0.05. 14% and 19% of GLOM & TI eQTLs, respectively, had > 1 independent signal associated with its expression. 12% and 26% of eQTLs were GLOM-specific and TI-specific, respectively. GLOM eQTLs were most significantly enriched in podocyte transcripts and TI eQTLs in proximal tubules. The IGAN TWAS identified significant GLOM & TI genes, primarily at the HLA region.\n\nIn this study of NS patients, we discovered GLOM & TI eQTLs, identified those that were tissue-specific, deconvoluted them into cell-specific signals, and used them to characterize known GWAS alleles. These data are publicly available for browsing and download at http://nephqtl.org.

genomics

High throughput characterization of genetic effects on DNA:protein binding and gene transcription

Many variants associated with complex traits are in non-coding regions, and contribute to phenotypes by disrupting regulatory sequences. To characterize these variants, we developed a streamlined protocol for a high-throughput reporter assay, BiT-STARR-seq (Biallelic Targeted STARR-seq), that identifies allele-specific expression (ASE) while accounting for PCR duplicates through unique molecular identifiers. We tested 75,501 oligos (43,500 SNPs) and identified 2,720 SNPs with significant ASE (FDR 10%). To validate disruption of binding as one of the mechanisms underlying ASE, we developed a new high throughput allele specific binding assay for NFKB-p50. We identified 2,951 SNPs with allele-specific binding (ASB) (FDR 10%); 173 of these SNPs also had ASE (OR=1.97, p-value=0.0006). Of variants associated with complex traits, 1,531 resulted in ASE and 1,662 showed ASB. For example, we characterized that the Crohns disease risk variant for rs3810936 increases NFKB binding and results in altered gene expression.

genetics

Genome-wide association study of 1 million people identifies 111 loci for atrial fibrillation

To understand the genetic variation underlying atrial fibrillation (AF), the most common cardiac arrhythmia, we performed a genome-wide association study (GWAS) of > 1 million people, including 60,620 AF cases and 970,216 controls. We identified 163 independent risk variants at 111 loci and prioritized 165 candidate genes likely to be involved in AF. Many of the identified risk variants fall near genes where more deleterious mutations have been reported to cause serious heart defects in humans or mice (MYH6, NKX2-5, PITX2, TBC1D32, TBX5),1,2 or near genes important for striated muscle function and integrity (e.g. MYH7, PKP2, SSPN, SGCA). Experiments in rabbits with heart failure and left atrial dilation identified a heterogeneous distributed molecular switch from MYH6 to MYH7 in the left atrium, which resulted in contractile and functional heterogeneity and may predispose to initiation and maintenance of atrial arrhythmia.

genetics

Environmental Perturbations Lead To Extensive Directional Shifts In RNA Processing

Environmental perturbations have large effects on both organismal and cellular traits, including gene expression, but the extent to which the environment affects RNA processing remains largely uncharacterized. Recent studies have identified a large number of genetic variants associated with variation in RNA processing that also have an important role in complex traits; yet we do not know in which contexts the different underlying isoforms are used. Here, we comprehensively characterized changes in RNA processing events across 89 environments in five human cell types and identified 15,300 event shifts (FDR = 15%) comprised of eight event types in over 4,000 genes. Many of these changes occur consistently in the same direction across conditions, indicative of global regulation by trans factors. Accordingly, we demonstrate that environmental modulation of splicing factor binding predicts shifts in intron retention, and that binding of transcription factors predicts shifts in AFE usage in response to specific treatments. We validated the mechanism hypothesized for AFE in two independent datasets. Using ATAC-seq, we found altered binding of 64 factors in response to selenium at sites of AFE shift, including ELF2 and other factors in the ETS family. We also performed AFE QTL mapping in 373 individuals and found an enrichment for SNPs predicted to disrupt binding of the ELF2 factor. Together, these results demonstrate that RNA processing is dramatically changed in response to environmental perturbations through specific mechanisms regulated by trans factors.\n\nAuthor SummaryChanges in a cells environment and genetic variation have been shown to impact gene expression. Here, we demonstrate that environmental perturbations also lead to extensive changes in alternative RNA processing across a large number of cellular environments that we investigated. These changes often occur in a non-random manner. For example, many treatments lead to increased intron retention and usage of the downstream first exon. We also show that the changes to first exon usage are likely dependent on changes in transcription factor binding. We provide support for this hypothesis by considering how first exon usage is affected by disruption of binding due to treatment with selenium. We further validate the role of a specific factor by considering the effect of genetic variation in its binding sites on first exon usage. These results help to shed light on the vast number of changes that occur in response to environmental stimuli and will likely aid in understanding the impact of compounds to which we are daily exposed.

genomics

QuASAR-MPRA: Accurate allele-specific analysis for massively parallel reporter assays

MotivationThe majority of the human genome is composed of non-coding regions containing regulatory elements such as enhancers, which are crucial for controlling gene expression. Many variants associated with complex traits are in these regions, and may disrupt gene regulatory sequences. Consequently, it is important to not only identify true enhancers but also to test if a variant within an enhancer affects gene regulation. Recently, allele-specific analysis in high-throughput reporter assays, such as massively parallel reporter assays (MPRA), have been used to functionally validate non-coding variants. However, we are still missing high-quality and robust data analysis tools for these datasets.\n\nResultsWe have further developed our method for allele-specific analysis QuASAR (quantitative allele-specific analysis of reads) to analyze allele-specific signals in barcoded read counts data from MPRA. Using this approach, we can take into account the uncertainty on the original plasmid proportions, over-dispersion, and sequencing errors. The provided allelic skew estimate and its standard error also simplifies meta-analysis of replicate experiments. Additionally, we show that a beta-binomial distribution better models the variability present in the allelic imbalance of these synthetic reporters and results in a test that is statistically well calibrated under the null. Applying this approach to the MPRA data by Tewhey et al. (2016), we found 602 SNPs with significant (FDR 10%) allele-specific regulatory function in LCLs. We also show that we can combine MPRA with QuASAR estimates to validate existing experimental and computational annotations of regulatory variants. Our study shows that with appropriate data analysis tools, we can improve the power to detect allelic effects in high throughput reporter assays.\n\nAvailabilityhttp://github.com/piquelab/QuASAR/tree/master/mpra\n\nContactfluca@wayne.edu; rpique@wayne.edu

bioinformatics