Search bioRxivSearch

Biology subjects

Zhan, L.

Publications and source records attributed to Zhan, L..

4 recordsLinked to original sources

ForestQC: quality control on genetic variants from next-generation sequencing data using random forest

Next-generation sequencing technology (NGS) enables discovery of nearly all genetic variants present in a genome. A subset of these variants, however, may have poor sequencing quality due to limitations in sequencing technology or in variant calling algorithms. In genetic studies that analyze a large number of sequenced individuals, it is critical to detect and remove those variants with poor quality as they may cause spurious findings. In this paper, we present a statistical approach for performing quality control on variants identified from NGS data by combining a traditional filtering approach and a machine learning approach. Our method uses information on sequencing quality such as sequencing depth, genotyping quality, and GC contents to predict whether a certain variant is likely to contain errors. To evaluate our method, we applied it to two whole-genome sequencing datasets where one dataset consists of related individuals from families while the other consists of unrelated individuals. Results indicate that our method outperforms widely used methods for performing quality control on variants such as VQSR of GATK by considerably improving the quality of variants to be included in the analysis. Our approach is also very efficient, and hence can be applied to large sequencing datasets. We conclude that combining a machine learning algorithm trained with sequencing quality information and the filtering approach is an effective approach to perform quality control on genetic variants from sequencing data.\n\nAuthor SummaryGenetic disorders can be caused by many types of genetic mutations, including common and rare single nucleotide variants, structural variants, insertions and deletions. Nowadays, next generation sequencing (NGS) technology allows us to identify various genetic variants that are associated with diseases. However, variants detected by NGS might have poor sequencing quality due to biases and errors in sequencing technologies and analysis tools. Therefore, it is critical to remove variants with low quality, which could cause spurious findings in follow-up analyses. Previously, people applied either hard filters or machine learning models for variant quality control (QC), which failed to filter out those variants accurately. Here, we developed a statistical tool, ForestQC, for variant QC by combining a filtering approach and a machine learning approach. We applied ForestQC to one family-based whole genome sequencing (WGS) dataset and one general case-control WGS dataset, to evaluate our method. Results show that ForestQC outperforms widely used methods for variant QC by considerably improving the quality of variants. Also, ForestQC is very efficient and scalable to large-scale sequencing datasets. Our study indicates that combining filtering approaches and machine learning approaches enables effective variant QC.

bioinformatics

Proximal recolonization by self-renewing microglia re-establishes microglial homeostasis in the adult mouse brain

Microglia are resident immune cells that play critical roles in maintaining normal physiology of the central nervous system. Remarkably, microglia have intrinsic capacity to replenish after being acutely ablated. However, the underlying mechanisms that drive such microglial restoration remain elusive. Here, we removed microglia via CSF1R inhibitor PLX5622 and characterized repopulation both spatially and temporally. We also investigated the cellular origin of repopulated microglia and report that microglia are replenished via self-renewal, with little contribution from non-microglial lineage progenitors, including nestin+ progenitors and the circulating myeloid population. Interestingly, spatial analyses with multi-color labeling reveal that newborn microglia recolonize the parenchyma by forming distinctive clusters that maintain stable territorial boundaries over time, indicating proximal expansion nature of adult microgliogenesis and stability of microglia tiling. Temporal transcriptome profiling from newborn microglia at different repopulation stages revealed that the adult newborn microglia gradually regain steady-state maturity from an immature state that is reminiscent of neonatal stage and follow a series of maturation programs that include NF-{kappa}B activation, interferon immune activation and apoptosis, etc. Importantly, we show that the restoration of microglial homeostatic density requires NF-{kappa}B signaling as well as apoptotic egress of excessive cells. In summary, our study reports key events that take place from microgliogenesis to homeostasis re-establishment.

neuroscience

High-Throughput Identification of Genetic Variation Impact on pre-mRNA Splicing Efficiency

AbstractUnderstanding the functional impact of genomic variants is a major goal of modern genetics and personalized medicine. Although many synonymous and non-coding variants act through altering the efficiency of pre-mRNA splicing, it is difficult to predict how these variants impact pre-mRNA splicing. Here, we describe a massively parallel approach we used to test the impact of 2,059 human genetic variants spanning 110 alternative exons on pre-mRNA splicing. This method yields data that reinforces known mechanisms of pre-mRNA splicing, can rapidly identify genomic variants that impact pre-mRNA splicing, and will be useful for increasing our understanding of genome function.

genomics

A Large-Scale Binding and Functional Map of Human RNA Binding Proteins

Genomes encompass all the information necessary to specify the development and function of an organism. In addition to genes, genomes also contain a myriad of functional elements that control various steps in gene expression. A major class of these elements function only when transcribed into RNA as they serve as the binding sites for RNA binding proteins (RBPs), which act to control post-transcriptional processes including splicing, cleavage and polyadenylation, RNA editing, RNA localization, stability, and translation. Despite the importance of these functional RNA elements encoded in the genome, they have been much less studied than genes and DNA elements. Here, we describe the mapping and characterization of RNA elements recognized by a large collection of human RBPs in K562 and HepG2 cells. These data expand the catalog of functional elements encoded in the human genome by addition of a large set of elements that function at the RNA level through interaction with RBPs.\n\nHighlightsO_LI223 eCLIP datasets for 150 RBPs reveal a wide variety of in vivo RNA target classes.\nC_LIO_LI472 knockdown/RNA-seq profiles of 263 RBPs reveal factor-responsive targets and integration with eCLIP indicates RNA expression and splicing regulatory patterns.\nC_LIO_LI78 RNA Bind-N-Seq profiles of in vitro binding motifs reveal links between in vitro and in vivo binding and indicate that eCLIP peaks that contain in vitro motifs are more strongly associated with regulation.\nC_LIO_LI274 maps of RBP subcellular localization by immunofluorescence indicate widespread organelle-specific RNA processing regulation.\nC_LIO_LI63 ChIP-seq profiles of DNA association suggest broad interconnectivity between chromatin association and RNA processing.\nC_LI

genomics