Search bioRxivSearch

Biology subjects

Wang, K.

Publications and source records attributed to Wang, K..

43 records · Page 3Linked to original sources

Accuracy, Reproducibility And Bias Of Next Generation Sequencing For Quantitative Small RNA Profiling: A Multiple Protocol Study Across Multiple Laboratories

Small RNA-seq is increasingly being used for profiling of small RNAs. Quantitative characteristics of long RNA-seq have been extensively described, but small RNA-seq involves fundamentally different methods for library preparation, with distinct protocols and technical variations that have not been fully and systematically studied. We report here the results of a study using common references (synthetic RNA pools of defined composition, as well as plasma-derived RNA) to evaluate the accuracy, reproducibility and bias of small RNA-seq library preparation for five distinct protocols and across nine different laboratories. We observed protocol-specific and sequence-specific bias, which was ameliorated using adapters for ligation with randomized end-nucleotides, and computational correction factors. Despite this technical bias, relative quantification using small RNA-seq was remarkably accurate and reproducible, even across multiple laboratories using different methods. These results provide strong evidence for the feasibility of reproducible cross-laboratory small RNA-seq studies, even those involving analysis of data generated using different protocols.

genomics

Reverse Engineering of Trascriptional Networks Uncovers Candidate Master Regulators Governing Neuropathology of Schizophrenia

Tissue-specific reverse engineering of transcriptional networks has uncovered master regulators (MRs) of cellular networks in various cancers, yet the application of this method to neuropsychiatric disorders is largely unexplored. Here, using RNA-Seq data on postmortem dorsolateral prefrontal cortex (DLPFC) from schizophrenia (SCZ) patients and control subjects, we deconvolved the transcriptional network to identify MRs that mediate expression of a large body of target genes. Together with an independent RNA-Seq data on cultured cells derived from olfactory neuroepithelium, we identified TCF4, a leading SCZ risk locus implicated by genome-wide association studies, as one of the top candidate MRs that may be potentially dysregulated in SCZ. We validated the dysregulated TCF4-related transcriptional network through examining the transcription factor binding footprints inferred from human induced pluripotent stem cell (hiPSC)-derived neuronal ATAC-Seq data, as well as direct binding sites obtained from ChIP-seq data in SH-SY5Y cells. The predicted TCF4 transcriptional targets were enriched for genes showing transcriptomic changes upon knockdown of TCF4 in hiPSC-derived neural progenitor cells (NPC) and glutamatergic neurons (Glut_N), based on observations from three separate cell lines. The altered TCF4 gene network perturbations in NPC, as compared to that in Glut_N, was more similar to the expression differences in the TCF4 gene network observed in the DLPFC of individuals with SCZ. Moreover, TCF4-associated gene expression changes in NPC were more enriched than Glut_N for pathways involved in neuronal activity, genome-wide significant SCZ risk genes, and SCZ-associated de novo mutations. Our results suggest that TCF4 may potentially serve as a MR of a gene network that confers susceptibility to SCZ at early stage of neurodevelopment, highlighting the importance of network dysregulation involving core genes and many hundreds of peripheral genes in conferring susceptibility to neuropsychiatric diseases.

neuroscience

Rapid Whole Brain Imaging Of Neural Activities In Freely Behaving Larval Zebrafish

The internal brain dynamics that link sensation and action are arguably better studied during natural animal behaviors. Here we report on a novel volume imaging and 3D tracking technique that monitors whole brain neural activity in freely swimming larval zebrafish (Danio rerio). We demonstrated the capability of our system through functional imaging of neural activity during visually evoked and prey capture behaviors in larval zebrafish.

neuroscience

The Calcineurin-FoxO-MuRF1 Signaling Pathway Regulates Myofibril Integrity in Cardiomyocytes

Altered Ca2+ handling is often present in diseased hearts undergoing structural remodeling and functional deterioration. The influences of Ca2+ signaling on cardiac function have been examined extensively, but whether Ca2+ directly regulates sarcomere structure has remained elusive. Using a mutant zebrafish model lacking NCX1 activity in the heart, we explored the impacts of impaired Ca2+ homeostasis on myofibril integrity. Gene expression profiling analysis revealed that the E3 ubiquitin ligase MuRF1 is upregulated in ncx1-deficient hearts. Intriguingly, knocking down MuRF1 activity or inhibiting proteasome activity preserved myofibril integrity in ncx1 deficient hearts, revealing a MuRF1-mediated proteasome degradation mechanism that is activated in response to abnormal Ca2+ homeostasis. Furthermore, we detected an accumulation of the MuRF1 regulator FoxO in the nuclei of ncx1-deficient cardiomyocytes. Overexpression of FoxO in wild type cardiomyocytes induced MuRF1 expression and caused myofibril disarray, whereas inhibiting Calcineurin activity attenuated FoxO-mediated MuRF1 expression and protected sarcomeres from degradation in ncx1-deficient hearts. Together, our findings reveal a novel mechanism by which Ca2+ overload disrupts the myofibril integrity in heart muscle cells by activating a Calcineurin-FoxO-MuRF1-proteosome signaling pathway.

cell biology

HadoopCNV: A Dynamic Programming Imputation Algorithm To Detect Copy Number Variants From Sequencing Data

BACKGROUNDWhole-genome sequencing (WGS) data may be used to identify copy number variations (CNVs). Existing CNV detection methods mostly rely on read depth or alignment characteristics (paired-end distance and split reads) to infer gains/losses, while neglecting allelic intensity ratios and cannot quantify copy numbers. Additionally, most CNV callers are not scalable to handle a large number of WGS samples.\n\nMETHODSTo facilitate large-scale and rapid CNV detection from WGS data, we developed a Dynamic Programming Imputation (DPI) based algorithm called HadoopCNV, which infers copy number changes through both allelic frequency and read depth information. Our implementation is built on the Hadoop framework, enabling multiple compute nodes to work in parallel.\n\nRESULTSCompared to two widely used tools - CNVnator and LUMPY, HadoopCNV has similar or better performance on both simulated data sets and real data on the NA12878 individual. Additionally, analysis on a 10-member pedigree showed that HadoopCNV has a Mendelian precision that is similar or better than other tools. Furthermore, HadoopCNV can accurately infer loss of heterozygosity (LOH), while other tools cannot. HadoopCNV requires only 1.6 hours for a human genome with 30X coverage, on a 32-node cluster, with a linear relationship between speed improvement and the number of nodes. We further developed a method to combine HadoopCNV and LUMPY result, and demonstrated that the combination resulted in better performance than any individual tools.\n\nCONCLUSIONSThe combination of high-resolution, allele-specific read depth from WGS data and Hadoop framework can result in efficient and accurate detection of CNVs.

bioinformatics

Genetic footprint of population fragmentation and contemporary collapse in a freshwater cetacean

Understanding demographic trends and patterns of gene flow in an endangered species is crucial for devising conservation strategies. Here, we examined the extent of population structure and recent evolution of the critically endangered Yangtze finless porpoise (Neophocaena asiaeorientalis asiaeorientalis). By analysing genetic variation at the mitochondrial and nuclear microsatellite loci for 148 individuals, we identified three populations along the Yangtze River, each one connected to a group of admixed ancestry. Each population displayed extremely low genetic diversity, consistent with extremely small effective size ([≤]92 individuals). Habitat degradation and distribution gaps correlated with highly asymmetric gene-flow that was inefficient in maintaining connectivity between populations. Genetic inferences of historical demography revealed that the populations in the Yangtze descended from a small number of founders colonizing the river from the sea during the last Ice Age. The colonization was followed by a rapid population split during the last millennium predating the Chinese Modern Economy Development. However, genetic diversity showed a clear footprint of population contraction over the last 50 years leaving only ~2% of the pre-collapsed size, consistent with the population collapses reported from field studies. This genetic perspective provides background information for devising mitigation strategies to prevent this species from extinction.

evolutionary biology

Evaluation on Efficient Detection of Structural Variants at Low Coverage by Long-Read Sequencing

BackgroundStructural variants (SVs) in human genomes are implicated in a variety of human diseases. Long-read sequencing delivers much longer read lengths than short-read sequencing and may greatly improve SV detection. However, due to the relatively high cost of long-read sequencing, it is unclear what coverage is needed and how to optimally use the aligners and SV callers.\n\nResultsIn this study, we developed NextSV, a meta-caller to perform SV calling from low coverage long-read sequencing data. NextSV integrates three aligners and three SV callers and generates two integrated call sets (sensitive/stringent) for different analysis purposes. We evaluated SV calling performance of NextSV under different PacBio coverages on two personal genomes, NA12878 and HX1. Our results showed that, compared with running any single SV caller, NextSV stringent call set had higher precision and balanced accuracy (F1 score) while NextSV sensitive call set had a higher recall. At 10X coverage, the recall of NextSV sensitive call set was 93.5% to 94.1% for deletions and 87.9% to 93.2% for insertions, indicating that ~10X coverage might be an optimal coverage to use in practice, considering the balance between the sequencing costs and the recall rates. We further evaluated the Mendelian errors on an Ashkenazi Jewish trio dataset.\n\nConclusionsOur results provide useful guidelines for SV detection from low coverage whole-genome PacBio data and we expect that NextSV will facilitate the analysis of SVs on long-read sequencing data.

genomics