Search bioRxivSearch

Biology subjects

Ding, L.

Publications and source records attributed to Ding, L..

6 recordsLinked to original sources

LRScaf: Improving Draft Genomes Using Long Noisy Reads

BackgroundThe advent of Third Generation Sequencing (TGS) technologies opens the door to improve genome assembly. Long reads are promised to enhance the quality of fragmental draft assemblies constructed from Next Generation Sequencing (NGS) technologies. To date, a few of algorithms, i.e., SSPACE-LongRead, OPERA-LG, SMIS, npScarf, DBG2OLC, Unicycler, and LINKS, have been released that are capable of improving draft assemblies. However, hybrid assembly on large genomes is still challenging.\n\nResultsWe develop a scalable and computationally efficient scaffolder, Long Reads Scaffolder (LRScaf), that is capable of boosting assembly contiguity to a large extent using long reads. In our experiment, our method significantly improves the contiguity of human draft assemblies, increasing the NG50 value of CHM1 from 127.5 Kb to 10.4 Mb using 20-fold coverage PacBio dataset and the NG50 value of NA12878 from 115.7 Kb to 17.4 Mb using 35-fold coverage Nanopore dataset. The run time for the scaffolding procedure using LRScaf is the shortest in all cases of our experiment. Compared with the run time of SSPACE-LongRead, LRScaf is faster 300 times for S. cerevisiae and 2,300 times for D. melanogaster. The peak RAM of LRScaf, by contrast, is more efficient than LINKS in our test. For the rice case, the peak RAM of LINKS (877.72 Gb) is about 196 times higher than LRScaf. For the experiment of human assembly, the peak RAM of LINKS is beyond the capacity of system memory (1 Tb) whereas LRScaf takes 20.28 and 41.20 Gb on CHM1 and NA12878 datasets.\n\nConclusionsThe new method, LRScaf, yields the best or at least moderate contiguity and accuracy of scaffolds in the shortest run time compared with the state-of-the-art methods. Furthermore, it offers a new opportunity for the hybrid assembly of large genomes.

bioinformatics

Concordance Of Genetic Variation That Increases Risk For Tourette Syndrome And That Influences Its Underlying Neurocircuitry

BACKGROUNDThere have been considerable recent advances in understanding the genetic architecture of Tourette Syndrome (TS) as well as its underlying neurocircuitry. However, the mechanisms by which genetic variations that increase risk for TS - and its main symptom dimensions - influence relevant brain regions are poorly understood. Here we undertook a genome-wide investigation of the overlap between TS genetic risk and genetic influences on the volume of specific subcortical brain structures that have been implicated in TS.\n\nMETHODSWe obtained summary statistics for the most recent TS genome-wide association study (GWAS) from the TS Psychiatric Genomics Consortium Working Group (4,644 cases and 8,695 controls) and GWAS of subcortical volumes from the ENIGMA consortium (30,717 individuals). We also undertook analyses using GWAS summary statistics of key symptom factors in TS, namely social disinhibition and symmetry behaviour. SNP Effect Concordance Analysis (SECA) was used to examine genetic pleiotropy - the same SNP affecting two traits - and concordance - the agreement in SNP effect directions across these two traits. In addition, a conditional false discovery rate (FDR) analysis was performed, conditioning the TS risk variants on each of the seven subcortical and the intracranial brain volume GWAS. Linkage Disequilibrium Score Regression (LDSR) was used as validation of SECA.\n\nRESULTSSECA revealed significant pleiotropy between TS and putaminal (p=2x10-4) and caudal (p=4x10-4) volumes, independent of direction of effect, and significant concordance between TS and lower thalamic volume (p=1x10-3). LDSR lent additional support for the association between TS and thalamic volume (p=5.85x10-2). Furthermore, SECA revealed significant evidence of concordance between the social disinhibition symptom dimension and lower thalamic volume (p=1x10-3), as well as concordance between symmetry behaviour and greater putaminal volume (p=7x10-4). Conditional FDR analysis further revealed novel variants significantly associated with TS (p<8x10-7) when conditioning on intracranial (rs2708146, q=0.046; and rs72853320, q=0.035 and hippocampal (rs1922786, q=0.001 volumes respectively.\n\nCONCLUSIONThese data indicate concordance for genetic variations involved in disorder risk and subcortical brain volumes in TS. Further work with larger samples is needed to fully delineate the genetic architecture of these disorders and their underlying neurocircuitry.

genomics

SupportNet: a novel incremental learning framework through deep learning and support data

MotivationIn most biological data sets, the amount of data is regularly growing and the number of classes is continuously increasing. To deal with the new data from the new classes, one approach is to train a classification model, e.g., a deep learning model, from scratch based on both old and new data. This approach is highly computationally costly and the extracted features are likely very different from the ones extracted by the model trained on the old data alone, which leads to poor model robustness. Another approach is to fine tune the trained model from the old data on the new data. However, this approach often does not have the ability to learn new knowledge without forgetting the previously learned knowledge, which is known as the catastrophic forgetting problem. To our knowledge, this problem has not been studied in the field of bioinformatics despite its existence in many bioinformatic problems.\n\nResultsHere we propose a novel method, SupportNet, to solve the catastrophic forgetting problem efficiently and effectively. SupportNet combines the strength of deep learning and support vector machine (SVM), where SVM is used to identify the support data from the old data, which are fed to the deep learning model together with the new data for further training so that the model can review the essential information of the old data when learning the new information. Two powerful consolidation regularizers are applied to ensure the robustness of the learned model. Comprehensive experiments on various tasks, including enzyme function prediction, subcellular structure classification and breast tumor classification, show that SupportNet drastically outperforms the state-of-the-art incremental learning methods and reaches similar performance as the deep learning model trained from scratch on both old and new data.\n\nAvailabilityOur program is accessible at: https://github.com/lykaust15/SupportNet.

bioinformatics

EZH2 co-opts gain-of-function p53 mutants to promote cancer growth and metastasis

With the unfolding of more and more cancer-driven gain-of-function (GOF) mutants of p53, it is important to define a common mechanism to systematically target different mutants rather than develop strategies tailored to inhibit each mutant individually. Here, using RNA immunoprecipitation sequencing (RIP-seq) we identified EZH2 as a p53 mRNA-binding protein. EZH2 bound to the internal ribosome entry site (IRES) in the 5 untranslated region (5UTR) of p53 mRNA and enhanced p53 protein translation in a methyltransferase-independent manner. EZH2 augmented p53 GOF mutant-mediated cancer growth and metastasis by increasing p53 GOF mutant protein level. EZH2 overexpression associated with the worse outcome only in patients with p53-mutated cancer. Depletion of EZH2 by antisense oligonucleotides inhibited p53 GOF mutant-mediated cancer growth. Our findings reveal a non-methyltransferase function of EZH2 that controls protein translation of p53 GOF mutants, inhibition of which causes synthetic lethality in cancer cells expressing p53 GOF mutants.

cancer biology

Ongoing, rational calibration of reward-driven perceptual biases

Decision-making is often interpreted in terms of normative computations that maximize a particular reward function for stable, average behaviors. Aberrations from the reward-maximizing solutions, either across subjects or across different sessions for the same subject, are often interpreted as reflecting poor learning or physical limitations. Here we show that such aberrations may instead reflect the involvement of additional satisficing and heuristic principles. For an asymmetric-reward perceptual decision-making task, three monkeys produced adaptive biases in response to changes in reward asymmetries and perceptual sensitivity. Their choices and response times were consistent with a normative accumulate-to-bound process. However, their context-dependent adjustments to this process deviated slightly but systematically from the reward-maximizing solutions. These adjustments were instead consistent with a rational process to find satisficing solutions based on the gradient of each monkeys reward-rate function. These results suggest new dimensions for assessing the rational and idiosyncratic aspects of flexible decision-making.

neuroscience

Scalable volumetric imaging for ultrahigh-speed brain mapping at synaptic resolution

We describe a new light-sheet microscopy method for fast, large-scale volumetric imaging. Combining synchronized scanning illumination and oblique imaging over cleared, thick tissue sections in smooth motion, our approach achieves high-speed 3D image acquisition of an entire mouse brain within 2 hours, at a resolution capable of resolving synaptic spines. It is compatible with immunofluorescence labeling, enabling flexible cell-type specific brain mapping, and is readily scalable for large biological samples such as primate brain.

neuroscience