Search bioRxivSearch

Biology subjects

Liu, C.

Publications and source records attributed to Liu, C..

At least 55 records · Page 3Linked to original sources

Contextual Regression: An Accurate and Conveniently Interpretable Nonlinear Model for Mining Discovery from Scientific Data

Machine learning algorithms such as linear regression, SVM and neural network have played an increasingly important role in the process of scientific discovery. However, none of them is both interpretable and accurate on nonlinear datasets. Here we present contextual regression, a method that joins these two desirable properties together using a hybrid architecture of neural network embedding and dot product layer. We demonstrate its high prediction accuracy and sensitivity through the task of predictive feature selection on a simulated dataset and the application of predicting open chromatin sites in the human genome. On the simulated data, our method achieved high fidelity recovery of feature contributions under random noise levels up to {+/-}200%. On the open chromatin dataset, the application of our method not only outperformed the state of the art method in terms of accuracy, but also unveiled two previously unfound open chromatin related histone marks. Our method fills in the gap of accurate and interpretable nonlinear modeling in scientific data mining tasks.

bioinformatics

Genetic Analysis of Deep Phenotyping Projects in Common Disorders

Several studies of complex psychotic disorders with large numbers of neurobiological phenotypes are currently under way, in living patients and controls, and on assemblies of brain specimens. Genetic analyses of such data typically present challenges, because of the choice of underlying hypotheses on genetic architecture of the studied disorders and phenotypes, large numbers of phenotypes, the appropriate multiple testing corrections, limited numbers of subjects, imputations required on missing phenotypes and genotypes, and the cross-disciplinary nature of the phenotype measures. Advances in genotype and phenotype imputation, and in genome-wide association (GWAS) methods, are useful in dealing with these challenges. As compared with the more traditional single-trait analyses, deep phenotyping with simultaneous genome-wide analyses serves as a discovery tool for previously unsuspected relationships of phenotypic traits with each other, and with specific molecular involvements.

genetics

Effects of mechanical loading on cortical defect repair using a novel mechanobiological model of bone healing

Mechanical loading is an important aspect of post-surgical care. The timing of load application relative to the injury event is thought to differentially regulate repair depending on the stage of healing. Here, we show using a novel mechanobiological model of cortical defect repair that daily loading (5 N peak load, 2 Hz, 60 cycles, 4 consecutive days) during hematoma consolidation and inflammation disrupts the injury site and activates cartilage formation on the periosteal surface adjacent to the defect. We also show that daily loading during the matrix deposition phase enhances both bone and cartilage formation at the defect site, while loading during the remodeling phase results in an enlarged woven bone regenerate. All loading regimens resulted in abundant cellular proliferation within the regenerate and at the periosteal surface and fibrous tissue formation directly above the defect. Stress was concentrated at the edges of the defect during exogenous loading, and finite element (FE)-modeled longitudinal strain ({varepsilon}zz) values along the anterior and posterior borders of the defect (~2200 {varepsilon}) were an order of magnitude larger than strain values on the proximal and distal borders (~50-100 {varepsilon}). These findings demonstrate that all phases of cortical defect healing are sensitive to physical stimulation. In addition, the proposed novel mechanobiological model offers several advantages including its technical simplicity and its well-characterized and spatially confined repair program, making effects of physical and biological interventions more easily assessed.

bioengineering

DeepLung: 3D Deep Convolutional Nets for Automated Pulmonary Nodule Detection and Classification

In this work, we present a fully automated lung CT cancer diagnosis system, DeepLung. DeepLung contains two parts, nodule detection and classification. Considering the 3D nature of lung CT data, two 3D networks are designed for the nodule detection and classification respectively. Specifically, a 3D Faster R-CNN is designed for nodule detection with a U-net-like encoder-decoder structure to effectively learn nodule features. For nodule classification, gradient boosting machine (GBM) with 3D dual path network (DPN) features is proposed. The nodule classification subnetwork is validated on a public dataset from LIDC-IDRI, on which it achieves better performance than state-of-the-art approaches, and surpasses the average performance of four experienced doctors. For the DeepLung system, candidate nodules are detected first by the nodule detection subnetwork, and nodule diagnosis is conducted by the classification subnetwork. Extensive experimental results demonstrate the DeepLung is comparable to the experienced doctors both for the nodule-level and patient-level diagnosis on the LIDC-IDRI dataset.

bioinformatics

The Early Diagnosis in Lung Cancer by the Detection of Circulating Tumor DNA

BackgroundRemarkable advances for clinical diagnosis and treatment in cancers including lung cancer involve cell-free circulating tumor DNA (ctDNA) detection through next generation sequencing. However, before the sensitivity and specificity of ctDNA detection can be widely recognized, the consistency of mutations in tumor tissue and ctDNA should be evaluated. The urgency of this consistency is extremely obvious in lung cancer to which great attention has been paid to in liquid biopsy field.\n\nMethodsWe have developed an approach named systematic error correction sequencing (Sec-Seq) to improve the evaluation of sequence alterations in circulating cell-free DNA. Averagely 10 ml preoperative blood samples were collected from 30 patients containing pulmonary space occupying pathological changes by traditional clinic diagnosis. cfDNA from plasma, genomic DNA from white blood cells, and genomic DNA from solid tumor of above patients were extracted and constructed as libraries for each sample before subjected to sequencing by a panel contains 50 cancer-associated genes encompassing 29 kb by custom probe hybridization capture with average depth >40000, 7000, or 6300 folds respectively.\n\nResultsDetection limit for mutant allele frequency in our study was 0.1%. The sequencing results were analyzed by bioinformatic expertise based on our previous studies on the baseline mutation profiling of circulating cell-free DNA and the clinicopathological data of these patients. Among all the lung cancer patients, 78% patients were predicted as positive by ctDNA sequencing when the shreshold was defined as at least one of the hotspot mutations detected in the blood (ctDNA) was also detected in tumor tissue. Pneumonia and pulmonary tuberculosis were detected as negative according to the above standard. When evaluating all hotspots in driver genes in the panel, 24% mutations detected in tumor tissue (tDNA) were also detected in patients blood (ctDNA). When evaluating all genetic variations in the panel, including all the driver genes and passenger genes, 28% detected in tumor tissue (tDNA) were also detected in patients blood (ctDNA). Positive detection rates of plasma ctDNA in stage I lung cancer patients is 85%, compared with 17% of tumor biomarkers.\n\nConclusionWe demonstrated the importance of sequencing both circulating cell-free DNA and genomic DNA in tumor tissue for ctDNA detection in lung cancer currently. We also determined and confirmed the consistency of ctDNA and tumor tissue through NGS according to the criteria explored in our studies. Our strategy can initially distinguish the lung cancer from benign lesions of lung. Our work shows that the consistency will be benefited from the optimization in sensitivity and specificity in ctDNA detection.

cancer biology

Complemented palindrome small RNAs first discovered from SARS coronavirus

In this study, we reported for the first time the existence of complemented palindrome small RNAs (cpsRNAs) and proposed cpsRNAs and palindrome small RNAs (psRNAs) as a novel class of small RNAs. The first discovered cpsRNA UCUUUAACAAGCUUGUUAAAGA from SARS coronavirus named SARS-CoV-cpsR-22 contained 22 nucleotides perfectly matching its reverse complementary sequence. Further sequence analysis supported that SARS-CoV-cpsR-22 originated from bat betacoronavirus. The results of RNAi experiments showed that one 19-nt segment of SARS-CoV-cpsR-22 significantly induced cell apoptosis. These results suggested that SARS-CoV-cpsR-22 could play a role in SARS-CoV infection or pathogenicity. The discovery of psRNAs and cpsRNAs paved the way to find new markers for pathogen detection and reveal the mechanisms in the infection or pathogenicity from a different point of view. The discovery of psRNAs and cpsRNAs also broaden the understanding of palindrome motifs in animal of plant genomes.

genomics

GVC: A superfast and universal genomic variant caller

Germline and somatic variant detection from human and cancer whole-genome sequencing data is a challenge task for genome-wide association study and cancer genomics in precision medicine. Many confounding factors contribute the difficulties including complexity of variant, sequencing and alignment error, tumor clonality and sample purity etc. Current genomic variant callers are too time-consuming to meet the requirement of clinical application in precision medicine. We developed superfast and universal Genomic Variant Caller (GVC), which can simultaneously detect various genomic variants including SNV, sINDEL and SV from personal and normal-cancer paired whole-genome/exome sequencing data within fifteen minutes. Whats more, it achieved higher sensitivity and precision than popular variant callers including GATK4, Mutect, NovoBreak in germline and somatic variant detection from NA12878 and ICGC-TCGA Dream Challenge Datasets respectuvely. It is worth mentioning that GVC achieved comparable performance in variant detection from NA12878 sequenced by three different high-throughput sequencing platforms including Illumina HiSeq2000, NovaSeq and BGISEQ-500.

bioinformatics

Accurate typing of class I human leukocyte antigen by Oxford nanopore sequencing

Oxford Nanopore Technologies MinION has expanded the current DNA sequencing toolkit by delivering long read lengths and extreme portability. The MinION has the potential to enable expedited point-of-care human leukocyte antigen (HLA) typing, an assay routinely used to assess the immunological compatibility between organ donors and recipients, but the platforms high error rate makes it challenging to type alleles with clinical-grade accuracy. Here, we developed and validated Athlon, an algorithm that iteratively scores nanopore reads mapped to a hierarchical database of HLA alleles to arrive at a consensus diploid genotype; Athlon achieved a 100% accuracy in class I HLA typing at high resolution.

bioinformatics

Common variants of NRXN1, LRP1B and RORA are associated with increased ventricular volumes in psychosis - GWAS findings from the B-SNIP deep phenotyping study

Schizophrenia, Schizoaffective, and Bipolar Disorders share common illness traits, intermediate phenotypes and a partially overlapping polygenic basis. We performed GWAS on deep phenotyping data, including structural MRI and DTI, clinical, and behavioral scales from 1,115 cases and controls. Significant associations were observed with two cerebrospinal fluid volumes: the temporal horn of left lateral ventricle was associated with NRXN1, and the volume of the cavum septum pellucidum was associated with LRP1B and RORA. Both volumes were associated with illness. Suggestive associations were observed with local gyrification indices, fractional anisotropy and age at onset. The deep phenotyping approach allowed unexpected genetic sharing to be found between phenotypes, including temporal horn of left lateral ventricle and age at onset.

genetics

Development of subcortical volumes across adolescence in males and females: A multisample study of longitudinal changes

The developmental patterns of subcortical brain volumes in males and females observed in previous studies have been inconsistent. To help resolve these discrepancies, we examined developmental trajectories using three independent longitudinal samples of participants in the age-span of 8-22 years (total 216 participants and 467 scans). These datasets, including Pittsburgh (PIT; University of Pittsburgh, USA), NeuroCognitive Development (NCD; University of Oslo, Norway), and Orygen Adolescent Development Study (OADS; The University of Melbourne, Australia), span three countries and were analyzed together and in parallel using mixed-effects modeling with both generalized additive models and general linear models. For all regions and across all samples, males were found to have significantly larger volumes as compared to females, and significant sex differences were seen in age trajectories over time. However, direct comparison of sample trajectories and sex differences identified within samples were not consistent. The trajectories for the amygdala, putamen, and nucleus accumbens were most consistent between the three samples. Our results suggest that even after using similar preprocessing and analytic techniques, additional factors, such as image acquisition or sample composition may contribute to some of the discrepancies in sex specific patterns in subcortical brain changes across adolescence, and highlight region-specific variations in congruency of developmental trajectories.

neuroscience

Identification of disease resistance genes from a Chinese wild grapevine (Vitis davidii) by analysing grape transcriptomes and transgenic Arabidopsis

HighlightTranscription profiles showed that 20 candidate genes were obviously co-expressed at 12 hpi in 30 Vitis davidii.\n\nVdWRKY53 trancription factor enhanced the resisitance in grapevine and Arabidopsis.\n\nAbstractThe molecular mechanisms underlying disease tolerance in grapevines remain uncharacterized, even though there are substantial differences in the resistance of grapevine species to fungal and bacterial diseases. In this study, we identified genes and genetic networks involved in disease resistance in grapevines by comparing the transcriptomes of a strongly resistant clone of Chinese wild grapevine (Vitis davidii cv. Ciputao 941, DAC) and a susceptible clone of European grapevine (Vitis vinifera cv. Manicure Finger, VIM) before and after infection with white rot disease (Coniella diplodiella). Disease resistance-related genes were triggered in DAC approximately 12 hours post infection (hpi) with C. diplodiella. Twenty candidate resistant genes were co-expressed in DAC. One of these candidate genes, VdWRKY53 (GenBank accession KY124243), was over-expressed in transgenic Arabidopsis thaliana plants and was found to provide these plants with enhanced resistance to C. diplodiella, Pseudomonas syringae pv tomato PDC3000, and Golovinomyces cichoracearum. This result indicates that VdWRKY53 may be involved in nonspecific resistance via interaction with fungal and oomycete elicitor signals and the activation of defence gene expression. These results provide potential gene targets for molecular breeding to develop resistant grape cultivars.

molecular biology

The pomegranate (Punica granatum L.) genome provides insights into fruit quality and ovule developmental biology

Pomegranate (Punica granatum L.) with an uncertain taxonomic status has an ancient cultivation history, and has become an emerging fruit due to its attractive features such as the bright red appearance and the high abundance of medicinally valuable ellagitannin-based compounds in its peel and aril. However, the absence of genomic resources has restricted further elucidating genetics and evolution of these interesting traits. Here we report a 274-Mb high-quality draft pomegranate genome sequence, which covers approximately 81.5% of the estimated 336 Mb genome, consists of 2,177 scaffolds with an N50 size of 1.7 Mb, and contains 30,903 genes. Phylogenomic analysis supported that pomegranate belongs to the Lythraceae family rather than the monogeneric Punicaceae family, and comparative analyses showed that pomegranate and Eucalyptus grandis shares the paleotetraploidy event. Integrated genomic and transcriptomic analyses provided insights into the molecular mechanisms underlying the biosynthesis of ellagitannin-based compounds, the color formation in both peels and arils during pomegranate fruit development, and the unique ovule development processes that are characteristic of pomegranate. This genome sequence represents the first reference in Lythraceae, providing an important resource to expand our understanding of some unique biological processes and to facilitate both comparative biology studies and crop breeding.

genomics

Positional effects revealed in Illumina Methylation Array and the impact on analysis

With the evolution of rapid epigenetic research, Illumina Infinium HumanMethylation BeadChips have been widely used to study DNA methylation. However, in evaluating the accuracy of this method, we found that the commonly used Illumina HumanMethylation BeadChips are substantially affected by positional effects; the DNA samples location in a chip affects the measured methylation levels. We analyzed three HumanMethylation450 and three HumanMethylation27 datasets by using four methods to prove the existence of positional effects. Three datasets were analyzed further for technical replicate analysis or differential methylation CpG sites analysis. The pre- and post-correction comparisons indicate that the positional effects could alter the measured methylation values and downstream analysis results. Nevertheless, ComBat, linear regression and functional normalization could all be used to minimize such artifact. We recommend performing ComBat to correct positional effects followed by the correction of batch effects in data preprocessing as this procedure slightly outperforms the others. In addition, randomizing the sample placement should be a critical laboratory practice for using such experimental platforms. Code for our method is freely available at: https://github.com/ChuanJ/posibatch.

bioinformatics

Near-Atomic Resolution Structure Determination in Over-Focus with Volta Phase Plate by Cs-corrected Cryo-EM

Volta phase plate (VPP) is a recently developed transmission electron microscope (TEM) apparatus that can significantly enhance the image contrast of biological samples in cryo-electron microscopy (cryo-EM) therefore impose the possibility to solve structures of relatively small macromolecules at high resolution. In this work, we performed theoretical analysis and found that using phase plate on objective lens spherical aberration (Cs)-corrected TEM may gain some interesting optical properties, including the over-focus imaging of macromolecules. We subsequently evaluated the imaging strategy of frozen-hydrated apo-ferritin with VPP on a Cs-corrected TEM and obtained the structure of apo-ferritin at near atomic resolution from both under- and over-focused dataset, illustrating the feasibility and new potential of combining VPP with Cs-corrected TEM for high resolution cryo-EM.\n\nHighlightsThe successful combination of volta phase plate and Cs-corrector in single particle cryo-EM.\n\nNear-atomic structure determined from over-focused images by cryo-EM. VPP-Cs-corrector coupled EM provides interesting optical properties.\n\nIn BriefWe took the unique advantage of the optical system by combining the volta phase plate and Cs-corrector in a modern TEM to collect high resolution micrographs of frozen-hydrated apo-ferritin in over-focus imaging conditions and determined the structure of apo-ferritin at 3.0 Angstrom resolution.

biophysics

Evaluation Of Chromatin Accessibility In Prefrontal Cortex Of Schizophrenia Cases And Controls

Schizophrenia genome-wide association (GWA) studies have identified over 150 regions of the genome that are associated with disease risk, yet there is little evidence that coding mutations contribute to this disorder. To explore the mechanism of non-coding regulatory elements in schizophrenia, we performed ATAC-seq on adult prefrontal cortex brain samples from 135 individuals with schizophrenia and 137 controls, and identified 118,152 ATAC-seq peaks. These accessible chromatin regions in brain are highly enriched for SNP-heritability for schizophrenia (10.6 fold enrichment, P=2.4x10-4, second only to genomic regions conserved in Eutherian mammals) and replicated in an independent dataset (9.0 fold enrichment, P=2.7x10-4). This degree of enrichment of schizophrenia heritability was higher than in open chromatin found in 138 different cell and tissue types. Brain open chromatin regions that overlapped highly conserved regions exhibited an even higher degree of heritability enrichment, indicating that conservation can identify functional subsets within regulatory elements active in brain. However, we did not identify chromatin accessibility differences between schizophrenia cases and controls, nor did we find an interaction of chromatin QTLs with case-control status. This indicates that although causal variants map within regulatory elements, mechanisms other than differential chromatin may govern the contribution of regulatory element variation to schizophrenia risk. Our results strongly implicate gene regulatory processes involving open chromatin in the pathogenesis of schizophrenia, and suggest a strategy to understand the hundreds of common variants emerging from large genomic studies of complex brain diseases.

genomics

The Proteolytic Landscape Of An Arabidopsis Separase-Deficient Mutant Reveals Novel Substrates Associated With Plant Development

Digestive proteolysis executed by the proteasome plays an important role in plant development. Yet, the role of limited proteolysis in this process is still obscured due to the absence of studies. Previously, we showed that limited proteolysis by the caspase-related protease separase (EXTRA SPINDLE POLES [ESP]) modulates development in plants through the cleavage of unknown substrates. Here we used a modified version of the positional proteomics method COmbined FRActional DIagonal Chromatography (COFRADIC) to survey the proteolytic landscape of wild-type and separase mutant RADIALLY SWOLLEN 4 (rsw4) root tip cells, as an attempt to identify targets of separase. We have discovered that proteins involved in the establishment of pH homeostasis and sensing, and lipid signalling in wild-type cells, suggesting novel potential roles for separase. We also observed significant accumulation of the protease PRX34 in rsw4 which negatively impacts growth. Furthermore, we observed an increased acetylation of N-termini of rsw4 proteins which usually comprise degrons identified by the ubiquitin-proteasome system, suggesting that separase intersects with additional proteolytic networks. Our results hint to potential pathways by which separase could regulate development suggesting also novel proteolytic functions.

plant biology

Genome-wide Association Study Of Plasma Proteins Identifies Putatively Causal Genes, Proteins, And Pathways For Cardiovascular Disease

Identifying genetic variants associated with circulating protein concentrations (pQTLs) and integrating them with variants from genome-wide association studies (GWAS) may illuminate the proteomes causal role in disease and bridge a GWAS knowledge gap for hitherto unexplained SNP-disease associations. We conducted GWAS of 71 high-value proteins for cardiovascular disease in 6,861 Framingham Heart Study participants followed by external replication. We comprehensively mapped thousands of pQTLs, including functional annotations and clinical-trait associations, and created an integrated plasma-protein-QTL searchable database. We next identified 15 proteins with pQTLs coinciding with coronary heart disease (CHD)-related variants from GWAS or tested causal for CHD by Mendelian randomization; most of these proteins were associated with new-onset cardiovascular disease events in Framingham participants with long-term follow-up. Identifying pQTLs and integrating them with GWAS results yields insights into genes, proteins, and pathways that may be causally associated with disease and can serve as therapeutic targets for treatment and prevention.

epidemiology

Direct Conversion Of Human Fibroblasts Into Osteoblasts And Osteocytes With Small Molecules And A Single Factor, Runx2

Human osteoblasts can be induced from somatic cells by introducing defined factors, however, the strategy limits cells therapeutic applications for its multi-factor and complicated genetic manipulations that may bring uncertainty into the genome. Another important cell type in bone metabolism, osteocytes, which play a central role in regulating the dynamic nature of bone in all its diverse functions, have not been obtained from transdifferetiation so far. Herein, we have established procedures to convert human fibroblast directly into osteocyte-like and osteoblast-like cells using a single transcription factor, Runx2 and chemical cocktails by activating Wnt and cAMP/PKA pathways. These induced osteoblast-like cells express osteogenic markers and generate mineralized nodule deposition. A good performance of bone formation from these cells was observed in subcutaneous site of mouse at 4 weeks post-transplantation. Moreover, further studies convert human fibroblasts into osteocyte-like cells by orchestrating timing of the aforementioned chemical cocktails exposure. These osteocyte-like cells express osteocyte-specific markers and display characteristic morphology features of osteocytes. In summary, this study provides a promising strategy for cell-based therapy in bone regenerative medicine by direct reprogramming of fibroblasts into osteocytes and osteoblasts.

cell biology