Search bioRxivSearch

Biology subjects

Tang, C.

Publications and source records attributed to Tang, C..

8 recordsLinked to original sources

Genetic control of the HDL proteome

High-density lipoproteins (HDL) are nanoparticles with >80 associated proteins, phospholipids, cholesterol and cholesteryl esters. A comprehensive genetic analysis of the regulation of proteome of HDL isolated from a panel of 100 diverse inbred strains of mice, Hybrid Mouse Diversity Panel (HMDP), revealed widely varied HDL protein levels across the strains. Some of this variation was explained by local, cis-acting regulation, termed cis-protein quantitative trait loci. Variations in apolipoprotein A-II and apolipoprotein C-3 affected the abundance of multiple HDL proteins indicating a coordinated regulation. We identified modules of co-varying proteins and define a protein-protein interaction network describing the protein composition of the naturally occurring subspecies of HDL in mice. Sterol efflux capacity varied up to 3-fold across the strains and HDL proteins displayed distinct correlation patterns with macrophage and ABCA1 specific cholesterol efflux capacity and cholesterol exchange, suggesting that subspecies of HDL participate in discrete functions. The baseline and stimulated sterol efflux capacity phenotypes associated with distinct QTLs with smaller effect size suggesting a multi genetic regulation. Our results highlight the complexity of HDL particles by revealing high degree of heterogeneity and intercorrelation, some of which is associated with functional variation, supporting the concept that HDL-cholesterol alone is not an accurate measure of HDLs properties such as protection against CAD.

systems biology

Toward deciphering developmental patterning with deep neural network

Complex biological functions are carried out by the interaction of genes and proteins. Uncovering the gene regulation network behind a function is one of the central themes in biology. Typically, it involves extensive experiments of genetics, biochemistry and molecular biology. In this paper, we show that much of the inference task can be accomplished by a deep neural network (DNN), a form of machine learning or artificial intelligence. Specifically, the DNN learns from the dynamics of the gene expression. The learnt DNN behaves like an accurate simulator of the system, on which one can perform in-silico experiments to reveal the underlying gene network. We demonstrate the method with two examples: biochemical adaptation and the gap-gene patterning in fruit fly embryogenesis. In the first example, the DNN can successfully find the two basic network motifs for adaptation - the negative feedback and the incoherent feed-forward. In the second and much more complex example, the DNN can accurately predict behaviors of essentially all the mutants. Furthermore, the regulation network it uncovers is strikingly similar to the one inferred from experiments. In doing so, we develop methods for deciphering the gene regulation network hidden in the DNN "black box". Our interpretable DNN approach should have broad applications in genotype-phenotype mapping. SignificanceComplex biological functions are carried out by gene regulation networks. The mapping between gene network and function is a central theme in biology. The task usually involves extensive experiments with perturbations to the system (e.g. gene deletion). Here, we demonstrate that machine learning, or deep neural network (DNN), can help reveal the underlying gene regulation for a given function or phenotype with minimal perturbation data. Specifically, after training with wild-type gene expression dynamics data and a few mutant snapshots, the DNN learns to behave like an accurate simulator for the genetic system, which can be used to predict other mutants behaviors. Furthermore, our DNN approach is biochemically interpretable, which helps uncover possible gene regulatory mechanisms underlying the observed phenotypic behaviors.

developmental biology

KS-Burden: Assessing distributional differences of rare variants in dichotomous traits

A number of rare variant tests have been developed to explore the effect of low frequency genetic variations on complex phenotypes. However, an often neglected aspect in these tests is the position of genetic variations. Here we are proposing a way to assess the differences in spatial organization of rare variants by assessing their distributional differences between affected and unaffected subjects. To do so, we have formulated an adaptation of the well know Kolmogorov-Smirnov (KS) test, combining both KS and a simple gene burden approach, called KS-Burden.\n\nThe performance of our test was evaluated under a comprehensive simulations framework using real data and various scenarios. Our results show that the KS-Burden test is able to outperform the commonly used SKAT-O test, as well as others, in the presents of clusters of causal variants within a genomic region. Furthermore, our test is able to maintain competitive statistical power in scenarios unfavorable to its original assumptions. Hence, the KS-Burden test is a valuable alternative to existing tests and provides better statistical power in the presents of causal clusters within a gene.

bioinformatics

Template switching causes artificial junction formation and false identification of circular RNAs

Hundreds of thousands of putative circular RNAs have been identified through deep sequencing and bioinformatic analyses. However, the circularity of these putative RNA circles has not been experimentally validated due to limited methodologies currently available. We reported here that the template-switching capability of commonly used reverse transcriptases (e.g., SuperScript II) leads to the formation of artificial junction sequences, and consequently misclassification of large linear RNAs as RNA circles. Use of reverse transcriptases without terminal transferase activity (e.g., MonsterScript) for cDNA synthesis is critical for the identification of physiological circular RNAs. We also report two methods, MonsterScript junction PCR and high-resolution melting curve analyses, which can reliably distinguish circular RNAs from their linear forms and thus, can be used to discover and validate true circular RNAs.\n\nSignificance StatementThe vast majority of circular RNAs were identified through computational detection of junction sequences in the deep sequencing reads because these unique fusion sequences represent back-splicing events. We found that artificial junction sequences could be formed through template switching (TS) when MMLV-derived reverse transcriptases, e.g., SuperScript II, are used to synthesize cDNAs. Thus, many of the reported circular RNAs may not be RNA circles, but rather experimental artifacts. Fake circular RNAs can be avoided by using reverse transcriptases without terminal transferase activity (e.g., MonsterScript) for cDNA synthesis. We developed two novel methods, MonsterScript junction PCR and high-resolution melting curve analyses, for distinguishing circular RNAs from their linear form.

molecular biology

Time-resolved analyses of elemental distribution and concentration in living plants: An example using manganese toxicity in cowpea leaves

O_LIKnowledge of elemental distribution and concentration within plant tissues is crucial in the understanding of almost every process that occurs within plants. However, analytical limitations have hindered the microscopic determination of changes over time in the location and concentration of nutrients and contaminants in living plant tissues.\nC_LIO_LIWe developed a novel method using synchrotron-based micro X-ray fluorescence (-XRF) that allows for laterally-resolved, multi-element, kinetic analyses of plant leaf tissues in vivo. To test the utility of this approach, we examined changes in the accumulation of Mn in unifoliate leaves of 7-d-old cowpea (Vigna unguiculata) plants grown for 48 h at 0.2 and 30 M Mn in solution.\nC_LIO_LIRepeated -XRF scanning did not damage leaf tissues demonstrating the validity of the method. Exposure to 30 M Mn for 48 h increased the initial number of small spots of localized high Mn and their concentration rose from 40 to 670 mg Mn kg-1 fresh mass. Extension of the two-dimensional -XRF scans to a three-dimensional geometry provided further assessment of Mn localization and concentration.\nC_LIO_LIThis method shows the value of synchrotron-based -XRF analyses for time-resolved in vivo analysis of elemental dynamics in plant sciences.\nC_LI

plant biology

The SAM domain of mouse SAMHD1 is critical for its activation and regulation

Human SAMHD1 (hSAMHD1) is a retroviral restriction factor that blocks HIV-1 infection by depleting the cellular nucleotides required for viral reverse transcription. SAMHD1 is allosterically activated by nucleotides that induce assembly of the active tetramer. Although the catalytic core of hSAMHD1 has been studied extensively, previous structures have not captured the regulatory SAM domain. In this study, we determined the first crystal structure of full-length SAMHD1 by capturing mouse SAMHD1 (mSAMHD1) structures in three different nucleotide bound states. Although mSAMHD1 and hSAMHD1 are highly similar in sequence and function, we found that mSAMHD1 possesses a more complex nucleotide-induced activation process, highlighting the regulatory role of the SAM domain. Our results provide new insights into the regulation of SAMHD1 activity, thereby will facilitate the improvement of HIV mouse models and the development of new therapies for certain cancers and autoimmune diseases.

biophysics

AASRA: An Anchor Alignment-Based Small RNA Annotation Pipeline

SncRNA-Seq has become a routine for sncRNA profiling; however, software packages currently available are either exclusively for miRNA or piRNA annotation (e.g., miRDeep, miRanalyzer, Shortstack, PIANO), or for direct mapping of the sequence reads to the genome (e.g., Bowtie 2, SOAP and BWA), which tend to generate inaccurate counting due to repetitive matches to the genome or sncRNA homologs. Moreover, novel sncRNA variants in the sequencing reads, including those bearing small overhangs or internal insertions, deletions or mutations, are totally excluded from counting by these algorithms, leading to potential quantification bias. To overcome these problems, a comprehensive software package that can annotate all known small RNA species with adjustable tolerance towards small mismatches is needed. AASRA is based on our unique anchor alignment algorithm, which not only avoids repetitive or ambiguous counting, but also distinguishes mature miRNA from precursor miRNA reads. Compared to all existing pipelines for small RNA annotation, AASRA is superior in the following aspects: 1) AASRA can annotate all known sncRNA species simultaneously with the capability of distinguishing mature and precursor miRNAs; 2) AASRA can identify and allow for inclusion of sncRNA variants with small overhangs and/or internal insertions/deletions into the final counts; 3) AASRA is the fastest among all small RNA annotation pipelines tested. AASRA represents an all-in-one sncRNA annotation pipeline, which allows for high-speed, simultaneous annotation of all known sncRNA species with the capability to distinguish mature from precursor miRNAs, and to identify novel sncRNA variants in the sncRNA-Seq sequencing reads.\n\nAvailability and ImplementationThe AASRA software is freely available at https://github.com/biogramming/AASRA.

bioinformatics

Optimal Growth Of Microbes On Mixed Carbon Sources

A classic problem in microbiology is that bacteria display two types of growth behavior when cultured on a mixture of two carbon sources: in certain mixtures the bacteria consume the two carbon sources sequentially (diauxie) and in other mixtures the bacteria consume both sources simultaneously (co-utilization). The search for the molecular mechanism of diauxie led to the discovery of the lac operon and gene regulation in general. However, why microbes would bother to have different strategies of taking up nutrients remained a mystery. Here we show that diauxie versus co-utilization can be understood from the topological features of the metabolic network. A model of optimal allocation of protein resources to achieve maximum growth quantitatively explains why and how the cell makes the choice when facing multiple carbon sources. Our work solves a long-standing puzzle and sheds light on microbes optimal growth in different nutrient conditions.

systems biology