Search bioRxivSearch

Biology subjects

Li, G.

Publications and source records attributed to Li, G..

34 records · Page 2Linked to original sources

Stratification of amyotrophic lateral sclerosis patients: a crowdsourcing approach

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease with substantial heterogeneity in clinical presentation with an urgent need for better stratification tools for clinical development and care. In this study we used a crowdsourcing approach to address the problem of ALS patient stratification. The DREAM Prize4Life ALS Stratification Challenge was a crowdsourcing initiative using data from >10,000 patients from completed ALS clinical trials and 1479 patients from community-based patient registers. Challenge participants used machine learning and clustering techniques to predict ALS progression and survival. By developing new approaches, the best performing teams were able to predict disease outcomes better than currently available methods. At the same time, the integration of clustering components across methods led to the emergence of distinct consensus clusters, separating patients into four consistent groups, each with its unique predictors for classification. This analysis reveals for the first time the potential of a crowdsourcing approach to uncover covert patient sub-populations, and to accelerate disease understanding and therapeutic development.

bioinformatics

Evolution analysis and expression divergence of the chitinase gene family against Leptosphaeria maculans and Sclerotinia sclerotiorum infection in Brassica napus

AbstractBlackleg and sclerotinia stem rot caused by Leptosphaeria maculans and Sclerotinia sclerotiorum respectively are two major diseases in rapeseed worldwide, which cause serious yield losses. Chitinases are pathogenesis-related proteins and play important roles in host resistance to various pathogens and abiotic stress responses. However, a systematic investigation of the chitinase gene family and its expression profile against L. maculans and S. sclerotiorum infection in rapeseed remains elusive. The recent release of assembled genome sequence of rapeseed allowed us to perform a genome-wide identification of the chitinase gene family. In this study, 68 chitinase genes were identified in Brassica napus genome. These genes were divided into five different classes and distributed among 15 chromosomes. Evolutionary analysis indicated that the expansion of the chitinase gene family was mainly attributed to segmental and tandem duplication. Moreover, the expression profiling of the chitinase gene family was investigated using RNA sequencing (RNA-Seq) and the results revealed that some chitinase genes were both induced while the other members exhibit distinct expression in response to L. maculans and S. sclerotiorum infection. This study presents a comprehensive survey of the chitinase gene family in B. napus and provides valuable information for further understanding the functions of the chitinase gene family.

plant biology

TFmapper: A tool for searching putative factors regulating gene expression using ChIP-seq data

BackgroundNext-generation sequencing coupled to chromatin immunoprecipitation (ChIP-seq), DNase I hypersensitivity (DNase-seq) and the transposase-accessible chromatin assay (ATAC-seq) has generated enormous amounts of data, markedly improved our understanding of the transcriptional and epigenetic control of gene expression. To take advantage of the availability of such datasets and provide clues on what factors, including transcription factors, epigenetic regulators and histone modifications, potentially regulates the expression of a gene of interest, a tool for simultaneous queries of multiple datasets using symbols or genomic coordinates as search terms is needed.\n\nResultsIn this study, we annotated the peaks of thousands of ChIP-seq datasets generated by ENCODE project, or ChIP-seq/DNase-seq/ATAC-seq datasets deposited in Gene Expression Omnibus and curated by CistromeDB; We built a MySQL database called TFmapper containing the annotations and associated metadata, allowing users without bioinformatics expertise to search across thousands of datasets to identify factors targeting a genomic region/gene of interest in a specified sample through a web interface. Users can also visualize multiple peaks in genome browsers and download the corresponding sequences.\n\nConclusionTFmapper will help users explore the vast amount of publicly available ChIP-seq/DNase-seq/ATAC-seq data, and perform integrative analyses to understand the regulation of a gene of interest. The web server is freely accessible at http://www.tfmapper.org/.

bioinformatics

Specific Eph Receptor-Cytoplasmic Effector Signaling Mediated by SAM-SAM Domain Interactions

The Eph receptor tyrosine kinase (RTK) family is the largest subfamily of RTKs playing critical roles in many developmental processes such as tissue patterning, neurogenesis and neuronal circuit formation, angiogenesis, etc. How the 14 Eph proteins, via their highly similar cytoplasmic domains, can transmit diverse and sometimes opposite cellular signals upon engaging ephrins is a major unresolved question. Here we systematically investigated the bindings of each SAM domain of Eph receptors to the SAM domains from SHIP2 and Odin, and uncover a highly specific SAM-SAM interaction-mediated cytoplasmic Eph-effector binding pattern. Comparative X-ray crystallographic studies of several SAM-SAM heterodimer complexes, together with biochemical and cell biology experiments, not only revealed the exquisite specificity code governing Eph/effector interactions but also allowed us to identify SAMD5 as a new Eph binding partner. Finally, these Eph/effector SAM heterodimer structures can explain numerous Eph SAM mutations identified in patients suffering from cancers and other diseases.

biophysics

Generation of a novel growth-enhanced and reduced environmental impact transgenic pig strain

In pig production, insufficient feed digestion causes excessive nutrients such as phosphorus and nitrogen, which are then released to the environment. To address the issue of environmental emissions, we have established transgenic pigs harboring a single-copy quad-cistronic transgene and simultaneously expressing three microbial enzymes, {beta}-glucanase, xylanase, and phytase in the salivary glands. All the transgenic enzymes were successfully expressed, and the digestion of non-starch polysaccharides (NSPs) and phytate in the feedstuff was enhanced. Fecal nitrogen and phosphate outputs were reduced by 23%-46%, and growth rate improved by 23.4% (gilts) and 24.4% (boars) when the pigs were fed on a corn and soybean-based diet and high-NSP diet. The transgenic pigs showed a 11.5%- 14.5% improvement in feed conversion rate compared to the age-matched wild-type littermates. These findings indicate that transgenic pigs are promising resources for improving feed efficiency and reducing nutrient emissions to the environment.

biochemistry

A bacterial GW-effector targets Arabidopsis AGO1 to promote pathogenicity and induces Effector-triggered immunity by disrupting AGO1 homeostasis

Pseudomonas syringae type III effectors were previously shown to suppress the Arabidopsis microRNA (miRNA) pathway through unknown mechanisms. Here, we first show that the HopT1-1 effector promotes bacterial growth by suppressing the Arabidopsis Argonaute 1 (AGO1)-dependent miRNA pathway. We further demonstrate that HopT1-1 interacts with Arabidopsis AGO1 through conserved glycine/tryptophan (GW) motifs, and in turn suppresses miRNA function. This process is not associated with a general decrease in miRNA accumulation. Instead, HopT1-1 reduces the level of AGO1-associated miRNAs in a GW-dependent manner. Therefore, HopT1-1 alters AGO1-miRISC activity, rather than miRNA biogenesis or stability. In addition, we show that the AGO1-binding platform of HopT1-1 is essential to suppress the production of reactive oxygen species (ROS) and of callose deposits during Pattern-triggered immunity (PTI). These data imply that the RNA silencing suppression activity of HopT1-1 is intimately coupled with its virulence function. Overall, these findings provide sound evidence that a bacterial effector has evolved to directly target a plant AGO protein to suppress PTI and cause disease.

plant biology

Coherent olfactory bulb gamma oscillations arise from coupling independent columnar oscillators

Spike timing-based representations of sensory information depend on embedded dynamical frameworks within neuronal networks that establish the rules of local computation and interareal communication. Here, we investigated the dynamical properties of olfactory bulb circuitry in mice of both sexes using microelectrode array recordings from slice and in vivo preparations. Neurochemical activation or optogenetic stimulation of sensory afferents evoked persistent gamma oscillations in the local field potential. These oscillations arose from slower, GABA(A) receptor-independent intracolumnar oscillators coupled by GABA(A)-ergic synapses into a faster, broadly coherent network oscillation. Consistent with the theoretical properties of coupled-oscillator networks, the spatial extent of zero-phase coherence was bounded in slices by the reduced density of lateral interactions. The intact in vivo network, however, exhibited long-range lateral interactions that suffice in simulation to enable zero-phase gamma coherence across the olfactory bulb. The timing of action potentials in a subset of principal neurons was phase-constrained with respect to evoked gamma oscillations. Coupled-oscillator dynamics in olfactory bulb thereby enable a common clock, robust to biological heterogeneities, that is capable of supporting gamma-band spike synchronization and phase coding across the ensemble of activated principal neurons. New & NoteworthyOdor stimulation evokes rhythmic gamma oscillations in the field potential of the olfactory bulb, but the dynamical mechanisms governing these oscillations have remained unclear. Establishing these mechanisms is important, as they determine the biophysical capacities of the bulbar circuit to, for example, maintain zero-phase coherence across a spatially extended network, or coordinate the timing of action potentials in principal neurons. These properties in turn constrain and suggest hypotheses of sensory coding.

neuroscience

Single-Cell RNAseq analysis of infiltrating neoplastic cells at the migrating front of human glioblastoma

Glioblastoma is the most common primary brain cancer in adults and is notoriously difficult to treat due to its diffuse nature. We performed single-cell RNAseq on 3589 cells in a cohort of four patients. We obtained cells from the tumor core as well as surrounding peripheral tissue. Our analysis revealed cellular variation in the tumors genome and transcriptome, We were able to identify infiltrating neoplastic cells in regions peripheral to the core lesions. Despite the existence of significant heterogeneity among neoplastic cells, we found that infiltrating GBM cells share a consistent gene signature between patients, suggesting a common mechanism of infiltration. Additionally, in investigating the immunological response to the tumors, we found transcriptionally distinct myeloid cell populations residing in the tumor core and the surrounding peritumoral space. Our data provide a detailed dissection of GBM cell types, revealing an abundance of novel information about tumor formation and migration.

cancer biology

Resequencing the Escherichia coli genome by GenoCare single molecule sequencing platform

Next generation sequencing (NGS) has revolutionized life sciences research. Recently, a new class of third-generation sequencing platforms has arrived to meet increasing demands in the clinic, capable of directly measuring DNA and RNA sequences at the single-molecule level without amplification. Here, we use the new GenoCare single molecule sequencing platform from Direct Genomics to resequence the E. coli genome and show comparable performance to the Illumina MiSeq system. Our platform detects single-molecule fluorescence by total internal reflection microscopy, with sequencing-by-synthesis chemistry. With a consensus sequence of 99.71% nucleotide identity to that of the Illumina MiSeq systems, GenoCare was determined to be a reliable platform for single-molecule sequencing, with strong potential for clinical applications.

genomics

Linearly changing stress environment causes cellular growth phenotype

Cells are exposed to changes in extracellular stimulus concentration that vary as a function of rate. However, the effect of stimulation rate on cell behavior and signaling remains poorly understood. Here, we examined how varying the rate of stress application alters budding yeast cell viability and mitogen-activated protein kinase (MAPK) signaling at the single-cell level. We show that cell survival and signaling depend on a rate threshold that operates in conjunction with a concentration threshold to determine the timing of MAPK signaling during rate-varying stimulus treatments. We also discovered that the stimulation rate threshold is sensitive to changes in the expression levels of the Ptp2 phosphatase, but not of another phosphatase that similarly regulates osmostress signaling during switch-like treatments. Our results demonstrate that stimulation rate is a regulated determinant of signaling output and provide a paradigm to guide the dissection of major stimulation rate-dependent mechanisms in other systems.

systems biology

Predicting single-cell transcription dynamics even when the central limit theorem fails

Despite substantial experimental and computational efforts, mechanistic modeling remains more predictive in engineering than in systems biology. The reason for this discrepancy is not fully understood. Although randomness and complexity of biological systems play roles in this concern, we hypothesize that significant and overlooked challenges arise due to specific features of single-molecule events that control crucial biological responses. Here we show that modern statistical tools to disentangle complexity and stochasticity, which assume normally distributed fluctuations or enormous datasets, don't apply to the discrete, positive, and non-symmetric distributions that characterize spatiotemporal mRNA fluctuations in single-cells. We demonstrate an alternate approach that fully captures discrete, non-normal effects within finite datasets. As an example, we integrate single-molecule measurements and these advanced computational analyses to explore Mitogen Activated Protein Kinase induction of multiple stress response genes. We discover and validate quantitatively precise, reproducible, and predictive understanding of diverse transcription regulation mechanisms, including gene activation, polymerase initiation, elongation, mRNA accumulation, spatial transport, and degradation. Our model-data integration approach extends to any discrete dynamic process with rare events and realistically limited data.\n\nSignificance StatementSystems biology seeks to combine experiments with computation to predict complex biological behaviors. However, despite tremendous data and knowledge, most biological models make terrible predictions. By analyzing single-cell-single-molecule measurements of mRNA in yeast during stress response, we explore how prediction accuracy is controlled by experimental distributions shapes. We find that asymmetric data distributions, which arise in measurements of positive quantities, can cause standard modeling approaches to yield excellent fits but make meaningless predictions. We demonstrate advanced computational tools that solve this dilemma and achieve predictive understanding of many spatiotemporal mechanisms of transcription control including RNA polymerase initiation and elongation and mRNA accumulation, transport and decay. Our approach extends to any discrete dynamic process with rare events and realistically limited data.

systems biology

Draft genome of the Reindeer (Rangifer tarandus)

AbstractO_ST_ABSBackgroundC_ST_ABSReindeer (Rangifer tarandus) is the only fully domesticated species in the Cervidae family, and is the only cervid with a circumpolar distribution. Unlike all other cervids, female reindeer regularly grow cranial appendages (antlers, the defining characteristics of cervids), as well as males. Moreover, reindeer milk contains more protein and less lactose than bovids milk. A high quality reference genome of this specie will assist efforts to elucidate these and other important features in the reindeer.\n\nFindingsWe obtained 723.2 Gb (Gigabase) of raw reads by an Illumina Hiseq 4000 platform, and a 2.64 Gb final assembly, representing 95.7% of the estimated genome (2.76 Gb according to k-mer analysis), including 92.6% of expected genes according to BUSCO analysis. The contig N50 and scaffold N50 sizes were 89.7 kilo base (kb) and 0.94 mega base (Mb), respectively. We annotated 21,555 protein-coding genes and 1.07 Gb of repetitive sequences by de novo and homology-based prediction. Homology-based searches detected 159 rRNA, 547 miRNA, 1,339 snRNA and 863 tRNA sequences in the genome of R. tarandus. The divergence time between R. tarandus, and ancestors of Bos taurus and Capra hircus, is estimated to be 29.55 million years ago (Mya).\n\nConclusionsOur results provide the first high-quality reference genome for the reindeer, and a valuable resource for studying evolution, domestication and other unusual characteristics of the reindeer.

genomics

Single Molecule Sequencing Of M13 Virus Genome Without Amplification

Third generation sequencing is a direct measurement of DNA/RNA sequences at the single molecule level without amplification. In this study, we report sequencing of the genome of the M13 virus by a new single molecule sequencing platform. Our platform detects single molecule fluorescence by the total internal reflection microscope technique, with sequencing-by-synthesis chemistry. We sequenced the genome of M13 to a depth of 316x and 100% coverage. The consensus sequence accuracy is 100%. We demonstrated that single molecule sequencing has no significant GC bias.

genomics

Molecular Mapping Of YrTZ2, A Stripe Rust Resistance Gene In Wild Emmer Accession TZ-2 And Its Comparative Analyses With Aegilops tauschii

Wheat stripe rust, caused by Puccinia striiformis f. sp. tritici (Pst), is a devastating disease that can cause severe yield losses. Identification and utilization of stripe rust resistance genes are essential for effective breeding against the disease. Wild emmer accession TZ-2, originally collected from Mount Hermon, Israel, confers near-immunity resistance against several prevailing Pst races in China. A set of 200 F6:7 recombinant inbred lines (RILs) derived from a cross between susceptible durum wheat cultivar Langdon and TZ-2 was used for stripe rust evaluation. Genetic analysis indicated that the stripe rust resistance of TZ-2 to Pst race CYR34 was controlled by a single dominant gene, temporarily designated YrTZ2. Through bulked segregant analysis (BSA) and SSR mapping, YrTZ2 was located on chromosome arm 1BS and flanked by SSR markers Xwmc230 and Xgwm413 with genetic distance of 0.8 cM (distal) and 0.3 cM (proximal), respectively. By applying wheat 90K iSelect SNP genotyping assay, 11 polymorphic loci (consist of 250 SNP markers) closely linked with YrTZ2 were identified. YrTZ2 was further delimited into a 0.8 cM genetic interval between SNP marker IWB19368 and SSR marker Xgwm413, and co-segregated with SNP marker IWB28744 (attached with 28 SNP markers). Comparative genomics analyses revealed high level of collinearity between the YrTZ2 genomic region and the orthologous region of Aegilops tauschii 1DS. The genomic region between loci IWB19368 and IWB31649 harboring YrTZ2 is orthologous to a 24.5 Mb genomic region between AT1D0112 and AT1D0150, spanning 15 contigs on chromosome 1DS. The genetic and comparative maps of YrTZ2 provide framework for map-based cloning and marker-assisted selection (MAS) of YrTZ2.

plant biology

An indicator cell assay for blood-based diagnostics

We have established proof of principle for the Indicator Cell Assay Platform (iCAP), a broadly applicable tool for blood-based diagnostics that uses specifically-selected, standardized cells as biosensors, relying on their innate ability to integrate and respond to diverse signals present in patients blood. To develop an assay, indicator cells are exposed in vitro to serum from case or control subjects and their global differential response patterns are used to train reliable, cost-effective disease classifiers based on a small number of features. In a feasibility study, the iCAP detected pre-symptomatic disease in a murine model of amyotrophic lateral sclerosis (ALS) with 94% accuracy (p-Value=3.81E-6) and correctly identified samples from a murine Huntingtons disease model as non-carriers of ALS. In a preliminary human disease assay, the iCAP detected early stage Alzheimers disease with 72% cross-validated accuracy (p-Value=3.10E-3). For both assays, iCAP features were enriched for disease-related genes, supporting the assays relevance for disease research.

systems biology

The Sequence of 1504 Mutants in the Model Rice Variety Kitaake Facilitates Rapid Functional Genomic Studies

The availability of a whole-genome sequenced mutant population and the cataloging of mutations of each line at a single-nucleotide resolution facilitates functional genomic analysis. To this end, we generated and sequenced a fast-neutron-induced mutant population in the model rice cultivar Kitaake (Oryza sativa L. ssp. japonica), which completes its life cycle in 9 weeks. We sequenced 1,504 mutant lines at 45-fold coverage and identified 91,513 mutations affecting 32,307 genes, 58% of all rice genes. We detected an average of 61 mutations per line. Mutation types include single base substitutions, deletions, insertions, inversions, translocations, and tandem duplications. We observed a high proportion of loss-of-function mutations. Using this mutant population, we identified an inversion affecting a single gene as the causative mutation for the short-grain phenotype in one mutant line with a small segregating population. This result reveals the usefulness of the resource for efficient identification of genes conferring specific phenotypes. To facilitate public access to this genetic resource, we established an open access database called KitBase that provides access to sequence data and seed stocks, enabling rapid functional genomic studies of rice.\n\nOne-sentence summaryWe have sequenced 1,504 mutant lines generated in the short life cycle rice variety Kitaake (9 weeks) and established a publicly available database, enabling rapid functional genomic studies of rice.

plant biology