Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Pediatric Brainstem Encephalitis Outbreak Investigation with Metagenomic Next-Generation Sequencing

In 2016, Catalonia experienced a pediatric brainstem encephalitis outbreak caused by enterovirus A71 (EV-A71). Conventional testing identified EV in peripheral body sites, but EV was rarely identified in cerebrospinal fluid (CSF). RNA was extracted from CSF (n=20), plasma (n=9), stool (n=15) and nasopharyngeal samples (n=16) from 10 children with brainstem encephalitis or encephalomyelitis and 10 contemporaneous pediatric controls with presumed viral meningitis or encephalitis. Unbiased complementary DNA libraries were sequenced, and microbial pathogens were identified using a custom bioinformatics pipeline. Full-length virus genomes were assembled for phylogenetic analyses. Metagenomic next-generation sequencing (mNGS) was concordant with qRT-PCR for all samples positive by PCR (n=25). In virus-negative samples (n=35), mNGS detected virus in 28.6% (n=10), including 5 CSF samples. mNGS co-detected EV-A71 and another EV in 5 patients. Overall, mNGS increased the proportion of EV-positive samples from 42% (25/60) to 57% (34/60) (McNemars test; p-value = 0.0077). For CSF, mNGS doubled the number of pathogen-positive samples (McNemars test; p-value = 0.074). Using phylogenetic analysis, the outbreak EV-A71 clustered with a neuroinvasive German EV-A71 isolate. Brainstem encephalitis specific, non-synonymous EV-A71 single nucleotide variants were not identified. mNGS demonstrated 100% concordance with clinical qRT-PCR of EV-related brainstem encephalitis and significantly increased the detection of enteroviruses. Our findings increase the probability that neurologic complications observed were virus-induced rather than para-infectious. A comprehensive genomic analysis confirmed that the EV-A71 outbreak strain was closely related to a neuroinvasive German EV-A71 isolate. There were no clear-cut viral genomic differences that discriminated between patients with differing neurologic phenotypes.

microbiology

RNAseq dataset describing transcriptional changes in cervical sensory ganglia after bilateral pyramidotomy and forelimb intramuscular gene therapy with AAV1 encoding human neurotrophin-3

Unilateral or bilateral corticospinal tract injury in the pyramids of adult rats causes changes in proprioceptive axon terminal arborization in the cervical spinal cord accompanied by hyperreflexia and abnormal movements including spasms [1, 2]. Treatment of affected forelimb muscles with an Adeno-Associated Viral Vector (AAV) encoding human neurotrophin-3 (NT3) normalizes many of these anatomical, neurophysiological and behavioural changes [1]. Interestingly, in several studies, neurotrophin-3 protein accumulates in cervical dorsal root ganglia (DRG) on the side ipsilateral to AAV injection [1, 3]. We hypothesize that neurotrophin-3 induces these changes (in proprioceptive axon wiring, proprioceptive reflex neurophysiology and sensorimotor behaviors involving proprioception) by modifying gene expression in affected cervical dorsal root ganglia (DRG). As a first step in testing this hypothesis, we analyzed the transcriptomes of cervical DRGs obtained during a previous study from naive rats and from rats after bilateral pyramidotomy (bPYX) with unilateral intramuscular injections of either AAV1-CMV-NT3 or AAV1-CMV-EGFP made 24h after injury [1]. Ten weeks after surgery, Poly(A) RNAs and small RNAs from C6 to C8 DRGs on the treated side were sequenced. We detected mRNAs or small RNAs that were significantly regulated under three conditions (bPYX+GFP vs naive; bPYX+NT3 versus naive; bPYX+NT3 vs bPYX+GFP). We identified mRNAs and small RNAs whose expression level was altered after pyramidotomy and normalized by neurotrophin-3 treatment. A bioinformatic analysis enabled us to identify genes that are likely to be expressed in proprioceptors after injury and which were regulated by neurotrophin-3 in the direction expected from other datasets involving knockout or overexpression of neurotrophin-3. This dataset will help us and others identify genes in sensory neurons whose expression levels are regulated by neurotrophin-3 treatment. This may help identify novel therapeutic targets to improve sensation and movement after neurological injury. Data has been deposited in the Gene Expression Omnibus (GSE82197).

neuroscience

Quantitative multi-locus metabarcoding and waggle dance interpretation reveal honey bee spring foraging patterns in Midwest agroecosystems

We explored the pollen foraging behavior of honey bee colonies situated in the corn and soybean dominated agroecosystems of central Ohio over a month-long period using both pollen metabarcoding and waggle dance inference of spatial foraging patterns. For molecular pollen analysis we developed simple and cost-effective laboratory and bioinformatics methods. Targeting four plant barcode loci (ITS2, rbcL, trnL and trnH), we implemented metabarcoding library preparation and dual-indexing protocols designed to minimize amplification biases and index mis-tagging events. We constructed comprehensive, curated reference databases for hierarchical taxonomic classification of metabarcoding data and used these databases to train the Metaxa2 DNA sequence classifier. Comparisons between morphological and molecular palynology provide strong support for the quantitative potential of multi-locus metabarcoding. Results revealed consistent foraging habits between locations and show clear trends in the phenological progression of honey bee spring foraging in these agricultural areas. Our data suggest that three key taxa, woody Rosaceae such as pome fruits and hawthorns, Salix, and Trifolium provided the majority of pollen nutrition during the study. Spatially, these foraging patterns were associated with a significant preference for forests and tree lines relative to crop fields and herbaceous land cover.

ecology

Beating Swords into Ploughshares: Domestication of a Phage Lysin for Housekeeping Function

Temperate phages constitute a potentially beneficial genetic reservoir for bacterial innovation despite being selfish entities encoding an infection cycle inherently at odds with bacterial fitness. These phages integrate their genomes into the bacterial host during infection, donating new, but deleterious, genetic material: the phage genome encodes toxic genes, such as lysins, that kill the bacterium during the phage infection cycle. Remarkably, some bacteria have exploited the destructive properties of phage genes for their own benefit by co-opting them as toxins for functions related to bacterial warfare, virulence, and secretion. However, do toxic phage genes ever become raw material for functional innovation? Here we report on a toxic phage gene whose product has lost its toxicity and has become a domain of a core cellular factor, SpmX, throughout the bacterial order Caulobacterales. Using a combination of phylogenetics, bioinformatics, structural biology, cell biology, and biochemistry, we have investigated the origin and function of SpmX and determined that its occurrence is the result of the detoxification of a phage peptidoglycan hydrolase gene. We show that the retained, attenuated activity of the phage-derived domain plays an important role in proper cell morphology and developmental regulation in representatives of this large bacterial clade. To our knowledge, this is the first observation of phage gene domestication in which a toxic phage gene has been co-opted for a housekeeping function.

microbiology

PolyA tracks and poly-lysine repeats are the Achilles heel of Plasmodium falciparum

Plasmodium falciparum, the causative agent of human malaria, is an apicomplexan parasite with a complex, multi-host life cycle. Sixty percent of transcripts from its extreme AT-rich (81%) genome possess coding polyadenosine (polyA) runs, distinguishing the parasite from its hosts and other sequenced organisms. Recent studies indicate that transcripts with polyA runs encoding poly-lysine are hot spots for ribosome stalling and frameshifting, eliciting mRNA surveillance pathways and attenuating protein synthesis in the majority of prokaryotic and eukaryotic organisms. Here, we show that the P. falciparum translational machinery is paradigm-breaking. Using bioinformatic and biochemical approaches, we demonstrate that both endogenous genes and reporter sequences containing long polyA runs are efficiently and accurately transcribed and translated in P. falciparum cells. Translation of polyA tracks in the parasite does not elicit any response from mRNA surveillance pathways usually seen in host human cells or organisms with similar AT content. The translation efficiency and accuracy of the parasite protein synthesis machinery reveals a unique role of ribosomes in the evolution and adaptation of P. falciparum to an AU-rich transcriptome and polybasic amino sequences. Finally, we show that the ability of P. falciparum to synthesize long poly-lysine repeats has given this parasite a unique protein exportome and an advantage in infectivity that can be suppressed by addition of exogenous poly-basic polymers.

microbiology

Genomic sequence capture of haemosporidian parasites: Methods and prospects for enhanced study of host-parasite evolution

Avian malaria and related haemosporidians (Plasmodium, [Para]Haemoproteus, and Leucocytoozoon) represent an exciting multi-host, multi-parasite system in ecology and evolution. Global research in this field accelerated after 1) the publication in 2000 of PCR protocols to sequence a haemosporidian mitochondrial (mtDNA) barcode, and 2) the development in 2009 of an open-access database to document the geographic and host ranges of parasite mtDNA haplotypes. Isolating haemosporidian nuclear DNA from bird hosts, however, has been technically challenging, slowing the transition to genomic-scale sequencing techniques. We extend a recently-developed sequence capture method to obtain hundreds of haemosporidian nuclear loci from wild bird samples, which typically have low levels of infection, or parasitemia. We tested 51 infected birds from Peru and New Mexico and evaluated locus recovery in light of variation in parasitemia, divergence from reference sequences, and pooling strategies. Our method was successful for samples with parasitemia as low as [~]0.03% (3 of 10,000 blood cells infected) and mtDNA divergence as high as 15.9% (one Leucocytozoon sample), and using the most cost-effective pooling strategy tested. Phylogenetic relationships estimated with >300 nuclear loci were well resolved, providing substantial improvement over the mtDNA barcode. We provide protocols for sample preparation and sequence capture including custom probe kit sequences, and describe our bioinformatics pipeline using aTRAM 2.0, PHYLUCE, and custom Perl and Python scripts. This approach can be applied to the tens of thousands of avian samples that have already been screened for haemosporidians, and greatly improve our understanding of parasite speciation, biogeography, and evolutionary dynamics.

genomics

Sexual selection rewires reproductive protein networks

Polyandry drives postcopulatory sexual selection (PCSS), resulting in rapid evolution of male ejaculate traits. Critical to male and female fitness, the ejaculate is known to contain rapidly evolving seminal fluid proteins (SFPs) produced by specialized male secretory accessory glands. The evidence that rapid evolution of some SFPs is driven by PCSS, however, is indirect, based on either plastic responses to changes in the sexual selection environment or correlative macroevolutionary patterns. Moreover, such studies focus on SFPs that represent but a small component of the accessory gland proteome. Neither how SFPs function with other reproductive proteins, nor how PCSS influences the underlying secretory tissue adaptations and content of the accessory gland, has been addressed at the level of the proteome. Here we directly test the hypothesis that PCSS results in rapid evolution of the entire male accessory gland proteome and protein networks by taking a system-level approach, combining divergent experimental evolution of PCSS in Drosophila pseudoobscura (Dpse), high resolution mass spectrometry (MS) and proteomic discovery, bioinformatics and population genetic analyses. We demonstrate that PCSS influences the abundance of over 200 accessory gland proteins, including SFPs. A small but significant number of these proteins display molecular signatures of positive selection. Divergent PCSS also results in fundamental and remarkably compartmentalized evolution of accessory gland protein networks in which males subjected to strong PCSS invest in protein networks that serve to increase protein production whereas males subjected to relaxed PCSS alters protein networks involved in protein surveillance and quality. These results directly demonstrate that PCSS is a key evolutionary driver that shapes not only individual reproductive proteins, but rewires entire reproductive protein networks.\n\nThe abbreviations used are

evolutionary biology

The UCSC Repeat Browser allows discovery and visualization of evolutionary conflict across repeat families

BackgroundNearly half the human genome consists of repeat elements, most of which are retrotransposons, and many of these sequences play important biological roles. However repeat elements pose several unique challenges to current bioinformatic analyses and visualization tools, as short repeat sequences can map to multiple genomic loci resulting in their misclassification and misinterpretation. In fact, sequence data mapping to repeat elements are often discarded from analysis pipelines. Therefore, there is a continued need for standardized tools and techniques to interpret genomic data of repeats. ResultsWe present the UCSC Repeat Browser, which consists of a complete set of human repeat reference sequences derived from the gold standard repeat database RepeatMasker. The UCSC Repeat Browser contains mapped annotations from the human genome to these references, and presents all of them as a comprehensive interface to facilitate work with repetitive elements. Furthermore, it provides processed tracks of multiple publicly available datasets of biological interest to the repeat community, including ChIP-SEQ datasets for KRAB Zinc Finger Proteins (KZNFs) - a family of proteins known to bind and repress certain classes of repeats. Here we show how the UCSC Repeat Browser in combination with these datasets, as well as RepeatMasker annotations in several non-human primates, can be used to trace the independent trajectories of species-specific evolutionary conflicts. ConclusionsThe UCSC Repeat Browser allows easy and intuitive visualization of genomic data on consensus repeat elements, circumventing the problem of multi-mapping, in which sequencing reads of repeat elements map to multiple locations on the human genome. By developing a reference consensus, multiple datasets and annotation tracks can easily be overlaid to reveal complex evolutionary histories of repeats in a single interactive window. Specifically, we use this approach to retrace the history of several primate specific LINE-1 families across apes, and discover several species-specific routes of evolution that correlate with the emergence and binding of KZNFs.

genomics

Identification of the bacterial biosynthetic gene clusters of the oral microbiome illuminates the unexplored social language of bacteria during health and disease

Small molecules are the primary communication media of the microbial world. Recent bioinformatics studies, exploring the biosynthetic gene clusters (BGCs) which produce many small molecules, have highlighted the incredible biochemical potential of the signaling molecules encoded by the human microbiome. Thus far, most research efforts have focused on understanding the social language of the gut microbiome, leaving crucial signaling molecules produced by oral bacteria, and their connection to health versus disease, in need of investigation. In this study, a total of 4,915 BGCs were identified across 461 genomes representing a broad taxonomic diversity of oral bacteria. Sequence similarity networking provided a putative product class for over 100 unclassified novel BGCs. The newly identified BGCs were cross-referenced against 254 metagenomes and metatranscriptomes derived from individuals with either good oral health, dental caries, or periodontitis. This analysis revealed 2,473 BGCs, which were differentially represented across the oral microbiomes associated with health versus disease. Co-abundance network analysis identified numerous inverse correlations between BGCs and specific oral taxa. These correlations were present in health, but greatly reduced in dental caries, which may suggest a defect in colonization resistance. Finally, corroborating mass spectrometry identified several compounds with homology to products of the predicted BGC classes. Together, these findings greatly expand the number of known biosynthetic pathways present in the oral microbiome and provide an atlas for experimental characterization of these abundant, yet poorly understood, molecules and socio-chemical relationships, which impact the development of caries and periodontitis, two of the worlds most common chronic diseases.\n\nIMPORTANCEThe healthy oral microbiome is symbiotic with the human host, importantly providing colonization resistance against potential pathogens. Dental caries and periodontitis are two of the worlds most common and costly chronic infectious diseases, and are caused by a localized dysbiosis of the oral microbiome. Bacterially produced small molecules, often encoded by BGCs, are the primary communication media of bacterial communities, and play a crucial, yet largely unknown, role in the transition from health to dysbiosis. This study provides a comprehensive mapping of the BGC repertoire of the human oral microbiome and identifies major differences in health compared to disease. Furthermore, BGC representation and expression is linked to the abundance of particular oral bacterial taxa in health versus dental caries and periodontitis. Overall, this study provides a significant insight into the chemical communication network of the healthy oral microbiome, and how it devolves in the case of two prominent diseases.

microbiology

GRID - Genomics of Rare Immune Disorders: a highly sensitive and specific diagnostic gene panel for patients with primary immunodeficiencies

Primary Immune disorders affect 15,000 new patients every year in Europe. Genetic tests are usually performed on a single or very limited number of genes leaving the majority of patients without a genetic diagnosis. We designed, optimised and validated a new clinical diagnostic platform called GRID, Genomics of Rare Immune Disorders, to screen in parallel 279 genes, including 2015 IUIS genes, known to be causative of Primary Immune disorders (PID). Validation to clinical standard using more than 58,000 variants in 176 PID patients shows an excellent sensitivity, specificity. The customised and automated bioinformatics pipeline prioritises and reports pertinent Single Nucleotide Variants (SNVs), INsertions and DELetions (INDELs) as well as Copy Number Variants (CNVs). An example of the clinical utility of the GRID panel, is represented by a patient initially diagnosed with X-linked agammaglobulinemia due to a missense variant in the BTK gene with severe inflammatory bowel disease. GRID results identified two additional compound heterozygous variants in IL17RC, potentially driving the altered phenotype.

genomics

Quantitative proteomics of the 2016 WHO Neisseria gonorrhoeae reference strains surveys vaccine candidates and antimicrobial resistance determinants

The sexually transmitted disease gonorrhea (causative agent: Neisseria gonorrhoeae) remains an urgent public health threat globally due to the repercussions on reproductive health, high incidence, widespread antimicrobial resistance (AMR), and absence of a vaccine. To mine gonorrhea antigens and enhance our understanding of gonococcal AMR at the proteome level, we performed the first large-scale proteomic profiling of a diverse panel (n=15) of gonococcal strains, including the 2016 World Health Organization (WHO) reference strains. These strains show all existing AMR profiles, previously described in regard to phenotypic and reference genome characteristics, and are intended for quality assurance in laboratory investigations. Herein, these isolates were subjected to subcellular fractionation and labeling with tandem mass tags coupled to mass spectrometry and multi-combinatorial bioinformatics. Our analyses detected 901 and 723 common proteins in cell envelope and cytoplasmic subproteomes, respectively. We identified nine novel gonorrhea vaccine candidates. Expression and conservation of new and previously selected antigens were investigated. In addition, established gonococcal AMR determinants were evaluated for the first time using quantitative proteomics. Six new proteins, WHO_F_00238, WHO_F_00635, WHO_F_00745, WHO_F_01139, WHO_F_01144, and WHO_F_01226, were differentially expressed in all strains, suggesting that they represent global proteomic AMR markers, indicate a predisposition toward developing or compensating gonococcal AMR, and/or act as new antimicrobial targets. Finally, phenotypic clustering based on the isolates defined antibiograms and common differentially expressed proteins yielded seven matching clusters between established and proteome-derived AMR signatures. Together, our investigations provide a reference proteomics databank for gonococcal vaccine and AMR research endeavors, which enables microbiological, clinical, or epidemiological projects and enhances the utility of the WHO reference strains.

microbiology

Direct PCR amplification of 16S rRNA genes offers accelerated bacterial identification using the MinION™ nanopore sequencer

Rapid identification of bacterial pathogens is crucial for appropriate and adequate antibiotic treatment, which significantly improves patient outcomes. 16S ribosomal RNA (rRNA) gene amplicon sequencing has proven to be a powerful strategy for diagnosing bacterial infections. We have recently established a sequencing method and bioinformatics pipeline for 16S rRNA gene analysis utilizing the Oxford Nanopore Technologies MinION sequencer. In combination with our taxonomy annotation analysis pipeline, the system enabled the molecular detection of bacterial DNA in a reasonable timeframe for diagnostic purposes. However, purification of bacterial DNA from specimens remains a rate-limiting step in the workflow. To further accelerate the process of sample preparation, we adopted a direct PCR strategy that amplifies 16S rRNA genes from bacterial cell suspensions without DNA purification. Our results indicate that differences in cell wall morphology significantly affect direct PCR efficiency and sequencing data. Notably, mechanical cell disruption preceding direct PCR was indispensable for obtaining an accurate representation of the specimen bacterial composition. Furthermore, 16S rRNA gene analysis of mock polymicrobial samples indicated that primer sequence optimization is required to avoid preferential detection of particular taxa and to cover a broad range of bacterial species. This study establishes a relatively simple workflow for rapid bacterial identification via MinIONTM sequencing, which reduces the turnaround time from sample to result, and provides a reliable method that may be applicable to clinical settings.

microbiology

Defining the RNA Interactome by Total RNA-Associated Protein Purification

UV crosslinking can be used to identify precise RNA targets for individual proteins, transcriptome-wide. We sought to develop a technique to generate reciprocal data, identifying precise sites of RNA-binding proteome-wide. The resulting technique, total RNA-associated protein purification (TRAPP), was applied to yeast (S. cerevisiae) and bacteria (E. coli). In all analyses, SILAC labelling was used to quantify protein recovery in the presence and absence of irradiation. For S. cerevisiae, we also compared crosslinking using 254 nm (UVC) irradiation (TRAPP) with 4-thiouracil (4tU) labelling combined with ~350 nm (UVA) irradiation (PAR-TRAPP). Recovery of proteins not anticipated to show RNA-binding activity was substantially higher in TRAPP compared to PAR-TRAPP. As an example of preferential TRAPP-crosslinking, we tested enolase (Eno1) and demonstrated its binding to tRNA loops in vivo. We speculate that many protein-RNA interactions have biophysical effects on localization and/or accessibility, by opposing or promoting phase separation for highly abundant protein. Homologous metabolic enzymes showed RNA crosslinking in S. cerevisiae and E. coli, indicating conservation of this property. TRAPP allows alterations in RNA interactions to be followed and we initially analyzed the effects of weak acid stress. This revealed specific alterations in RNA-protein interactions; for example, during late 60S ribosome subunit maturation. Precise sites of crosslinking at the level of individual amino acids (iTRAPP) were identified in 395 peptides from 155 unique proteins, following phospho-peptide enrichment combined with a bioinformatics pipeline (Xi). TRAPP is quick, simple and scalable, allowing rapid characterization of the RNA-bound proteome in many systems.

cell biology

Illumina-based sequencing framework for accurate detection and mapping of influenza virus defective interfering particle-associated RNAs

The mechanisms and consequences of defective interfering particle (DIP) formation during influenza virus infection remain poorly understood. The development of next generation sequencing (NGS) technologies has made it possible to identify large numbers of DIP-associated sequences, providing a powerful tool to better understand their biological relevance. However, NGS approaches pose numerous technical challenges including the precise identification and mapping of deletion junctions in the presence of frequent mutation and base-calling errors, and the potential for numerous experimental and computational artifacts. Here we detail an Illumina-based sequencing framework and bioinformatics pipeline capable of generating highly accurate and reproducible profiles of DIP-associated junction sequences. We use a combination of simulated and experimental control datasets to optimize pipeline performance and demonstrate the absence of significant artifacts. Finally, we use this optimized pipeline to generate a high-resolution profile of DIP-associated junctions produced during influenza virus infection and demonstrate how this data can provide insight into mechanisms of DIP formation. This work highlights the specific challenges associated with NGS-based detection of DIP-associated sequences, and details the computational and experimental controls required for such studies.

microbiology

Galectin-1 promotes the invasion of bladder cancer urothelia through their matrix milieu

The progression of carcinoma of the urinary bladder involves migration of cancer epithelia through their surrounding tissue matrix microenvironment. This was experimentally confirmed when a gender- and grade-diverse set of bladder cancer cell lines were cultured in pathomimetic three-dimensional laminin-rich environments. The high-grade cells, particularly female, formed multicellular invasive morphologies in 3D. In comparison, low- and intermediate-grade counterparts showed growth-restricted phenotypes. A proteomic approach combining mass spectrometry and bioinformatics analysis identified the estrogen-driven lactose-binding lectin Galectin-1 (GAL-1) as a putative candidate that could drive this invasion. Expression of LGALS1, the gene encoding GAL-1 showed an association with tumor grade progression in bladder cell lines. Immunohisto- and cyto-chemical experiments suggested greater extracellular levels of GAL-1 in 3D cultures of high-grade bladder cells and cancer tissues. High levels of GAL-1 associated with increased proliferation- and adhesion- of bladder cancer cells when grown on laminin-rich matrices. Pharmacological inhibition and Gal-1 knockdown in high-grade female cells decreased their adhesion to, and viability on, laminin-rich substrata. Higher GAL-1 also correlated with reduced E-cadherin and increased N-cadherin levels in consonance with a mesenchymal-like phenotype that we observed in 3D culture. The inhibition of GAL-1 reversed the stellate invasive phenotype to a more growth-restricted one in high-grade cells embedded within both basement-membrane-like and stromal collagenous matrix scaffolds. Finally, inhibition of GAL-1 specifically altered cell surface sialic acids, suggesting the mechanism by which the levels of GAL-1 may underlie the aggression and poor prognosis of invasive bladder cancer, especially in women.

cancer biology

Robust Estimation of the Phylogenetic Origin of Plastids Using a tRNA-Based Phyloclassifier

The trait of oxygenic photosynthesis was acquired by the last common ancestor of Archaeplastida through endosymbiosis of the cyanobacterial progenitor of modern-day plastids. Although a single origin of plastids by endosymbiosis is broadly supported, recent phylogenomic studies report contradictory evidence that plastids branch either early or late within the cyanobacterial Tree of Life. Here we describe CYANO-MLP, a general-purpose phyloclassifier of cyanobacterial genomes implemented using a Multi-Layer Perceptron. CYANO-MLP exploits consistent phylogenetic signals in bioinformatically estimated structure-function maps of tRNAs. CYANO-MLP accurately classifies cyanobacterial genomes into one of eight well-supported cyanobacterial clades in a manner that is robust to missing data, unbalanced data and variation in model specification. CYANO-MLP supports a late-branching origin of plastids: we classify 99.32% of 440 plastid genomes into one of two late-branching cyanobacterial clades with strong statistical support, and confidently assign 98.41% of plastid genomes to one late-branching clade containing unicellular starch-producing marine/freshwater diazotrophic Cyanobacteria. CYANO-MLP correctly classifies the chromatophore of Paulinella chromatophora and rejects a sister relationship between plastids and the early-branching cyanobacterium Gloeomargarita lithophora. We show that recently applied phylogenetic models and character recoding strategies fit cyanobacterial/plastid phylogenomic datasets poorly, because of heterogeneity both in substitution processes over sites and compositions over lineages.

evolutionary biology

Metatranscriptome profiling of the dynamic transcription of mRNA and sRNA of a probiotic Lactobacillus strain in human gut

Metatranscriptomic sequencing has recently been applied to study how pathogens and probiotics affect human gastrointestinal (GI) tract microbiota, which provides new insights into their mechanisms of action. In this study, metatranscriptomic sequencing was applied to deduce the in vivo expression patterns of an ingested Lactobacillus casei strain, which was compared with its in vitro growth transcriptomes. Extraction of the strain-specific reads revealed that transcripts from the ingested L. casei were increased, while those from the resident L. paracasei strains remained unchanged. Mapping of all metatranscriptomic reads and transcriptomic reads to L. casei genome showed that gene expression in vitro and in vivo differed dramatically. About 39% (1163) mRNAs and 45% (93) sRNAs of L. casei well-expressed were repressed after ingested into human gut. Expression of ABC transporter genes and amino acid metabolism genes was induced at day-14 of ingestion; and genes for sugar and SCFA metabolisms were activated at day-28 of ingestion. Moreover, expression of sRNAs specific to the in vitro log phase was more likely to be activated in human gut. Expression of rli28c sRNA with peaked expression during the in vitro stationary phase was also activated in human gut; this sRNA repressed L. casei growth and lactic acid production in vitro. These findings implicate that the ingested L. casei might have to successfully change its transcription patterns to survive in human gut, and the time-dependent activation patterns indicate a highly dynamic cross-talk between the probiotic and human gut including its microbe community.\n\nImportanceProbiotic bacteria are important in food industry and as model microorganisms in understanding bacterial gene regulation. Although probiotic functions and mechanisms in human gastrointestinal tract are linked to the unique probiotic gene expression, it remains elusive how transcription of probiotic bacteria is dynamically regulated after being ingested. Previous study of probiotic gene expression in human fecal samples has been restricted due to its low abundance and the presence of of closely related species. In this study, we took the advantage of the good depth of metatranscriptomic sequencing reads and developed a strain-specific read analysis method to discriminate the transcription of the probiotic Lactobacillus casei and those of its resident relatives. This approach and additional bioinformatics analysis allowed the first study of the dynamic transcriptome profiles of probiotic L casei in vivo. The novel findings indicate a highly regulated repression and dynamic activation of probiotic genome in human GI tract.

microbiology

A novel signature derived from immunoregulatory and hypoxia genes predicts prognosis in liver and five other cancers

BackgroundDespite much progress in cancer research, its incidence and mortality continue to rise. A robust biomarker that would predict tumor behavior is highly desirable and could improve patient treatment and prognosis.\n\nMethodsIn a retrospective bioinformatics analysis involving patients with liver cancer (n=839), we developed a prognostic signature consisting of 45 genes associated with tumor-infiltrating lymphocytes and cellular responses to hypoxia. From this gene set, we were able to identify a second prognostic signature comprised of 8 genes. Its performance was further validated in five other cancers: head and neck (n=520), renal papillary cell (n=290), lung (n=515), pancreas (n=178) and endometrial (n=370).\n\nFindingsThe 45-gene signature predicted overall survival in three liver cancer cohorts: hazard ratio (HR)=1.82, P=0.006; HR=1.84, P=0.008 and HR=2.67, P=0.003. Additionally, the reduced 8-gene signature was sufficient and effective in predicting survival in liver and five other cancers: liver (HR=2.36, P=0.0003; HR=2.43, P=0.0002 and HR=3.45, P=0.0007), head and neck (HR=1.64, P=0.004), renal papillary cell (HR=2.31, P=0.04), lung (HR=1.45, P=0.03), pancreas (HR=1.96, P=0.006) and endometrial (HR=2.33, P=0.003). Receiver operating characteristic analyses demonstrated both signatures superior performance over current tumor staging parameters. Multivariate Cox regression analyses revealed that both 45-gene and 8-gene signatures were independent of other clinicopathological features in these cancers. Combining the gene signatures with somatic mutation profiles increased their prognostic ability.\n\nConclusionsThis study, to our knowledge, is the first to identify a gene signature uniting both tumor hypoxia and lymphocytic infiltration as a prognostic determinant in six cancer types (n=2,712). The 8-gene signature can be used for patient risk stratification by incorporating hypoxia information to aid clinical decision making.

cancer biology