Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Genome-wide DNA methylome analysis reveals novel epigenetically dysregulated non-coding RNAs in human breast cancer

The development of human breast cancer is driven by changes in the genetic and epigenetic landscape of the cell. Despite growing appreciation of the importance of epigenetics in breast cancers, our knowledge of epigenetic alterations of non-coding RNAs (ncRNAs) in breast cancers remains limited. Here, we explored the epigenetic patterns of ncRNAs in breast cancers via a sequencing-based comparative methylome analysis, mainly focusing on two most popular ncRNA biotypes, long non-coding RNAs (lncRNAs) and miRNAs. Besides global hypomethylation and extensive CpG islands (CGIs) hypermethylation, we observed widely aberrant methylation in the promoters of ncRNAs, which was higher than that of protein-coding genes. Specifically, intergenic ncRNAs were observed to contribute a large slice of the aberrantly methylated ncRNA promoters. Moreover, we summarized five patterns of ncRNA promoter aberrant methylation in the context of genomic CGIs, where aberrant methylation occurred not only on the CGIs, but also flanking regions and CGI sparse promoters. Integration with transcriptional datasets, we found that the ncRNA promoter methylation events were associated with transcriptional changes. Furthermore, a panel of ncRNAs were identified as biomarkers that were able to discriminate between disease phenotypes (AUCs>0.90). Finally, the potential functions for aberrantly methylated ncRNAs were predicted based on similar patterns, adjacency and/or target genes, highlighting that ncRNAs and coding genes coordinately mediated pathways dysregulation in the development and progression of breast cancers. This study presents the aberrant methylation patterns of ncRNAs, which will be a highly valuable resource for investigations at understanding epigenetic regulation of breast cancers.\n\n[Supplemental material is available online at www.genome.org.]

Bioinformatics

Modeling bi-modality improves characterization of cell cycle on gene expression in single cells

Advances in high-throughput, single cell gene expression are allowing interrogation of cell heterogeneity. However, there is concern that the cell cycle phase of a cell might bias characterizations of gene expression at the single-cell level. We assess the effect of cell cycle phase on gene expression in single cells by measuring 333 genes in 930 cells across three phases and three cell lines. We determine each cells phase non-invasively without chemical arrest and use it as a covariate in tests of differential expression. We observe bi-modal gene expression, a previously-described phenomenon, wherein the expression of otherwise abundant genes is either strongly positive, or undetectable within individual cells. This bi-modality is likely both biologically and technically driven. Irrespective of its source, we show that it should be modeled to draw accurate inferences from single cell expression experiments. To this end, we propose a semi-continuous modeling framework based on the generalized linear model, and use it to characterize genes with consistent cell cycle effects across three cell lines. Our new computational framework improves the detection of previously characterized cell-cycle genes compared to approaches that do not account for the bi-modality of single-cell data. We use our semi-continuous modelling framework to estimate single cell gene co-expression networks. These networks suggest that in addition to having phase-dependent shifts in expression (when averaged over many cells), some, but not all, canonical cell cycle genes tend to be co-expressed in groups in single cells. We estimate the amount of single cell expression variability attributable to the cell cycle. We find that the cell cycle explains only 5%-17% of expression variability, suggesting that the cell cycle will not tend to be a large nuisance factor in analysis of the single cell transcriptome.

Bioinformatics

SNP-guided identification of monoallelic DNA-methylation events from enrichment-based sequencing data

Monoallelic gene expression is typically initiated early in the development of an organism. Dysregulation of monoallelic gene expression has already been linked to several non-Mendelian inherited genetic disorders. In humans, DNA-methylation is deemed to be an important regulator of monoallelic gene expression, but only few examples are known. One important reason is that current, cost-affordable truly genome-wide methods to assess DNA-methylation are based on sequencing post enrichment. Here, we present a new methodology that combines methylomic data from MethylCap-seq with associated SNP profiles to identify monoallelically methylated loci. Using the Hardy-Weinberg theorem for each SNP locus, it could be established whether the observed frequency of samples featured by biallelic methylation was lower than randomly expected. Applied on 334 MethylCap-seq samples of very diverse origin, this resulted in the identification of 80 genomic regions featured by monoallelic DNA-methylation. Of these 80 loci, 49 are located in genic regions of which 25 have already been linked to imprinting. Further analysis revealed statistically significant enrichment of these loci in promoter regions, further establishing the relevance and usefulness of the method. Additional validation of the found loci was done using 14 whole-genome bisulfite sequencing data sets. Importantly, the developed approach can be easily applied to other enrichment-based sequencing technologies, such as the ChIP-seq-based identification of monoallelic histone modifications.

Bioinformatics

A phase diagram for gene selection and disease classification

Identifying a small subset of discriminate genes is important for predicting clinical outcomes and facilitating disease diagnosis. Based on the model population analysis framework, we present a method, called PHADIA, which is able to output a phase diagram displaying the predictive ability of each variable, which provides an intuitive way for selecting informative variables. Using two publicly available microarray datasets, its demonstrated that our method can selects a few informative genes and achieves significantly better or comparable classification accuracy compared to the reported results in the literature. The source codes are freely available at: www.libpls.net.

Bioinformatics

Automated ensemble assembly and validation of microbial genomes

BackgroundThe continued democratization of DNA sequencing has sparked a new wave of development of genome assembly and assembly validation methods. As individual research labs, rather than centralized centers, begin to sequence the majority of new genomes, it is important to establish best practices for genome assembly. However, recent evaluations such as GAGE and the Assemblathon have concluded that there is no single best approach to genome assembly. Instead, it is preferable to generate multiple assemblies and validate them to determine which is most useful for the desired analysis; this is a labor-intensive process that is often impossible or unfeasible.\n\nResultsTo encourage best practices supported by the community, we present iMetAMOS, an automated ensemble assembly pipeline; iMetAMOS encapsulates the process of running, validating, and selecting a single assembly from multiple assemblies. iMetAMOS packages several leading open-source tools into a single binary that automates parameter selection and execution of multiple assemblers, scores the resulting assemblies based on multiple validation metrics, and annotates the assemblies for genes and contaminants. We demonstrate the utility of the ensemble process on 225 previously unassembled Mycobacterium tuberculosis genomes as well as a Rhodobacter sphaeroides benchmark dataset. On these real data, iMetAMOS reliably produces validated assemblies and identifies potential contamination without user intervention. In addition, intelligent parameter selection produces assemblies of R. sphaeroides that exceed the quality of those from the GAGE-B evaluation, affecting the relative ranking of some assemblers.\n\nConclusionsEnsemble assembly with iMetAMOS provides users with multiple, validated assemblies for each genome. Although computationally limited to small or mid-sized genomes, this approach is the most effective and reproducible means for generating high-quality assemblies and enables users to select an assembly best tailored to their specific needs.

Bioinformatics

Sashimi plots: Quantitative visualization of alternative isoform expression from RNA-seq data

To the Editor: To the Editor: Software documentation and... References Analysis of RNA sequencing (RNA-Seq) data revealed that the vast majority of human genes express multiple mRNA isoforms, produced by alternative pre-mRNA splicing and other mechanisms, and that most alternative isoforms vary in expression between human tissues (Pan et al., 2008; Wang et al., 2008). As RNA-Seq datasets grow in size, it remains challenging to visualize isoform expression across multiple samples. We present Sashimi plots, a quantitative multi-sample visualization of RNA-Seq reads aligned to gene annotations, which enables quantitative comparison of isoform usage across samples or experimental conditions. Given an input annotation and spliced alignments of reads from a sample, a region of interest is visualized in a Sashimi plot as follows: (i) alignments in ...

Bioinformatics

A reassessment of consensus clustering for class discovery

Consensus clustering (CC) is an unsupervised class discovery method widely used to study sample heterogeneity in high-dimensional datasets. It calculates \"consensus rate\" between any two samples as how frequently they are grouped together in repeated clustering runs under a certain degree of random perturbation. The pairwise consensus rates form a between-sample similarity matrix, which has been used (1) as a visual proof that clusters exist, (2) for comparing stability among clusters, and (3) for estimating the optimal number (K) of clusters. However, the sensitivity and specificity of CC have not been systemically studied. To assess its performance, we investigated the most common implementations of CC; and compared CC with other popular methods that also focus on cluster stability and estimation of K. We evaluated these methods using simulated datasets with either known structure or known absence of structure. Our results showed that (1) CC was able to divide randomly generated unimodal data into pre-specified numbers of clusters, and was able to show apparent stability of these chance partitions of known cluster-less data; (2) for data with known structure, the proportion of ambiguously clustered (PAC) pairs infers the known number of clusters more reliably than several commonly used K estimating methods; and (3) validation of the optimal K by choosing the most discriminant genes from the discovery cohort and applying them in an independent cohort often exaggerates the confidence in K due to inherent gene-gene correlations among the selected genes. While these results do not yet prove that any of the published studies using CC has generated false positive findings, they show that datasets with subtle or no structure are fully capable of producing strong evidence of consensus clustering. We therefore recommend caution is using CC in class discovery and validation.

Bioinformatics

Spectacle: Faster and more accurate chromatin state annotation using spectral learning

Recently, a wealth of epigenomic data has been generated by biochemical assays and next-generation sequencing (NGS) technologies. In particular, histone modification data generated by the ENCODE project and other large-scale projects show specific patterns associated with regulatory elements in the human genome. It is important to build a unified statistical model to decipher the patterns of multiple histone modifications in a cell type to annotate chromatin states such as transcription start sites, enhancers and transcribed regions rather than to map histone modifications individually to regulatory elements.\n\nSeveral genome-wide statistical models have been developed based on hidden Markov models (HMMs). These methods typically use the Expectation-Maximization (EM) algorithm to estimate the parameters of the model. Here we used spectral learning, a state-of-the-art parameter estimation algorithm in machine learning. We found that spectral learning plus a few (up to five) iterations of local optimization of the likelihood outper-forms the standard EM algorithm. We also evaluated our software implementation called Spectacle on independent biological datasets and found that Spectacle annotated experimentally defined functional elements such as enhancers significantly better than a previous state-of-the-art method.\n\nSpectacle can be downloaded from https://github.com/jiminsong/Spectacle.

Bioinformatics

Large-scale non-targeted metabolomic profiling in three human population-based studies

Metabolomic profiling is an emerging technique in life sciences. Human studies using these techniques have been performed in a small number of individuals or have been targeted at a restricted number of metabolites. In this article, we propose a data analysis workflow to perform non-targeted metabolomic profiling in large human population-based studies using ultra performance liquid chromatography-mass spectrometry (UPLC-MS). We describe challenges and propose solutions for quality control, statistical analysis and annotation of metabolic features. Using the data analysis workflow, we detected more than 8,000 metabolic features in serum samples from 2,489 fasting individuals. As an illustrative example, we performed a non-targeted metabolome-wide association analysis of high-sensitive C-reactive protein (hsCRP) and detected 407 metabolic features corresponding to 90 unique metabolites that could be replicated in an external population. Our results reveal unexpected biological associations, such as metabolites identified as monoacylphosphorylcholines (LysoPC) being negatively associated with hsCRP. R code and fragmentation spectra for all metabolites are made publically available. In conclusion, the results presented here illustrate the viability and potential of non-targeted metabolomic profiling in large population-based studies.

Bioinformatics

HTSeq - A Python framework to work with high-throughput sequencing data

MotivationA large choice of tools exists for many standard tasks in the analysis of high-throughput sequencing (HTS) data. However, once a project deviates from standard work flows, custom scripts are needed.\n\nResultsWe present HTSeq, a Python library to facilitate the rapid development of such scripts. HTSeq offers parsers for many common data formats in HTS projects, as well as classes to represent data such as genomic coordinates, sequences, sequencing reads, alignments, gene model information, variant calls, and provides data structures that allow for querying via genomic coordinates. We also present htseq-count, a tool developed with HTSeq that preprocesses RNA-Seq data for differential expression analysis by counting the overlap of reads with genes.\n\nAvailabilityHTSeq is released as open-source software under the GNU General Public Licence and available from http://www-huber.embl.de/HTSeq or from the Python Package Index https://pypi.python.org/pypi/HTSeq.\n\nContactsanders@fs.tum.de

Bioinformatics

Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2

In comparative high-throughput sequencing assays, a fundamental task is the analysis of count data, such as read counts per gene in RNA-seq, for evidence of systematic changes across experimental conditions. Small replicate numbers, discreteness, large dynamic range and the presence of outliers require a suitable statistical approach. We present DESeq2, a method for differential analysis of count data, using shrinkage estimation for dispersions and fold changes to improve stability and interpretability of estimates. This enables a more quantitative analysis focused on the strength rather than the mere presence of differential expression. The DESeq2 package is available at http://www.bioconductor.org/packages/release/bioc/html/DESeq2.html.

Bioinformatics

A GWAS platform built on iPlant cyber-infrastructure

We demonstrated a flexible Genome-Wide Association Study (GWAS) platform built upon the iPlant Collaborative Cyber-infrastructure. The platform supports big data management, sharing, and large scale study of both genotype and phenotype data on clusters. End users can add their own analysis tools, and create customized analysis workflows through the graphical user interfaces in both iPlant Discovery Environment and BioExtract server.

Bioinformatics

Functional normalization of 450k methylation array data improves replication in large cancer studies

We propose an extension to quantile normalization which removes unwanted technical variation using control probes. We adapt our algorithm, functional normalization, to the Illumina 450k methylation array and address the open problem of normalizing methylation data with global epigenetic changes, such as human cancers. Using datasets from The Cancer Genome Atlas and a large case-control study, we show that our algorithm outperforms all existing normalization methods with respect to replication of results between experiments, and yields robust results even in the presence of batch effects. Functional normalization can be applied to any microarray platform, provided suitable control probes are available.

Bioinformatics

T-lex2: genotyping, frequency estimation and re-annotation of transposable elements using single or pooled next-generation sequencing data

Transposable elements (TEs) are the most active, diverse and ancient component in a broad range of genomes. As such, a complete understanding of genome function and evolution cannot be achieved without a thorough understanding of TE impact and biology. However, in-depth analyses of TEs still represent a challenge due to the repetitive nature of these genomic entities. In this work, we present a broadly applicable and flexible tool: T-lex2. T-lex2 is the only available software that allows routine,automatic, and accurate genotyping of individual TE insertions and estimation of their population frequencies both using individual strain and pooled next-generation sequencing (NGS) data. Furthermore, T-lex2 also assesses the quality of the calls allowing the identification of miss-annotated TEs and providing the necessary information to re-annotate them. Although we tested the fidelity of T-lex2 using the high quality Drosophila melanogaster genome, the flexible and customizable design of T-lex2 allows running it in any genome and for any type of TE insertion. Overall, T-lex2 represents a significant improvement in our ability to analyze the contribution of TEs to genome function and evolution as well as learning about the biology of TEs. T-lex2 is freely available online at http://petrov.stanford.edu/cgi-bin/Tlex.html.

Bioinformatics

Efficient synergistic single-cell genome assembly

As the vast majority of all microbes are unculturable, single-cell sequencing has become a significant method to gain insight into microbial physiology. Single-cell sequencing methods, currently powered by multiple displacement genome amplification (MDA), have passed important milestones such as finishing and closing the genome of a prokaryote. However, the quality and reliability of genome assemblies from single cells are still unsatisfactory due to uneven coverage depth and the absence of scattered chunks of the genome in the final collection of reads caused by MDA bias. In this work, our new algorithm Hybrid De novo Assembler (HyDA) demonstrates the power of co-assembly of multiple single-cell genomic data sets through significant improvement of the assembly quality in terms of predicted functional elements and length statistics. Co-assemblies contain significantly more base pairs and protein coding genes, cover more subsystems, and consist of longer contigs compared to individual assemblies by the same algorithm as well as state-of-the-art single-cell assemblers SPAdes and IDBA-UD. Hybrid De novo Assembler (HyDA) is also able to avoid chimeric assemblies by detecting and separating shared and exclusive pieces of sequence for input data sets. By replacing one deep single-cell sequencing experiment with a few single-cell sequencing experiments of lower depth, the co-assembly method can hedge against the risk of failure and loss of the sample, without significantly increasing sequencing cost. Application of the single-cell coassembler HyDA to the study of three uncultured members of an alkane-degrading methanogenic community validated the usefulness of the co-assembly concept.

Bioinformatics

Genomic Repeat Element Analyzer for Mammals (GREAM)

Background: Understanding the mechanism behind the transcriptional regulation of genes is still a challenge. Recent findings indicate that the genomic repeat elements (such as LINES, SINES and LTRs) could play an important role in the transcription control. Hence, it is important to further explore the role of genomic repeat elements in the gene expression regulation, and perhaps in other molecular processes. Although many computational tools exists for repeat element analysis, almost all of them simply identify and/or classifying the genomic repeat elements within query sequence(s); none of them facilitate identification of repeat elements that are likely to have a functional significance, particularly in the context of transcriptional regulation.\n\nResult: We developed the Genomic Repeat Element Analyzer for Mammals (GREAM) to allow gene-centric analysis of genomic repeat elements in 17 mammalian species, and validated it by comparing with some of the existing experimental data. The output provides a categorized list of the specific type of transposons, retro-transposons and other genome-wide repeat elements that are statistically over-represented across specific neighborhood regions of query genes. The position and frequency of these elements, within the specified regions, are displayed as well. The tool also offers queries for position-specific distribution of repeat elements within chromosomes. In addition, GREAM facilitates the analysis of repeat element distribution across the neighborhood of orthologous genes.\n\nConclusion: GREAM allows researchers to short-list the potentially important repeat elements, from the genomic neighborhood of genes, for further experimental analysis. GREAM is free and available for all at http://resource.ibab.ac.in/GREAM/

Bioinformatics

Detecting translational regulation by change point analysis of ribosome profiling datasets

Ribo-Seq maps the location of translating ribosomes on mature mRNA transcripts. While ribosome density is constant along the length of the mRNA coding region, it can be altered by translational regulatory events. In this study, we developed a method to detect translational regulation of individual mRNAs from their ribosome profiles, utilizing changes in ribosome density. We used mathematical modelling to show that changes in ribosome density should occur along the mRNA at the point of regulation. We analyzed a Ribo-Seq dataset obtained for mouse embryonic stem cells and showed that normalization by corresponding RNA-Seq can be used to improve the Ribo-Seq quality by removing bias introduced by deep-sequencing and alignment artefacts. After normalization, we applied a change point algorithm to detect changes in ribosome density present in individual mRNA ribosome profiles. Additional sequence and gene isoform information obtained from the UCSC Genome Browser allowed us to further categorize the detected changes into different mechanisms of regulation. In particular, we detected several mRNAs with known post-transcriptional regulation, e.g. premature termination for selenoprotein mRNAs and translational control of Atf4, but also several more mRNAs with hitherto unknown translational regulation. Additionally, our approach proved useful for identification of new gene isoforms.

Bioinformatics

Alignathon: A competitive assessment of whole genome alignment methods.

BackgroundMultiple sequence alignments (MSAs) are a prerequisite for a wide variety of evolutionary analyses. Published assessments and benchmark datasets for protein and, to a lesser extent, global nucleotide MSAs are available, but less effort has been made to establish benchmarks in the more general problem of whole genome alignment (WGA).\n\nResultsUsing the same model as the successful Assemblathon competitions, we organized a competitive evaluation in which teams submitted their alignments, and assessments were performed collectively after all the submissions were received. Three datasets were used: two of simulated primate and mammalian phylogenies, and one of 20 real fly genomes. In total 35 submissions were assessed, submitted by ten teams using 12 different alignment pipelines.\n\nConclusionsWe found agreement between independent simulation-based and statistical assessments, indicating that there are substantial accuracy differences between contemporary alignment tools. We saw considerable difference in the alignment quality of differently annotated regions, and found few tools aligned the duplications analysed. We found many tools worked well at shorter evolutionary distances, but fewer performed competitively at longer distances. We provide all datasets, submissions and assessment programs for further study, and provide, as a resource for future benchmarking, a convenient repository of code and data for reproducing the simulation assessments.

Bioinformatics