Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Genome ARTIST: a robust, high-accuracy aligner tool for mapping transposon insertions and self-insertions

A critical topic of insertional mutagenesis experiments performed on model organisms is mapping the hits of artificial transposons (ATs) at nucleotide level accuracy. Obviously, mapping errors may occur when sequencing artifacts or mutations as SNPs and small indels are present very close to the junction between a genomic sequence and a transposon inverted repeat (TIR). Another particular item of insertional mutagenesis is mapping of the transposon self-insertions and, to our best knowledge, there is no publicly available mapping tool designed to analyze such molecular events. We developed Genome ARTIST, a pairwise gapped aligner tool which works out both issues by means of an original, robust mapping strategy. Genome ARTIST is not designed to use NGS data but to analyze ATs insertions obtained in small to medium-scale mutagenesis experiments. Genome ARTIST employs a heuristic approach to find DNA sequence similarities and harnesses a multi-step implementation of a Smith-Waterman adapted algorithm to compute the mapping alignments. The experience is enhanced by easily customizable parameters and a user-friendly interface that describes the genomic landscape surrounding the insertion. Genome ARTIST deals with many genomes of bacteria and eukaryotes available in Ensembl and GenBank repositories. Our tool specifically harnesses/exploits the sequence annotation data provided by FlyBase for Drosophila melanogaster (the fruit fly), which enables mapping of insertions relative to various genomic features such as natural transposons. Genome ARTIST was tested against other alignment tools using relevant query sequences derived from the D. melanogaster and Mus musculus (mouse) genomes. Real and simulated query sequences were also comparatively inquired, revealing that Genome ARTIST is a very robust solution for mapping transposon insertions.\n\nGenome ARTIST is a stand-alone user-friendly application, designed for high-accuracy mapping of transposon insertions and self-insertions. The tool is also useful for routine aligning assessments like detection of SNPs or checking the specificity of primers and probes. Genome ARTIST is an open source software and is available for download at www.genomeartist.ro and at www.bioinformatics.org.

Bioinformatics

SFS_CODE: More Efficient and Flexible Forward Simulations

SUMMARYModern implementations of forward population genetic simulations are efficient and flexible, enabling the exploration of complex models that may otherwise be intractable. Here we describe an updated version of SFS_CODE, which has increased efficiency and includes many novel features. Among these features is an arbitrary model of dominance, the ability to simulate partial and soft selective sweeps, as well as track the trajectories of mutations and/or ancestries across multiple populations under complex models that are not possible under a coalescent framework. We also release sfs_coder, a Python wrapper to SFS_CODE allowing the user to easily generate command lines for common models of demography, selection, and human genome structure, as well as parse and simulate phenotypes from SFS_CODE output.\n\nAvailability and ImplementationOur open source software is written in C and Python, and are available under the GNU General Public License at http://sfscode.sourceforge.net.\n\nContactryan.hernandez@ucsf.edu\n\nSupplementary informationDetailed usage information is available from the project website at http://sfscode.sourceforge.net.

Bioinformatics

A Graph Theoretical Approach to Data Fusion

The rapid development of high throughput experimental techniques has resulted in a growing diversity of genomic datasets being produced and requiring analysis. A variety of computational techniques allow us to analyse such data and to model the biological processes behind them. However, it is increasingly being recognised that we can gain deeper understanding by combining the insights obtained from multiple, diverse datasets. We therefore require scalable computational approaches for data fusion.\n\nWe propose a novel methodology for scalable unsupervised data fusion. Our technique exploits network representations of the data in order to identify (and quantify) similarities among the datasets. We may work within the Bayesian formalism, using Bayesian nonparametric approaches to model each dataset; or (for fast, approximate, and massive scale data fusion) can naturally switch to more heuristic modelling techniques. An advantage of the proposed approach is that each dataset can initially be modelled independently (and therefore in parallel), before applying a fast post-processing step in order to perform data fusion. This allows us to incorporate new experimental data in an online fashion, without having to rerun all of the analysis. The methodology can be applied to genomic scale datasets and we demonstrate its applicability on examples from the literature, using a broad range of genomic datasets, and also on a recent gene expression dataset from Sporadic inclusion body myositis Availability. Example R code and instructions are available from https://sites.google.com/site/gtadatafusion/.

Bioinformatics

Differential transcript usage from RNA-seq data: isoform pre-filtering improves performance of count-based methods

Large-scale sequencing of cDNA (RNA-seq) has been a boon to the quantitative analysis of transcriptomes. A notable application is the detection of changes in transcript usage between experimental conditions. For example, discovery of pathological alternative splicing may allow the development of new treatments or better management of patients. From an analysis perspective, there are several ways to approach RNA-seq data to unravel differential transcript usage, such as annotation-based exon-level counting, differential analysis of the percent spliced in measure or quantitative analysis of assembled transcripts. The goal of this research is to compare and contrast current state-of-the-art methods, as well as to suggest improvements to commonly used workflows.\n\nWe assess the performance of representative workflows using synthetic data and explore the effect of using non-standard counting bin definitions as input to a state-of-the-art inference engine (DEXSeq). Although the canonical counting provided the best results overall, several non-canonical approaches were as good or better in specific aspects and most counting approaches outperformed the evaluated event- and assembly-based methods. We show that an incomplete annotation catalog can have a detrimental effect on the ability to detect differential transcript usage in transcriptomes with few isoforms per gene and that isoform-level pre-filtering can considerably improve false discovery rate (FDR) control.\n\nCount-based methods generally perform well in detection of differential transcript usage. Controlling the FDR at the imposed threshold is difficult, mainly in complex organisms, but can be improved by pre-filtering of the annotation catalog.

Bioinformatics

Modelizing Drosophila melanogaster longevity curves using a new discontinuous 2-Phases of Aging model

Aging is commonly described as being a continuous process affecting progressively organisms as time passes. This process results in a progressive decrease in individuals fitness through a wide range of both organismal - decreased motor activity, fertility, resistance to stress - and molecular phenotypes - decreased protein and energy homeostasis, impairment of insulin signaling. In the past 20 years, numerous genes have been identified as playing a major role in the aging process, yet little is known about the events leading to that loss of fitness. We recently described an event characterized by a dramatic increase of intestinal permeability to a blue food dye in aging flies committed to die within a few days. Importantly, flies showing this so called Smurf phenotype are the only ones, among a population, to show various age-related changes and exhibit a high-risk of impending death whatever their chronological age. Thus, these observations suggest that instead of being one continuous phenomenon, aging may be a discontinuous process well described by at least two distinguishable phases. In this paper we addressed this hypothesis by implementing a new 2-Phases of Aging mathematiCal model (2PAC model) to simulate longevity curves based on the simple hypothesis of two consecutive phases of lifetime presenting different properties. We first present a unique equation for each phase and discuss the biological significance of the 3 associated parameters. Then we evaluate the influence of each parameter on the shape of survival curves. Overall, this new mathematical model, based on simple biological observations, is able to reproduce many experimental longevity curves, supporting the existence of 2-phases of aging exhibiting specific properties and separated by a dramatic transition that remains to be characterized. Moreover, it indicates that Smurf survival can be approximated by one single constant parameter for a broad range of genotypes that we have tested under our environmental conditions.\n\nAuthor SummaryThe perception we can have of a process directly affects the way we study it. In the literature, aging is generally described as being a continuous process, progressively affecting organisms through a broad range of molecular and physiological changes ultimately leading to a dramatic decrease of individuals life expectancy. As such, aging studies focus on changes occurring in groups of individuals through time, considering individuals taken at a given time as being all equivalent. Instead, the recently described Smurf phenotype [1] suggested that any given time, a population could be divided in two subpopulations each characterized by a significantly different risk of impending death.\n\nBy formalizing here the concept of a discontinuous aging process using a mathematical model based on simple experimental observations, we propose a theoretical framework in which aging is actually separated in two consecutive phases characterized by three parameters easily quantifiable in vivo. Thus, the model we present here brings new tools to assess the events occurring during aging using a novel angle that we hope will open a better understanding of the very processes driving aging.

Bioinformatics

OEFinder: A user interface to identify and visualize ordering effects in single-cell RNA-seq data

A recent paper identified an artifact in multiple single-cell RNA-seq (scRNA-seq) data sets generated by the Fluidigm C1 platform. Specifically, Leng* et al. showed significantly increased gene expression in cells captured from sites with small or large plate output IDs. We refer to this artifact as an ordering effect (OE). Including OE genes in downstream analyses could lead to biased results. To address this problem, we developed a statistical method and software called OEFinder to identify a sorted list of OE genes. OEFinder is available as an R package along with user-friendly graphical interface implementations that allows users to check for potential artifacts in scRNA-seq data generated by the Fluidigm C1 platform.\n\nAvailability and ImplementationOEFinder is freely available at https://github.com/lengning/OEFinder\n\nContactrstewart@morgridge.org

Bioinformatics

Genome scaffolding with PE-contaminated mate-pair libraries

Scaffolding is often an essential step in a genome assembly process, in which contigs are ordered and oriented using read pairs from a combination of paired-ends libraries and longer-range mate-pair libraries. Although a simple idea, scaffolding is unfortunately hard to get right in practice. One source of problem is so-called PE-contamination in mate-pair libraries, in which a non-negligible fraction of the read pairs get the wrong orientation and a much smaller insert size than what is expected. This contamination has been discussed in previous work on integrated scaffolders in end-to-end assemblers such as Allpaths-LG and MaSuRCA but the methods relies on the fact that the orientation is observable, e.g., by finding the junction adapter sequence in the reads. This is not always the case, making orientation and insert size of a read pair stochastic. Furthermore, work on modeling PE-contamination has so far been disregarded in stand-alone scaffolders and the effect that PE-contamination has on scaffolding quality has not been examined before.\n\nWe have addressed PE-contamination in an update of our scaffolder BESST. We formulate the problem as an Integer Linear Program (ILP) and use characteristics of the problem, such as contig lengths and insert size, to efficiently solve the ILP using a linear amount (with respect to the number of contigs) of Linear Programs. Our results show significant improvement over both integrated and standalone scaffolders. The impact of modeling PE-contamination is quantified by comparison with the previous BESST model. We also show how other scaffolders are vulnerable to PE-contaminated libraries, resulting in increased number of misassemblies, more conservative scaffolding, and inflated assembly sizes.\n\nThe model is implemented in BESST. Source code and usage instructions are found at https://github.com/ksahlin/BESST. BESST can also be downloaded using PyPI.

Bioinformatics

Simultaneously inferring T cell fate and clonality from single cell transcriptomes

The heterodimeric T cell receptor (TCR) comprises two protein chains that pair to determine the antigen specificity of each T lymphocyte. The enormous sequence diversity within TCR repertoires allows specific TCR sequences to be used as lineage markers for T cells that derive from a common progenitor. We have developed a computational method, called TraCeR, to reconstruct full-length, paired TCR sequences from T lymphocyte single-cell RNA-seq by combining existing assembly and alignment programs with a \"synthetic genome\" library comprising all possible TCR sequences. We validate this method with PCR to quantify its accuracy and sensitivity, and compare to other TCR sequencing methods. Our inferred TCR sequences reveal clonal relationships between T cells, which we put into the context of each cells functional state from the complete transcriptional landscape quantified from the remaining RNA-seq data. This provides a powerful tool to link T cell specificity with functional response in a variety of normal and pathological conditions. We demonstrate this by determining the distribution of members of expanded T cell clonotypes in response to Salmonella infection in the mouse. We show that members of the same clonotype span early activated CD4+ T cells, as well as mature effector and memory cells.

Bioinformatics

Modeling of RNA-seq fragment sequence bias reduces systematic errors in transcript abundance estimation

RNA-seq technology is widely used in biomedical and basic science research. These studies rely on complex computational methods that quantify expression levels for observed transcripts. We find that current computational methods can lead to hundreds of false positive results related to alternative isoform usage. This flaw in the current methodology stems from a lack of modeling sample-specific bias that leads to drops in coverage and is related to sequence features like fragment GC content and GC stretches. By incorporating features that explain this bias into transcript expression models, we greatly increase the specificity of transcript expression estimates, with more than a four-fold reduction in the number of false positives for reported changes in expression. We introduce alpine, a method for estimation of bias-corrected transcript abundance. The method is available as a Bioconductor package that includes data visualization tools useful for bias discovery.

Bioinformatics

Teaser: Individualized benchmarking and optimization of read mapping results for NGS data

Mapping reads to a genome remains challenging, especially for non-model organisms with poorer quality assemblies, or for organisms with higher rates of mutations. While most research has focused on speeding up the mapping process, little attention has been paid to optimize the choice of mapper and parameters for a users dataset. Here we present Teaser, which assists in these choices through rapid automated benchmarking of different mappers and parameter settings for individualized data. Within minutes, Teaser completes a quantitative evaluation of an ensemble of mapping algorithms and parameters. Using Teaser, we demonstrate how Bowtie2 can be optimized for different data.

Bioinformatics

ProtAnnot: an App for Integrated Genome Browser to display how alternative splicing and transcription affect proteins

SummaryOne gene can produce multiple transcript variants encoding proteins with different functions. To facilitate visual analysis of transcript variants, we developed ProtAnnot, which shows protein annotations in the context of genomic sequence. ProtAnnot searches InterPro and displays profile matches (protein annotations) alongside gene models, exposing how alternative promoters, splicing, and 3 end processing add, remove, or remodel functional motifs. To draw attention to these effects, ProtAnnot color-codes exons by frame and displays a cityscape graphic summarizing exonic sequence at each position. These techniques make visual analysis of alternative transcripts faster and more convenient for biologists. Availability and ImplementationProtAnnot is a plug-in App for Integrated Genome Browser, an open source desktop genome browser available from http://www.bioviz.org. Contactaloraine@uncc.edu

Bioinformatics

Privacy-Preserving Microbiome Analysis Using Secure Computation

MotivationDeveloping targeted therapeutics and identifying biomarkers relies on large amounts of patient data. Beyond human DNA, researchers now investigate the DNA of micro-organisms inhabiting the human body. An individuals collection of microbial DNA consistently identifies that person and could be used to link a real-world identity to a sensitive attribute in a research dataset. Unfortunately, the current suite of DNA-specific privacy-preserving analysis tools does not meet the requirements for microbiome sequencing studies.\n\nResultsWe augment an existing categorization of genomic-privacy attacks to incorporate microbiome sequencing and provide an implementation of metagenomic analyses using secure computation. Our implementation allows researchers to perform analysis over combined data without revealing individual patient attributes. We implement three metagenomic analyses and perform an evaluation on real datasets for comparative analysis. We use our implementation to simulate sharing data between four policy-domains and measure the increase in significant discoveries. Additionally, we describe an application of our implementation to form patient pools of data to allow drug companies to query against and compensate patients for the analysis.\n\nAvailabilityThe software is freely available for download at: http://cbcb.umd.edu/[~]hcorrada/projects/secureseq.html

Bioinformatics

A Comparison of Methods: Normalizing High-Throughput RNA Sequencing Data

As RNA-Seq and other high-throughput sequencing grow in use and remain critical for gene expression studies, technical variability in counts data impedes studies of differential expression studies, data across samples and experiments, or reproducing results. Studies like Dillies et al. (2013) compare several between-lane normalization methods involving scaling factors, while Hansen et al. (2012) and Risso et al. (2014) propose methods that correct for sample-specific bias or use sets of control genes to isolate and remove technical variability. This paper evaluates four normalization methods in terms of reducing intra-group, technical variability and facilitating differential expression analysis or other research where the biological, inter-group variability is of interest. To this end, the four methods were evaluated in differential expression analysis between data from Pickrell et al. (2010) and Montgomery et al. (2010) and between simulated data modeled on these two datasets. Though the between-lane scaling factor methods perform worse on real data sets, they are much stronger for simulated data. We cannot reject the recommendation of Dillies et al. to use TMM and DESeq normalization, but further study of power to detect effects of different size under each normalization method is merited.

Bioinformatics

Revisiting inconsistency in large pharmacogenomic studies

BackgroundIn 2012, two large pharmacogenomic studies, the Genomics of Drug Sensitivity in Cancer (GDSC) and Cancer Cell Line Encyclopedia (CCLE), were published, each reported gene expression data and measures of drug response for a large number of drugs and hundreds of cell lines. In 2013, we published a comparative analysis that reported gene expression profiles for the 471 cell lines profiled in both studies and dose response measurements for the 15 drugs characterized in the common cell lines by both studies. While we found good concordance in gene expression profiles, there was substantial inconsistency in the drug responses reported by the GDSC and CCLE projects. Our paper was widely discussed and we received extensive feedback on the comparisons that we performed. This feedback, along with the release of new data, prompted us to revisit our initial analysis. Here we present a new analysis using these expanded data in which we address the most significant suggestions for improvements on our published analysis: that drugs with different response characteristics should have been treated differently, that targeted therapies and broad cytotoxic drugs should have been treated differently in assessing consistency, that consistency of both molecular profiles and drug sensitivity measurements should both be compared across cell lines to accurately assess differences in the studies, that we missed some biomarkers that are consistent between studies, and that the software analysis tools we provided with our analysis should have been easier to run, particularly as the GDSC and CCLE released additional data.\n\nMethodsFor each drug, we used published sensitivity data from the GDSC and CCLE to separately estimate drug dose-response curves. We then used two statistics, the area between drug dose-response curves (ABC) and the Matthews correlation coefficient (MCC), to robustly estimate the consistency of continuous and discrete drug sensitivity measures, respectively. We also used recently released RNA-seq data together with previously published gene expression microarray data to assess inter-platform reproducibility of cell line gene expression profiles.\n\nResultsThis re-analysis supports our previous finding that gene expression data are significantly more consistent than drug sensitivity measurements. The use of new statistics to assess data consistency allowed us to identify two broad effect drugs -- 17-AAG and PD-0332901 -- and three targeted drugs -- PLX4720, nilotinib and crizotinib -- with moderate to good consistency in drug sensitivity data between GDSC and CCLE. Not enough sensitive cell lines were screened in both studies to robustly assess consistency for three other targeted drugs, PHA-665752, erlotinib, and sorafenib. Concurring with our published results, we found evidence of inconsistencies in pharmacological phenotypes for the remaining eight drugs. Further, to discover \"consistency\" between studies required the use of multiple statistics and the selection of specific measures on a case-by-case basis.\n\nConclusionOur results reaffirm our initial findings of an inconsistency in drug sensitivity measures for eight of fifteen drugs screened both in GDSC and CCLE, irrespective of which statistical metric was used to assess correlation. Taken together, our findings suggest that the phenotypic data on drug response in the GDSC and CCLE continue to present challenges for robust biomarker discovery. This re-analysis provides additional support for the argument that experimental standardization and validation of pharmacogenomic response will be necessary to advance the broad use of large pharmacogenomic screens.

Bioinformatics

Evolution of genes neighborhood within reconciled phylogenies: an ensemble approach

ContextThe reconstruction of evolutionary scenarios for whole genomes in terms of genome rearrangements is a fundamental problem in evolutionary and comparative genomics. The DeCo algorithm, recently introduced by Berard et al., computes parsimonious evolutionary scenarios for gene adjacencies, from pairs of reconciled gene trees. However, as for many combinatorial optimization algorithms, there can exist many co-optimal, or slightly sub-optimal, evolutionary scenarios that deserve to be considered.\n\nContributionWe extend the DeCo algorithm to sample evolutionary scenarios from the whole solution space under the Boltzmann distribution, and also to compute Boltzmann probabilities for specific ancestral adjacencies.\n\nResultsWe apply our algorithms to a dataset of mammalian gene trees and adjacencies, and observe a significant reduction of the number of syntenic conflicts observed in the resulting ancestral gene adjacencies.

Bioinformatics

Integrated Genome Browser: visual analytics platform for genomics

MotivationGenome browsers that support fast navigation and interactive visual analytics can help scientists achieve deeper insight into large-scale genomic data sets more quickly, thus accelerating the discovery process. Toward this end, we developed Integrated Genome Browser (IGB), a highly configurable, interactive and fast open source desktop genome browser.\n\nResultsHere we describe multiple updates to IGB, including all-new capability to display and interact with data from high-throughput sequencing experiments. To demonstrate, we describe example visualizations and analyses of data sets from RNA-Seq, ChIP-Seq, and bisulfite sequencing experiments. Understanding results from genome-scale experiments requires viewing the data in the context of reference genome annotations and other related data sets. To facilitate this, we enhanced IGBs ability to consume data from diverse sources, including Galaxy, Distributed Annotation, and IGB-specific Quickload servers. To support future visualization needs as new genome-scale assays enter wide use, we transformed the IGB codebase into a modular, extensible platform for developers to create and deploy all-new visualizations of genomic data.\n\nAvailabilityIGB is open source and is freely available from http://bioviz.org/igb.\n\nContactaloraine@uncc.edu

Bioinformatics

pcaReduce: Hierarchical Clustering of Single Cell Transcriptional Profiles

MotivationAdvances in single cell genomics provides a way of routinely generating transcriptomics data at the single cell level. A frequent requirement of single cell expression experiments is the identification of novel patterns of heterogeneity across single cells that might explain complex cellular states or tissue composition. To date, classical statistical analysis tools have being routinely applied to single cell data, but there is considerable scope for the development of novel statistical approaches that are better adapted to the challenges of inferring cellular hierarchies.\n\nResultsHere, we present a novel integration of principal components analysis and hierarchical clustering to create a framework for characterising cell state identity. Our methodology uses agglomerative clustering to generate a cell state hierarchy where each cluster branch is associated with a principal component of variation that can be used to differentiate two cellular states. We demonstrate that using real single cell datasets this approach allows for consistent clustering of single cell transcriptional profiles across multiple scales of interpretation.\n\nAvailabilityR implementation of pcaReduce algorithm is available from https://github.com/JustinaZ/pcaReduce

Bioinformatics

Using Cell line and Patient samples to improve Drug Response Prediction

BackgroundRecent advances in high-throughput technologies have facilitated the profiling of large panels of cancer cell lines with responses measured for thousands of drugs. The computational challenge is now to realize the potential of these data in predicting patients responses to these drugs in the clinic.\n\nMethodsWe address this issue by examining the spectrum of prediction models of patient response: models predicting directly from cell lines, those predicting directly from patients, and those trained on cell lines and patients at the same time. We tested 21 classification models on four drugs, that are bortezomib, erlotinib, docetaxel and epirubicin, for which clinical trial data were available.\n\nResultsOur integrative models consistently outperform cell line-based predictors, indicating that there are limitations to the predictive potential of in vitro data alone. Furthermore, these integrative models achieve better predictive accuracy and require substantially fewer patients than would be the case if only patient data were available.\n\nConclusionsThe integration of in vitro and ex vivo genomic data results in more accurate predictors using only a fraction of the patient information, which can help optimize the development of personalized predictors of therapy response. Altogether our results support the relevance of preclinical data for therapy prediction in clinical trials, enabling more efficient and cost-effective trial design.

Bioinformatics