Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

Fuzzy-FishNet: A highly precise distribution-free network approach for feature selection in clinical proteomics

Network-based analysis methods can help resolve coverage and inconsistency issues in proteomics data. Previously, it was demonstrated that a suite of rank-based network approaches (RBNAs) provides unparalleled consistency and reliable feature selection. However, reliance on the t-statistic/t-distribution and hypersensitivity (coupled to a relatively flat p-value distribution) makes feature prioritization for validation difficult. To address these concerns, a refinement based on the fuzzified Fisher exact test, Fuzzy-FishNet was developed. Fuzzy-FishNet is highly precise (providing probability values that allows exact ranking of features). Furthermore, feature ranks are stable, even in small sample size scenario. Comparison of features selected by genomics and proteomics data respectively revealed that in spite of relative feature stability, cross-platform overlaps are extremely limited, suggesting that networks may not be the answer towards bridging the proteomics-genomics divide.

Bioinformatics

Toward high-throughput predictive modeling of protein binding/unbinding kinetics

One of the unaddressed challenges in drug discovery is that drug potency determined in vitro is not a reliable indicator of drug activity in humans. Accumulated evidences suggest that in vivo activity is more strongly correlated with the binding/unbinding kinetics than the equilibrium thermodynamics of protein-ligand interactions (PLI). However, existing experimental and computational techniques are insufficient in studying the molecular details of kinetics process of PLI. Consequently, we not only have limited mechanistic understanding of the kinetic process but also lack a practical platform for the high-throughput screening and optimization of drug leads based on their kinetic properties. Here we address this unmet need by integrating energetic and conformational dynamic features derived from molecular modeling with multi-task learning. To test our method, HIV-1 protease is used as a model system. Our integrated model provides us with new insights into the molecular determinants of kinetics of PLI. We find that the coherent coupling of conformational dynamics between protein and ligand may play a critical role in determining the kinetic rate constants of PLI. Furthermore, we demonstrate that the relative movement of normal nodes of amino acids upon ligand binding is an important feature to capture conformational dynamics of the binding/unbinding kinetics. Coupled with the multi-task learning, we can predict combined kon and koff accurately with an accuracy of 74.35%. Thus, it is possible to screen and optimize compounds based on their binding/unbinding kinetics. The further development of such computational tools will bridge one of the critical missing links in drug discovery.

Bioinformatics

Cookiecutter: a tool for kmer-based read filtering and extraction

MotivationKmer-based analysis is a powerful method used in read error correction and implemented in various genome assembly tools. A number of read processing routines include extracting or removing sequence reads from the results of high-throughput sequencing experiments prior to further analysis. Here we present a new approach to sorting or filtering of raw reads based on a provided list of kmers.\n\nResultsWe developed Cookiecutter -- a computational tool for rapid read extraction or removing according to a provided list of k-mers generated from a FASTA file. Cookiecutter is based on the implementation of the Aho-Corasik algorithm and is useful in routine processing of high-throughput sequencing datasets. Cookiecutter can be used for both removing undesirable reads and read extraction from a user-defined region of interest.\n\nAvailabilityThe open-source implementation with user instructions can be obtained from GitHub: https://github.com/ad3002/Cookiecutter.

Bioinformatics

Optimal Point Process Filtering and Estimation of the Coalescent Process

The coalescent process is an important and widely used model for inferring the dynamics of biological populations from samples of genetic diversity. Coalescent analysis typically involves applying statistical methods to either samples of genetic sequences or an estimated genealogy in order to estimate the demographic history of the population from which the samples originated. Several parametric and non-parametric estimation techniques, employing diverse methods, such as Gaussian processes and Monte Carlo particle filtering, already exist. However, these techniques often trade estimation accuracy and sophistication for methodological flexibility and ease of use. Thus, there is room for new coalescent estimation techniques that can be easily implemented for a range of inference problems while still maintaining some sense of statistical optimality.\n\nHere we introduce the Bayesian Snyder filter as a natural, easily implementable and flexible minimum mean square error estimator for parametric demographic functions. By reinterpreting the coalescent as a self-correcting inhomogeneous Poisson process, we show that the Snyder filter can be applied to both isochronous (sampled at one time point) and heterochronous (serially sampled) estimation problems. We test the estimation performance of the filter on both standard, simulated demographic models and on a well-studied empirical dataset comprising hepatitis= C virus sequences from Egypt. Additionally, we provide some analytical insight into the relationship between the Snyder filter and popular maximum likelihood and skyline plot techniques for coalescent inference. The Snyder filter is an exact and direct Bayesian estimation method that provides optimal mean square error estimates. It has the potential to become as a useful, alternative technique for coalescent inference.

Bioinformatics

Data science identifies novel drug interactions that prolong the QT interval

Drug-induced prolongation of the QT interval on the electrocardiogram (long QT syndrome, LQTS) can lead to a potentially fatal ventricular arrhythmia called torsades de pointes (TdP). 180 drugs with both cardiac and non-cardiac indications have been found to increase risk for TdP, but drug-drug interactions contributing to LQTS (QT-DDIs) remain poorly characterized. Traditional methods for mining observational healthcare data are poorly equipped to detect QT- DDI signals due to low reporting numbers and a lack of direct evidence for LQTS. In this study we present an integrative data science pipeline that addresses these limitations by identifying latent signals for QT-DDIs in the FDAs Adverse Event Reporting System and retrospectively validating these predictions using electrocardiogram data in electronic health records. We present 26 novel QT-DDIs flagged using this method that warrant further investigation.\n\nKey Points- Drug-induced long QT syndrome (LQTS) can lead to potentially fatal arrhythmias (torsades de pointes, TdP). Drug-drug interactions that prolong the QT interval (QT- DDIs) can be clinically significant but remain poorly characterized.\n- Observational health data (such as adverse event spontaneous reporting systems and electronic health records) offer an opportunity to mine for new QT-DDIs, but when used individually these datasets have a number of limitations that prevent identification of true signals.\n- We present an integrative data science approach that combines mining for latent QT- DDI signals in the FDA Adverse Event Reporting System and retrospective analysis of electrocardiogram lab results in electronic health records at Columbia University Medical Center to identify 26 novel QT-DDIs.

Bioinformatics

Genome ARTIST: a robust, high-accuracy aligner tool for mapping transposon insertions and self-insertions

A critical topic of insertional mutagenesis experiments performed on model organisms is mapping the hits of artificial transposons (ATs) at nucleotide level accuracy. Obviously, mapping errors may occur when sequencing artifacts or mutations as SNPs and small indels are present very close to the junction between a genomic sequence and a transposon inverted repeat (TIR). Another particular item of insertional mutagenesis is mapping of the transposon self-insertions and, to our best knowledge, there is no publicly available mapping tool designed to analyze such molecular events. We developed Genome ARTIST, a pairwise gapped aligner tool which works out both issues by means of an original, robust mapping strategy. Genome ARTIST is not designed to use NGS data but to analyze ATs insertions obtained in small to medium-scale mutagenesis experiments. Genome ARTIST employs a heuristic approach to find DNA sequence similarities and harnesses a multi-step implementation of a Smith-Waterman adapted algorithm to compute the mapping alignments. The experience is enhanced by easily customizable parameters and a user-friendly interface that describes the genomic landscape surrounding the insertion. Genome ARTIST deals with many genomes of bacteria and eukaryotes available in Ensembl and GenBank repositories. Our tool specifically harnesses/exploits the sequence annotation data provided by FlyBase for Drosophila melanogaster (the fruit fly), which enables mapping of insertions relative to various genomic features such as natural transposons. Genome ARTIST was tested against other alignment tools using relevant query sequences derived from the D. melanogaster and Mus musculus (mouse) genomes. Real and simulated query sequences were also comparatively inquired, revealing that Genome ARTIST is a very robust solution for mapping transposon insertions.\n\nGenome ARTIST is a stand-alone user-friendly application, designed for high-accuracy mapping of transposon insertions and self-insertions. The tool is also useful for routine aligning assessments like detection of SNPs or checking the specificity of primers and probes. Genome ARTIST is an open source software and is available for download at www.genomeartist.ro and at www.bioinformatics.org.

Bioinformatics

SFS_CODE: More Efficient and Flexible Forward Simulations

SUMMARYModern implementations of forward population genetic simulations are efficient and flexible, enabling the exploration of complex models that may otherwise be intractable. Here we describe an updated version of SFS_CODE, which has increased efficiency and includes many novel features. Among these features is an arbitrary model of dominance, the ability to simulate partial and soft selective sweeps, as well as track the trajectories of mutations and/or ancestries across multiple populations under complex models that are not possible under a coalescent framework. We also release sfs_coder, a Python wrapper to SFS_CODE allowing the user to easily generate command lines for common models of demography, selection, and human genome structure, as well as parse and simulate phenotypes from SFS_CODE output.\n\nAvailability and ImplementationOur open source software is written in C and Python, and are available under the GNU General Public License at http://sfscode.sourceforge.net.\n\nContactryan.hernandez@ucsf.edu\n\nSupplementary informationDetailed usage information is available from the project website at http://sfscode.sourceforge.net.

Bioinformatics

A Graph Theoretical Approach to Data Fusion

The rapid development of high throughput experimental techniques has resulted in a growing diversity of genomic datasets being produced and requiring analysis. A variety of computational techniques allow us to analyse such data and to model the biological processes behind them. However, it is increasingly being recognised that we can gain deeper understanding by combining the insights obtained from multiple, diverse datasets. We therefore require scalable computational approaches for data fusion.\n\nWe propose a novel methodology for scalable unsupervised data fusion. Our technique exploits network representations of the data in order to identify (and quantify) similarities among the datasets. We may work within the Bayesian formalism, using Bayesian nonparametric approaches to model each dataset; or (for fast, approximate, and massive scale data fusion) can naturally switch to more heuristic modelling techniques. An advantage of the proposed approach is that each dataset can initially be modelled independently (and therefore in parallel), before applying a fast post-processing step in order to perform data fusion. This allows us to incorporate new experimental data in an online fashion, without having to rerun all of the analysis. The methodology can be applied to genomic scale datasets and we demonstrate its applicability on examples from the literature, using a broad range of genomic datasets, and also on a recent gene expression dataset from Sporadic inclusion body myositis Availability. Example R code and instructions are available from https://sites.google.com/site/gtadatafusion/.

Bioinformatics

Differential transcript usage from RNA-seq data: isoform pre-filtering improves performance of count-based methods

Large-scale sequencing of cDNA (RNA-seq) has been a boon to the quantitative analysis of transcriptomes. A notable application is the detection of changes in transcript usage between experimental conditions. For example, discovery of pathological alternative splicing may allow the development of new treatments or better management of patients. From an analysis perspective, there are several ways to approach RNA-seq data to unravel differential transcript usage, such as annotation-based exon-level counting, differential analysis of the percent spliced in measure or quantitative analysis of assembled transcripts. The goal of this research is to compare and contrast current state-of-the-art methods, as well as to suggest improvements to commonly used workflows.\n\nWe assess the performance of representative workflows using synthetic data and explore the effect of using non-standard counting bin definitions as input to a state-of-the-art inference engine (DEXSeq). Although the canonical counting provided the best results overall, several non-canonical approaches were as good or better in specific aspects and most counting approaches outperformed the evaluated event- and assembly-based methods. We show that an incomplete annotation catalog can have a detrimental effect on the ability to detect differential transcript usage in transcriptomes with few isoforms per gene and that isoform-level pre-filtering can considerably improve false discovery rate (FDR) control.\n\nCount-based methods generally perform well in detection of differential transcript usage. Controlling the FDR at the imposed threshold is difficult, mainly in complex organisms, but can be improved by pre-filtering of the annotation catalog.

Bioinformatics

Modelizing Drosophila melanogaster longevity curves using a new discontinuous 2-Phases of Aging model

Aging is commonly described as being a continuous process affecting progressively organisms as time passes. This process results in a progressive decrease in individuals fitness through a wide range of both organismal - decreased motor activity, fertility, resistance to stress - and molecular phenotypes - decreased protein and energy homeostasis, impairment of insulin signaling. In the past 20 years, numerous genes have been identified as playing a major role in the aging process, yet little is known about the events leading to that loss of fitness. We recently described an event characterized by a dramatic increase of intestinal permeability to a blue food dye in aging flies committed to die within a few days. Importantly, flies showing this so called Smurf phenotype are the only ones, among a population, to show various age-related changes and exhibit a high-risk of impending death whatever their chronological age. Thus, these observations suggest that instead of being one continuous phenomenon, aging may be a discontinuous process well described by at least two distinguishable phases. In this paper we addressed this hypothesis by implementing a new 2-Phases of Aging mathematiCal model (2PAC model) to simulate longevity curves based on the simple hypothesis of two consecutive phases of lifetime presenting different properties. We first present a unique equation for each phase and discuss the biological significance of the 3 associated parameters. Then we evaluate the influence of each parameter on the shape of survival curves. Overall, this new mathematical model, based on simple biological observations, is able to reproduce many experimental longevity curves, supporting the existence of 2-phases of aging exhibiting specific properties and separated by a dramatic transition that remains to be characterized. Moreover, it indicates that Smurf survival can be approximated by one single constant parameter for a broad range of genotypes that we have tested under our environmental conditions.\n\nAuthor SummaryThe perception we can have of a process directly affects the way we study it. In the literature, aging is generally described as being a continuous process, progressively affecting organisms through a broad range of molecular and physiological changes ultimately leading to a dramatic decrease of individuals life expectancy. As such, aging studies focus on changes occurring in groups of individuals through time, considering individuals taken at a given time as being all equivalent. Instead, the recently described Smurf phenotype [1] suggested that any given time, a population could be divided in two subpopulations each characterized by a significantly different risk of impending death.\n\nBy formalizing here the concept of a discontinuous aging process using a mathematical model based on simple experimental observations, we propose a theoretical framework in which aging is actually separated in two consecutive phases characterized by three parameters easily quantifiable in vivo. Thus, the model we present here brings new tools to assess the events occurring during aging using a novel angle that we hope will open a better understanding of the very processes driving aging.

Bioinformatics

OEFinder: A user interface to identify and visualize ordering effects in single-cell RNA-seq data

A recent paper identified an artifact in multiple single-cell RNA-seq (scRNA-seq) data sets generated by the Fluidigm C1 platform. Specifically, Leng* et al. showed significantly increased gene expression in cells captured from sites with small or large plate output IDs. We refer to this artifact as an ordering effect (OE). Including OE genes in downstream analyses could lead to biased results. To address this problem, we developed a statistical method and software called OEFinder to identify a sorted list of OE genes. OEFinder is available as an R package along with user-friendly graphical interface implementations that allows users to check for potential artifacts in scRNA-seq data generated by the Fluidigm C1 platform.\n\nAvailability and ImplementationOEFinder is freely available at https://github.com/lengning/OEFinder\n\nContactrstewart@morgridge.org

Bioinformatics

Genome scaffolding with PE-contaminated mate-pair libraries

Scaffolding is often an essential step in a genome assembly process, in which contigs are ordered and oriented using read pairs from a combination of paired-ends libraries and longer-range mate-pair libraries. Although a simple idea, scaffolding is unfortunately hard to get right in practice. One source of problem is so-called PE-contamination in mate-pair libraries, in which a non-negligible fraction of the read pairs get the wrong orientation and a much smaller insert size than what is expected. This contamination has been discussed in previous work on integrated scaffolders in end-to-end assemblers such as Allpaths-LG and MaSuRCA but the methods relies on the fact that the orientation is observable, e.g., by finding the junction adapter sequence in the reads. This is not always the case, making orientation and insert size of a read pair stochastic. Furthermore, work on modeling PE-contamination has so far been disregarded in stand-alone scaffolders and the effect that PE-contamination has on scaffolding quality has not been examined before.\n\nWe have addressed PE-contamination in an update of our scaffolder BESST. We formulate the problem as an Integer Linear Program (ILP) and use characteristics of the problem, such as contig lengths and insert size, to efficiently solve the ILP using a linear amount (with respect to the number of contigs) of Linear Programs. Our results show significant improvement over both integrated and standalone scaffolders. The impact of modeling PE-contamination is quantified by comparison with the previous BESST model. We also show how other scaffolders are vulnerable to PE-contaminated libraries, resulting in increased number of misassemblies, more conservative scaffolding, and inflated assembly sizes.\n\nThe model is implemented in BESST. Source code and usage instructions are found at https://github.com/ksahlin/BESST. BESST can also be downloaded using PyPI.

Bioinformatics

Simultaneously inferring T cell fate and clonality from single cell transcriptomes

The heterodimeric T cell receptor (TCR) comprises two protein chains that pair to determine the antigen specificity of each T lymphocyte. The enormous sequence diversity within TCR repertoires allows specific TCR sequences to be used as lineage markers for T cells that derive from a common progenitor. We have developed a computational method, called TraCeR, to reconstruct full-length, paired TCR sequences from T lymphocyte single-cell RNA-seq by combining existing assembly and alignment programs with a \"synthetic genome\" library comprising all possible TCR sequences. We validate this method with PCR to quantify its accuracy and sensitivity, and compare to other TCR sequencing methods. Our inferred TCR sequences reveal clonal relationships between T cells, which we put into the context of each cells functional state from the complete transcriptional landscape quantified from the remaining RNA-seq data. This provides a powerful tool to link T cell specificity with functional response in a variety of normal and pathological conditions. We demonstrate this by determining the distribution of members of expanded T cell clonotypes in response to Salmonella infection in the mouse. We show that members of the same clonotype span early activated CD4+ T cells, as well as mature effector and memory cells.

Bioinformatics

Modeling of RNA-seq fragment sequence bias reduces systematic errors in transcript abundance estimation

RNA-seq technology is widely used in biomedical and basic science research. These studies rely on complex computational methods that quantify expression levels for observed transcripts. We find that current computational methods can lead to hundreds of false positive results related to alternative isoform usage. This flaw in the current methodology stems from a lack of modeling sample-specific bias that leads to drops in coverage and is related to sequence features like fragment GC content and GC stretches. By incorporating features that explain this bias into transcript expression models, we greatly increase the specificity of transcript expression estimates, with more than a four-fold reduction in the number of false positives for reported changes in expression. We introduce alpine, a method for estimation of bias-corrected transcript abundance. The method is available as a Bioconductor package that includes data visualization tools useful for bias discovery.

Bioinformatics

Teaser: Individualized benchmarking and optimization of read mapping results for NGS data

Mapping reads to a genome remains challenging, especially for non-model organisms with poorer quality assemblies, or for organisms with higher rates of mutations. While most research has focused on speeding up the mapping process, little attention has been paid to optimize the choice of mapper and parameters for a users dataset. Here we present Teaser, which assists in these choices through rapid automated benchmarking of different mappers and parameter settings for individualized data. Within minutes, Teaser completes a quantitative evaluation of an ensemble of mapping algorithms and parameters. Using Teaser, we demonstrate how Bowtie2 can be optimized for different data.

Bioinformatics

ProtAnnot: an App for Integrated Genome Browser to display how alternative splicing and transcription affect proteins

SummaryOne gene can produce multiple transcript variants encoding proteins with different functions. To facilitate visual analysis of transcript variants, we developed ProtAnnot, which shows protein annotations in the context of genomic sequence. ProtAnnot searches InterPro and displays profile matches (protein annotations) alongside gene models, exposing how alternative promoters, splicing, and 3 end processing add, remove, or remodel functional motifs. To draw attention to these effects, ProtAnnot color-codes exons by frame and displays a cityscape graphic summarizing exonic sequence at each position. These techniques make visual analysis of alternative transcripts faster and more convenient for biologists. Availability and ImplementationProtAnnot is a plug-in App for Integrated Genome Browser, an open source desktop genome browser available from http://www.bioviz.org. Contactaloraine@uncc.edu

Bioinformatics

Privacy-Preserving Microbiome Analysis Using Secure Computation

MotivationDeveloping targeted therapeutics and identifying biomarkers relies on large amounts of patient data. Beyond human DNA, researchers now investigate the DNA of micro-organisms inhabiting the human body. An individuals collection of microbial DNA consistently identifies that person and could be used to link a real-world identity to a sensitive attribute in a research dataset. Unfortunately, the current suite of DNA-specific privacy-preserving analysis tools does not meet the requirements for microbiome sequencing studies.\n\nResultsWe augment an existing categorization of genomic-privacy attacks to incorporate microbiome sequencing and provide an implementation of metagenomic analyses using secure computation. Our implementation allows researchers to perform analysis over combined data without revealing individual patient attributes. We implement three metagenomic analyses and perform an evaluation on real datasets for comparative analysis. We use our implementation to simulate sharing data between four policy-domains and measure the increase in significant discoveries. Additionally, we describe an application of our implementation to form patient pools of data to allow drug companies to query against and compensate patients for the analysis.\n\nAvailabilityThe software is freely available for download at: http://cbcb.umd.edu/[~]hcorrada/projects/secureseq.html

Bioinformatics

A Comparison of Methods: Normalizing High-Throughput RNA Sequencing Data

As RNA-Seq and other high-throughput sequencing grow in use and remain critical for gene expression studies, technical variability in counts data impedes studies of differential expression studies, data across samples and experiments, or reproducing results. Studies like Dillies et al. (2013) compare several between-lane normalization methods involving scaling factors, while Hansen et al. (2012) and Risso et al. (2014) propose methods that correct for sample-specific bias or use sets of control genes to isolate and remove technical variability. This paper evaluates four normalization methods in terms of reducing intra-group, technical variability and facilitating differential expression analysis or other research where the biological, inter-group variability is of interest. To this end, the four methods were evaluated in differential expression analysis between data from Pickrell et al. (2010) and Montgomery et al. (2010) and between simulated data modeled on these two datasets. Though the between-lane scaling factor methods perform worse on real data sets, they are much stronger for simulated data. We cannot reject the recommendation of Dillies et al. to use TMM and DESeq normalization, but further study of power to detect effects of different size under each normalization method is merited.

Bioinformatics