Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

MuCor: Mutation Aggregation and Correlation

MotivationThere are many tools for variant calling and effect prediction, but little to tie together large sample groups. Aggregating, sorting, and summarizing variants and effects across a cohort is often done with ad hoc scripts that must be re-written for every new project. In response, we have written MuCor, a tool to gather variants from a variety of input formats (including multiple files per sample), perform database lookups and frequency calculations, and write many report types. In addition to use in large studies with numerous samples, MuCor can also be employed to directly compare variant calls from the same sample across two or more platforms, parameters, or pipelines. A companion utility, DepthGauge, measures coverage at regions of interest to increase confidence in calls.\n\nAvailabilitySource code is freely available at https://github.com/blachlylab\n\nContactjames.blachly@osumc.edu\n\nSupplementary dataSupplementary data, including detailed documentation, are available online.

Bioinformatics

CUA: a Flexible and Comprehensive Codon Usage Analyzer

Codon usage bias (CUB) is pervasive in genomes. Studying its patterns and causes is fundamental for understanding genome evolution. Rapidly emerging large-scale RNA and DNA sequences make studying CUB in many species feasible. Existing software however is limited in incorporating the new data resources. Therefore, I release the software CUA which can compute all popular CUB metrics, including CAI, tAI, Fop, ENC. More importantly, CUA allows users to incorporate user-specific data, such as tRNA abundance and highly expressed genes from considered tissues; this flexibility enables computing CUB metrics for any species with improved accuracy. In sum, CUA eases codon usage studies and establishes a platform for incorporating new metrics in future. CUA is available at http://search.cpan.org/dist/Bio-CUA/ with help documentation and tutorial.

Bioinformatics

Ancestral gene synteny reconstruction improves extant species scaffolding

We exploit the methodological similarity between ancestral genome reconstruction and extant genome scaffolding. We present a method, called ARO_SCPCAPTC_SCPCAP-DO_SCPCAPEC_SCPCAPCO_SCPCAPOC_SCPCAP that constructs neighborhood relationships between genes or contigs, in both ancestral and extant genomes, in a phylogenetic context. It is able to handle dozens of complete genomes, including genes with complex histories, by using gene phylogenies reconciled with a species tree, that is, annotated with speciation, duplication and loss events. Reconstructed ancestral or extant synteny comes with a support computed from an exhaustive exploration of the solution space. We compare our method with a previously published one that follows the same goal on a small number of genomes with universal unicopy genes. Then we test it on the whole Ensembl database, by proposing partial ancestral genome structures, as well as a more complete scaffolding for many partially assembled genomes on 69 eukaryote species. We carefully analyze a couple of extant adjacencies proposed by our method, and show that they are indeed real links in the extant genomes, that were missing in the current assembly. On a reduced data set of 39 eutherian mammals, we estimate the precision and sensitivity of ARO_SCPCAPTC_SCPCAP-DO_SCPCAPEC_SCPCAPCO_SCPCAPOC_SCPCAP by simulating a fragmentation in some well assembled genomes, and measure how many adjacencies are recovered. We find a very high precision, while the sensitivity depends on the quality of the data and on the proximity of closely related genomes.

Bioinformatics

destiny – diffusion maps for large-scale single-cell data in R

SummaryDiffusion maps are a spectral method for non-linear dimension reduction and have recently been adapted for the visualization of single cell expression data. Here we present destiny, an efficient R implementation of the diffusion map algorithm. Our package includes a single-cell specific noise model allowing for missing and censored values. In contrast to previous implementations, we further present an efficient nearest-neighbour approximation that allows for the processing of hundreds of thousands of cells and a functionality for projecting new data on existing diffusion maps. We exemplarily apply destiny to a recent time-resolved mass cytometry dataset of cellular reprogramming.\n\nAvailability and implementationdestiny is an open-source R/Bioconductor package http://bioconductor.org/packages/ destiny also available at https://www.helmholtz-muenchen.de/icb/destiny. A detailed vignette describing functions and workflows is provided with the package.\n\nContactcarsten.marr@helmholtz-muenchen.de, f.buettner@helmholtz-muenchen.de

Bioinformatics

Circlator: automated circularization of genome assemblies using long sequencing reads

The assembly of DNA sequence data into finished genomes is undergoing a renais-sance thanks to emerging technologies producing reads of tens of kilobases. Assembling complete bacterial and small eukaryotic genomes is now possible, but the final step of circularizing sequences remains unsolved. Here we present Circlator, the first tool to automate assembly circularization and produce accurate linear rep-resentations of circular sequences. Using Pacific Biosciences and Oxford Nanopore data, Circlator correctly circularized 26 of 27 circularizable sequences, comprising 11 chromosomes and 12 plasmids from bacteria, the apicoplast and mitochondrion of Plasmodium falciparum and a human mitochondrion. Circlator is available at http://sanger-pathogens.github.io/circlator/.

Bioinformatics

Dissecting the roles of local packing density and longer-range effects in protein sequence evolution

What are the structural determinants of protein sequence evolution? A number of site-specific structural characteristics have been proposed, most of which are broadly related to either the density of contacts or the solvent accessibility of individual residues. Most importantly, there has been disagreement in the literature over the relative importance of solvent accessibility and local packing density for explaining site-specific sequence variability in proteins. We show here that this discussion has been confounded by the definition of local packing density. The most commonly used measures of local packing, such as the contact number and the weighted contact number, represent by definition the combined effects of local packing density and longer-range effects. As an alternative, we here propose a truly local measure of packing density around a single residue, based on the Voronoi cell volume. We show that the Voronoi cell volume, when calculated relative to the geometric center of amino-acid side chains, behaves nearly identically to the relative solvent accessibility, and both can explain, on average, approximately 34% of the site-specific variation in evolutionary rate in a data set of 209 enzymes. An additional 10% of variation can be explained by non-local effects that are captured in the weighted contact number. Consequently, evolutionary variation at a site is determined by the combined action of the immediate amino-acid neighbors of that site and of effects mediated by more distant amino acids. We conclude that instead of contrasting solvent accessibility and local packing density, future research should emphasize the relative importance of immediate contacts and longer-range effects on evolutionary variation.

Bioinformatics

Inverting proteomics analysis provides powerful insight into the peptide/protein conundrum

In proteomics, a large proportion of mass spectrometry (MS) data is ignored due to the lack of, or insufficient statistical evidence for mappable peptides. In reality, only a small fraction of features are expected to be differentially relevant anyway. Mapping spectra to peptides and subsequently, proteins, produces uncertainty at several levels. We propose it is better to analyze proteomic profiling data directly at MS level, and then relate these features to peptides/proteins. In a renal cancer data comprising 12 normal and 12 cancer subjects, we demonstrate that a simple rule-based binning approach can give rise to informative features. We note that the peptides associated with significant spectral bins gave rise to better class separation than the corresponding proteins, suggesting a loss of signal in the peptide-to-protein transition. Additionally, the binning approach sharpens focus on relevant protein splice forms rather than just canonical sequences. Taken together, the inverted raw spectra analysis paradigm, which is realised by the MZ-Bin method described in this article, provides new possibilities and insights, in how MS-data can be interpreted.

Bioinformatics

Elucidating Variants in Systemic JIA from Analysis of Illumina & Ion Torrent Exome Data

MotivationChildren with systemic juvenile idiopathic arthritis (SJIA) are affected by a wide-range of complications. Partially arising from the difficulty of diagnosis due to the idiopathic nature of the indication. There may be a genetic basis for SJIA, which could help in both diagnosis, and treatment.\n\nResultsTwo mutations in the Fc epsilon RI pathway, including PIK3CD, were detected in low-coverage Ion Torrent data. Variants of unknown significance were detected within HLA regions on standard Illumina exomes. CSF2RA, which could account for pulmonary observations, had insignificant coverage on both datasets.\n\nAvailabilityftp.systemicjia.com\n\nContactmo@genedrop.com

Bioinformatics

Perturbative formulation of general continuous-time Markov model of sequence evolution via insertions/deletions, Part III: Algorithm for first approximation

BackgroundInsertions and deletions (indels) account for more nucleotide differences between two related DNA sequences than substitutions do, and thus it is imperative to develop a stochastic evolutionary model that enables us to reliably calculate the probability of the sequence evolution through indel processes. In a separate paper (Ezawa, Graur and Landan 2015a), we established an ab initio perturbative formulation of a continuous-time Markov model of the evolution of an entire sequence via insertions and deletions. And we showed that, under a certain set of conditions, the ab initio probability of an alignment can be factorized into the product of an overall factor and contributions from regions (or local alignments) separated by gapless columns. Moreover, in another separate paper (Ezawa, Graur and Landan 2015b), we performed concrete perturbation analyses on all types of local pairwise alignments (PWAs) and some typical types of local multiple sequence alignments (MSAs). The analyses indicated that even the fewest-indel terms alone can quite accurately approximate the probabilities of local alignments, as long as the segments and the branches in the tree are of modest lengths.\n\nResultsTo examine whether or not the fewest-indel terms alone can well approximate the alignment probabilities of more general types of local MSAs as well, and as a first step toward the automatic application of our ab initio perturbative formulation, we developed an algorithm that calculates the first approximation of the probability of a given MSA under a given parameter setting including a phylogenetic tree. The algorithm first chops the MSA into gapped and gapless segments, second enumerates all parsimonious indel histories potentially responsible for each gapped segment, and finally calculates their contributions to the MSA probability. We performed validation analyses using more than ten million local MSAs. The results indicated that even the first approximation can quite accurately estimate the probability of each local MSA, as long as the gaps and tree branches are at most moderately long.\n\nConclusionsThe newly developed algorithm, called LOLIPOG, brought our ab initio perturbation formulation at least one step closer to a practically useful method to quite accurately calculate the probability of a MSA under a given biologically realistic parameter setting.\n\n[This paper and three other papers (Ezawa, Graur and Landan 2015a,b,c) describe a series of our efforts to develop, apply, and extend the ab initio perturbative formulation of a general continuous-time Markov model of indels.]

Bioinformatics

Fast Bayesian Inference of Copy Number Variants using Hidden Markov Models with Wavelet Compression

By combining Haar wavelets with Bayesian Hidden Markov Models, we improve detection of genomic copy number variants (CNV) in array CGH experiments compared to the state-of-the-art, including standard Gibbs sampling. At the same time, we achieve drastically reduced running times, as the method concentrates computational effort on chromosomal segments which are difficult to call, by dynamically and adaptively recomputing consecutive blocks of observations likely to share a copy number. This makes routine diagnostic use and re-analysis of legacy data collections feasible; to this end, we also propose an effective automatic prior. An open source software implementation of our method is available at http://bioinformatics.rutgers.edu/Software/HaMMLET/. The web supplement is at http://bioinformatics.rutgers.edu/Supplements/HaMMLET/

Bioinformatics

D3M: Detection of differential distributions of methylation levels

Motivation: DNA methylation is an important epigenetic modification related to a variety of diseases including cancers. We focus on the methylation data from Illuminas Infinium HumanMethylation450 BeadChip. One of the key issues of methylation analysis is to detect the differential methylation sites between case and control groups. Previous approaches describe data with simple summary statistics and kernel function, and then use statistical tests to determine the difference. However, a summary statistics-based approach cannot capture complicated underlying structure, and a kernel functions-based approach lacks interpretability of results.\n\nResults: We propose a novel method D3M, for detection of differential distribution of methylation, based on distribution-valued data. Our method can detect high-order moments, such as shapes of underlying distributions in methylation profiles, based on the Wasserstein metric. We test the significance of the difference between case and control groups and provide an interpretable summary of the results. The simulation results show that the proposed method achieves promising accuracy and shows favorable results compared with previous methods. Glioblastoma multiforme and lower grade glioma data from The Cancer Genome Atlas show that our method supports recent biological advances and suggests new insights.\n\nAvailability: R implemented code is freely available from\n\nhttps://github.com/ymatts/D3M/\n\nhttps://cran.r-project.org/package=D3M.\n\nContact: ymatsui@med.nagoya-u.ac.jp

Bioinformatics

The impact of partitioning on phylogenomic accuracy

Several strategies have been proposed to assign substitution models in phylogenomic datasets, or partitioning. The accuracy of these methods, and most importantly, their impact on phylogenetic estimation has not been thoroughly assessed using computer simulations. We simulated multiple partitioning scenarios to benchmark two a priori partitioning schemes (one model for the whole alignment, one model for each data block), and two statistical approaches (hierarchical clustering and greedy) implemented in PartitionFinder and in our new program, PartitionTest. Most methods were able to identify optimal partitioning schemes closely related to the true one. Greedy algorithms identified the true partitioning scheme more frequently than the clustering algorithms, but selected slightly less accurate partitioning schemes and tended to underestimate the number of partitions. PartitionTest was several times faster than PartitionFinder, with equal or better accuracy. Importantly, maximum likelihood phylogenetic inference was very robust to the partitioning scheme. Best-fit partitioning schemes resulted in optimal phylogenetic performance, without appreciable differences compared to the use of the true partitioning scheme. However, accurate trees were also obtained by a \"simple\" strategy consisting of assigning independent GTR+G models to each data block. On the contrary, leaving the data unpartitioned always diminished the quality of the trees inferred, to a greater or lesser extent depending on the simulated scenario. The analysis of empirical data confirmed these trends, although suggesting a stronger influence of the partitioning scheme. Overall, our results suggests that statistical partitioning, but also the a priori assignment of independent GTR+G models, maximize phylogenomic performance.

Bioinformatics

DADA2: High resolution sample inference from amplicon data

Microbial communities are commonly characterized by amplifying and sequencing target genes, but errors limit the precision of amplicon sequencing. We present DADA2, a software package that models and corrects amplicon errors. DADA2 identified more real variants than other methods in Illumina-sequenced mock communities, some differing by a single nucleotide, while outputting fewer spurious sequences. DADA2 analysis of vaginal samples revealed a diversity of Lactobacillus crispatus strains undetected by OTU methods.

Bioinformatics

Fuzzy-FishNet: A highly precise distribution-free network approach for feature selection in clinical proteomics

Network-based analysis methods can help resolve coverage and inconsistency issues in proteomics data. Previously, it was demonstrated that a suite of rank-based network approaches (RBNAs) provides unparalleled consistency and reliable feature selection. However, reliance on the t-statistic/t-distribution and hypersensitivity (coupled to a relatively flat p-value distribution) makes feature prioritization for validation difficult. To address these concerns, a refinement based on the fuzzified Fisher exact test, Fuzzy-FishNet was developed. Fuzzy-FishNet is highly precise (providing probability values that allows exact ranking of features). Furthermore, feature ranks are stable, even in small sample size scenario. Comparison of features selected by genomics and proteomics data respectively revealed that in spite of relative feature stability, cross-platform overlaps are extremely limited, suggesting that networks may not be the answer towards bridging the proteomics-genomics divide.

Bioinformatics

Toward high-throughput predictive modeling of protein binding/unbinding kinetics

One of the unaddressed challenges in drug discovery is that drug potency determined in vitro is not a reliable indicator of drug activity in humans. Accumulated evidences suggest that in vivo activity is more strongly correlated with the binding/unbinding kinetics than the equilibrium thermodynamics of protein-ligand interactions (PLI). However, existing experimental and computational techniques are insufficient in studying the molecular details of kinetics process of PLI. Consequently, we not only have limited mechanistic understanding of the kinetic process but also lack a practical platform for the high-throughput screening and optimization of drug leads based on their kinetic properties. Here we address this unmet need by integrating energetic and conformational dynamic features derived from molecular modeling with multi-task learning. To test our method, HIV-1 protease is used as a model system. Our integrated model provides us with new insights into the molecular determinants of kinetics of PLI. We find that the coherent coupling of conformational dynamics between protein and ligand may play a critical role in determining the kinetic rate constants of PLI. Furthermore, we demonstrate that the relative movement of normal nodes of amino acids upon ligand binding is an important feature to capture conformational dynamics of the binding/unbinding kinetics. Coupled with the multi-task learning, we can predict combined kon and koff accurately with an accuracy of 74.35%. Thus, it is possible to screen and optimize compounds based on their binding/unbinding kinetics. The further development of such computational tools will bridge one of the critical missing links in drug discovery.

Bioinformatics

Cookiecutter: a tool for kmer-based read filtering and extraction

MotivationKmer-based analysis is a powerful method used in read error correction and implemented in various genome assembly tools. A number of read processing routines include extracting or removing sequence reads from the results of high-throughput sequencing experiments prior to further analysis. Here we present a new approach to sorting or filtering of raw reads based on a provided list of kmers.\n\nResultsWe developed Cookiecutter -- a computational tool for rapid read extraction or removing according to a provided list of k-mers generated from a FASTA file. Cookiecutter is based on the implementation of the Aho-Corasik algorithm and is useful in routine processing of high-throughput sequencing datasets. Cookiecutter can be used for both removing undesirable reads and read extraction from a user-defined region of interest.\n\nAvailabilityThe open-source implementation with user instructions can be obtained from GitHub: https://github.com/ad3002/Cookiecutter.

Bioinformatics

Optimal Point Process Filtering and Estimation of the Coalescent Process

The coalescent process is an important and widely used model for inferring the dynamics of biological populations from samples of genetic diversity. Coalescent analysis typically involves applying statistical methods to either samples of genetic sequences or an estimated genealogy in order to estimate the demographic history of the population from which the samples originated. Several parametric and non-parametric estimation techniques, employing diverse methods, such as Gaussian processes and Monte Carlo particle filtering, already exist. However, these techniques often trade estimation accuracy and sophistication for methodological flexibility and ease of use. Thus, there is room for new coalescent estimation techniques that can be easily implemented for a range of inference problems while still maintaining some sense of statistical optimality.\n\nHere we introduce the Bayesian Snyder filter as a natural, easily implementable and flexible minimum mean square error estimator for parametric demographic functions. By reinterpreting the coalescent as a self-correcting inhomogeneous Poisson process, we show that the Snyder filter can be applied to both isochronous (sampled at one time point) and heterochronous (serially sampled) estimation problems. We test the estimation performance of the filter on both standard, simulated demographic models and on a well-studied empirical dataset comprising hepatitis= C virus sequences from Egypt. Additionally, we provide some analytical insight into the relationship between the Snyder filter and popular maximum likelihood and skyline plot techniques for coalescent inference. The Snyder filter is an exact and direct Bayesian estimation method that provides optimal mean square error estimates. It has the potential to become as a useful, alternative technique for coalescent inference.

Bioinformatics

Data science identifies novel drug interactions that prolong the QT interval

Drug-induced prolongation of the QT interval on the electrocardiogram (long QT syndrome, LQTS) can lead to a potentially fatal ventricular arrhythmia called torsades de pointes (TdP). 180 drugs with both cardiac and non-cardiac indications have been found to increase risk for TdP, but drug-drug interactions contributing to LQTS (QT-DDIs) remain poorly characterized. Traditional methods for mining observational healthcare data are poorly equipped to detect QT- DDI signals due to low reporting numbers and a lack of direct evidence for LQTS. In this study we present an integrative data science pipeline that addresses these limitations by identifying latent signals for QT-DDIs in the FDAs Adverse Event Reporting System and retrospectively validating these predictions using electrocardiogram data in electronic health records. We present 26 novel QT-DDIs flagged using this method that warrant further investigation.\n\nKey Points- Drug-induced long QT syndrome (LQTS) can lead to potentially fatal arrhythmias (torsades de pointes, TdP). Drug-drug interactions that prolong the QT interval (QT- DDIs) can be clinically significant but remain poorly characterized.\n- Observational health data (such as adverse event spontaneous reporting systems and electronic health records) offer an opportunity to mine for new QT-DDIs, but when used individually these datasets have a number of limitations that prevent identification of true signals.\n- We present an integrative data science approach that combines mining for latent QT- DDI signals in the FDA Adverse Event Reporting System and retrospective analysis of electrocardiogram lab results in electronic health records at Columbia University Medical Center to identify 26 novel QT-DDIs.

Bioinformatics