Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

A de novo DNA Sequencing and Variant Calling Algorithm for Nanopores

The single-molecule accuracy of nanopore sequencing has been an area of rapid academic and commercial advancement, but remains insufficient for the de novo analysis of genomes. We introduce here a novel algorithm for the error correction of nanopore data, utilizing statistical models of the physical system in order to obtain high accuracy de novo sequences at a range of coverage depths. We demonstrate the technique by sequencing M13 bacteriophage DNA to 99% accuracy at moderate coverage as well as its use in an assembly pipeline by sequencing {lambda} DNA at a range of coverages. We also show the algorithms ability to accurately classify sequence variants at far lower coverage than existing methods.

Bioinformatics

Diagnosis of coronary heart diseases using gene expression profiling; stable coronary artery disease, cardiac ischemia with and without myocardial necrosis

Cardiovascular disease including coronary artery disease and myocardial infarction is one of the leading causes of death in Europe, and is influenced by both environmental and genetic factors. With the advancements in genomic tools and technologies there is potential to predict and diagnose heart disease using molecular data from analysis of blood cells. We analyzed gene expression data from blood samples taken from normal people (n=21), non-significant coronary artery disease (n=93), patients with unstable angina (n=16), stable coronary artery disease (n=14) and myocardial infarction (MI; n=207). We used a feature selection approach to identify a set of gene expression variables which successfully differentiate different cardiovascular diseases. The initial features were discovered by fitting a linear model for each probe set across all arrays of normal individuals and patients with myocardial infarction. Three different feature optimisation algorithms were devised which identified two most discriminating sets of genes one using MI and normal controls (total genes=8) and another one using MI and unstable angina patients (total genes=17). The results proved the diagnostic robustness of the final feature sets in discriminating not only patients with myocardial infraction from healthy controls but also from patients with clinical symptoms of cardiac ischemia with myocardial necrosis and stable coronary artery disease despite the influence of batch effects and different microarray gene chips and platforms. selection approach to identify a set of gene expression variables which successfully differentiate different cardiovascular diseases. The initial features were discovered by fitting a linear model for each probe set across all arrays of normal individuals and patients with myocardial infarction. Three different feature optimisation algorithms were devised which identified two most discriminating sets of genes one using MI and normal controls (total genes=8) and another one using MI and unstable angina patients (total genes=17). The results proved the diagnostic robustness of the final feature sets in discriminating not only patients with myocardial infraction from healthy controls but also from patients with clinical symptoms of cardiac ischemia with myocardial necrosis and stable coronary artery disease despite the influence of batch effects and different microarray gene chips and platforms.

Bioinformatics

GenoWAP: Post-GWAS Prioritization Through Integrated Analysis of Genomic Functional Annotation

MotivationGenome-wide association study (GWAS) has been a great success in the past decade. However, significant challenges still remain in both identifying new risk loci and interpreting results. Bonferroni-corrected significance level is known to be conservative, leading to insufficient statistical power when the effect size is moderate at risk locus. Complex structure of linkage disequilibrium also makes it challenging to separate causal variants from nonfunctional ones in large haplotype blocks.\n\nResultsWe describe GenoWAP, a post-GWAS prioritization method that integrates genomic functional annotation and GWAS test statistics. The effectiveness of GenoWAP is demonstrated through its applications to Crohns disease and schizophrenia using the largest studies available, where highly ranked loci show substantially stronger signals in the whole dataset after prioritization based on a subset of samples. At the single nucleotide polymorphism (SNP) level, top ranked SNPs after prioritization have both higher replication rates and consistently stronger enrichment of eQTLs. Within each risk locus, GenoWAP is also able to distinguish functional sites from groups of correlated SNPs.\n\nAvailability and ImplementationGenoWAP is freely available on the web at http://genocanyon.med.yale.edu/GenoWAP

Bioinformatics

Re-Annotator: Annotation Pipeline for Microarrays

BackgroundMicroarray technologies are established approaches for high throughput gene expression, methylation and genotyping analysis. An accurate mapping of the array probes is essential to generate reliable biological findings. Manufacturers typically provide incomplete and outdated annotation tables, which often rely on older genome and transcriptome versions differing substantially from up-to-date sequence databases.\n\nResultsHere, we present the Re-Annotator, a re-annotation pipeline for microarrays. It is primarily designed for gene expression microarrays but can be adapted to other types of microarrays. The Re-Annotator is based on a custom-built mRNA reference, used to identify the positions of gene expression array probe sequences. A comparison of our re-annotation of the Human-HT12-v4 microarray to the manufacturers annotation led to over 25% differently interpreted probes.\n\nConclusionsA thorough re-annotation of probe information is crucial to any microarray analysis. The Re-Annotator pipeline consists of Perl and Shell scripts, freely available at http://sourceforge.net/projects/reannotator. Re-annotation files for Illumina microarrays Human HT-12 v3/v4 and MouseRef-8 v2 are available as well.

Bioinformatics

A simple model-based approach to inferring and visualizing cancer mutation signatures

Recent advances in sequencing technologies have enabled the production of massive amounts of data on somatic mutations from cancer genomes. These data have led to the detection of characteristic patterns of somatic mutations or \"mutation signatures\" at an unprecedented resolution, with the potential for new insights into the causes and mechanisms of tumorigenesis.\n\nHere we present new methods for modelling, identifying and visualizing such mutation signatures. Our methods greatly simplify mutation signature models compared with existing approaches, reducing the number of parameters by orders of magnitude even while increasing the contextual factors (e.g. the number of flanking bases) that are accounted for. This improves both sensitivity and robustness of inferred signatures. We also provide a new intuitive way to visualize the signatures, analogous to the use of sequence logos to visualize transcription factor binding sites.\n\nWe illustrate our new method on somatic mutation data from urothelial carcinoma of the upper urinary tract, and a larger dataset from 30 diverse cancer types. The results illustrate several important features of our methods, including the ability of our new visualization tool to clearly highlight the key features of each signature, the improved robustness of signature inferences from small sample sizes, and more detailed inference of signature characteristics such as strand biases and sequence context effects at the base two positions 5 to the mutated site.\n\nThe overall framework of our work is based on probabilistic models that are closely connected with \"mixed-membership models\" which are widely used in population genetic admixture analysis, and in machine learning for document clustering. We argue that recognizing these relationships should help improve understanding of mutation signature extraction problems, and suggests ways to further improve the statistical methods.\n\nOur methods are implemented in an R package pmsignature (https://github.com/friend1ws/pmsignature) and a web application available at https://friend1ws.shinyapps.io/pmsignature_shiny/.\n\nAuthor SummarySomatic (non-inherited) mutations are acquired throughout our lives in cells throughout our body. These mutations can be caused, for example, by DNA replication errors or exposure to environmental mutagens such as tobacco smoke. Some of these mutations can lead to cancer.\n\nDifferent cancers, and even different instances of the same cancer, can show different distinctive patterns of somatic mutations. These distinctive patterns have become known as \"mutation signatures\". For example, C > A mutations are frequent in lung caners whereas C > T and CC > TT mutations are frequent in skin cancers. Each mutation signature may be associated with a specific kind of carcinogen, such as tobacco smoke or ultraviolet light. Identifying mutation signatures therefore has the potential to identify new carcinogens, and yield new insights into the mechanisms and causes of cancer,\n\nIn this paper, we introduce new statistical tools for tackling this important problem. These tools provide more robust and interpretable mutation signatures compared to previous approaches, as we demonstrate by applying them to large-scale cancer genomic data.

Bioinformatics

SAM/BAM format v1.5 extensions for de novo assemblies

SummaryThe plain text Sequence Alignment/Map (SAM) file format and its companion binary form (BAM) are a generic alignment format for storing read alignments against reference sequences (and unmapped reads) together with structured meta-data (Li et al., 2009). Driven by the needs of the 1000 Genomes Project which sequenced many individual human genomes, early SAM/BAM usage focused on pairwise alignments of reads to a reference. However, through the CIGAR P operator multiple sequence alignments can also be preserved. Herein we describe clarifications and additions in version 1.5 of the specification to facilitate storing de novo sequence alignments: Padded reference sequences (with gap characters), annotation of reads or regions of the reference, and the option of embedding the reference sequence within the file.\n\nAvailabilityThe latest public release of the specification is at http://samtools.sourceforge.net/SAM1.pdf, with in development drafts at https://github.com/samtools/hts-specs/ under version control.\n\nContactpeter.cock@hutton.ac.uk

Bioinformatics

The GL Service: Web Service to Exchange GL String Encoded HLA & KIR Genotypes With Complete and Accurate Allele and Genotype Ambiguity

Genotype List (GL) Strings use a set of hierarchical character delimiters to represent allele and genotype ambiguity in HLA and KIR genotypes in a complete and accurate fashion. A RESTful web service called Genotype List Service was created to allow users to register a GL String and receive a unique identifier for that string in the form of a URI. By exchanging URIs and dereferencing them through the GL Service, users can easily transmit HLA genotypes in a variety of useful formats. The GL Service was developed to be secure, scalable, and persistent. An instance of the GL Service is configured with a nomenclature and can be run in strict or non-strict modes. Strict mode requires alleles used in the GL String to be present in the allele database using the fully qualified nomenclature. Non-strict mode allows any GL String to be registered as long as it is syntactically correct. The GL Service source code is free and open source software, distributed under the GNU Lesser General Public License (LGPL) version 3 or later.\n\nAbbreviations\n\nAbbreviations

Bioinformatics

Optimizing error correction of RNAseq reads

MotivationThe correction of sequencing errors contained in Illumina reads derived from genomic DNA is a common pre-processing step in many de novo genome assembly pipelines, and has been shown to improved the quality of resultant assemblies. In contrast, the correction of errors in transcriptome sequence data is much less common, but can potentially yield similar improvements in mapping and assembly quality. This manuscript evaluates several popular read-correction tools ability to correct sequence errors commonplace to transcriptome derived Illumina reads.\n\nResultsI evaluated the efficacy of correction of transcriptome derived sequencing reads using using several metrics across a variety of sequencing depths. This evaluation demonstrates a complex relationship between the quality of the correction, depth of sequencing, and hardware availability which results in variable recommendations depending on the goals of the experiment, tolerance for false positives, and depth of coverage. Overall, read error correction is an important step in read quality control, and should become a standard part of analytical pipelines.\n\nAvailabilityResults are non-deterministically repeatable using AMI:ami-3dae4956 (MacManes_EC_2015) and the Makefile available here: https://goo.gl/oVIuE0\n\nContactmatthew.macmanes@unh.edu and @PeroMHC

Bioinformatics

Learning quantitative sequence-function relationships from high-throughput biological data

Understanding the transcriptional regulatory code, as well as other types of information encoded within biomolecular sequences, will require learning biophysical models of sequence-function relationships from high-throughput data. Controlling and characterizing the noise in such experiments, however, is notoriously difficult. The unpredictability of such noise creates problems for standard likelihood-based methods in statistical learning, which require that the quantitative form of experimental noise be known precisely. However, when this unpredictability is properly accounted for, important theoretical aspects of statistical learning which remain hidden in standard treatments are revealed. Specifically, one finds a close relationship between the standard inference method, based on likelihood, and an alternative inference method based on mutual information. Here we review and extend this relationship. We also describe its implications for learning sequence-function relationships from real biological data. Finally, we detail an idealized experiment in which these results can be demonstrated analytically.

Bioinformatics

OPERA-LG: Efficient and exact scaffolding of large, repeat-rich eukaryotic genomes with performance guarantees

The assembly of large, repeat-rich eukaryotic genomes continues to represent a significant challenge in genomics. While long-read technologies have made the high-quality assembly of small, microbial genomes increasingly feasible, data generation can be prohibitively expensive for larger genomes. Advances in assembly algorithms are thus essential to exploit the characteristics of short and long-read sequencing technologies to consistently and reliably provide high-quality assemblies in a cost-efficient manner. OPERA-LG is a scalable, exact algorithm for the scaffold assembly of large, repeat-rich genomes, with consistent improvement over state-of-the-art programs for scaffold correctness and contiguity. It provides a rigorous framework for scaffolding of repetitive sequences and a systematic approach for combining data from different second-generation (Illumina, Ion Torrent) and third-generation (PacBio, ONT) sequencing technologies. OPERA-LG efficiently scaffolds large genomes with provable scaffold properties, providing an avenue for systematic augmentation and improvement of 1000s of existing draft eukaryotic genome assemblies.

Bioinformatics

The next 20 years of genome research

The last 20 years have been a remarkable era for biology and medicine. One of the most significant achievements has been the sequencing of the first human genomes, which has laid the foundation for profound insights into human genetics, the intricacies of regulation and development, and the forces of evolution. Incredibly, as we look into the future over the next 20 years, we see the very real potential for sequencing more than one billion genomes, bringing with it even deeper insights into human genetics as well as the genetics of millions of other species on the planet. Realizing this great potential, though, will only be achieved through the integration and development of highly scalable computational and quantitative approaches can keep pace with the rapid improvements to biotechnology. In this perspective, we aim to chart out these future technologies, anticipate the major themes of research, and call out the challenges ahead. One of the largest shifts will be in the training used to prepare the class of 2035 for their highly interdisciplinary world.

Bioinformatics

Triplex Domain Finder: Detection of Triple Helix Binding Domains in Long Non-Coding RNAs

Long (>200vbps) non-coding RNAs (lncRNA) can act as a scaffold promoting the interaction of several proteins, RNA and DNA. Some lncRNAs interact with the DNA via a triple helix formation. Triple helices are formed by a single stranded RNA/DNA molecule, which binds to the major groove of a double helix following a canonical code. Recently, sequence analysis methods have been proposed to detect triple helices for a given RNA and DNA sequences. We propose the Triplex Domain Finder (TDF) to detect DNA binding domains in RNA molecules. For a candidate lncRNA and potential target DNA regions, i.e. promoter of genes differentially regulated after the knockdown of the lncRNA, TDF evaluates whether particular RNA regions are likely to form DNA binding domains (DBD). Moreover, the DNA binding sites from the predicted DBDs are used to indicate potential target DNA regions, i.e. genes with high binding site coverage in their promoter. The command line tool provides results on a user friendly and graphical html interface. A case study on FENDRR, an lncRNA known to form triple helices, demonstrates that TDF is able to recover both previously discovered DBDs and DNA binding sites. Source code, tutorial and case studies are available at www.regulatory-genomics.org/tdf.

Bioinformatics

A semi-supervised approach uncovers thousands of intragenic enhancers differentially activated in human cells

BackgroundTranscriptional enhancers are generally known to regulate gene transcription from afar. Their activation involves a series of changes in chromatin marks and recruitment of protein factors. These enhancers may also occur inside genes, but how many may be active in human cells and their effects on the regulation of the host gene remains unclear.\n\nResultsWe describe a novel semi-supervised method based on the relative enrichment of chromatin signals between 2 conditions to predict active enhancers. We applied this method to the tumoral K562 and the normal GM12878 cell lines to predict enhancers that are differentially active in one cell type. These predictions show enhancer-like properties according to positional distribution, correlation with gene expression and production of enhancer RNAs. Using this model, we predict 10,365 and 9,777 intragenic active enhancers in K562 and GM12878, respectively, and relate the differential activation of these enhancers to expression and splicing differences of the host genes.\n\nConclusionsWe propose that the activation or silencing of intragenic transcriptional enhancers modulate the regulation of the host gene by means of a local change of the chromatin and the recruitment of enhancer-related factors that may interact with the RNA directly or through the interaction with RNA binding proteins. Predicted enhancers are available at http://regulatorygenomics.upf.edu/Projects/enhancers.html

Bioinformatics

Segmenting Microarrays with Deep Neural Networks

Microarray images consist of thousands of spots, e ach of which corresponds to a different biological material. The microarray segmentation problem is to work out which pixels belong to which spots, even in presence of noise and corruption. We propose a solution based on deep neural networks, which achieves excellent results both on simulated and experimental data. We have made the source code for our solution available on Github under a permissive license.

Bioinformatics

DISSECT: A new tool for analyzing extremely large genomic datasets

Computational tools are quickly becoming the main bottleneck to analyze large-scale genomic and genetic data. This big-data problem, affecting a wide range of fields, is becoming more acute with the fast increase of data available. To address it, we developed DISSECT, a new, easy to use, and freely available software able to exploit the parallel computer architectures of supercomputers to perform a wide range of genomic and epidemiologic analyses which currently can only be carried out on reduced sample sizes or in restricted conditions. We showcased our new tool by addressing the challenge of predicting phenotypes from genotype data in human populations using Mixed Linear Model analysis. We analyzed simulated traits from half a million individuals genotyped for 590,004 SNPs using the combined computational power of 8,400 processor cores. We found that prediction accuracies in excess of 80% of the theoretical maximum could be achieved with large numbers of training individuals.

Bioinformatics

Resolving microsatellite genotype ambiguity in populations of allopolyploid and diploidized autopolyploid organisms using negative correlations between allelic variables

A major limitation in the analysis of genetic marker data from polyploid organisms is non-Mendelian segregation, particularly when a single marker yields allelic signals from multiple, independently segregating loci (isoloci). However, with markers such as microsatellites that detect more than two alleles, it is sometimes possible to deduce which alleles belong to which isoloci. Here we describe a novel mathematical property of codominant marker data when it is recoded as binary (presence/absence) allelic variables: under random mating in an infinite population, two allelic variables will be negatively correlated if they belong to the same locus, but uncorrelated if they belong to different loci. We present an algorithm to take advantage of this mathematical property, sorting alleles into isoloci based on correlations, then refining the allele assignments after checking for consistency with individual genotypes. We demonstrate the utility of our method on simulated data, as well as a real microsatellite dataset from a natural population of octoploid white sturgeon (Acipenser transmontanus). Our methodology is implemented in the R package O_SCPLOWPOLYSATC_SCPLOW version 1.5.

Bioinformatics

Hi-Cpipe: a pipeline for high-throughput chromosome capture

HiC-inspector is a toolkit for the analysis and visualization of data generated by high-throughput chromatin conformation capture (HiC). The analysis module comprises steps of data formatting, genome alignment, quality control and filtering, identification of genome-wide chromatin interactions, and statistics. The interactive browser enables visual inspection of interaction data generated and analysis results.

Bioinformatics

WaspAtlas: A Nasonia vitripennis gene database

SummaryWaspAtlas is a new integrated gene database for the emerging model organism Nasonia vitripennis, which combines annotation data from all available annotation releases with original analyses to form the most comprehensive N. vitripennis resource to date. WaspAtlas allows users to browse and search for gene information in a clear and co-ordinated fashion providing detailed illustrations and easy to understand summaries. The database provides a platform for integrating gene expression and DNA methylation data. WaspAtlas also functions as an archive for empirical data relating to genes, allowing users to easily browse published data relating to their gene(s) of interest.\n\nAvailabilityFreely available on the web at http://waspatlas.com. Website implemented in Catalyst, MySQL and Apache, with all major browsers supported.\n\nContactnjd23@le.ac.uk

Bioinformatics