Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

PatternMarkers and Genome-Wide CoGAPS Analysis in Parallel Sets (GWCoGAPS) for data-driven detection of novel biomarkers via whole transcriptome Non-negative matrix factorization (NMF)

SummaryNon-negative Matrix Factorization (NMF) algorithms associate gene expression with biological processes (e.g., time-course dynamics or disease subtypes). Compared with univariate associations, the relative weights of NMF solutions can obscure biomarkers. Therefore, we developed a novel PatternMarkers statistic to extract genes for biological validation and enhanced visualization of NMF results. Finding novel and unbiased gene markers with PatternMarkers requires whole-genome data. However, NMF algorithms typically do not converge for the tens of thousands of genes in genome-wide profiling. Therefore, we also developed Genome-Wide CoGAPS Analysis in Parallel Sets (GWCoGAPS), the first robust whole genome Bayesian NMF using the sparse, MCMC algorithm, CoGAPS. This software contains analytic and visualization tools including a Shiny web application, patternMatcher, which are generalized for any NMF. Using these tools, we find granular brain-region and cell-type specific signatures with corresponding biomarkers in GTex data, illustrating GWCoGAPS and patternMarkers ascertainment of data-driven biomarkers from whole-genome data.\n\nAvailabilityPatternMarkers & GWCoGAPS are in the CoGAPS Bioconductor package (3.5) under the GPL license.\n\nContactgsteinobrien@jhmi.edu; ccolantu@jhmi.edu; ejfertig@jhmi.edu

bioinformatics

Minimal-assumption inference from population-genomic data

Samples of multiple complete genome sequences contain vast amounts of information about the evolutionary history of populations, much of it in the associations among polymorphisms at different loci. Current methods that take advantage of this linkage information rely on models of recombination and coalescence, limiting the sample sizes and populations that they can analyze. We introduce a method, Minimal-Assumption Genomic Inference of Coalescence (MAGIC), that reconstructs key features of the evolutionary history, including the distribution of coalescence times, by integrating information across genomic length scales without using an explicit model of recombination, demography or selection. Using simulated data, we show that MAGICs performance is comparable to PSMC on single diploid samples generated with standard coalescent and recombination models. More importantly, MAGIC can also analyze arbitrarily large samples and is robust to changes in the coalescent and recombination processes. Using MAGIC, we show that the inferred coalescence time histories of samples of multiple human genomes exhibit inconsistencies with a description in terms of an effective population size based on single-genome data.

evolutionary biology

Predicting Enhancer-Promoter Interaction from Genomic Sequence with Deep Neural Networks

In the human genome, distal enhancers are involved in regulating target genes through proxi-mal promoters by forming enhancer-promoter interactions. Although recently developed high-throughput experimental approaches have allowed us to recognize potential enhancer-promoter interactions genome-wide, it is still largely unclear to what extent the sequence-level information encoded in our genome help guide such interactions. Here we report a new computational method (named \"SPEID\") using deep learning models to predict enhancer-promoter interactions based on sequence-based features only, when the locations of putative enhancers and promoters in a particular cell type are given. Our results across six different cell types demonstrate that SPEID is effective in predicting enhancer-promoter interactions as compared to state-of-the-art methods that only use information from a single cell type. As a proof-of-principle, we also applied SPEID to identify somatic non-coding mutations in melanoma samples that may have reduced enhancer-promoter interactions in tumor genomes. This work demonstrates that deep learning models can help reveal that sequence-based features alone are sufficient to reliably predict enhancer-promoter interactions genome-wide.

bioinformatics

On the (im)possibility to reconstruct plasmids from whole genome short-read sequencing data

Plasmids are autonomous extra-chromosomal elements in bacterial cells that can carry genes that are important for bacterial survival. To benchmark algorithms for automated plasmid sequence reconstruction from short read sequencing data, we selected 42 publicly available complete bacterial genome sequences which were assembled by a combination of long- and short-read data. The selected bacterial genome sequence projects span 12 genera, containing 148 plasmids. We predicted plasmids from short-read data with four different programs (PlasmidSPAdes, Recycler, cBar and PlasmidFinder) and compared the outcome to the reference sequences.\n\nPlasmidSPAdes reconstructs plasmids based on coverage differences in the assembly graph. It reconstructed most of the reference plasmids (recall = 0.82) but approximately a quarter of the predicted plasmid contigs were false positives (precision = 0.76). PlasmidSPAdes merged 83 % of the predictions from genomes with multiple plasmids in a single bin. Recycler searches the assembly graph for sub-graphs corresponding to circular sequences and correctly predicted small plasmids but failed with long plasmids (recall = 0.12, precision = 0.30). cBar, which applies pentamer frequency composition analysis to detect plasmid-derived contigs, showed an overall recall and precision of 0.78 and 0.64. However, cBar only categorizes contigs as plasmid-derived and does not bin the different plasmids correctly within a bacterial isolate. PlasmidFinder, which searches for matches in a replicon database, had the highest precision (1.0) but was restricted by the contents of its database and the contig length obtained from de novo assembly (recall = 0.36).\n\nSurprisingly, PlasmidSPAdes and Recycler detected single isolated components corresponding to putative novel small plasmids (<10 kbp) which were also predicted as plasmids by cBar.\n\nThis study shows that it is possible to automatically predict plasmid sequences, but only for small plasmids. The reconstruction of large plasmids (>50 kbp) containing repeated sequences remains challenging and limits the high-throughput analysis of WGS data.\n\nAuthor SummaryShort read sequencing of the DNA of bacteria is often used to understand characteristics such as antibiotic resistance. However the assembly of short read sequencing data with the goal of reconstructing a complete genome is often fragmented and leaves gaps. Therefore independently replicating DNA fragments called plasmids cannot easily be identified from an assembly. Lately a number of programs have been developed to enable the automated prediction of the sequences of plasmids. Here we tested these programs by comparing their outcomes with complete genome sequences. None of the tested programs were able to fully and unambiguously predict distinct plasmid sequences. All programs performed best with the prediction of plasmids smaller than 50 kbp. Larger plasmids were only correctly predicted if they were present as a single contig in the assembly. While predictions by PlasmidSPAdes and cBar contained most of the plasmids, they were merged with or indistinguishable from other plasmids and sometimes chromosome sequences. PlasmidFinder missed most plasmids but all its predictions were correct. Without manual steps or long-read sequencing information, plasmid reconstruction from short read sequencing data remains challenging.

microbiology

Millstone: Software for Multiplex Microbial Genome Analysis and Engineering

Inexpensive DNA sequencing and advances in genome editing have made computational analysis a major rate-limiting step in adaptive laboratory evolution and microbial genome engineering. We describe Millstone, a web-based platform which automates genotype comparison and visualization for projects with up to hundreds of genomic samples. To enable iterative genome engineering, Millstone allows users to design oligonucleotide libraries and create successive versions of reference genomes. Millstone is open source and easily deployable to a cloud platform, local cluster, or desktop, making it a scalable solution for any lab.

synthetic biology

Locus-specific ChIP combined with NGS analysis reveals genomic regulatory regions that physically interact with the Pax5 promoter in a chicken B cell line

Chromosomal interactions regulate genome functions, such as transcription, via dynamic chromosomal organization in the nucleus. In this study, we identified genomic regions that physically bind to the promoter region of the Pax5 gene in the chicken B-cell line DT40, with the goal of obtaining mechanistic insight into transcriptional regulation through chromosomal interaction. Using insertional chromatin immunoprecipitation (iChIP) in combination with next-generation sequencing (NGS) (iChIP-Seq), we found that the Pax5 promoter bound to multiple genomic regions. The identified chromosomal interactions were independently confirmed by in vitro engineered DNA-binding molecule-mediated ChIP (in vitro enChIP) in combination with NGS (in vitro enChIP-Seq). Comparing chromosomal interactions in wild-type DT40 with those in a macrophage-like counterpart, we found that some of the identified chromosomal interactions were organized in a B cell-specific manner. In addition, deletion of a B cell-specific interacting genomic region in chromosome 11, which was marked by active enhancer histone modifications, resulted in moderate but significant down-regulation of Pax5 transcription. Together, these results suggested that Pax5 transcription in DT40 cells is regulated by inter-chromosomal interactions. Moreover, these analyses showed that iChIP-Seq and in vitro enChIP-Seq are useful for non-biased identification of functional genomic regions that physically interact with a locus of interest.

molecular biology

Non-random inversion landscapes in prokaryotic genomes are shaped by heterogeneous selection pressures

Inversions are a major contributor to structural genome evolution in prokaryotes. Here, using a novel alignment-based method, we systematically compare 1651 bacterial and 98 archaeal genomes to show that inversion landscapes are frequently biased towards (symmetric) inversions around the origin-terminus axis. However, symmetric inversion bias is not a universal feature of prokaryotic genome evolution but varies considerably across clades. At the extremes, inversion landscapes in Bacillus-Clostridium and Actinobacteria are dominated by symmetric inversions, while there is little or no systematic bias favouring symmetric rearrangements in archaea with a single origin of replication. Within clades, we find strong but clade-specific relationships between symmetric inversion bias and different features of adaptive genome architecture, including the distance of essential genes to the origin of replication and the preferential localization of genes on the leading strand. We suggest that heterogeneous selection pressures have converged to produce similar patterns of structural genome evolution across prokaryotes.

evolutionary biology

Novel insights into karyotype evolution and whole genome duplications in legumes

Legumes (family Fabaceae) are globally important crops due to their nitrogen fixing ability. Papilionoideae, the best-studied subfamily, have undergone a Whole Genome Duplication (WGD) around 59 million years ago. Recent study found varying WGD ages in subfamilies Mimosoideae and Caesalpinioideae and proposed multiple occurrences of WGD across the family based on gene duplication patterns. Despite that, the genome evolution of legume ancestor into modern legumes after the WGD is not well-understood. We aimed to study genome evolution at the subfamily level using gene-based linkage maps for Acacia auriculiformis and A. mangium (Mimosoideae) and we discovered evidence for a WGD event in Acacia. In additional to synonymous substitution rate (Ks) analysis, we used ancestral karyotype prediction to further corroborate this WGD and elucidate underlying mechanisms of karyotype evolution in Fabaceae. Using publicly available transcriptome resources from 25 species across the family Fabaceae and 2 species from order Fabales, we found that the variations in WGD ages highly correlate (R=0.8606, p-value<0.00001) with the divergence age of Vitis vinifera as an outgroup. If the variation of Ks is corrected, the age of WGDs of the family Fabaceae should be the same and therefore, parsimony would favour a single WGD near the base of Fabaceae over multiple independent WGDs across Fabaceae. In addition, we demonstrated that genome comparison of Papilionoideae with other subfamily provide important insights in understanding genome evolution in legumes.

evolutionary biology

Fragile dynamics enable diverse genomic determinants to influence arrhythmia propensity

Genotype-phenotype relationships are determinants of human diseases. Often, we know little about why so many genes are involved in complex common diseases. We hypothesized that this multigene effect arises from the relationship between genes and physiological dynamics. We tested this hypothesis for arrhythmias as physiological dynamics define this disease. We integrated graph theory analysis of genomic and protein-protein interaction networks with dynamical models of ion channel function to identify the physiological dynamics of genome wide variation for five different arrhythmias. Regulatory networks for the cardiac conduction system and arrhythmias were constructed from GWAS and known disease genes. Electrophysiological models of myocyte action potentials were used to conduct extensive parameter variations to identify robust and fragile kinetic parameters that were then, using regulatory networks, associated with genomic determinants. We find that genome-wide determinants of arrhythmias that represent many cellular processes are selectively associated with fragile physiological dynamics of ion channel kinetics. This association predicts disease propensity. Deep RNA sequencing from human left ventricular tissue of arrhythmia and control subjects confirmed the predictive relationship. Taken together these studies show that the varied multigene effects of arrhythmias arises because of associations with fragile kinetic parameters of cardiac electrophysiology.\n\nSignificance StatementOur understanding of the genetics of common diseases has advanced exponentially over the past decade. We now know that differences and variation in multiple genes contribute to disease susceptibility with significant heterogeneity in the phenotype. However, how genetic variation contributes to disease phenotypes remains unknown. We hypothesized that the relationships between physiological dynamics and genetic architecture is a fundamental determinant of disease susceptibility and genetic heterogeneity. To test our hypothesis, we integrated mathematical models of cardiac electrophysiology with genetic network models of cardiac arrhythmias. We found that disease related genome variants were selectively associated with fragile kinetic parameters that predict disease propensity and identified several novel cellular processes associated with arrhythmogenesis.

systems biology

A Flow Procedure for the Linearization of Genome Sequence Graphs.

1Efforts to incorporate human genetic variation into the reference human genome have converged on the idea of a graph representation of genetic variation within a species, a genome sequence graph. A sequence graph represents a set of individual haploid reference genomes as paths in a single graph. When that set of reference genomes is sufficiently diverse, the sequence graph implicitly contains all frequent human genetic variations, including translocations, inversions, deletions, and insertions.\n\nIn representing a set of genomes as a sequence graph one encounters certain challenges. One of the most important is the problem of graph linearization, essential both for efficiency of storage and access, as well as for natural graph visualization and compatibility with other tools. The goal of graph linearization is to order nodes of the graph in such a way that operations such as access, traversal and visualization are as efficient and effective as possible.\n\nA new algorithm for the linearization of sequence graphs, called the flow procedure, is proposed in this paper. Comparative experimental evaluation of the flow procedure against other algorithms shows that it outperforms its rivals in the metrics most relevant to sequence graphs.

bioinformatics

A scalable Bayesian method for integrating functional information in genome-wide association studies

Although genome-wide association studies (GWASs) have identified many risk loci for complex traits and common diseases, most of the identified associations reside in noncoding regions and have unknown biological functions. Recent genomic sequencing studies have produced a rich resource of annotations that help characterize the function of genetic variants. Integrative analysis that incorporates these functional annotations into GWAS can help elucidate the biological mechanisms underlying the identified associations and help prioritize causal-variants. Here, we develop a novel, flexible Bayesian variable selection model with efficient computational techniques for such integrative analysis. Different from previous approaches, our method models the effect-size distribution and probability of causality for variants with different annotations and jointly models genome-wide variants to account for linkage disequilibrium (LD), thus prioritizing associations based on the quantification of the annotations and allowing for multiple causal-variants per locus. Our efficient computational algorithm dramatically improves both computational speed and posterior sampling convergence by taking advantage of the block-wise LD structures of human genomes. With simulations, we show that our method accurately quantifies the functional enrichment and performs more powerful for identifying true causal-variants than several competing methods. The power gain brought up by our method is especially apparent in cases when multiple causal-variants in LD reside in the same locus. We also apply our method for an in-depth GWAS of age-related macular degeneration with 33,976 individuals and 9,857,286 variants. We find the strongest enrichment for causality among non-synonymous variants (54x more likely to be causal, 1.4x larger effect-sizes) and variants in active promoter (7.8x more likely, 1.4x larger effect-sizes), as well as identify 5 potentially novel loci in addition to the 32 known AMD risk loci. In conclusion, our method is shown to efficiently integrate functional information in GWASs, helping identify causal variants and underlying biology.\n\nAuthor summaryWe propose a novel Bayesian hierarchical model to account for linkage disequilibrium (LD) and multiple functional annotations in GWAS, paired with an expectation-maximization Markov chain Monte Carlo (EM-MCMC) computational algorithm to jointly analyze genome-wide variants. Our method improves the MCMC convergence property to ensure accurate Bayesian inference of the quantifications of the functional enrichment pattern and fine-mapped association results. By applying our method to the real GWAS of age-related macular degeneration (AMD) with various functional annotations (i.e., gene-based, regulatory, and chromatin states), we find that the variants of non-synonymous, coding, and active promoter annotations have the highest causal probability and the largest effect-sizes. In addition, our method produces fine-mapped association results in the identified risk loci, two of which are shown as examples (C2/CFB/SKIV2L and C3) with justifications by haplotype analysis, model comparison, and conditional analysis. Therefore, we believe our integrative method will be useful for quantifying the enrichment pattern of functional annotations in GWAS, and then prioritizing associations with respect to the learned functional enrichment pattern.

genetics

Insights into the genetic determinism andevolution of recombination rates fromcombining multiple genome-wide datasets inSheep

Recombination is a complex biological process that results from a cascade of multiple events during meiosis. Understanding the genetic determinism of recombination can help to understand if and how these events are interacting. To tackle this question, we studied the patterns of recombination in sheep, using multiple approaches and datasets. We constructed male recombination maps in a dairy breed from the south of France (the Lacaune breed) at a fine scale by combining meiotic recombination rates from a large pedigree genotyped with a 50K SNP array and historical recombination rates from a sample of unrelated individuals genotyped with a 600K SNP array. This analysis revealed recombination patterns in sheep similar to other mammals but also genome regions that have likely been affected by directional and diversifying selection. We estimated the average recombination rate of Lacaune sheep at 1.5 cM/Mb, identified about 50,000 crossover hotspots on the genome and found a high correlation between historical and meiotic recombination rate estimates. A genome-wide association study revealed two major loci affecting inter-individual variation in recombination rate in Lacaune, including the RNF212 and HEI10 genes and possibly 2 other loci of smaller effects including the KCNJ15 and FSHR genes. Finally, we compared our results to those obtained previously in a distantly related population of domestic sheep, the Soay. This comparison revealed that Soay and Lacaune males have a very similar distribution of recombination along the genome and that the two datasets can be combined to create more precise male meiotic recombination maps in sheep. Despite their similar recombination maps, we show that Soay and Lacaune males exhibit different heritabilities and QTL effects for inter-individual variation in genome-wide recombination rates.

genetics

Millennia of genomic stability within the invasive Para C Lineage of Salmonella enterica

Salmonella enterica serovar Paratyphi C is the causative agent of enteric (paratyphoid) fever. While today a potentially lethal infection of humans that occurs in Africa and Asia, early 20th century observations in Eastern Europe suggest it may once have had a wider-ranging impact on human societies. We recovered a draft Paratyphi C genome from the 800-year-old skeleton of a young woman in Trondheim, Norway, who likely died of enteric fever. Analysis of this genome against a new, significantly expanded database of related modern genomes demonstrated that Paratyphi C is descended from the ancestors of swine pathogens, serovars Choleraesuis and Typhisuis, together forming the Para C Lineage. Our results indicate that Paratyphi C has been a pathogen of humans for at least 1,000 years, and may have evolved after zoonotic transfer from swine during the Neolithic period.\n\nOne Sentence SummaryThe combination of an 800-year-old Salmonella enterica Paratyphi C genome with genomes from extant bacteria reshapes our understanding of this pathogens origins and evolution.

microbiology

Validation and Implementation of CLIA-Compliant Whole Genome Sequencing (WGS) in Public Health Laboratory

BackgroundPublic health microbiology laboratories (PHL) are at the cusp of unprecedented improvements in pathogen identification, antibiotic resistance detection, and outbreak investigation by using whole genome sequencing (WGS). However, considerable challenges remain due to the lack of common standards.\n\nObjectives1) Establish the performance specifications of WGS applications used in PHL to conform with CLIA (Clinical Laboratory Improvements Act) guidelines for laboratory developed tests (LDT), 2) Develop quality assurance (QA) and quality control (QC) measures, 3) Establish reporting language for end users with or without WGS expertise, 4) Create a validation set of microorganisms to be used for future validations of WGS platforms and multi-laboratory comparisons and, 5) Create modular templates for the validation of different sequencing platforms.\n\nMethodsMiSeq Sequencer and Illumina chemistry (Illumina, Inc.) were used to generate genomes for 34 bacterial isolates with genome sizes from 1.8 to 4.7 Mb and wide range of GC content (32.1%-66.1%). A customized CLCbio Genomics Workbench - shell script bioinformatics pipeline was used for the data analysis.\n\nResultsWe developed a validation panel comprising ten Enterobacteriaceae isolates, five gram-positive cocci, five gram-negative non-fermenting species, nine Mycobacterium tuberculosis, and five miscellaneous bacteria; the set represented typical workflow in the PHL. The accuracy of MiSeq platform for individual base calling was >99.9% with similar results shown for reproducibility/repeatability of genome-wide base calling. The accuracy of phylogenetic analysis was 100%. The specificity and sensitivity inferred from MLST and genotyping tests were 100%. A test report format was developed for the end users with and without WGS knowledge.\n\nConclusionWGS was validated for routine use in PHL according to CLIA guidelines for LDTs. The validation panel, sequencing analytics, and raw sequences will be available for future multi-laboratory comparisons of WGS in PHL. Additionally, the WGS performance specifications and modular validation template are likely to be adaptable for the validation of other platforms and reagents kits.

microbiology

Genomic contingencies and beak shape variation in a hybrid species

Hybridization is increasingly recognized as a potent evolutionary force. Though additive genetic variation and novel combinations of parental genes theoretically increase the potential for hybrid species to adapt, few empirical studies have investigated the adaptive potential within a hybrid species. Here, we investigate factors promoting phenotypic divergence using genomically diverged island populations of the homoploid hybrid Italian sparrow Passer italiae from Crete, Corsica, and Sicily. We address whether genomic contingencies, adaptation to climate or diet best explain divergence in beak morphology. Populations vary significantly in beak morphology, both between and within islands of origin. Temperature seasonality best explains population divergence in beak size. Interestingly, beak shape along all significant dimensions of variation was best explained by annual precipitation, genomic composition and their interaction, suggesting a role for contingencies. Moreover, beak shape similarity to a parent species correlates with proportion of the genome inherited from that species, consistent with the presence of contingencies. In conclusion, adaptation to local conditions and genomic contingencies arising from putatively independent hybridization events jointly explain beak morphology in the Italian sparrow. Hence, hybridization may induce contingencies and restrict evolution in certain directions dependent on the genetic background.

evolutionary biology

Intervene: a tool for intersection and visualization of multiple gene or genomic region sets

BackgroundA common task for scientists relies on comparing lists of genes or genomic regions derived from high-throughput sequencing experiments. While several tools exist to intersect and visualize sets of genes, similar tools dedicated to the visualization of genomic region sets are currently limited.\n\nResultsTo address this gap, we have developed the Intervene tool, which provides an easy and automated interface for the effective intersection and visualization of genomic region or list sets, thus facilitating their analysis and interpretation. Intervene contains three modules: venn to generate Venn diagrams of up to six sets, upset to generate UpSet plots of multiple sets, and pairwise to compute and visualize intersections of multiple sets as clustered heat maps. Intervene, and its interactive web ShinyApp companion, generate publication-quality figures for the interpretation of genomic region and list sets.\n\nConclusionsIntervene and its web application companion provide an easy command line, and an interactive web interface to compute intersections of multiple genomic and list sets. They also have the capacity to plot intersections using easy-to-interpret visual approaches. Intervene is developed and designed to meet the needs of both computer scientists and biologists. The source code is freely available at https://bitbucket.org/CBGR/intervene, with the web application available at https://asntech.shinyapps.io/intervene.

bioinformatics

Genome-wide regulatory model from MPRA data predicts functional regions, eQTLs, and GWAS hits

Massively-parallel reporter assays (MPRA) enable unprecedented opportunities to test for regulatory activity of thousands of regulatory sequences. However, MPRA only assay a subset of the genome thus limiting their applicability for genome-wide functional annotations. To overcome this limitation, we have used existing MPRA datasets to train a machine learning model that uses DNA sequence information, regulatory motif annotations, evolutionary conservation, and epigenomic information to predict genomic regions that show enhancer activity when tested in MPRA assays. We used the resulting model to generate global predictions of regulatory activity at single-nucleotide resolution across 14 million common variants. We find that genetic variants with stronger predicted regulatory activity show significantly lower minor allele frequency, indicative of evolutionary selection within the human population. They also show higher over-lap with eQTL annotations across multiple tissues relative to the background SNPs, indicating that their perturbations in vivo more frequently result in changes in gene expression. In addition, they are more frequently associated with trait-associated SNPs from genome-wide association studies (GWAS), enabling us to prioritize genetic variants that are more likely to be causal based on their predicted regulatory activity. Lastly, we use our model to compare MPRA inferences across cell types and platforms and to prioritize the assays most predictive of MPRA assay results, including cell-dependent DNase hypersensitivity sites and transcription factors known to be active in the tested cell types. Our results indicate that high-throughput testing of thousands of putative regions, coupled with regulatory predictions across millions of sites, presents a powerful strategy for systematic annotation of genomic regions and genetic variants.

bioinformatics

REPARATION: Ribosome Profiling Assisted (Re-)Annotation of Bacterial genomes.

Prokaryotic genome annotation is highly dependent on automated methods, as manual curation cannot keep up with the exponential growth of sequenced genomes. Current automated methods depend heavily on sequence context and often underestimate the complexity of the proteome. We developed REPARATION (RibosomeE Profiling Assisted (Re-)AnnotaTION), a de novo algorithm that takes advantage of experimental protein translation evidence from ribosome profiling (Ribo-seq) to delineate translated open reading frames (ORFs) in bacteria, independent of genome annotation. REPARATION evaluates all possible ORFs in the genome and estimates minimum thresholds based on a growth curve model to screen for spurious ORFs. We applied REPARATION to three annotated bacterial species to obtain a more comprehensive mapping of their translation landscape in support of experimental data. In all cases, we identified hundreds of novel (small) ORFs including variants of previously annotated ORFs. Our predictions were supported by matching mass spectrometry (MS) proteomics data, sequence composition and conservation analysis. REPARATION is unique in that it makes use of experimental translation evidence to perform de novo ORF delineation in bacterial genomes irrespective of the sequence context of the reading frame.

microbiology