Search bioRxivSearch

Biology subjects

Tatiana Tatarinova

Publications and source records attributed to Tatiana Tatarinova.

4 recordsLinked to original sources

Seqping: Gene Prediction Pipeline for Plant Genomes using Self-Trained Gene Models and Transcriptomic Data

SummaryAlthough various software are available for gene prediction, none of the currently available gene-finders have a universal Hidden Markov Models (HMM) that can perform gene prediction for all organisms equally well in an automatic fashion. Here, we report an automated pipeline that performs gene prediction using selftrained HMM models and transcriptomic data. The program processes the genome and transcriptome sequences of a target species through GlimmerHMM, SNAP, and AUGUSTUS training pipeline that ends with the program MAKER2 combining the predictions from the three models in association with the transcriptomic evidence. The pipeline generates species-specific HMMs and is able to predict genes that are not biased to other model organisms. Our evaluation of the program revealed that it performed better than the use of the closest related HMM from a standalone program.\n\nAvailability and ImplementationDistributed under the GNU license with free download at http://sourceforge.net/projects/seqping and http://genomsawit.mpob.gov.my.\n\nContactchankl@mpob.gov.my\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Whole-genome modeling accurately predicts quantitative traits, as revealed in plants.

Many adaptive events in natural populations, as well as response to artificial selection, are caused by polygenic action. Under selective pressure, the adaptive traits can quickly respond via small allele frequency shifts spread across numerous loci. We hypothesize that a large proportion of current phenotypic variation between individuals may be best explained by population admixture.\n\nWe thus consider the complete, genome-wide universe of genetic variability, spread across several ancestral populations originally separated. We experimentally confirmed this hypothesis by predicting the differences in quantitative disease resistance levels among accessions in the wild legume Medicago truncatula. We discovered also that variation in genome admixture proportion explains most of phenotypic variation for several quantitative functional traits, but not for symbiotic nitrogen fixation. We shown that positive selection at the species level might not explain current, rapid adaptation.\n\nThese findings prove the infinitesimal model as a mechanism for adaptation of quantitative phenotypes. Our study produced the first evidence that the whole-genome modeling of DNA variants is the best approach to describe an inherited quantitative trait in a higher eukaryote organism and proved the high potential of admixture-based analyses. This insight contribute to the understanding of polygenic adaptation, and can accelerate plant and animal breeding, and biomedicine research programs.

Genetics

The mysterious orphans of Mycoplasmataceae

BackgroundThe length of a protein sequence is largely determined by its function, i.e. each functional group is associated with an optimal size. However, comparative genomics revealed that proteins length may be affected by additional factors. In 2002 it was shown that in bacterium Escherichia coli and the archaeon Archaeoglobus fulgidus, protein sequences with no homologs are, on average, shorter than those with homologs [1]. Most experts now agree that the length distributions are distinctly different between protein sequences with and without homologs in bacterial and archaeal genomes. In this study, we examine this postulate by a comprehensive analysis of all annotated prokaryotic genomes and focusing on certain exceptions.\n\nResultsWe compared lengths distributions of \"having homologs proteins\" (HHPs) and \"non-having homologs proteins\" (orphans or ORFans) in all currently annotated completely sequenced prokaryotic genomes. As expected, the HHPs and ORFans have strikingly different length distributions in almost all genomes. As previously established, the HHPs, indeed, are, on average, longer than the ORFans, and the length distributions for the ORFans have a relatively narrow peak, in contrast to the HHPs, whose lengths spread over a wider range of values. However, about thirty genomes do not obey these rules. Practically all genomes of Mycoplasma and Ureaplasma have atypical ORFans distributions, with the mean lengths of ORFan larger than the mean lengths of HHPs. These genera constitute over 80% of atypical genomes.\n\nConclusionsWe confirmed on a ubiquitous set of genomes the previous observation that HHPs and ORFans have different gene length distributions. We also showed that Mycoplasmataceae genomes have very distinctive distributions of ORFans lengths. We offer several possible biological explanations of this phenomenon.

Evolutionary Biology

benchNGS : An approach to benchmark short reads alignment tools

In the last decade a number of algorithms and associated software were developed to align next generation sequencing (NGS) reads to relevant reference genomes. The results of these programs may vary significantly, especially when the NGS reads are contain mutations not found in the reference genome. Yet there is no standard way to compare these programs and assess their biological relevance.\n\nWe propose a benchmark to assess accuracy of the short reads mapping based on the pre-computed global alignment of closely related genome sequences. In this paper we outline the method and also present a short report of an experiment performed on five popular alignment tools.

Bioinformatics