Search bioRxivSearch

Biology subjects

Domanska, D.

Publications and source records attributed to Domanska, D..

3 recordsLinked to original sources

MirGeneDB2.0: the curated microRNA Gene Database

Small non-coding RNAs have gained substantial attention due to their roles in animal development and human disorders. Among them, microRNAs are unique because individual gene sequences are conserved across the animal kingdom. In addition, unique and mechanistically well understood features can clearly distinguish bona fide miRNAs from the myriad other small RNAs generated by cells. However, making this separation is not a common practice and, thus, not surprisingly, the heterogeneous quality of available miRNA complements has become a major concern in microRNA research. We addressed this by extensively expanding our curated microRNA gene database MirGeneDB to 45 organisms that represent the full taxonomic breadth of Metazoa. By consistently annotating and naming more than 10,900 microRNA genes in these organisms, we show that previous microRNA annotations contained not only many false positives, but surprisingly lacked more than 2,100 bona fide microRNAs. Indeed, curated microRNA complements of closely related organisms are very similar and can be used to reconstruct Metazoan evolution. MirGeneDB represents a robust platform for microRNA-based research, providing deeper and more significant insights into the biology and evolution of miRNAs but also biomedical and biomarker research. MirGeneDB is publicly and freely available at http://mirgenedb.org/.

genomics

Mind your gaps: Overlooking assembly gaps confounds statistical testing in genome analysis

BackgroundThe difficulties associated with sequencing and assembling some regions of the DNA sequence result in gaps in the reference genomes that are typically represented as stretches of Ns. Although the presence of assembly gaps causes a slight reduction in the mapping rate in many experimental settings, that does not invalidate the typical statistical testing comparing read count distributions across experimental conditions. However, we hypothesize that not handling assembly gaps in the null model may confound statistical testing of co-localization of genomic features.\n\nResultsFirst, we performed a series of explorative analyses to understand whether and how the public genomic tracks intersect the assembly gaps track (hg19). The findings rightly confirm that the genomic regions in public genomic tracks intersect very little with assembly gaps and the intersection was observed only at the beginning and end regions of the assembly gaps rather than covering the whole gap sizes. Further, we simulated a set of query and reference genomic tracks in a way that nullified any dependence between them to test our hypothesis that not avoiding assembly gaps in the null model would result in spurious inflation of statistical significance. We then contrasted the distributions of test statistics and p-values of Monte Carlo simulation-based permutation tests that either avoided or not avoided assembly gaps in the null model when testing for significant co-localization between a pair of query and reference tracks. We observed that the statistical tests that did not account for the assembly gaps in the null model resulted in a distribution of the test statistic that is shifted to the right and a distribu tion of p-values that is shifted to the left (leading to inflated significance).\n\nConclusionOur results shows that not accounting for assembly gaps in statistical testing of co-localization analysis may lead to false positives and over-optimistic findings.

bioinformatics

Genome Build Information Is An Essential Part Of Genomic Track Files

Genomic locations are represented as coordinates on a specific genome build version, but the build information is frequently missing when coordinates are provided. It is essential to correctly interpret and analyse the genomic intervals contained in genomic track files. Here, we demonstrate that this crucial metadatum (or rather datum) is often isolated from the genomic track files in public repositories and journal articles, which could be a major time thief. We propose best practices to ensure that genome build version is always carried along with genomic track files. Although not a substitute to the best practices, we also provide a tool to predict the genome build version of genomic track files.

bioinformatics