Search bioRxiv⌕ Search

Biology subjects

Mareboina, M.

Publications and source records attributed to Mareboina, M..

2 recordsLinked to original sources

Nucleic Quasi-Primes: Identification of the Shortest Unique Oligonucleotide Sequences in a Species

Despite the exponential increase in sequencing information driven by massively parallel DNA sequencing technologies, universal and succinct genomic fingerprints for each organism are still missing. Identifying the shortest species-specific nucleic sequences offers insights into species evolution and holds potential practical applications in agriculture, wildlife conservation, and healthcare. We propose a new method for sequence analysis termed nucleic "quasi-primes", the shortest occurring sequences in each of 45,785 organismal reference genomes, present in one genome and absent from every other examined genome. In the human genome, we find that the genomic loci of nucleic quasi-primes are most enriched for genes associated with brain development and cognitive function. In a single-cell case study focusing on the human primary motor cortex, nucleic quasi-prime genes account for a significantly larger proportion of the variation based on average gene expression. Non-neuronal cell-types, including astrocytes, endothelial cells, microglia perivascular-macrophages, oligodendrocytes, and vascular and leptomeningeal cells, exhibited significant activation of quasi-prime containing gene associations related to cancer, while simultaneously suppressing quasi-prime containing genes were associated with cognitive, mental, and developmental disorders. We also show that human disease-causing variants, eQTLs, mQTLs and sQTLs are 4.43-fold, 4.34-fold, 4.29-fold and 4.21-fold enriched at human quasi-prime loci, respectively. These findings indicate that nucleic quasi-primes are genomic loci linked to the evolution of species-specific traits and in humans they provide insights in the development of cognitive traits and human diseases, including neurodevelopmental disorders.

genomics↗

The determinants of the rarity of nucleic and peptide short sequences in nature

The prevalence of nucleic and peptide short sequences across organismal genomes and proteomes has not been thoroughly investigated. Here we examined 45,785 reference genomes and 21,871 reference proteomes, spanning archaea, bacteria, viruses and eukaryotes to calculate the rarity of short sequences in them. To capture this, we developed a metric of the rarity of each sequence in nature, the Anti-Kardashian index. We find that the frequency of certain dipeptides in rare oligopeptide sequences is hundreds of times lower than expected, which is not the case for any dinucleotides. We also generate predictive regression models that infer the rarity of nucleic and proteomic sequences in nature. For six-mer peptide kmers the R2 performance of the regression models based on amino acid and dipeptide content is 0.816, whereas models based on physicochemical features achieve an R2 of 0.788. For twelve-mer nucleic kmers the R2 performance of our models based on mono and dinucleotides is 0.481. Our results indicate that the mono and dinucleotide composition of nucleic sequences and the amino acids, dipeptides and physicochemical properties of peptide sequences can explain a significant proportion of the variance in their frequencies between organisms in nature.

genomics↗