Search bioRxiv⌕ Search

Biology subjects

Ghaddar, F.

Publications and source records attributed to Ghaddar, F..

2 recordsLinked to original sources

Random and natural non-coding RNA have similar structural motif patterns but can be distinguished by bulge, loop, and bond counts

An important question in evolutionary biology is whether and in what ways genotype-phenotype (GP) map biases can influence evolutionary trajectories. Untangling the relative roles of natural selection and biases (and other factors) in shaping phenotypes can be difficult. Because RNA secondary structure (SS) can be analysed in detail mathematically and computationally, is biologically relevant, and a wealth of bioinformatic data is available, it offers a good model system for studying the role of bias. For quite short RNA (length L [≤] 126), it has recently been shown that natural and random RNA are structurally very similar, suggesting that bias strongly constrains evolutionary dynamics. Here we extend these results with emphasis on much larger RNA with length up to 3000 nucleotides. By examining both abstract shapes and structural motif frequencies (ie the numbers of helices, bonds, bulges, junctions, and loops), we find that large natural and random structures are also very similar, especially when contrasted to typical structures sampled from the space of all possible RNA structures. Our motif frequency study yields another result, that the frequencies of different motifs can be used in machine learning algorithms to classify random and natural RNA with quite high accuracy, especially for longer RNA (eg ROC AUC 0.86 for L = 1000). The most important motifs for classification are found to be the number of bulges, loops, and bonds. This finding may be useful in using SS to detect candidates for functional RNA within junk DNA regions.

evolutionary biology↗

Phenotype bias determines how RNA structures occupy the morphospace of all possible shapes

Morphospaces representations of phenotypic characteristics are often populated unevenly, leaving large parts unoccupied. Such patterns are typically ascribed to contingency, or else to natural selection disfavouring certain parts of the morphospace. The extent to which developmental bias, the tendency of certain phenotypes to preferentially appear as potential variation, also explains these patterns is hotly debated. Here we demonstrate quantitatively that developmental bias is the primary explanation for the occupation of the morphospace of RNA secondary structure (SS) shapes. Upon random mutations, some RNA SS shapes (the frequent ones) are much more likely to appear than others. By using the RNAshapes method to define coarse-grained SS classes, we can directly compare the frequencies that non-coding RNA SS shapes appear in the RNAcentral database to frequencies obtained upon random sampling of sequences. We show that: a) Only the most frequent structures appear in nature; the vast majority of possible structures in the morphospace have not yet been explored. b) Remarkably small numbers of random sequences are needed to produce all the RNA SS shapes found in nature so far. c) Perhaps most surprisingly, the natural frequencies are accurately predicted, over several orders of magnitude in variation, by the likelihood that structures appear upon uniform random sampling of sequences. The ultimate cause of these patterns is not natural selection, but rather strong phenotype bias in the RNA genotype-phenotype map, a type of developmental bias or "findability constraint", which limits evolutionary dynamics to a hugely reduced subset of structures that are easy to "find".

evolutionary biology↗