Search bioRxiv⌕ Search

Biology subjects

Badelt, S.

Publications and source records attributed to Badelt, S..

2 recordsLinked to original sources

DrTransformer: Heuristic cotranscriptional RNA folding using the nearest neighbor energy model

BackgroundFolding during transcription can have an important influence on the structure and function of [R]NA molecules, as regions closer to the 5 end can fold into metastable structures before potentially stronger interactions with the 3 end become available. Thermodynamic [R]NA folding models are not suitable to analyze this problem, as they can only calculate properties of the equilibrium distribution. Other software packages that simulate the kinetic process of [R]NA folding during transcription exist, but they are mostly applicable for short sequences. ResultsWe present a new algorithm that tracks changes to the [R]NA secondary structure ensemble during transcription. At every transcription step, new representative local minima are identified, a neighborhood relation is defined and transition rates are estimated for kinetic simulations. After every simulation, a part of the ensemble is removed and the remainder is used to search for new potentially relevant structures. The presented algorithm is deterministic (up to numeric instabilities of simulations), fast (in comparison with existing methods), and it is capable of folding [R]NAs much longer than 200 nucleotides. AvailabilityThis software is open-source and available at https://github.com/ViennaRNA/drtransformer.

bioinformatics↗

Caveats to deep learning approaches to RNA secondary structure prediction

Machine learning (ML) and in particular deep learning techniques have gained popularity for predicting structures from biopolymer sequences. An interesting case is the prediction of RNA secondary structures, where well established biophysics based methods exist. The accuracy of these classical methods is limited due to lack of experimental parameters and certain simplifying assumptions and has seen little improvement over the last decade. This makes RNA folding an attractive target for machine learning and consequently several deep learning models have been proposed in recent years. However, for ML approaches to be competitive for de-novo structure prediction, the models must not just demonstrate good phenomenological fits, but be able to learn a (complex) biophysical model. In this contribution we discuss limitations of current approaches, in particular due to biases in the training data. Furthermore, we propose to study capabilities and limitations of ML models by first applying them on synthetic data (obtained from a simplified biophysical model) that can be generated in arbitrary amounts and where all biases can be controlled. We assume that a deep learning model that performs well on these synthetic, would also perform well on real data, and vice versa. We apply this idea by testing several ML models of varying complexity. Finally, we show that the best models are capable of capturing many, but not all, properties of RNA secondary structures. Most severely, the number of predicted base pairs scales quadratically with sequence length, even though a secondary structure can only accommodate a linear number of pairs.

bioinformatics↗