Search bioRxiv⌕ Search

Biology subjects

Bini, G.

Publications and source records attributed to Bini, G..

3 recordsLinked to original sources

Decoding RNA-RNA Interactions: The Role of Low-Complexity Repeats and a Deep Learning Framework for Sequence-Based Prediction

RNA-RNA interactions (RRIs) are fundamental to gene regulation and RNA processing, yet their molecular determinants remain unclear. In this work, we analyze several large-scale RRI datasets and identify low-complexity repeats (LCRs), including simple tandem repeats, as key drivers of RRIs. Our findings reveal that LCRs enable thermodynamically stable interactions with multiple partners, positioning them as key hubs in RNA-RNA interaction networks. These RRIs appear to be important for several aspects of RNA metabolism. Sequencing-based analysis of the lncRNA Lhx1os interactors validates the importance of LCRs in shaping contacts potentially involved in neuronal development. Recognizing the pivotal role of sequence determinants, we develop RIME, a deep learning model that predicts RRIs by leveraging embeddings from a nucleic acid language model. RIME outperforms traditional thermodynamics-based tools, successfully captures the role of LCRs and prioritizes high-confidence interactions, including those established by lncRNAs. RIME is freely available at https://tools.tartaglialab.com/rna_rna.

biochemistry↗

Accurate Predictions of Phase Separating Proteins at Single Amino Acid Resolution

Liquid-liquid phase separation (LLPS) is a molecular mechanism that leads to the formation of membraneless organelles inside the cell. Despite recent advances in the experimental probing and computational prediction of proteins involved in this process, the identification of the protein regions driving LLPS and the prediction of the effect of mutations on LLPS are lagging behind. Here, we introduce catGRANULE 2.0 ROBOT (R - Ribonucleoprotein, O - Organization, in B - Biocondensates, O - Organelle, T - Types), an advanced algorithm for predicting protein LLPS at single amino acid resolution. Integrating physico-chemical properties of the proteins and structural features derived from AlphaFold models, catGRANULE 2.0 ROBOT significantly surpasses traditional sequence-based and state-of-the-art structure-based methods in performance, achieving an Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.76 or higher. We present a comprehensive evaluation of the algorithm across multiple organisms and cellular components, demonstrating its effectiveness in predicting LLPS propensities at the single amino acid level and the impacts of mutations on LLPS. Our results are robustly supported by experimental validations, including immunofluorescence microscopy images from the Human Protein Atlas. catGRANULE 2.0 ROBOTs potential in protein design and mutation control can improve our understanding of proteins propensity to form subcellular compartments and help develop strategies to influence biological processes through LLPS. catGRANULE 2.0 ROBOT is freely available at https://tools.tartaglialab. com/catgranule2.

biochemistry↗

Machine learning methods applied to classify complex diseases using genomic data

Complex diseases pose challenges in disease prediction due to their multifactorial and polygenic nature. In this work, we explored the prediction of two complex diseases, multiple sclerosis (MS) and Alzheimers disease (AD), using machine learning (ML) methods and genomic data from UK Biobank. Different ML methods were applied, including logistic regressions (LR), gradient boosting decision trees (GB), extremely randomized trees (ET), random forest (RF), feedforward networks (FFN), and convolutional neural networks (CNN). The primary goal of this research was to investigate the variability of ML models in classifying complex diseases based on genomic risk. LR was the most robust method across folds and diseases, whereas deep learning methods (FFN and CNN) exhibited high variability. When comparing the performance of polygenic risk scores (PRS) with ML methods, PRS consistently performed at an average level. However, PRS still offers several practical advantages over ML methods. Despite implementing feature selection techniques to exclude non-informative and correlated predictors, the performance of ML models did not improve significantly, underscoring the ability of ML methods to achieve optimal performance even in the presence of correlated features due to linkage disequilibrium. Upon applying explainability tools to extract information about the genomic features contributing most to the classification task, the results confirmed the polygenicity of MS. The prevalence of HLA gene annotations among the top genomic features on chromosome 6 aligns with their significance in the context of MS. Overall, the highest-prioritized genomic variants were identified as expression or splicing quantitative trait loci (eQTL or sQTL) located in non-coding regions within or near genes associated with the immune response and MS. In summary, this research offers deeper insights into how ML models discern genomic patterns related to complex diseases.

bioinformatics↗