Search bioRxiv⌕ Search

Biology subjects

Bagherian, M.

Publications and source records attributed to Bagherian, M..

2 recordsLinked to original sources

Start right to end right: authentic open reading frame selection matters

Accurate annotation of open reading frames (ORFs) is fundamental for understanding gene function and post-transcriptional regulation. A critical but often overlooked aspect of transcriptome annotation is the selection of authentic translation start sites. Many genome annotation pipelines identify the longest possible ORF in alternatively spliced transcripts, using internal methionine codons as putative start sites. However, this computational approach ignores the biological reality that ribosomes select start codons based on sequence context, not ORF length. Here, we demonstrate that this practice leads to systematic misannotation of nonsense-mediated decay (NMD) targets in the Arabidopsis thaliana Araport11 reference transcriptome. Using TranSuite software to identify authentic start codons, we reanalyzed transcriptomic data from an NMD-deficient mutant and found that correct ORF annotation more than doubles the number of identifiable NMD targets with premature termination codons followed by downstream exon junctions, from 203 to 426 transcripts. Furthermore, we show that incorrect ORF annotations can lead to erroneous protein structure predictions, potentially introducing computational artifacts into protein databases. Our findings underscore the importance of biologically informed ORF annotation for accurate assessment of post-transcriptional regulation and proteome prediction, with implications for all eukaryotic genome annotation projects.

genomics↗

Data Fusion by Matrix Completion for Exposome Target Interaction Prediction

Human exposure to toxic chemicals presents a huge health burden and disease risk. Key to understanding chemical toxicity is knowledge of the molecular target(s) of the chemicals. Because a comprehensive safety assessment for all chemicals is infeasible due to limited resources, a robust computational method for discovering targets of environmental exposures is a promising direction for public health research. In this study, we implemented a novel matrix completion algorithm named coupled matrix-matrix completion (CMMC) for predicting exposome-target interactions, which exploits the vast amount of accumulated data regarding chemical exposures and their molecular targets. Our approach achieved an AUC of 0.89 on a benchmark dataset generated using data from the Comparative Toxicogenomics Database. Our case study with bisphenol A (BPA) and its analogues shows that CMMC can be used to accurately predict molecular targets of novel chemicals without any prior bioactivity knowledge. Overall, our results demonstrate the feasibility and promise of computational predicting environmental chemical-target interactions to efficiently prioritize chemicals for further study.

pharmacology and toxicology↗