Search bioRxiv⌕ Search

Biology subjects

Anglada-Girotto, M.

Publications and source records attributed to Anglada-Girotto, M..

2 recordsLinked to original sources

Evolutionary conservation of A/T-ending codons reflects co-regulation of expression and complex formation.

BackgroundIn a wide variety of organisms, synonymous codons are used with different frequencies, a phenomenon known as codon bias that plays an important role in determining expression levels. However, the importance of codon bias to facilitate the simultaneous turnover of thousands of protein-coding transcripts to bring about phenotypic changes in cellular programs such as development, has not yet been investigated in detail. ResultsHere, we discover that genes with A/T-ending codon preferences are expressed coordinately and display a high codon conservation in mammals. This feature is not observed in genes enriched in G/C-ending codons. A paradigmatic case of this phenomenon is KRAS, from the RAS family, an A/T-rich gene with a high codon conservation (95%) in comparison to HRAS (76%). Also, we find that genes with similar codon composition are more likely to be part of the same protein complex, and that genes with A/T-ending codons are more prone to form protein complexes than those rich in G/C. The codon preferences of genes with A/T-ending codons are conserved among vertebrates. We propose that codon conservation, a feature of expression-coordinated transcripts, is linked to the high expression variation and coordination of tRNA isoacceptors reading A/T-ending codons. ConclusionsOur data indicate that cells exploit A/T-ending codons to generate coordinated, fine-tuned changes of protein-coding transcripts. We suggest that this orchestration contributes to tissue-specific and ontogenetic-specific expression, which can facilitate, for instance, timely protein complex formation.

evolutionary biology↗

robustica: customizable robust independent component analysis

BackgroundIndependent Component Analysis (ICA) allows the dissection of omic datasets into modules that help to interpret global molecular signatures. The inherent randomness of this algorithm can be overcome by clustering many iterations of ICA together to obtain robust components. Existing algorithms for robust ICA are dependent on the choice of clustering method and on computing a potentially biased and large Pearson distance matrix. ResultsWe present robustica, a Python-based package to compute robust independent components with a fully customizable clustering algorithm and distance metric. Here, we exploited its customizability to revisit and optimize robust ICA systematically. From the 6 popular clustering algorithms considered, DBSCAN performed the best at clustering independent components across ICA iterations. After confirming the bias introduced with Pearson distances, we created a subroutine that infers and corrects the components signs across ICA iterations to enable using Euclidean distance. Our subroutine effectively corrected the bias while simultaneously increasing the precision, robustness, and memory efficiency of the algorithm. Finally, we show the applicability of robustica by dissecting over 500 tumor samples from low-grade glioma (LGG) patients, where we define a new gene expression module with the key modulators of tumor aggressiveness downregulated upon IDH1 mutation. Conclusionrobustica brings precise, efficient, and customizable robust ICA into the Python toolbox. Through its customizability, we explored how different clustering algorithms and distance metrics can further optimize robust ICA. Then, we showcased how robustica can be used to discover gene modules associated with combinations of features of biological interest. Taken together, given the broad applicability of ICA for omic data analysis, we envision robustica will facilitate the seamless computation and integration of robust independent components in large pipelines. Contactmiquel.anglada@crg.eu

systems biology↗