Search bioRxiv⌕ Search

Biology subjects

Cleynen, A.

Publications and source records attributed to Cleynen, A..

3 recordsLinked to original sources

FracFixR: A compositional statistical framework for absolute proportion estimation between fractions in RNA sequencing data

MotivationRNA fractionation followed by sequencing is widely used to study RNA localization, translation, and subcellular compartmentalization. Interpreting fractionated RNA-seq data poses a fundamental compositional challenge: library preparation and sequencing depth obscure the original proportions of RNA fractions, which can bias comparisons - particularly when biological changes shift RNA distribution across fractions. This bias compromises comparisons of fraction-specific RNA profiles and limits the utility of standard differential expression methods. Existing approaches using transcript frequency ratios or standard normalization fail to account for the compositional nature of fractionated samples and cannot estimate the unrecoverable "lost" fraction. ResultsWe developed FracFixR, a statistical framework that reconstructs original fraction proportions by modeling the compositional relationship between whole and fractionated RNA samples. Using non-negative linear regression on carefully selected transcripts, FracFixR estimates global fraction weights, corrects individual transcript frequencies, and quantifies unrecoverable material. The framework includes methods for differential proportion testing between conditions using binomial GLM, logit, or beta-binomial models. We rigorously validated FracFixR using synthetic data with known ground truth and real polysome profiling data from multiple cell lines, demonstrating accurate reconstruction of fraction weights (Pearson correlation > 0.85) and enabling detection of differentially translated transcripts between cancer subtypes. Availability and implementationFracFixR is implemented as an R package freely available on GitHub at https://github.com/Arnaroo/FracFixR.

bioinformatics↗

High-Accuracy RNA Integrity Definition for Unbiased Transcriptome Comparisons with INDEGRA

RNA sample integrity variability introduces biases and obscures natural RNA degradation, posing a significant challenge in transcriptomics. To address this, we developed the Direct Transcriptome Integrity (DTI) measure, a universal and robust RNA integrity metric based on nanopore sequencing. By accurately modeling RNA fragmentation, DTI provides a reliable assessment of sample quality. Integrated into the INDEGRA package (freely available at https://github.com/Arnaroo/INDEGRA), we provide tools to correct false discoveries and enable precise differential expression and RNA degradation analyses, even for challenging sample types. INDEGRA software can be used to accurately measure RNA DTI stability metric, isolate biological component of RNA degradation from technical biases, compare biological RNA stability transcriptome-wide and suppress false degradation-induced differential gene expression hits to allow broad comparisons across samples of different quality DTI offers a straightforward and accurate method for assessing RNA degradation, characterizing both overall sample integrity and transcript-specific degradation rates using direct RNA sequencing (DRS) data. Calculated through INDEGRA, DTI reveals inter- and intra-transcript variability in degradation, while INDEGRA separates RNA degradation from mapping inaccuracies, and connects degradation profiles to RNA fragmentation rates. By leveraging INDEGRA, researchers can minimize false differential transcript abundance findings caused by variations in overall sample integrity, while preserving genuine transcript-specific differences in stability and degradation. INDEGRA supports integration with widely used differential transcript abundance tools like DESeq2, limma-voom, and edgeR, enabling seamless analysis pipelines. INDEGRA enhances the accuracy and reliability of RNA quantification in high-throughput data and simplifies comparisons across diverse transcriptomic datasets, including those derived from different tissues, species, or experimental protocols.

bioinformatics↗

Unsupervized identification of prognostic copy-numberalterations using segmentation and lasso regularization

Identifying copy-number alteration with prognostic impact is typically done in a supervised approach, were candidate regions are user-selected (chomosome arms, oncogenes, etc). Yet CNA events may range from whole chromosome alterations to small focal amplifications or deletions, with no available approach to combine the potential prognostic impact of different aberration ranges. We propose and compare different statistical models to integrate the effects of multi-scale CNA events by exploiting the longitudinal structure of the genome, and assume that the survival distribution follows a Cox-proportional hazard model. These methods are adaptable to any cohorts screened for CNA by genome-wide assays such as CGH-array or whole-genome sequencing technologies, and with sufficient follow-up time. We show that combining a segmentation in the survival odds strategy with a lasso-regularization selection approach provides the best results in terms of recovering the true significant CNA regions as well as predicting survival outcomes. In particular, as shown on a 551 Multiple Myeloma patient cohort, this method allows to refine previously identified regions to exhibit potential novel driver genes.

bioinformatics↗