Search bioRxiv⌕ Search

Biology subjects

Passemiers, A.

Publications and source records attributed to Passemiers, A..

4 recordsLinked to original sources

When batch correction corrupts gene expression: uncovering distortions in correlation structures

Batch correction is essential for integrating datasets and enabling population-level insights into health and disease. Embedding-based approaches are among the most widely used solutions, but here we highlight a critical, overlooked limitation: these methods can distort feature-to-feature (e.g., gene-gene) relationships, potentially undermining downstream analyses. We investigate this issue and introduce a novel metric to quantify it.

bioinformatics↗

geneRNIB: a living benchmark for gene regulatory network inference

Gene regulatory networks (GRNs) underpin cellular identity and function, playing a key role in health and disease. GRN inference has received substantial attention, motivating systematic benchmarking. Despite various benchmarking efforts, existing studies remain limited in the number of methods, datasets, and metrics, fail to capture the context-specific nature of regulatory interactions across biological conditions, and are constrained by the absence of a reliable ground truth. Here, we introduce geneRNIB, a comprehensive GRN inference benchmarking framework built on three key principles: continuous integration, context-specific evaluation, and holistic assessment in the absence of a true reference network. geneRNIB enables the seamless incorporation of new algorithms, datasets, and evaluation metrics to reflect ongoing developments. In the current version, we systematically integrated and assessed 12 GRN inference methods, spanning single- and multiomics approaches across 11 datasets including thousands of perturbation scenarios. We introduced complementary metrics specifically designed to assess context-specific inference. Our findings indicate that simple models with fewer assumptions often outperform more complex pipelines across several perturbation-informed and predictive metrics. Notably, gene expression-based algorithms yielded better results than more advanced multimodal approaches. In addition, we identify several potential factors that influence the performance of GRN inference and offer actionable guidelines for the future development of the method. By addressing these critical limitations in existing benchmarks, geneRNIB advances GRN inference research and fosters progress toward personalized medicine.

bioinformatics↗

Critical issues found in "Dissecting cell identity via network inference and in silico gene perturbation"

1In the 2023 Nature publication "Dissecting cell identity via network inference and in silico gene perturbation" [1], the authors introduced CellOracle (CO), a novel method leveraging mRNA-seq and ATAC-seq data to construct gene regulatory networks (GRNs), which are subsequently used for gene perturbation. They designed CO to account for the role of distal cis-regulatory elements, e.g. enhancers, as well as proximal promoters in the gene regulation system. For this purpose, they employed Cicero to determine the co-accessibility scores between peaks, provided by ATAC-seq data. These scores are then used to identify the interaction of distal regions with the target gene. Using CO, they have conducted multiple perturbation studies on different organisms and identified novel phenotypes resulting from transcriptional factor (TF) perturbation. In addition, they benchmarked COs performance using ChIP-seq data as ground truth against other state-of-the-art GRN methods across multiple mouse tissue samples. However, our evaluation reveals critical limitations in the implementation of their methodology, both in terms of ATAC-seq data integration as well as benchmarking. In this report, we first explain the limitations in their approach of integrating ATAC-seq data. We show that the proposed algorithm fails to account for distal regulatory interactions. After, we present the issues associated with their benchmarking algorithm and the data used for benchmarking. We show that their findings regarding the comparative performance of CO against other GRN inference methods is invalid and requires further evaluation. In conclusion, we detect multiple inaccuracies in this paper which undermine the validity of their published protocol and the results. The materials supporting our findings are accessible on GitHub1.

bioinformatics↗

Alleviating cell-free DNA sequencing biases with optimal transport

Cell-free DNA (cfDNA) is a rich source of biomarkers for various (patho)physiological conditions. Recent developments have used Machine Learning on large cfDNA data sets to enhance the detection of cancers and immunological diseases. Preanalytical variables, such as the library preparation protocol or sequencing platform, are major confounders that influence such data sets and lead to domain shifts (i.e., shifts in data distribution as those confounders vary across time or space). Here, we present a domain adaptation method that builds on the concept of optimal transport, and explicitly corrects for the effect of such preanalytical variables. Our approach can be used to merge cohorts representative of the same population but separated by technical biases. Moreover, we also demonstrate that it improves cancer detection via Machine Learning by alleviating the sources of variation that are not of biological origin. Our method also improves over the widely used GC-content bias correction, both in terms of bias removal and cancer signal isolation. These results open perspectives for the downstream analysis of larger data sets through the integration of cohorts produced by different sequencing pipelines or collected in different centers. Notably, the approach is rather general with the potential for application to many other genomic data analysis problems.

bioinformatics↗