Search bioRxiv⌕ Search

Biology subjects

Sonder, E.

Publications and source records attributed to Sonder, E..

3 recordsLinked to original sources

On the identification of differentially-active transcription factors from ATAC-seq data

ATAC-seq has emerged as a rich epigenome profiling technique, and is commonly used to identify Transcription Factors (TFs) underlying given phenomena. A number of methods can be used to identify differentially-active TFs through the accessibility of their DNA-binding motif, however little is known on the best approaches for doing so. Here we benchmark several such methods using a combination of curated datasets with various forms of short-term perturbations on known TFs, as well as semi-simulations. We include both methods specifically designed for this type of data as well as some that can be repurposed for it. We also investigate variations to these methods, and identify three particularly promising approaches (a chromVAR-limma workflow with critical adjustments, monaLisa and a combination of GC smooth quantile normalization and multivariate modeling). We further investigate the specific use of nucleosome-free fragments, the combination of top methods, and the impact of technical variation. Finally, we illustrate the use of the top methods on a novel dataset to characterize the impact on DNA accessibility of TRAnscription Factor TArgeting Chimeras (TRAFTAC), which can deplete TFs - in our case NFkB - at the protein level. Author summaryTranscription factors regulate gene expression by binding sites in the genome that often harbor a specific DNA motif. The collective accessibility of these motif-matching regions, measured by technologies such as ATAC-seq, can be used to infer the activity of the corresponding transcription factors. Here we use curated datasets of 11 TF-specific perturbations as well as 116 semi-simulated datasets to benchmark various methods for identifying factors that differ in activity between experimental conditions. We investigate important variations in the analysis and make recommendations pertaining to such analysis. Finally, we illustrate the application of the top methods to characterize the effects of a novel method for perturbing transcription factors at the protein level.

bioinformatics↗

Validation of hypermethylated DNA regions found in colorectal cancers as potential aging-independent biomarkers of precancerous colorectal lesions

BackgroundWe previously identified 16,772 colorectal cancer-associated hypermethylated DNA regions that were also detectable in precancerous colorectal lesions (preCRCs) and unrelated to normal mucosal aging. We have now conducted a study to validate 990 of these differentially methylated DNA regions (DMR) in a new series of preCRCs. MethodsWe used targeted bisulfite sequencing to validate these 990 potential biomarkers in 59 preCRC tissue samples (41 conventional adenomas, 18 sessile serrated lesions), each with a patient-matched normal mucosal sample. Based on differential DNA methylation tests, a panel of (candidate) DMRs was chosen on a subset of the (our) cohort and validated on the remaining part of our cohort and (two) further publicly available datasets with respect to their stratifying potential between preCRCs and normal mucosa. ResultsStrong statistical significance for the difference in methylation levels was observed across the full set of 990 investigated DMRs. From these, a selected candidate panel of 30 DMRs correctly identified 58/59 tumors (area under the receiver operating curve: 0.998). ConclusionsThese validated DNA hypermethylation markers can be exploited to develop more accurate noninvasive colorectal tumor screening assays.

genomics↗

Meta-analysis of (single-cell method) benchmarks reveals the need for extensibility and interoperability

Computational methods represent the lifeblood of modern molecular biology. Benchmarking is important for all methods, but with a focus here on computational methods, benchmarking is critical to dissect important steps of analysis pipelines, formally assess performance across common situations as well as edge cases, and ultimately guide users on what tools to use. Benchmarking can also be important for community building and advancing methods in a principled way. We conducted a meta-analysis of recent single-cell benchmarks to summarize the scope, extensibility, neutrality, as well as technical features and whether best practices in open data and reproducible research were followed. The results highlight that while benchmarks often make code available and are in principle reproducible, they remain difficult to extend, for example, as new methods and new ways to assess methods emerge. In addition, embracing containerization and workflow systems would enhance reusability of intermediate benchmarking results, thus also driving wider adoption.

bioinformatics↗