Search bioRxiv⌕ Search

Biology subjects

Wall, B. P. G.

Publications and source records attributed to Wall, B. P. G..

2 recordsLinked to original sources

Beyond Blacklists: A Critical Assessment of Exclusion Set Generation Strategies and Alternative Approaches

Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Tools like the widely used Blacklist software facilitate this process, but their algorithmic details and parameter choices are not always clearly documented, affecting reproducibility and biological relevance. We examined the Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to variability in input data, aligner choice, and read length. We also identified and addressed a coding issue that led to over-annotation of high-signal regions. We further explored the use of "sponge" sequences--unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA--as an alternative approach. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets and suggest that sponge sequences offer a flexible, alignment-guided strategy for reducing artifacts and improving functional genomics analyses.

bioinformatics↗

scHiCcompare: an R package for differential analysis of single-cell Hi-C data

Changes in the three-dimensional (3D) structure of the human genome are key indicators of cancer and developmental disorders. Techniques like chromatin conformation capture (Hi-C) have been developed to study these global 3D structures, typically requiring millions of cells and an extremely high sequencing depth (around 1 billion reads per sample) for bulk Hi-C. In contrast, single-cell Hi-C (scHi-C) captures 3D structures at the individual cell level but faces significant data sparsity, marked by a high proportion of zeros. While differential analysis methods exist for bulk Hi-C data, they are limited for scHi-C data. To address this, we developed a method for differential scHi-C analysis, building on existing techniques in the HiCcompare R package. Our approach imputes sparse scHi-C data by considering genomic distances and creates pseudo-bulk Hi-C matrices by summing condition-specific data. The data are normalized using LOESS regression, and differential chromatin interactions are detected via Gaussian Mixture Model (GMM) clustering. Our workflow outperforms existing methods in identifying differential chromatin interactions across various genomic distances, fold changes, resolutions, and sample sizes in both simulated and experimental contexts. This allows for effective detection of cell type-specific differences in chromatin structure, which has meaningful associations with biological and epigenetic features. Our method is implemented in the scHiCcompare R package, available at https://github.com/dozmorovlab/scHiCcompare. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=74 SRC="FIGDIR/small/622369v1_ufig1.gif" ALT="Figure 1"> View larger version (32K): org.highwire.dtl.DTLVardef@b94f54org.highwire.dtl.DTLVardef@72fddorg.highwire.dtl.DTLVardef@1d77702org.highwire.dtl.DTLVardef@c63088_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗