Search bioRxiv⌕ Search

Biology subjects

Diekema, M. H.

Publications and source records attributed to Diekema, M. H..

2 recordsLinked to original sources

miDGD: a multi-modal deep generative model predicts miRNA expression from bulk or single-cell mRNA expression

MicroRNAs (miRNAs) are key post-transcriptional regulators, yet standard bulk and single-cell RNA-seq do not capture them, leaving this regulatory layer invisible in most transcriptomic data. We present miDGD, a deep generative decoder that jointly models paired mRNA and miRNA profiles through a shared latent representation, enabling miRNA expression to be predicted from mRNA alone. Trained on tumors (TCGA), healthy tissues (GTEx), and human cell lines, miDGD recovers hundreds of miRNAs in held-out tumors (mean Spearman {rho} = 0.56), captures both tissue-specific and ubiquitous miRNAs, and preserves known miRNA--target repression and host-gene co-expression. Without label supervision, its latent space separates 32 cancer types (80% accuracy). Predictions remain stable at single-cell-like sparsity and transfer across datasets--from tumors to healthy tissues and from bulk to single cells--where miDGD outperforms existing supervised and activity-inference methods. miDGD thus unlocks miRNA regulation in the vast body of existing mRNA-only data, including single-cell datasets.

bioinformatics↗

SNV and indel error modeling of deep targeted cell-free DNA sequencing data for sensitive detection of circulating tumor DNA in colorectal cancer

Circulating tumor DNA (ctDNA) is a promising biomarker for cancer detection, but low tumor burden makes it difficult to distinguish true signal from background noise. To aggregate and better evaluate weak mutational signals, we propose PyDREAMS, which incorporates both single-nucleotide variants (SNVs) and insertion and deletions (indels) for ctDNA detection and quantification. To distinguish signal from noise, a neural network background error model is learned from healthy controls. It captures the joint effects of cell-free DNA (cfDNA)-specific lesions and sequencing errors, accounting for both genomic context and read-level features. Finally, a statistical test is used to evaluate the presence of mutational signals. We evaluate the method in a tumor-informed setting, using cohorts of colorectal cancer samples with deep targeted plasma cfDNA sequencing across 12 cancer driver genes. We trained PyDREAMS on 46 healthy controls, with feature analysis revealing that both SNV and indel error rates were lowest at mononucleosomal fragment lengths, suggesting that nucleosomes protect cfDNA and reduce lesion accumulation during circulation and sample handling. In the validation cohort, combining SNVs with indels improved detection, with indels contributing approximately 1.5-fold more evidence per mutation than SNVs. On a test cohort of 209 stage I-III colorectal cancer (CRC) patients and 24 healthy controls, PyDREAMS outperformed a Shearwater-based caller (AUC 0.917 vs 0.909). In stage III post-operative (post-OP) samples (n = 26), where ctDNA was expected only in non-cured patients, PyDREAMS detected ctDNA in 5 patients, including 3/9 with later recurrence, while Shearwater detected none. Together, these results show that PyDREAMS improves evaluation of ultra-low-frequency tumor signals through unified read-level modelling of SNV and indel background error.

bioinformatics↗