Search bioRxivSearch

Biology subjects

Schmitt, A.

Publications and source records attributed to Schmitt, A..

3 recordsLinked to original sources

Morphologic intratumoral heterogeneity from routine whole-slide histopathology is prognostic for survival in primary central nervous system lymphoma: development in the LOC Network and international external validation

Background: Clinical scores incompletely capture outcomes in primary central nervous system lymphoma (PCNSL). We quantified morphologic heterogeneity in pretreatment hematoxylin and eosin (H\&E) whole slides. Patients and methods: Three independent cohorts of immunocompetent, HIV- and EBV-negative patients treated recently were analyzed: LOC 2023 (122 slides), phase III BLOCAGE-01 (245 slides; NCT02313389), and external Barcelona (BCN; 41 slides). UNI embeddings, prototype learning, spatial metrics, and elastic-net Cox regression defined ITH-C. Results: Models achieved bootstrap-corrected concordance of 0.797--0.834. Age-, sex-, and KPS-adjusted ITH-C HRs were 1.29 (95\% CI 1.01--1.64), 1.27 (1.07--1.51), and 2.13 (1.35--3.37), respectively. Adding ITH-C increased MSKCC C-index from 0.671 to 0.717, 0.560 to 0.593, and 0.588 to 0.706. Spatial transcriptomics linked ITH-C to immune programs. Conclusions: Routine H\&E encodes prognostic spatial heterogeneity in PCNSL. ITH-C complements clinical scores, supporting prospective risk stratification.

bioinformatics

Integration of phased Hi-C and molecular phenotype data to study genetic and epigenetic effects on chromatin looping

While genetic variation at chromatin loops is relevant for human disease, the relationships between loop strength, genetics, gene expression, and epigenetics are unclear. Here, we quantitatively interrogate this relationship using Hi-C and molecular phenotype data across cell types and haplotypes. We find that chromatin loops consistently form across multiple cell types and quantitatively vary in strength, instead of exclusively forming within only one cell type. We show that large haplotype loop imbalance is primarily associated with imprinting and copy number variation, rather than genetically driven traits such as allele-specific expression. Finally, across cell types and haplotypes, we show that subtle changes in chromatin loop strength are associated with large differences in other molecular phenotypes, with a 2-fold change in looping corresponding to a 100-fold change in gene expression. Our study suggests that regulatory genetic variation could mediate its effects on gene expression through subtle modification of chromatin loop strength.

genomics

Integrating Hi-C links with assembly graphs for chromosome-scale assembly

Long-read sequencing and novel long-range assays have revolutionized de novo genome assembly by automating the reconstruction of reference-quality genomes. In particular, Hi-C sequencing is becoming an economical method for generating chromosome-scale scaffolds. Despite its increasing popularity, there are limited open-source tools available. Errors, particularly inversions and fusions across chromosomes, remain higher than alternate scaffolding technologies. We present a novel open-source Hi-C scaffolder that does not require an a priori estimate of chromosome number and minimizes errors by scaffolding with the assistance of an assembly graph. We demonstrate higher accuracy than the state-of-the-art methods across a variety of Hi-C library preparations and input assembly sizes. The Python and C++ code for our method is openly available at https://github.com/machinegun/SALSA\n\nAuthor summaryHi-C technology was originally proposed to study the 3D organization of a genome. Recently, it has also been applied to assemble large eukaryotic genomes into chromosome-scale scaffolds. Despite this, there are few open source methods to generate these assemblies. Existing methods are also prone to small inversion errors due to noise in the Hi-C data. In this work, we address these challenges and develop a method, named SALSA2. SALSA2 uses sequence overlap information from an assembly graph to correct inversion errors and provide accurate chromosome-scale assemblies.

bioinformatics