Search bioRxivSearch

Biology subjects

Abdel-Wahab, O.

Publications and source records attributed to Abdel-Wahab, O..

3 recordsLinked to original sources

High throughput droplet single-cell Genotyping of Transcriptomes (GoT) reveals the cell identity dependency of the impact of somatic mutations

Defining the transcriptomic identity of clonally related malignant cells is challenging in the absence of cell surface markers that distinguish cancer clones from one another or from admixed non-neoplastic cells. While single-cell methods have been devised to capture both the transcriptome and genotype, these methods are not compatible with droplet-based single-cell transcriptomics, limiting their throughput. To overcome this limitation, we present single-cell Genotyping of Transcriptomes (GoT), which integrates cDNA genotyping with high-throughput droplet-based single-cell RNA-seq. We further demonstrate that multiplexed GoT can interrogate multiple genotypes for distinguishing subclonal transcriptomic identity. We apply GoT to 26,039 CD34+ cells across six patients with myeloid neoplasms, in which the complex process of hematopoiesis is corrupted by CALR-mutated stem and progenitor cells. We define high-resolution maps of malignant versus normal hematopoietic progenitors, and show that while mutant cells are comingled with wildtype cells throughout the hematopoietic progenitor landscape, their frequency increases with differentiation. We identify the unfolded protein response as a predominant outcome of CALR mutations, with significant cell identity dependency. Furthermore, we identify that CALR mutations lead to NF-{kappa}B pathway upregulation specifically in uncommitted early stem cells. Collectively, GoT provides high-throughput linkage of single-cell genotypes with transcriptomes and reveals that the transcriptional output of somatic mutations is heavily dependent on the native cell identity.

cancer biology

Impaired hematopoiesis and leukemia development in mice with a \"knock-in\" allele of U2af1(S34F)

Mutations affecting the spliceosomal protein U2AF1 are commonly found in myelodysplastic syndromes (MDS) and secondary acute myeloid leukemia (sAML). We have generated mice that carry Cre-dependent \"knock-in\" alleles of U2af1(S34F), the murine version of the most common mutant allele of U2AF1 encountered in human cancers. Cre-mediated recombination in murine hematopoietic lineages caused changes in RNA splicing, as well as multilineage cytopenia, macrocytic anemia, decreased hematopoietic stem and progenitor cells, low-grade dysplasias, and impaired transplantability, but without lifespan shortening or leukemia development. In an attempt to identify U2af1(S34F)-cooperating changes that promote leukemogenesis, we combined U2af1(S34F) with Runx1 deficiency in mice and further treated the mice with a mutagen, N-Ethyl-N-Nitrosourea (ENU). Overall, three of sixteen ENU-treated compound transgenic mice developed AML. However, AML did not arise in mice with other genotypes or without ENU treatment. Sequencing DNA from the three AMLs revealed somatic mutations homologous to those considered to be drivers of human AML, including predicted loss-or gain-of-function mutations in Tet2, Gata2, Idh1, and Ikzfl. However, the engineered U2af1(S34F) missense mutation reverted to wild type (WT) in two of the three AML cases, implying that U2af1(S34F) is dispensable, or even selected against, once leukemia is established.\n\nSIGNIFICANCE STATEMENTSomatic mutations in four splicing factor genes (U2AF1, SRSF2, SF3B1, and ZRSR2) are found in MDS and MDS-related AML, blood cancers with few effective treatment options. However, the pathophysiological effects of these mutations remain poorly characterized, in part due to the paucity of disease-relevant models. Here, we report the establishment of mouse models to study the most common U2AF1 mutation, U2af1(S34F). Production of the mutant protein specifically in the murine hematopoietic compartment disrupts hematopoiesis in ways resembling human MDS. We further identified deletion of the Runx1 gene and other known oncogenic mutations as changes that might collaborate with U2af1(S34F) to give rise to frank AML in mice.

cancer biology

ProteomeGenerator: A framework for comprehensive proteomics based on de novo transcriptome assembly and high-accuracy peptide mass spectral matching

Modern mass spectrometry now permits genome-scale and quantitative measurements of biological proteomes. However, analyses of specific specimens are currently hindered by the incomplete representation of biological variability of protein sequences in canonical reference proteomes, and the technical demands for their construction. Here, we report ProteomeGenerator, a framework for de novo and reference-assisted proteogenomic database construction and analysis based on sample-specific transcriptome sequencing and high-resolution and high-accuracy mass spectrometry proteomics. This enables assembly of proteomes encoded by actively transcribed genes, including sample-specific protein isoforms resulting from non-canonical mRNA transcription, splicing, or editing. To improve the accuracy of protein isoform identification in non-canonical proteomes, ProteomeGenerator relies on statistical target-decoy database matching augmented with spectral-match calibrated sample-specific controls. We applied this method for the proteogenomic discovery of splicing factor SRSF2-mutant leukemia cells, demonstrating high-confidence identification of non-canonical protein isoforms arising from alternative transcriptional start sites, intron retention, and cryptic exon splicing, as well as improved accuracy of genome-scale proteome discovery. Additionally, we report proteogenomic performance metrics for the current state-of-the-art implementations of SEQUEST HT, Proteome Discoverer, MaxQuant, Byonic, and PEAKS mass spectral analysis algorithms. Finally, ProteomeGenerator is implemented as a Snakemake workflow, enabling open, scalable, and facile discovery of sample-specific, non-canonical and neomorphic biological proteomes (https://github.com/jtpoirier/proteomegenerator).

bioinformatics