Search bioRxiv⌕ Search

Biology subjects

Ge, H. G.

Publications and source records attributed to Ge, H. G..

6 recordsLinked to original sources

A Human Genetics Framework for De-risking Gene Editing Targets for Hematopoietic Cell and Gene Therapy

Developing novel therapeutics requires robust early-stage target de-risking to ensure safety and efficacy. We developed a scalable proteogenomic framework integrating population-scale human genetics and plasma proteomics to identify genes tolerant of inactivation (i.e., dispensable) within hematopoietic compartments, thereby enabling safer targeted immunotherapies. Using CD33 as a validated benchmark, we observed that naturally occurring loss-of-function (LoF) variants lead to concordant RNA and protein depletion, supporting functional gene inactivation. Early clinical results from the Trem-Cel trial (NCT05945849) further provide proof of concept that deletion of dispensable antigens can enable safe and effective immunotherapy in humans. We extended this approach genome-wide in the UK Biobank and identified 237 candidate dispensable genes, filtered by plasma proteomic data and hematopoietic expression, highlighting LY75 (CD205) as a novel candidate with strong proteogenomic evidence of LoF tolerance. This work establishes a generalizable, quantitative proteogenomic framework for systematic prioritization of dispensable gene targets for editing, providing a foundation for next-generation cell and gene therapies that minimize on-target, off-tumor toxicities.

genetics↗

Cutting Through the Artifacts: Dissecting gRNA Impurities with FUSS-seq

CRISPR-based therapeutics rely on guide RNAs (gRNAs) and the Cas9 endonuclease for precise gene editing. Ensuring gRNA purity and base-level sequence integrity is essential for clinical translation. While industry-standard practice relies on liquid chromatography-high-resolution mass spectrometry to assess oligonucleotide identity and purity, more recent FDA guidance recommends complementary base-by-base sequence analysis (FDA CBER Webinar, 2024). In this study, we evaluated next-generation sequencing (NGS) strategies for characterizing chemically synthesized gRNAs. We found that the widely used SMARTer assay, while capable of producing sequenceable libraries, introduced substantial artifacts during library preparation. These included truncated scaffold species at oligo(A) stretches in the scaffold region and 5'(n-1) deletions within the spacer sequence. Although absent in the original gRNA, these artifacts accounted for over 10% of the sequencing reads, creating the false appearance of impurities. Through experimental and computational approaches, we traced these artifacts to mispriming by template-switching oligonucleotides (TSOs). Importantly, these artifacts occur during sequencing, and although they do not reflect real gRNA impurities, they compromise assay accuracy and can obscure true sequence impurities. To overcome these limitations, we developed FUSS-seq (Full-length Uncoupled Second-strand Synthesis followed by sequencing), a novel assay that integrates principles from 5' RACE with a modified TSO bearing a 3' polymerase-blocking moiety. FUSS-seq markedly reduced artifacts and increased full-length gRNA recovery, providing a more accurate and lower-bias method for gRNA purity assessment. This approach supports improved Chemistry Manufacturing and Controls (CMC) characterization of gRNAs and strengthens the analytical toolkit needed for reliable CRISPR-based therapeutic development.

molecular biology↗

Integrating Human Genetics and Protective Genome Editing to Enable ADGRE2-Directed AML Therapy

Acute myeloid leukemia (AML) remains a major therapeutic challenge due to extensive disease heterogeneity and lack of cancer-specific antigens. ADGRE2 has emerged as a promising AML target with broad expression in AML patient blast and leukemic stem cell-enriched populations. However, comparable expression in healthy hematopoietic stem and progenitor cells (HSPCs) and myeloid lineages suggests a high susceptibility to on-target, off-tumor myelotoxicity with ADGRE2-targeted therapies. Guided by human genetics data identifying loss-of-function variants, we evaluated whether ADGRE2 is dispensable in hematopoietic stem cells as a protective approach for transplant-based shielding from ADGRE2-directed therapies. Using CRISPR-Cas9 and adenine base editors, we achieved high-efficiency ADGRE2 knockout (>94%) in HSPCs with corresponding protein loss without impairing cell viability, differentiation, and cytokine release in vitro, or long-term engraftment, multilineage differentiation, and persistence of gene editing in mouse xenografts. We also developed novel ADGRE2-specific chimeric antigen receptor (CAR) T cells that demonstrated potent cytotoxicity against AML cells, even at low antigen levels. Together, these findings establish ADGRE2 as a compelling AML target and provide a framework for hematopoietic stem cell transplant with protective gene editing to enable ADGRE2-directed immunotherapies while minimizing myelotoxicity.

cancer biology↗

A resource and computational approach for quantifying gene editing allelism at single-cell resolution

CRISPR-Cas9-based gene editing is a powerful approach to developing gene and cell therapies for several diseases. Engineering cell therapies requires accurate assessment of gene editing allelism because editing patterns can vary across cells leading to genotypic heterogeneity. This can hinder development of robust cell therapies. Droplet-based targeted single-cell DNA sequencing (scDNAseq) has been used to genotype targeted loci across thousands of cells enabling high-throughput assessment of gene editing efficiency. Here, we constructed a "ground truth" gene editing single-cell DNAseq atlas, along with an artifact-aware computational workflow called GUMM (Genotyping Using Mixture Models) to systematically infer single-cell allelism from these data. This resource was created by expanding CRISPR-Cas9-edited HL-60 clones that harbored distinct insertion-deletion (indel) profiles in CLEC12A and mixing them at pre-defined ratios to create artificial cocktails that mimic the potential editing diversity of a CRISPR-Cas9 experiment. This enabled assessment of technical artifacts that confound interpretation of allelism in the readouts of gene edited cells. GUMM was able to accurately genotype cells and infer the original clonal composition of the artificial cocktails even in the presence of artifacts.

bioinformatics↗

SANTON: Sequencing Analysis Toolkits for Off-target Nomination

BackgroundGenome-wide off-target nomination and screening sequencing methods, such as synthetic oligonucleotide-based sequencing and whole genome-based sequencing, evaluate the precision and safety of gene-editing technologies. However, there remains a lack of comprehensive bioinformatics tools for analyzing the sequencing data generated by off-target nomination assays for various gene editors, including Cas9, Cas12a, and base editors. ResultsWe introduce a Sequencing Analysis Toolkits for Off-target Nomination (SANTON) for the identification and quantification of potential off-target sites using datasets generated from synthetic oligonucleotide-based sequencing and whole genome-based sequencing (e.g., Digenome-seq) methods. By applying SANTON to a Cas9 treated oligo synthesization-based sequencing data, a comprehensive set of potential off-target sites are evaluated in which the top off-target sites displayed highly constituency with a published study using in vivo method (e.g, GUIDE-seq). Utilizing SANTON to previously published Digenome-seq datasets, we identified more potential off-target regions than previous studies, including ones showing significantly higher cleavage level than the on-target site. To the best of our knowledge, SANTON is the first public package capable of analyzing synthetic oligonucleotide-based sequencing data and the first to specific optimized for the analysis of whole genome-based sequencing analysis for Cas12a and base editors. We demonstrated the capability of SANTON to effectively analyze complex cleavage patterns by a diverse editing system. ConclusionsSANTON provides a powerful computational solution to support genome editing research, facilitating more reliable and comprehensive off-target profiling.

bioinformatics↗

scGenAI: A generative AI platform with biological context embedding of multimodal features enhances single cell state classification

SummarySingle-cell sequencing has advanced the understanding of cellular heterogeneity, yet traditional cell type annotation tools struggle with increasingly complex datasets and multimodal integration. Recent large language models (LLMs) offer improved accuracy but rely on pre-trained models with limited gene vocabularies and lack of biological contextualization. These constraints make it challenging to effectively fine-tune pre-trained models, especially for novel datasets, non-human species, or disease-specific studies where unique gene expression patterns are critical. To address these constraints, we developed scGenAI, an LLM-based tool that supports straightforward and flexible de novo training, allowing researchers to incorporate genomic and biofunctional contexts to improve accuracy and interpretability on the prediction of cell states. Here, we demonstrate that scGenAI outperforms conventional models in tasks such as cell type prediction and acute myeloid leukemia (AML) malignant cell states identification, making it a powerful tool for single-cell analysis workflows. Availability and implementationscGenAI (DOI: 10.5281/zenodo.14927611) is distributed as an open-source Python package, with source code and comprehensive documentation accessible on GitHub at https://github.com/VOR-Quantitative-Biology/scGenAI.

bioinformatics↗