Search bioRxiv⌕ Search

Biology subjects

Boyle, I. A.

Publications and source records attributed to Boyle, I. A..

2 recordsLinked to original sources

Dissecting context-dependent cancer vulnerabilities using Perturb-seq

Background CRISPR-mediated viability assays in diverse cancer cell lines have informed cancer biology and precision medicine, but cell fitness is not the only cancer-relevant phenotype. Gene expression profiling provides insight into cellular stress, inflammation, and differential state, while still identifying activation of cell-death pathways. Perturb-seq allows scalable functional genomics screening of expression phenotypes at single-cell resolution, however existing datasets cover only a small number of work-horse cell lines. Results We produced a proof-of-concept Perturb-seq dataset targeting 100 genes in 16 diverse cancer cell lines. In the process, we established methods to address single-cell technical artifacts, identified Cas9-mediated chromosomal aberrations and assessed screen quality. Even with a limited library, we observed common signatures of deleting essential genes as well as context-specific responses based on intrinsic genomic properties of the models. For example, we inferred a previously undescribed relationship between dependence on the ER-golgi transport gene immediate early response 3 interacting protein 1 (IER3IP1) and oxidative stress, demonstrating the potential of integrated Perturb-seq for hypothesis generation. Conclusions We established a framework for building a comprehensive map of post-perturbational transcriptional phenotypes using parallel Perturb-seq experiments across multiple cell lines. We demonstrated that integrated Perturb-seq experiments spanning diverse contexts enable hypotheses about gene function specific to tissue types or cancer subtypes - suggesting large-scale, genome-wide datasets would offer invaluable insight into the highly context-dependent nature of cancer biology.

cancer biology↗

Natural language processing of gene descriptions for overrepresentation analysis with GeneTEA

Overrepresentation analysis (ORA) is used to identify the biological relationships in a list of genes by testing gene sets for enrichment in the query. However, the inconsistent definition and highly overlapping nature of gene set databases can make interpreting ORA results difficult. Here, we introduce GeneTEA, a model that takes in free-text gene descriptions and incorporates several natural language processing methods to learn a sparse gene-by-term embedding, which can be treated as a de novo gene set database. We benchmark performance against other popular ORA tools and find that only GeneTEA properly controls false discovery while consistently surfacing relevant biology. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=75 SRC="FIGDIR/small/646026v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@145a76aorg.highwire.dtl.DTLVardef@1f221d3org.highwire.dtl.DTLVardef@18acdacorg.highwire.dtl.DTLVardef@1c5106a_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗