Search bioRxivSearch

Biology subjects

Pinello, L.

Publications and source records attributed to Pinello, L..

12 recordsLinked to original sources

Analysis and comparison of genome editing using CRISPResso2

Genome editing technologies are rapidly evolving, and analysis of deep sequencing data from target or off-target regions is necessary for measuring editing efficiency and evaluating safety. However, no software exists to analyze base editors, perform allele-specific quantification or that incorporates biologically-informed and scalable alignment approaches. Here, we present CRISPResso2 to fill this gap and illustrate its functionality by experimentally measuring and analyzing the editing properties of six genome editing agents.

bioinformatics

CRISPR-SURF: Discovering regulatory elements by deconvolution of CRISPR tiling screen data

Tiling screens using CRISPR-Cas technologies provide a powerful approach to map regulatory elements to phenotypes of interest, but computational methods that effectively model these experimental approaches for different CRISPR technologies are not readily available. Here we present CRISPR-SURF, a deconvolution framework to identify functional regulatory regions in the genome from data generated by CRISPR-Cas nuclease, CRISPR interference (CRISPRi), or CRISPR activation (CRISPRa) tiling screens. We validated CRISPR-SURF on previously published and new data, identifying both experimentally validated and new potential regulatory elements. With CRISPR tiling screens now being increasingly used to elucidate the regulatory architecture of the non-coding genome, CRISPRSURF provides a generalizable and accessible solution for the discovery of regulatory elements.

bioinformatics

CRISPRO Identifies Functional Protein Coding Sequences Based on Genome Editing Dense Mutagenesis

CRISPR/Cas9 pooled screening permits parallel evaluation of comprehensive guide RNA libraries to systematically perturb protein coding sequences in situ and correlate with functional readouts. For the analysis and visualization of the resulting datasets we have developed CRISPRO, a computational pipeline that maps functional scores associated with guide RNAs to genome, transcript, and protein coordinates and structure. No available tool has similar functionality. The ensuing genotype-phenotype linear and 3D maps raise hypotheses about structure-function relationships at discrete protein regions. Machine learning based on CRISPRO features improves prediction of guide RNA efficacy. The CRISPRO tool is freely available at gitlab.com/bauerlab/crispro.

bioinformatics

STREAM: Single-cell Trajectories Reconstruction, Exploration And Mapping of omics data

Single-cell transcriptomic assays have enabled the de novo reconstruction of lineage differentiation trajectories, along with the characterization of cellular heterogeneity and state transitions. Several methods have been developed for reconstructing developmental trajectories from single-cell transcriptomic data, but efforts on analyzing single-cell epigenomic data and on trajectory visualization remain limited. Here we present STREAM, an interactive pipeline capable of disentangling and visualizing complex branching trajectories from both single-cell transcriptomic and epigenomic data.

genomics

AmpUMI: Design and analysis of unique molecular identifiers for deep amplicon sequencing

MotivationUnique molecular identifiers (UMIs) are added to DNA fragments before PCR amplification to discriminate between alleles arising from the same genomic locus and sequencing reads produced by PCR amplification. While computational methods have been developed to take into account UMI information in genome-wide and single-cell sequencing studies, they are not designed for modern amplicon based sequencing experiments, especially in cases of high allelic diversity. Importantly, no guidelines are provided for the design of optimal UMI length for amplicon-based sequencing experiments.\n\nResultsBased on the total number of DNA fragments and the distribution of allele frequencies, we present a model for the determination of the minimum UMI length required to prevent UMI collisions and reduce allelic distortion. We also introduce a user-friendly software tool called AmpUMI to assist in the design and the analysis of UMI-based amplicon sequencing studies. AmpUMI provides quality control metrics on frequency and quality of UMIs, and trims and deduplicates amplicon sequences with user specified parameters for use in downstream analysis. AmpUMI is open-source and freely available at http://github.com/pinellolab/AmpUMI.\n\nContactIpinello@mgh.harvard.edu

bioinformatics

High-precision CRISPR-Cas9 base editors with minimized bystander and off-target mutations

Recently described base editor (BE) technology, which uses CRISPR-Cas9 to direct cytidine deaminase enzymatic activity to specific genomic loci, enables the highly efficient introduction of precise cytidine-to-thymidine (C [->] T) DNA alterations in many different cell types and organisms1-6. In contrast to genome-editing nucleases7-9, BEs avoid the need to introduce double-strand breaks or exogenous donor DNA templates and induce lower levels of unwanted variable-length insertion/deletion mutations (indels)1,2,10. However, existing BEs can also efficiently create unwanted C to T alterations when more than one C is present within the five base pair \"editing window\" of these proteins, a lack of precision that can cause potentially deleterious by ...

molecular biology

In vivo CRISPR-Cas gene editing with no detectable genome-wide off-target mutations

CRISPR-Cas genome-editing nucleases hold substantial promise for human therapeutics1-5 but identifying unwanted off-target mutations remains an important requirement for clinical translation6, 7. For ex vivo therapeutic applications, previously published cell-based genome-wide methods provide potentially useful strategies to identify and quantify these off-target mutation sites8-12. However, a well-validated method that can reliably identify off-targets in vivo has not been described to date, leaving the question of whether and how frequently these types of mutations occur. Here we describe Verification of In Vivo Off-targets (VIVO), a highly sensitive, unbiased, and generalizable strategy that we show can robustly identify genome-wide CRISPR-Cas nuclease off-target effects in vivo. To our knowledge, these studies provide the first demonstration that CRISPR-Cas nucleases can induce substantial off-target mutations in vivo, a result we obtained using a deliberately promiscuous guide RNA (gRNA). More importantly, we used VIVO to show that appropriately designed gRNAs can direct efficient in vivo editing without inducing detectable off-target mutations. Our findings provide strong support for and should encourage further development of in vivo genome editing therapeutic strategies.

molecular biology

Detecting genome-wide directional effects of transcription factor binding on polygenic disease risk

Biological interpretation of GWAS data frequently involves analyzing unsigned genomic annotations comprising SNPs involved in a biological process and assessing enrichment for disease signal. However, it is often possible to generate signed annotations quantifying whether each SNP allele promotes or hinders a biological process, e.g., binding of a transcription factor (TF). Directional effects of such annotations on disease risk enable stronger statements about causal mechanisms of disease than enrichments of corresponding unsigned annotations. Here we introduce a new method, signed LD profile regression, for detecting such directional effects using GWAS summary statistics, and we apply the method using 382 signed annotations reflecting predicted TF binding. We show via theory and simulations that our method is well-powered and is well-calibrated even when TF binding sites co-localize with other enriched regulatory elements, which can confound unsigned enrichment methods. We further validate our method by showing that it recovers known transcriptional regulators when applied to molecular QTL in blood. We then apply our method to eQTL in 48 GTEx tissues, identifying 651 distinct TF-tissue expression associations at per-tissue FDR < 5%, including 30 associations with robust evidence of tissue specificity. Finally, we apply our method to 46 diseases and complex traits (average N = 289,617) and identify 77 annotation-trait associations at per-trait FDR < 5% representing 12 independent TF-trait associations, and we conduct gene-set enrichment analyses to characterize the underlying transcriptional programs. Our results implicate new causal disease genes (including causal genes at known GWAS loci), and in some cases suggest a detailed mechanism for a causal genes effect on disease. Our method provides a new way to leverage functional data to draw inferences about disease etiology.

genetics

Haystack: systematic analysis of the variation of epigenetic states and cell-type specific regulatory elements

MotivationWith the increasing amount of genomic and epigenomic data in the public domain, a pressing challenge is how to integrate these data to investigate the role of epigenetic mechanisms in regulating gene expression and maintenance of cell-identity. To this end, we have implemented a computational pipeline to systematically study epigenetic variability and uncover regulatory DNA sequences that play a role in gene regulation.\n\nResultsHaystack is a bioinformatics pipeline to characterize hotspots of epigenetic variability across different cell-types as well as cell-type specific cis-regulatory elements along with their corresponding transcription factors. Our approach is generally applicable to any epigenetic mark and provides an important tool to investigate cell-type identity and the mechanisms underlying epigenetic switches during development. Additionally, we make available a set of precomputed tracks for a number of epigenetic marks across several cell types. These precomputed results may be used as an independent resource for functional annotation of the human genome.\n\nAvailabilityThe Haystack pipeline is implemented as an open-source, multiplatform, Python package called haystack_bio available at https://github.com/pinellolab/haystack_bio.\n\nContactlpinello@mgh.harvard.edu, gcyuan@jimmy.harvard.edu

bioinformatics

"Unexpected mutations after CRISPR-Cas9 editing in vivo" are most likely pre-existing sequence variants and not nuclease-induced mutations

Schaefer et al. recently advanced the provocative conclusion that CRISPR-Cas9 nuclease can induce off-target alterations at genomic loci that do not resemble the intended on-target site.1 Using high-coverage whole genome sequencing (WGS), these authors reported finding SNPs and indels in two CRISPR-Cas9-treated mice that were not present in a single untreated control mouse. On the basis of this association, Schaefer et al. concluded that these sequence variants were caused by CRISPR-Cas9. This new proposed CRISPR-Cas9 off-target activity runs contrary to previously published work2-8 and, if the authors are correct, could have profound implications for research and therapeutic applications. Here, we demonstrate that the simplest interpretation of Schaefer et al.s data is that the two CRISPR-Cas9-treated mice are actually more closely related genetically to each other than to the control mouse. This strongly suggests that the so-called \"unexpected mutations\" simply represent SNPs and indels shared in common by these mice prior to nuclease treatment. In addition, given the genomic and sequence distribution profiles of these variants, we show that it is challenging to explain how CRISPR-Cas9 might be expected to induce such changes. Finally, we argue that the lack of appropriate controls in Schaefer et al.s experimental design precludes assignment of causality to CRISPR-Cas9. Given these substantial issues, we urge Schaefer et al. to revise or re-state the original conclusions of their published work so as to avoid leaving misleading and unsupported statements to persist in the literature.

molecular biology

Integrated Computational Guide Design, Execution, And Analysis Of Arrayed And Pooled CRISPR Genome Editing Experiments

CRISPR genome editing experiments offer enormous potential for the evaluation of genomic loci using arrayed single guide RNAs (sgRNAs) or pooled sgRNA libraries. Numerous computational tools are available to help design sgRNAs with optimal on-target efficiency and minimal off-target potential. In addition, computational tools have been developed to analyze deep sequencing data resulting from genome editing experiments. However, these tools are typically developed in isolation and oftentimes not readily translatable into laboratory-based experiments. Here we present a protocol that describes in detail both the computational and benchtop implementation of an arrayed and/or pooled CRISPR genome editing experiment. This protocol provides instructions for sgRNA design with CRISPOR, experimental implementation, and analysis of the resulting high-throughput sequencing data with CRISPResso. This protocol allows for design and execution of arrayed and pooled CRISPR experiments in 4-5 weeks by non-experts as well as computational data analysis in 1-2 days that can be performed by both computational and non-computational biologists alike.

molecular biology

The role of Cdx2 as a lineage specific transcriptional repressor for pluripotent network during trophectoderm and inner cell mass specification

The first cellular differentiation event in mouse development leads to the formation of the blastocyst consisting of the inner cell mass (ICM) and an outer functional epithelium called trophectoderm (TE). The lineage specific transcription factor CDX2 is required for proper TE specification, where it promotes expression of TE genes, and represses expression of Pou5f1 (OCT4) by inhibiting OCT4 from promoting its own expression. However its downstream network in the developing early embryo is not fully characterized. Here, we performed high-throughput single embryo qPCR analysis in Cdx2 null embryos to identify components of the CDX2-regulated network in vivo. To identify genes likely to be regulated by CDX2 directly, we performed CDX2 ChIP-Seq on trophoblast stem (TS) cells, derived from the TE. In addition, we examined the dynamics of gene expression changes using an inducible CDX2 embryonic stem (ES) cell system, so that we could predict which CDX2-bound genes are activated or repressed by CDX2 binding. By integrating these data with observations of chromatin modifications, we were able to identify novel regulatory elements that are likely to repress gene expression in a lineage-specific manner. Interestingly, we found CDX2 binding sites within regulatory elements of key pluripotent genes such as Pou5f1 and Nanog, pointing to the existence of a novel mechanism by which CDX2 maintains repression of OCT4 in trophoblast. Our study proposes a general mechanism in regulating lineage segregation during mammalian development.

cell biology