Search bioRxivSearch

Biology subjects

Luscombe, N. M.

Publications and source records attributed to Luscombe, N. M..

12 recordsLinked to original sources

Target-specific precision of CRISPR-mediated genome editing

The CRISPR-Cas9 system has successfully been adapted to edit the genome of various organisms. However, our ability to predict editing accuracy, efficacy and outcome at specific sites is limited by an incomplete understanding of how the bacterial system interacts with eukaryotic genomes and DNA repair machineries. Here, we performed the largest comparison of indel profiles to date, examining over one thousand sites in the genome of human cells, and uncovered general principles guiding CRISPR-mediated DNA editing. We find that precision of DNA editing varies considerably among sites, with some targets showing one highly-preferred indel and others displaying a wide range of infrequent indels. Editing precision correlates with editing efficiency, homology-associated end-joining for both insertions and deletions, and a preference for single-nucleotide insertions. Precise targets and the identity of their preferred indel can be predicted based on simple rules that mainly depend on the fourth nucleotide upstream of the PAM sequence. Regardless of precision, site-specific indel profiles are highly robust and depend on both DNA sequence and chromatin features. Our findings have important implications for clinical applications of CRISPR technology and reveal general patterns of broken end-joining that can inform us on DNA repair mechanisms in human cells.

molecular biology

Heteromeric RNP assembly at LINEs controls lineage-specific RNA processing

It is challenging for RNA processing machineries to select exons within long intronic regions. We find that intronic LINE repeat sequences (LINEs) contribute to this selection by recruiting dozens of RNA-binding proteins (RBPs). This includes MATR3, which promotes binding of PTBP1 to multivalent binding sites in LINEs. Both RBPs repress splicing and 3 end processing within and around LINEs, as demonstrated in cultured human cells and mouse brain. Notably, repressive RBPs preferentially bind to evolutionarily young LINEs, which are confined to deep intronic regions. These RBPs insulate both LINEs and surrounding regions from RNA processing. Upon evolutionary divergence, gradual loss of insulation diversifies the roles of LINEs. Older LINEs are located closer to exons, are a common source of tissue-specific exons, and increasingly bind to RBPs that enhance RNA processing. Thus, LINEs are hubs for assembly of repressive RBPs, and contribute to evolution of new, lineage-specific transcripts in mammals.

genomics

Identifying the genetic basis of variation in cell behaviour in human iPS cell lines from healthy donors

Large cohorts of human iPSCs from healthy donors are potentially a powerful tool for investigating the relationship between genetic variants and cellular phenotypes. Here we integrate high content imaging, gene expression and DNA sequence datasets for over 100 human iPSC lines to identify the genetic basis of inter-individual variability in cell behaviour. By applying a dimensionality reduction approach, Probabilistic Estimation of Expression Residuals (PEER), we identified genes that correlated in expression with intrinsic (genetic) and extrinsic (ECM) factors. However, variation in mRNA levels could not account for outlier cell behaviour. Instead, we identified rare, deleterious SNVs in the coding sequence of genes involved in ECM adhesion that occurred in cell lines that were outliers for one or more phenotypes such as cell spreading. These also correlated with altered germ layer differentiation on micropatterned surfaces. Our study thus establishes a strategy for integrating genetic and cell biological measurements for high-throughput analysis.

cell biology

Generative adversarial networks uncover epidermal regulators and predict single cell perturbations

Recent advances have enabled gene expression profiling of single cells at lower cost. As more data is produced there is an increasing need to integrate diverse datasets and better analyse underutilised data to gain biological insights. However, analysis of single cell RNA-seq data is challenging due to biological and technical noise which not only varies between laboratories but also between batches. Here for the first time, we apply a new generative deep learning approach called Generative Adversarial Networks (GAN) to biological data. We apply GANs to epidermal, neural and hematopoietic scRNA-seq data spanning different labs and experimental protocols. We show that it is possible to integrate diverse scRNA-seq datasets and in doing so, our generative model is able to simulate realistic scRNA-seq data that covers the full diversity of cell types. In contrast to many machine-learning approaches, we are able to interpret internal parameters in a biologically meaningful manner. Using our generative model we are able to obtain a universal representation of epidermal differentiation and use this to predict the effect of cell state perturbations on gene expression at high time-resolution. We show that our trained neural networks identify biological state-determining genes and through analysis of these networks we can obtain inferred gene regulatory relationships. Finally, we use internal GAN learned features to perform dimensionality reduction. In combination these attributes provide a powerful framework to progress the analysis of scRNA-seq data beyond exploratory analysis of cell clusters and towards integration of multiple datasets regardless of origin.

genomics

Machine learning models in electronic health records can outperform conventional survival models for predicting patient mortality in coronary artery disease

Prognostic modelling is important in clinical practice and epidemiology for patient management and research. Electronic health records (EHR) provide large quantities of data for such models, but conventional epidemiological approaches require significant researcher time to implement. Expert selection of variables, fine-tuning of variable transformations and interactions, and imputing missing values in datasets are time-consuming and could bias subsequent analysis, particularly given that missingness in EHR is both high, and may carry meaning.\n\nUsing a cohort of over 80,000 patients from the CALIBER programme, we performed a systematic comparison of several machine-learning approaches in EHR. We used Cox models and random survival forests with and without imputation on 27 expert-selected variables to predict all-cause mortality. We also used Cox models, random forests and elastic net regression on an extended dataset with 586 variables to build prognostic models and identify novel prognostic factors without prior expert input.\n\nWe observed that data-driven models used on an extended dataset can outperform conventional models for prognosis, without data preprocessing or imputing missing values, and with no need to scale or transform continuous data. An elastic net Cox regression based with 586 unimputed variables with continuous values discretised achieved a C-index of 0.801 (bootstrapped 95% CI 0.799 to 0.802), compared to 0.793 (0.791 to 0.794) for a traditional Cox model comprising 27 expert-selected variables with imputation for missing values.\n\nWe also found that data-driven models allow identification of novel prognostic variables; that the absence of values for particular variables carries meaning, and can have significant implications for prognosis; and that variables often have a nonlinear association with mortality, which discretised Cox models and random forests can elucidate.\n\nThis demonstrates that machine-learning approaches applied to raw EHR data can be used to build reliable models for use in research and clinical practice, and identify novel predictive variables and their effects to inform future research.

epidemiology

Regionalization of the nervous system requires axial allocation prior to neural lineage commitment

SummaryNeural induction in vertebrates generates a central nervous system that extends the rostral-caudal length of the body. The prevailing view is that neural cells are initially induced with anterior (forebrain) identity, with caudalising signals then converting a proportion to posterior fates (spinal cord). To test this model, we used chromatin accessibility assays to define how cells adopt region-specific neural fates. Together with genetic and biochemical perturbations this identified a developmental time window in which genome-wide chromatin remodeling events preconfigure epiblast cells for neural induction. Contrary to the established model, this revealed that cells commit to a regional identity before acquiring neural identity. This \"primary regionalization\" allocates cells to anterior or posterior regions of the nervous system, explaining how cranial and spinal neurons are generated at appropriate axial positions. These findings prompt a revision to models of neural induction and support the proposed dual evolutionary origin of the vertebrate central nervous system.

developmental biology

Data Science Issues in Understanding Protein-RNA Interactions

An interplay of experimental and computational methods is required to achieve a comprehensive understanding of protein-RNA interactions. Crosslinking and immunoprecipitation (CLIP) identifies endogenous interactions by sequencing RNA fragments that co-purify with a selected RBP under stringent conditions. Here we focus on approaches for the analysis of resulting data and appraise the methods for peak calling, visualisation, analysis and computational modelling of protein-RNA binding sites. We advocate a combined assessment of cDNA complexity and specificity for data quality control. Moreover, we demonstrate the value of analysing sequence motif enrichment in peaks assigned from CLIP data, and of visualising RNA maps, which examine the positional distribution of peaks around regulated landmarks in transcripts. We use these to assess how variations in CLIP data quality, and in different peak calling methods, affect the insights into regulatory mechanisms. We conclude by discussing future opportunities for the computational analysis of protein-RNA interaction experiments.

bioinformatics

Identifying developmentally important genes with single-cell RNA-seq from an embryo

Single-cell RNA-seq has been established as a reliable and accessible technique enabling new types of analyses, such as identifying cell types and studying spatial and temporal gene expression variation and change at single-cell resolution. Recently, single-cell RNA-seq has been applied to developing embryos, which offers great potential for finding and characterising genes controlling the course of development along with their expression patterns. In this study, we applied single-cell RNA-seq to the 16-cell stage of the Ciona embryo, a marine chordate and performed a computational search for cell-specific gene expression patterns. We recovered many known expression patterns from our single-cell RNA-seq data and despite extensive previous screens, we succeeded in finding new cell-specific patterns, which we validated by in situ and single-cell qPCR.

developmental biology

Integrated analysis sheds light on evolutionary trajectories of young transcription start sites in the human genome

Previous studies revealed widespread transcription initiation and fast turnover of transcription start sites (TSSs) in mammalian genomes. Yet how new TSSs originate and how they evolve over time remain poorly understood. To address these questions, we analyzed [~]200,000 human TSSs by integrating evolutionary and functional genomic data, particularly focusing on TSSs that emerged in the primate lineages. We found that intrinsic factors of repetitive sequences and their proximity to established regulatory modules (extrinsic factors) contribute significantly to origin of new TSSs. In early periods, young TSSs experience rapid sequence evolution driven by endogenous mutational mechanisms that reduce the instability of associated repetitive sequences. In later periods, the regulatory functions of young TSSs are gradually modified, and with evolutionary changes subject to temporal (fewer regulatory changes in younger TSSs) and spatial constraints (fewer regulatory changes in more isolated TSSs). These findings advance our understanding of how regulatory innovations arise in the genome throughout evolution and highlight the roles of repetitive sequences in these processes.

genomics

Post-transcriptional remodelling is temporally deregulated during motor neurogenesis in human ALS models

Mutations causing amyotrophic lateral sclerosis (ALS) strongly implicate regulators of RNA-processing that are ubiquitously expressed throughout development. To understand the molecular impact of ALS-causing mutations on early neuronal development and disease, we performed transcriptomic analysis of differentiated human control and VCP-mutant induced pluripotent stem cells (iPSCs) during motor neurogenesis. We identify intron retention (IR) as the predominant splicing change affecting early stages of wild-type neural differentiation, targeting key genes involved in the splicing machinery. Importantly, IR occurs prematurely in VCP-mutant cultures compared with control counterparts; these events are also observed in independent RNAseq datasets from SOD1- and FUS-mutant motor neurons (MNs). Together with related effects on 3UTR length variation, these findings implicate alternative RNA-processing in regulating distinct stages of lineage restriction from iPSCs to MNs, and reveal a temporal deregulation of such processing by ALS mutations. Thus, ALS-causing mutations perturb the same post-transcriptional mechanisms that underlie human motor neurogenesis.\n\nHIGHLIGHTSO_LIIntron retention is the main mode of alternative splicing in early differentiation.\nC_LIO_LIThe ALS-causing VCP mutation leads to premature intron retention.\nC_LIO_LIIncreased intron retention is seen with multiple ALS-causing mutations.\nC_LIO_LITranscriptional programs are unperturbed despite post-transcriptional defects.\nC_LI\n\neTOC BLURBLuisier et al. identify post-transcriptional changes underlying human motor neurogenesis: extensive variation in 3 UTR length and intron retention (IR) are the early predominant modes of splicing. The VCP mutation causes IR to occur prematurely during motor neurogenesis and these events are validated in other ALS-causing mutations, SOD1 and FUS.

neuroscience

3’UTR Remodelling of Axonal Transcripts in Sympathetic Neurons

The 3 untranslated regions (3UTRs) of messenger RNAs (mRNA) are non-coding sequences that regulate several aspects of mRNA metabolism, including intracellular localisation and translation. Here, we show that in sympathetic neuron axons, the 3UTRs of many transcripts undergo cleavage, generating both translatable isoforms expressing a shorter 3UTR, and 3UTR fragments. 3end RNA sequencing indicated that 3UTR cleavage is a potentially widespread event in axons, which is mediated by a protein complex containing the endonuclease Ago2 and the RNA binding protein HuD. Analysis of the Inositol monophosphatase 1 (Impa1) mRNA revealed that a stem loop structure within the 3UTR is necessary for Ago2 cleavage. Thus, remodeling of the 3UTR provides an alternative mechanism that simultaneously regulates local protein synthesis and generates a new class of 3UTR RNAs with yet unknown functions.

neuroscience

Epidermal Wnt signalling regulates transcriptome heterogeneity and proliferative fate in neighbouring cells

Canonical Wnt/beta-catenin signalling regulates self-renewal and lineage selection within the mouse epidermis. Although the transcriptional response of keratinocytes that receive a Wnt signal is well characterised, little is known about the mechanism by which keratinocytes in proximity to the Wntreceiving cell are co-opted to undergo a change in cell fate. To address this, we performed single-cell mRNA-Seq on mouse keratinocytes co-cultured with and without the presence of beta-catenin activated neighbouring cells. We identified seven distinct cell states in cultures that had not been exposed to the beta-catenin stimulus and show that the stimulus redistributes wild type subpopulation proportions. Using temporal single-cell analysis we reconstruct the cell fate changes induced by neighbour Wnt activation. Gene expression heterogeneity was reduced in neighbouring cells and this effect was most dramatic for protein synthesis associated genes. The changes in gene expression were accompanied by a shift from a quiescent to a more proliferative stem cell state. By integrating imaging and reconstructed sequential gene expression changes during the state transition we identified transcription factors, including Smad4 and Bcl3, that were responsible for effecting the transition in a contact-dependent manner. Our data indicate that non cell-autonomous Wnt/beta-catenin signalling decreases transcriptional heterogeneity and further our understanding of how epidermal Wnt signalling orchestrates regeneration and self-renewal.

genomics