Search bioRxivSearch

Biology subjects

Regev, A.

Publications and source records attributed to Regev, A..

At least 19 recordsLinked to original sources

Single-cell transcriptomics of the aged mouse brain reveals convergent, divergent and unique aging signatures

The mammalian brain is complex, with multiple cell types performing a variety of diverse functions, but exactly how the brain is affected with aging remains largely unknown. Here we performed a single-cell transcriptomic analysis of young and old mouse brains. We provide a comprehensive dataset of aging-related genes, pathways and ligand-receptor interactions in nearly all brain cell types. Our analysis identified gene signatures that vary in a coordinated manner across cell types and gene sets that are regulated in a cell type specific manner, even at times in opposite directions. Thus, our data reveals that aging, rather than inducing a universal program drives a distinct transcriptional course in each cell population. These data provide an important resource for the aging community and highlight key molecular processes, including ribosomal biogenesis, underlying aging. We believe that this large-scale dataset, which is publicly accessible online (aging-mouse-brain), will facilitate additional discoveries directed towards understanding and modifying the aging process.

neuroscience

Molecular Classification and Comparative Taxonomics of Foveal and Peripheral Cells in Primate Retina

High acuity vision in primates, including humans, is mediated by a small central retinal region called the fovea. As more accessible model organisms lack a fovea, its specialized function and dysfunction in ocular diseases remain poorly understood. We used 165,000 single-cell RNA-seq profiles to generate and validate comprehensive cellular taxonomies of macaque fovea and peripheral retina. More than 80% of >65 cell types match between the two regions, but exhibit substantial differences in proportions and gene expression, some of which we relate to functional differences. Comparison of macaque retinal types with those of mice reveals that interneuron types are tightly conserved, but that projection neuron types and programs diverge, despite conserved transcription factor codes. Key macaque types are conserved in humans, allowing mapping of cell-type and region-specific expression of >190 genes associated with 6 human retinal diseases. Our work provides a framework for comparative single-cell analysis across tissue regions and species.

neuroscience

Regulatory network controlling tumor-promoting inflammation in human cancers

Using an inducible, inflammatory model of breast cellular transformation, we describe the transcriptional regulatory network mediated by STAT3, NF-{kappa}B, and AP-1 factors on a genomic scale. These regulators form transcriptional complexes that directly regulate the expression of hundreds of genes in oncogenic pathways via a positive feedback loop. This inflammatory feedback loop, which functions to various extents in many types of cancer cells and patient tumors, is the basis for an \"inflammation\" index that defines cancer types by functional criteria. We identify a network of non-inflammatory genes whose expression is well correlated with the cancer inflammatory index. Conversely, the inflammation index is negatively correlated with expression of genes involved in DNA metabolism, and transformation is associated with genome instability. Inflammatory tumors are preferentially associated with infiltrating immune cells that might be recruited to the site of the tumor via inflammatory molecules produced by the cancer cells.

cancer biology

A single cell-based atlas of human microglial states reveals associations with neurological disorders and histopathological features of the aging brain

Recent studies of bulk microglia have provided insights into the role of this immune cell type in central nervous system development, homeostasis and dysfunction. Nonetheless, our understanding of the diversity of human microglial cell states remains limited; microglia are highly plastic and have multiple different roles, making the extent of phenotypic heterogeneity a central question, especially in light of the development of therapies targeting this cell type. Here, we investigated the population structure of human microglia by single-cell RNA-sequencing. Using surgical- and autopsy-derived cortical brain samples, we identified 14 human microglial subpopulations and noted substantial intra- and inter-individual heterogeneity. These putative subpopulations display divergent associations with Alzheimers disease, multiple sclerosis, and other diseases. Several states show enrichment for genes found in disease-associated mouse microglial states, suggesting additional diversity among human microglia. Overall, human microglia appear to exist in different functional states with varying levels of involvement in different brain pathologies.

neuroscience

Conservation and divergence in modules of the transcriptional programs of the human and mouse immune systems

Studies in mouse have shed important light on human hematopoietic differentiation and disease. However, substantial differences between the two species often limit the translation of findings from mouse to human.\n\nHere, we compare previously defined modules of co-expressed genes in human and mouse immune cells based on compendia of genome-wide profiles. We show that the overall modular organization of the transcriptional program is conserved. We highlight modules of co-expressed genes in one species that dissolve or split in the other species. Many of the associated regulatory mechanisms - as reflected by computationally inferred trans regulators, or enriched cis-regulatory elements - are conserved between the species. Nevertheless, the degree of conservation in regulatory mechanism is lower than that of expression, suggesting that distinct regulation may underlie some of the conserved transcriptional responses.

bioinformatics

A quantitative model for characterizing the evolutionary history of mammalian gene expression

Characterizing the evolutionary history of a genes expression profile is a critical component for understanding the relationship between genotype, expression, and phenotype. However, it is not well-established how best to distinguish the different evolutionary forces acting on gene expression. Here, we use RNA-seq across 7 tissues from 17 mammalian species to show that expression evolution across mammals is accurately modeled by the Ornstein-Uhlenbeck (OU) process. This stochastic process models expression trajectories across time as Gaussian distributions whose variance is parameterized by the rate of genetic drift and strength of stabilizing selection. We use these mathematical properties to identify expression pathways under neutral, stabilizing, and directional selection, and quantify the extent of selective pressure on a genes expression. We further detect deleterious expression levels outside expected evolutionary distributions in expression data from individual patients. Our work provides a statistical framework for interpreting expression data across species and in disease.\n\nOne Sentence SummaryWe demonstrate the power of a stochastic model for quantifying selective pressure on expression and estimating evolutionary distributions of optimal gene expression.

genomics

Deciphering cis-regulatory logic with 100 million synthetic promoters

Deciphering cis-regulation, the code by which transcription factors (TFs) interpret regulatory DNA sequence to control gene expression levels, is a long-standing challenge. Previous studies of native or engineered sequences have remained limited in scale. Here, we use random sequences as an alternative, allowing us to measure the expression output of over 100 million synthetic yeast promoters. Random sequences yield a broad range of reproducible expression levels, indicating that the fortuitous binding sites in random DNA are functional. From these data we learn models of transcriptional regulation that predict over 94% of the expression driven from independent test data and nearly 89% from sequences from yeast promoters. These models allow us to characterize the activity of TFs and their interactions with chromatin, and help refine cis-regulatory motifs. We find that strand, position, and helical face preferences of TFs are widespread and depend on interactions with neighboring chromatin. Such massive-throughput regulatory assays of random DNA provide the diverse examples necessary to learn complex models of cis-regulatory logic.

genomics

T helper cells modulate intestinal stem cell renewal and differentiation

In the small intestine, a cellular niche of diverse accessory cell types supports the rapid generation of mature epithelial cell types through self-renewal, proliferation, and differentiation of intestinal stem cells (ISCs). However, not much is known about interactions between immune cells and ISCs, and it is unclear if and how immune cell dynamics affect eventual ISC fate or the balance between self-renewal and differentiation. Here, we used single-cell RNA-seq (scRNA-Seq) of intestinal epithelial cells (IECs) to identify new mechanisms for ISC-immune cell interactions. Surprisingly, MHC class II (MHCII) is enriched in two distinct subsets of Lgr5+ crypt base columnar ISCs, which are also distinguished by higher proliferation rates. Using co-culture of T cells with intestinal organoids, cytokine stimulations, and in vivo mouse models, we confirm that CD4+ T helper (Th) cells communicate with ISCs and affect their differentiation, in a manner specific to the Th subtypes and their signature cytokines and dependent on MHCII expression by ISCs. Specific inducible knockout of MHCII in intestinal epithelial cells in mice in vivo results in expansion of the ISC pool. Mice lacking T cells have expanded ISC pools, whereas specific depletion of Treg cells in vivo results in substantial reduction of ISC numbers. Our findings show that interactions between Th cells and ISCs mediated via MHCII expressed in intestinal epithelial stem cells help orchestrate tissue-wide responses to external signals.

immunology

A molecular network of the aging brain implicates INPPL1 and PLXNB1 in Alzheimer’s disease

The fact that only symptomatic therapies of small effect are available for Alzheimers disease (AD) today highlights the need for new therapeutic targets with which to prevent a major contributor to aging-related cognitive decline. Here, we report the construction and validation of a molecular network of the aging human frontal cortex. Using RNA sequence data from 478 individuals, we first identify the role of modules of coexpressed genes, and then confirm them in independent AD datasets. Then, we prioritize influential genes in AD-related modules and test our predictions in human model systems. We functionally validate two putative regulator genes in human astrocytes: INPPL1 and PLXNB1, whose activity in AD may be related to semaphorin signalling and type II diabetes, which have both been implicated in AD. This arc of network identification followed by statistical and experimental validation provides specific new targets for therapeutic development and illustrates a network approach to a complex disease.\n\nOne sentence summaryMolecular network analysis of RNA sequencing data from the aging human cortex identifies new Alzheimers and cognitive decline genes.

systems biology

A unified web platform for network-based analyses of genomic data

Functional genomics networks are widely used to identify unexpected pathway relationships in large genomic datasets. However, it is challenging to quantitatively compare the signal-to-noise ratio of different networks, the biology they describe, and to identify the optimal network to interpret a particular genetic dataset. Via GeNets users can train a machine-learning model (Quack) to make such comparisons; and they can execute, store, and share analyses of genetic and RNA sequencing datasets.

genomics

Reconstruction of developmental landscapes by optimal-transport analysis of single-cell gene expression sheds light on cellular reprogramming.

Understanding the molecular programs that guide cellular differentiation during development is a major goal of modern biology. Here, we introduce an approach, WADDINGTON-OT, based on the mathematics of optimal transport, for inferring developmental landscapes, probabilistic cellular fates and dynamic trajectories from large-scale single-cell RNA-seq (scRNA-seq) data collected along a time course. We demonstrate the power of WADDINGTON-OT by applying the approach to study 65,781 scRNA-seq profiles collected at 10 time points over 16 days during reprogramming of fibroblasts to iPSCs. We construct a high-resolution map of reprogramming that rediscovers known features; uncovers new alternative cell fates including neuraland placental-like cells; predicts the origin and fate of any cell class; highlights senescent-like cells that may support reprogramming through paracrine signaling; and implicates regulatory models in particular trajectories. Of these findings, we highlight Obox6, which we experimentally show enhances reprogramming efficiency. Our approach provides a general framework for investigating cellular differentiation.

bioinformatics

Genetic analysis of isoform usage in the human anti-viral response reveals influenza-specific regulation of ERAP2 transcripts under balancing selection

While the impact of common genetic variants on gene expression response to cellular stimuli has been analyzed in depth, less is known about how stimulation modulates the genetic control of isoform usage. Analyzing RNA-seq profiles of monocyte-derived dendritic cells from 243 individuals, we uncovered thousands of unannotated isoforms synthesized in response to viral infection and stimulation with type I interferon. We identified more than a thousand single nucleotide polymorphisms associated with isoform usage (isoQTLs), > 40% of which are independent of expression QTLs for the same gene. Compared to eQTLs, isoQTLs are enriched for splice sites and untranslated regions, and depleted of sequences upstream of annotated transcription start sites. Both eQTLs and isoQTLs in stimulated cells explain a significant proportion of the disease heritability attributed to common genetic variants. At the IRF7 locus, we found alternative promoter usage in response to influenza as a possible mechanism by which DNA variants previously associated with immune-related disorders mediate disease risk. At the ERAP2 locus, we shed light on the function of the major haplotype that has been maintained under long-term balancing selection. At baseline and following type 1 interferon stimulation, the major haplotype is associated with absence of ERAP2 expression while the minor haplotype, known to increase Crohns disease risk, is associated with high ERAP2 expression. Surprisingly, in response to influenza infection, the major haplotype results in the expression of two uncharacterized, alternatively transcribed, spliced and translated short isoforms. Thus, genetic variants at a single locus could modulate independent gene regulatory processes in the innate immune response, and in the case of ERAP2, may confer a historical fitness advantage in response to virus.

genomics

Heterogeneous Responses of Hematopoietic Stem Cells to Inflammatory Stimuli are Altered with Age

Long-term hematopoietic stem cells (LT-HSCs) maintain hematopoietic output throughout an animal's lifespan. With age, however, they produce a myeloid-biased output that may lead to poor immune responses to infectious challenge and the development of myeloid leukemias. Here, we show that young and aged LT-HSCs respond differently to inflammatory stress, such that aged LT-HSCs produce a cell-intrinsic, myeloid-biased expression program. Using single-cell RNA-seq, we identify a myeloid-biased subset within the LT-HSC population (mLT-HSCs) that is much more common amongst aged LT-HSCs and is uniquely primed to respond to acute inflammatory challenge. We predict several transcription factors to regulate differentially expressed genes between mLT-HSCs and other LT-HSC subsets. Among these, we show that Klf5, Ikzf1 and Stat3 play important roles in age-related inflammatory myeloid bias. These factors may regulate myeloid versus lymphoid balance with age, and can potentially mitigate the long-term deleterious effects of inflammation that lead to hematopoietic pathologies.\n\nHighlightsO_LILT-HSCs from young and aged mice have differential responses to acute inflammatory challenge.\nC_LIO_LIHSPCs directly sense inflammatory stimuli in vitro and have a robust transcriptional response.\nC_LIO_LIAged LT-HSCs demonstrate a cell-intrinsic myeloid bias during inflammatory challenge.\nC_LIO_LISingle-cell RNA-seq unmasked the existence of two subsets within the LT-HSC population that was apparent upon stimulation but not steady-state. One of the LT-HSC subsets is more prevalent in young and the other in aged mice.\nC_LIO_LIKlf5, Ikzf1 and Stat3 regulate age- and inflammation-related LT-HSC myeloid-bias.\nC_LI\n\nOne sentence summaryMurine hematopoietic stem cells display transcriptional heterogeneity that is quantitatively altered with age and leads to the age-dependent myeloid bias evident after inflammatory challenge.

immunology

The Human Cell Atlas

The recent advent of methods for high-throughput single-cell molecular profiling has catalyzed a growing sense in the scientific community that the time is ripe to complete the 150-year-old effort to identify all cell types in the human body, by undertaking a Human Cell Atlas Project as an international collaborative effort. The aim would be to define all human cell types in terms of distinctive molecular profiles (e.g., gene expression) and connect this information with classical cellular descriptions (e.g., location and morphology). A comprehensive reference map of the molecular state of cells in healthy human tissues would propel the systematic study of physiological states, developmental trajectories, regulatory circuitry and interactions of cells, as well as provide a framework for understanding cellular dysregulation in human disease. Here we describe the idea, its potential utility, early proofs-of-concept, and some design considerations for the Human Cell Atlas.

cell biology

Deciphering Variance In Epigenomic Regulators By k-mer Factorization

BackgroundVariation in chromatin organization across single cells can help shed important light on the mechanisms controlling gene expression, but scale, noise, and sparsity pose significant challenges for interpretation of single cell chromatin data. Here, we develop BROCKMAN (Brockman Representation Of Chromatin by K-mers in Mark-Associated Nucleotides), an approach to infer variation in transcription factor (TF) activity across samples through unsupervised analysis of the variation in DNA sequences associated with an epigenomic mark.\n\nResultsBROCKMAN represents each sample as a vector of epigenomic-mark-associated DNA word frequencies, and decomposes the resulting matrix to find hidden structure in the data, followed by unsupervised grouping of samples and identification of the TFs that distinguish groups. Applied to single cell ATAC-seq, BROCKMAN readily distinguished cell types, treatments, batch effects, experimental artifacts, and cycling cells. We show that each variable component in the k-mer landscape reflects a set of co-varying TFs, which are often known to physically interact. For example, in K562 cells, AP-1 TFs were central determinant of variability in chromatin accessibility through their variable expression levels and diverse interactions with other TFs. We provide a theoretical basis for why cooperative TF binding - and any associated epigenomic mark - is inherently more variable than non-cooperative binding.\n\nConclusionsBROCKMAN and related approaches will help gain a mechanistic understanding of the trans determinants of chromatin variability between cells, treatments, and individuals.

genomics

STAR-Fusion: Fast and Accurate Fusion Transcript Detection from RNA-Seq

MotivationFusion genes created by genomic rearrangements can be potent drivers of tumorigenesis. However, accurate identification of functionally fusion genes from genomic sequencing requires whole genome sequencing, since exonic sequencing alone is often insufficient. Transcriptome sequencing provides a direct, highly effective alternative for capturing molecular evidence of expressed fusions in the precision medicine pipeline, but current methods tend to be inefficient or insufficiently accurate, lacking in sensitivity or predicting large numbers of false positives. Here, we describe STAR-Fusion, a method that is both fast and accurate in identifying fusion transcripts from RNA-Seq data.\n\nResultsWe benchmarked STAR-Fusions fusion detection accuracy using both simulated and genuine Illumina paired-end RNA-Seq data, and show that it has superior performance compared to popular alternative fusion detection methods.\n\nAvailability and implementationSTAR-Fusion is implemented in Perl, freely available as open source software at http://star-fusion.github.io, and supported on Linux.\n\nContactbhaas@broadinstitute.org

bioinformatics

DroNc-Seq: Deciphering cell types in human archived brain tissues by massively-parallel single nucleus RNA-seq

Single nucleus RNA-Seq (sNuc-Seq) profiles RNA from tissues that are preserved or cannot be dissociated, but does not provide the throughput required to analyse many cells from complex tissues. Here, we develop DroNc-Seq, massively parallel sNuc-Seq with droplet technology. We profile 29,543 nuclei from mouse and human archived brain samples to demonstrate sensitive, efficient and unbiased classification of cell types, paving the way for charting systematic cell atlases.

genomics

Composite measurements and molecular compressed sensing for highly efficient transcriptomics

RNA profiling is an excellent phenotype of cellular responses and tissue states, but can be costly to generate at the massive scale required for studies of regulatory circuits, genetic states or perturbation screens. Here, we draw on a series of advances over the last decade in the field of mathematics to establish a rigorous link between biological structure, data compressibility, and efficient data acquisition. We propose that very few random composite measurements - in which gene abundances are combined in a random linear combination - are needed to approximate the high-dimensional similarity between any pair of gene abundance profiles. We then show how finding latent, sparse representations of gene expression data would enable us to \"decompress\" a small number of random composite measurements and recover high-dimensional gene expression levels that were not measured (unobserved). We present a new algorithm for finding sparse, modular structure, which improves the ability to interpret samples in terms of small numbers of active modules, and show that the modular structure we find is sufficient to recover gene expression profiles from composite measurements (with ~100-fold fewer composite measurements than genes). Moreover, the knowledge that sparse, modular structures exist allows us to recover expression profiles from composite measurements, even without access to any training data. Finally, we present a proof-of-concept experiment for making composite measurements in the laboratory, involving the measurement of linear combinations of RNA abundances. Altogether, our results suggest new compressive modalities in experimental biology that can form a foundation for massive scaling in high-throughput measurements, while also offering new insights into the interpretation of high-dimensional data.

systems biology