Search bioRxivSearch

Biology subjects

Nolan, G. P.

Publications and source records attributed to Nolan, G. P..

4 recordsLinked to original sources

Continuous visualization of differences between biological conditions in single-cell data

In high-dimensional single cell data, comparing changes in functional markers between conditions is typically done across manual or algorithm-derived partitions based on population-defining markers. Visualizations of these partitions is commonly done on low-dimensional embeddings (eg. t-SNE), colored by per-partition changes. Here, we provide an analysis and visualization tool that performs these comparisons across overlapping k-nearest neighbor (KNN) groupings. This allows one to color low-dimensional embeddings by marker changes without hard boundaries imposed by partitioning. We devised an objective optimization of k based on minimizing functional marker KNN imputation error. Proof-of-concept work visualized the exact location of an IL-7 responsive subset in a B cell developmental trajectory on a t-SNE map independent of clustering. Per-condition cell frequency analysis revealed that KNN is sensitive to detecting artifacts due to marker shift, and therefore can also be valuable in a quality control pipeline. Overall, we found that KNN groupings lead to useful multiple condition visualizations and efficiently extract a large amount of information from mass cytometry data. Our software is publicly available through the Bioconductor package Sconify.

systems biology

Evolutionary Origin of the Mammalian Hematopoietic System Found in a Colonial Chordate

Hematopoiesis is an essential process that evolved in multicellular animals. At the heart of this process are hematopoietic stem cells (HSCs), which are multipotent, self-renewing and generate the entire repertoire of blood and immune cells throughout life. Here we studied the hematopoietic system of Botryllus schlosseri, a colonial tunicate that has vasculature, circulating blood cells, and interesting characteristics of stem cell biology and immunity. Self-recognition between genetically compatible B. schlosseri colonies leads to the formation of natural parabionts with shared circulation, whereas incompatible colonies reject each other. Using flow-cytometry, whole-transcriptome sequencing of defined cell populations, and diverse functional assays, we identified HSCs, progenitors, immune-effector cells, the HSC niche, and demonstrated that self-recognition inhibits cytotoxic reaction. Our study implies that the HSC and myeloid lineages emerged in a common ancestor of tunicates and vertebrates and suggests that hematopoietic bone marrow and the B. schlosseri endostyle niche evolved from the same origin.

immunology

Meta-analysis of Cytometry Data Reveals Racial Differences in Immune Cells

While meta-analysis has demonstrated increased statistical power and more robust estimations in studies, the application of this commonly accepted methodology to cytometry data has been challenging. Different cytometry studies often involve diverse sets of markers. Moreover, the detected values of the same marker are inconsistent between studies due to different experimental designs and cytometer configurations. As a result, the cell subsets identified by existing auto-gating methods cannot be directly compared across studies. We developed MetaCyto for automated meta-analysis of both flow and mass cytometry (CyTOF) data. By combining clustering methods with a silhouette scanning method, MetaCyto is able to identify commonly labeled cell subsets across studies, thus enabling meta-analysis. Applying MetaCyto across a set of 10 heterogeneous cytometry studies totaling 2926 samples enabled us to identify multiple cell populations exhibiting differences in abundance between White and Asian adults. Software is released to the public through GitHub (github.com/hzc363/MetaCyto).

bioinformatics

Scalable Multi-Sample Single-Cell Data Analysis by Partition-Assisted Clustering and Multiple Alignments of Networks

Mass cytometry (CyTOF) has greatly expanded the capability of cytometry. It is now easy to generate multiple CyTOF samples in a single study, with each sample containing single-cell measurement on 50 markers for more than hundreds of thousands of cells. Current methods do not adequately address the issues concerning combining multiple samples for subpopulation discovery, and these issues can be quickly and dramatically amplified with increasing number of samples. To overcome this limitation, we developed Partition-Assisted Clustering and Multiple Alignments of Networks (PAC-MAN) for the fast automatic identification of cell populations in CyTOF data closely matching that of expert manual-discovery, and for alignments between subpopulations across samples to define dataset-level cellular states. PAC-MAN is computationally efficient, allowing the management of very large CyTOF datasets, which are increasingly common in clinical studies and cancer studies that monitor various tissue samples for each subject.\n\nAuthor SummaryRecently, the cytometry field has experienced rapid advancement in the development of mass cytometry (CyTOF). CyTOF enables a significant increase in the ability to monitor 50 or more cellular markers for millions of cells at the single-cell level. Initial studies with CyTOF focused on few samples, in which expert manual discovery of cell types were acceptable. As the technology matures, it is now feasible to collect more samples, which enables systematic studies of cell types across multiple samples. However, the statistical and computational issues surrounding multi-sample analysis have not been previously examined in detail. Furthermore, it was not clear how the data analysis could be scaled for hundreds of samples, such as those in clinical studies. In this work, we present a scalable analysis pipeline that is grounded in strong statistical foundation. Partition-Assisted Clustering (PAC) offers fast and accurate clustering and Multiple Alignments of Networks (MAN) utilizes network structures learned from each homogeneous cluster to organize the data into data-set level clusters. PAC-MAN thus enables the analysis of a large CyTOF dataset that was previously too large to be analyzed systematically; this pipeline can be extended to the analysis of similarly large or larger datasets.

bioinformatics