Search bioRxivSearch

Biology subjects

Cooper, S. J.

Publications and source records attributed to Cooper, S. J..

3 recordsLinked to original sources

powerTCR: a model-based approach to comparative analysis of the clone size distribution of the T cell receptor repertoire

Sequencing of the T cell receptor repertoire is a powerful tool for deeper study of immune response, but the unique structure of this type of data makes its meaningful quantification challenging. We introduce a new method, the Gamma-GPD spliced threshold model, to address this difficulty. This biologically interpretable model captures the distribution of the TCR repertoire, demonstrates stability across varying sequencing depths, and permits comparative analysis across any number of sampled individuals. We apply our method to several datasets and obtain insights regarding the differentiating features in the T cell receptor repertoire among sampled individuals across conditions. We have implemented our method in the open-source R package powerTCR.\n\nAuthor summaryA more detailed understanding of the immune response can unlock critical information concerning diagnosis and treatment of disease. Here, in particular, we study T cells through T cell receptor sequencing, as T cells play a vital role in immune response. One important feature of T cell receptor sequencing data is the frequencies of each receptor in a given sample. These frequencies harbor global information about the landscape of the immune response. We introduce a flexible method that extracts this information by modeling the distribution of these frequencies, and show that it can be used to quantify differences in samples from individuals of different biological conditions.

immunology

R2DGC: Threshold-free peak alignment and identification for 2D gas chromatography mass in R

Summary: Comprehensive two dimensional gas chromatography-mass spectrometry is a powerful method for analyzing complex mixtures of volatile compounds. This method produces a large amount of raw data that requires downstream processing to align signals of interest (peaks) across multiple samples and match peak characteristics to reference standard libraries prior to downstream statistical analysis. To address the paucity of applications addressing this need, we have developed an R package that implements retention time and mass spectra similarity threshold-free alignments, seamlessly integrates retention time standards for universally reproducible alignments, performs common ion filtering, and provides compatibility with multiple peak quantification methods. We demonstrate the packages utility on a controlled mix of metabolite standards separated under variable chromatography conditions and data generated from cell lines.\n\nAvailability and documentation: R2DGC can be downloaded at https://github.com/rramaker/R2DGC or installed via the Comprehensive R Archive Network (CRAN).\n\nContact: sjcooper@hudsonalpha.org\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

A genome-wide interactome of DNA-associated proteins in the human liver

Large-scale efforts like the Encyclopedia of DNA Elements (ENCODE) Project have made tremendous progress in cataloging the genomic binding patterns of DNA-associated proteins (DAPs), such as transcription factors (TFs). However most chromatin immunoprecipitation-sequencing (ChIP-seq) analyses have focused on a few immortalized cell lines whose activities and physiology deviate in important ways from endogenous cells and tissues. Consequently, binding data from primary human tissue are essential to improving our understanding of in vivo gene regulation. Here we analyze ChIP-seq data for 20 DAPs assayed in two healthy human liver tissue samples, identifying more than 450,000 binding sites. We integrated binding data with transcriptome and phased whole genome data to investigate allelic DAP interactions and the impact of heterozygous sequence variation on the expression of neighboring genes. We find our tissue-based dataset demonstrates binding patterns more consistent with liver biology than cell lines, and describe uses of these data to better prioritize impactful non-coding variation. Collectively, our rich dataset offers novel insights into genome function in healthy liver tissue and provides a valuable research resource for assessing disease-related disruptions.

genomics