Search bioRxivSearch

Biology subjects

Gottardo, R.

Publications and source records attributed to Gottardo, R..

3 recordsLinked to original sources

DataPackageR: Reproducible data preprocessing, standardization and sharing using R/Bioconductor for collaborative data analysis.

A central tenet of reproducible research is that scientific results are published along with the underlying data and software code necessary to reproduce and verify the findings. A host of tools and software have been released that facilitate such work-flows and scientific journals have increasingly demanded that code and primary data be made available with publications. There has been little practical advice on implementing reproducible research work-flows for large omics or systems biology data sets used by teams of analysts working in collaboration. In such instances it is important to ensure all analysts use the same version of a data set for their analyses. Yet, instantiating relational databases and standard operating procedures can be unwieldy, with high \"startup\" costs and poor adherence to procedures when they deviate substantially from an analysts usual work-flow. Ideally a reproducible research work-flow should fit naturally into an individuals existing work-flow, with minimal disruption. Here, we provide an overview of how we have leveraged popular open source tools, including Bioconductor, Rmarkdown, git version control, R, and specifically Rs package system combined with a new tool DataPackageR, to implement a lightweight reproducible research work-flow for preprocessing large data sets, suitable for sharing among small-to-medium sized teams of computational scientists. Our primary contribution is the DataPackageR tool, which decouples time-consuming data processing from data analysis while leaving a traceable record of how raw data is processed into analysis-ready data sets. The software ensures packaged data objects are properly documented and performs checksum verification of these along with basic package version management, and importantly, leaves a record of data processing code in the form of package vignettes. Our group has implemented this work-flow to manage, analyze and report on pre-clinical immunological trial data from multi-center, multi-assay studies for the past three years.

bioinformatics

cytometree: a binary tree algorithm for automatic gating in cytometry analysis

MotivationFlow cytometry is a powerful technology that allows the high-throughput quantification of dozens of surface and intracellular proteins at the single-cell level. It has become the most widely used technology for immunophenotyping of cells over the past three decades. Due to the increasing complexity of cytometry experiments (more cells and more markers), traditional manual flow cytometry data analysis has become untenable due to its subjectivity and time-consuming nature.\n\nResultsWe present a new unsupervised algorithm called \"cytometree\" to perform automated population discovery (aka gating) in flow cytometry. cytometree is based on the construction of a binary tree, the nodes of which are subpopulations of cells. At each node, the marker distributions are modeled by mixtures of normal distribution. Node splitting is done according to a normalized difference of Akaike information criteria (AIC) between the two models. Post-processing of the tree structure and derived populations allows us to complete the annotation of the derived populations. The algorithm is shown to perform better than the state-of-the-art unsupervised algorithms previously proposed on panels introduced by the Flow Cytometry: Critical Assessment of Population Identification Methods (FlowCAP I) project. The algorithm is also applied to a T-cell panel proposed by the Human Immunology Project Consortium (HIPC) program; it also outperforms the best unsupervised open-source available algorithm while requiring the shortest computation time.\n\nAvailabilityAn R package named \"cytometree\" is available on the CRAN repository.\n\nContactdaniel.commenges@u-bordeaux.fr; rodolphe.thiebaut@u-bordeaux.fr\n\nSupplementary informationSupplementary data are available.

bioinformatics

CD32+ and PD-1+ Lymph Node CD4 T Cells Support Persistent HIV-1 Transcription in Treated Aviremic Individuals

A recent study conducted in blood has proposed CD32 as the marker identifying the elusive HIV reservoir. We have investigated the distribution of CD32+ CD4 T cells in blood and lymph nodes(LNs) of healthy HIV-1 uninfected, viremic untreated and long-term treated HIV-1 infected individuals and their relationship with PD-1+ CD4 T cells. The frequency of CD32+ CD4 T cells was increased in viremic as compared to treated individuals in LNs and a large proportion(up to 50%) of CD32+ cells co-expressed PD-1 and were enriched within T follicular helper cells(Tfh) cells. We next investigated the role of LN CD32+ CD4 T cells in the HIV reservoir. Total HIV DNA was enriched in CD32+ and PD-1+ CD4 T cells as compared to CD32- and PD-1- cells in both viremic and treated individuals but there was no difference between CD32+ and PD-1+ cells. There was not enrichment of latently infected cells with inducible HIV-1 in CD32+ versus PD-1+ cells in ART treated individuals. HIV-1 transcription was then analyzed in LN memory CD4 T cell populations sorted on the basis of CD32 and PD-1 expression. CD32+PD-1+ CD4 T cells were significantly enriched in cell associated HIV RNA as compared to CD32-PD-1-(average 5.2 fold in treated and 86.6 fold in viremics), to CD32+PD-1-(2.2 fold in treated and 4.3 fold in viremics) and to CD32-PD-1+ cell populations(2.2 fold in ART treated and 4.6 fold in viremics). Similar levels of HIV-1 transcription were found in CD32+PD-1- and CD32-PD-1+ CD4 T cells. Interestingly, the proportion of CD32+ and PD-1+ CD4 T cells negatively correlated with CD4 T cell counts and length of therapy while positively correlated with viremia. Therefore, the expression of CD32 identifies, independently of PD-1, a CD4 T cell population with persistent HIV-1 transcription and CD32 and PD-1 co-expression the CD4 T cell population with the highest levels of HIV-1 transcription in both viremic and treated individuals.\n\nImportanceThe existence of long-lived latently infected resting memory CD4 T cells represents a major obstacle to the eradication of HIV infection. Identifying cell markers defining latently infected cells containing replication competent virus is important in order to determine the mechanisms of HIV persistence and to develop novel therapeutic strategies to cure HIV infection. We provide evidence that PD-1 and CD32 may have a complementary role in better defining CD4 T cell populations infected with HIV-1. Furthermore, CD4 T cells co-expressing CD32 and PD-1 identify a CD4 T cell population with high levels of persistent HIV-1 transcription.

immunology