Search bioRxiv⌕ Search

Biology subjects

Blood Service Biobank,

Publications and source records attributed to Blood Service Biobank,.

2 recordsLinked to original sources

Blood donor biobank pipeline to collect genome-based samples for research

The integration of genome data with electronic health records, driven by large biobank studies, has advanced human genetics by allowing systematic exploration of genotype-phenotype links. Regular donation enables large, longitudinal sample cohorts. Because blood donors are generally healthy, disease treatments or progression do not disturb interpretations in functional studies. We describe here a pipeline on how to collect blood donors high quality plasma, serum, and living cell samples for multi-omics studies. Peripheral blood mononuclear cells (PBMC) were frozen and, after thawing, contained standard levels of immune cell subpopulations, responded to immune activation, and were of good quality starting material for multi-omics and cell imaging studies. We demonstrate that most genetic variants of interest to the major genomics study in Finland, FinnGen, could be found by random collection of samples during the standard blood donation without recall. Probing simple associations in the multi-omics data confirmed expected associations with e.g. age and sex, demonstrating good sample quality. As an example of interesting findings, we observed a significant association between frequent blood donation and lower levels of per- and polyfluoroalkyl substances (PFAS). The study demonstrates that regular blood donors are a suitable target population for high-quality, cost-effective sample collections.

molecular biology↗

Accurate multi-population imputation of MICA, MICB, HLA-E, HLA-F and HLA-G alleles from genome SNP data

In addition to the classical HLA genes, the major histocompatibility complex (MHC) harbors a high number of other polymorphic genes with less established roles in disease associations and transplantation matching. To facilitate studies of the non-classical and non-HLA genes in large patient and biobank cohorts, we trained imputation models for MICA, MICB, HLA-E, HLA-F and HLA-G alleles on genome SNP array data. We show, using both population-specific and multi-population 1000 Genomes references, that the alleles of these genes can be accurately imputed for screening and research purposes. The best imputation model for MICA, MICB, HLA-E, -F and -G achieved a mean accuracy of 99.3% (min, max: 98.6, 99.9). Furthermore, validation of the 1000 Genomes exome short-read sequencing-based allele calling against a clinical-grade reference data showed an average accuracy of 99.8%, testifying for the quality of the 1000 Genomes data as an imputation reference. The imputation models, trained using the HIBAG algorithm, are available at GitHub (https://github.com/FRCBS/HLA_EFG_MICAB_imputation) and can be run locally, thus avoiding the need of sending sensitive genome data to remote portals.

bioinformatics↗