Search bioRxivSearch

Biology subjects

Hovig, E.

Publications and source records attributed to Hovig, E..

6 recordsLinked to original sources

Rationally designed mimotope library for profiling of the human IgM repertoire

Specific antibody reactivities are routinely used as biomarkers but the use of antibody repertoire profiles is still awaiting recognition. Here we suggest to expedite the adoption of this class of system level biomarkers by rationally designing a peptide array as an efficient probe for an appropriately chosen repertoire compartment. Most IgM antibodies are characterized by few somatic mutations, polyspecificity and physiological autoreactivity with housekeeping function. Previously, probing this repertoire with a set of immunodominant self-proteins provided only coarse information on repertoire profiles. In contrast, here we describe the rational selection of a peptide mimotope set, appropriately sized as a potential diagnostic, that also represents optimally the diversity of the human public IgM reactivities. A 7-mer random peptide phage display library was panned on pooled human IgM. Next generation sequencing of the selected phage yielded a non-exhaustive set of 224087 mimotopes which clustered in 790 sequence clusters. A set of 594 mimotopes, representative of the most significant clusters, was used to demonstrate that this approach samples symmetrically the space of IgM reactivities. When probed with diverse patients sera in an oriented peptide array, this set produced a higher and more dynamic signal as compared to 1) random peptides, 2) random peptides purged of mimotope-like sequences and 3) mimotopes from a small subset of clusters. In this respect, the representative library is an optimized probe of the human IgM diversity. Proof of principle predictors for randomly selected diagnoses based on the optimized library demonstrated that it contains more than 1070 different profiles with the capacity to correlate with diverse pathologies. Thus, an optimized small library of IgM mimotopes is found to address very efficiently the dynamic diversity of the human IgM repertoire providing informationally dense and structurally interpretable IgM reactivity profiles. Author SummaryThe presence in the blood of antibodies specific for a particular infectious agent is used routinely as a diagnostic tool. The overall profile of available antibody reactivities (or their repertoire) in an individual has been studied much less. As an omics approach to immunity it can be a rich source of information about the system beyond just the individual history of antigenic exposure. Using a subset of antibodies - IgM, which are involved also in housekeeping functions like removing dead cells, and bacteriophage based techniques for selection of specific peptides, we managed to define a non-exhaustive set of 224087 peptides recognized by IgM antibodies present in most individuals. They were found to group naturally in 790 structural groups. Limiting these to the most outstanding 594 groups, we used one representative from each group to assemble a reasonably small set of peptides that extracts the maximum information from the antibody repertoire at a minimum cost per test. We demonstrate, that this representative peptide library is a better probe of the human IgM diversity than comparably sized libraries constructed on other principles. The optimized library contains more than 1070 different potentially profiles useful for the diagnosis, prognosis or monitoring of inflammatory and infectious conditions, tumors, neurodegenerative diseases, etc.

immunology

MirGeneDB2.0: the curated microRNA Gene Database

Small non-coding RNAs have gained substantial attention due to their roles in animal development and human disorders. Among them, microRNAs are unique because individual gene sequences are conserved across the animal kingdom. In addition, unique and mechanistically well understood features can clearly distinguish bona fide miRNAs from the myriad other small RNAs generated by cells. However, making this separation is not a common practice and, thus, not surprisingly, the heterogeneous quality of available miRNA complements has become a major concern in microRNA research. We addressed this by extensively expanding our curated microRNA gene database MirGeneDB to 45 organisms that represent the full taxonomic breadth of Metazoa. By consistently annotating and naming more than 10,900 microRNA genes in these organisms, we show that previous microRNA annotations contained not only many false positives, but surprisingly lacked more than 2,100 bona fide microRNAs. Indeed, curated microRNA complements of closely related organisms are very similar and can be used to reconstruct Metazoan evolution. MirGeneDB represents a robust platform for microRNA-based research, providing deeper and more significant insights into the biology and evolution of miRNAs but also biomedical and biomarker research. MirGeneDB is publicly and freely available at http://mirgenedb.org/.

genomics

Sample-Index Misassignment Impacts Tumor Exome Sequencing

Sample pooling enabled by dedicated indexes is a common and cost-effective strategy used in high-throughput DNA sequencing. Index misassignment leading to cross-sample contamination has however been described as a general problem of sequencing instruments which utilize exclusion amplification. Using real-life data from multiple tumor sequencing projects, we demonstrate that co-multiplexed samples can induce artifactual calls closely resembling high-quality somatic variant calls, and argue that dual indexing is the most reliable countermeasure.

genomics

Personal Cancer Genome Reporter: Variant Interpretation Report For Precision Oncology

SummaryIndividual tumor genomes pose a major challenge for clinical interpretation due to their unique sets of acquired mutations. There is a general scarcity of tools that can i) systematically interrogate cancer genomes in the context of diagnostic, prognostic, and therapeutic biomarkers, ii) prioritize and highlight the most important findings, and iii) present the results in a format accessible to clinical experts. We have developed a stand-alone, open-source software package for somatic variant annotation that integrates a comprehensive set of knowledge resources related to tumor biology and therapeutic biomarkers, both at the gene and variant level. Our application generates a tiered report that will aid the interpretation of individual cancer genomes in a clinical setting.\n\nAvailability and ImplementationThe software is implemented in Python/R, and is freely available through Docker technology. Documentation, example reports, and installation instructions are accessible via the project GitHub page: https://github.com/sigven/pcgr)\n\nContactsigven@ifi.uio.no

bioinformatics

Genome Build Information Is An Essential Part Of Genomic Track Files

Genomic locations are represented as coordinates on a specific genome build version, but the build information is frequently missing when coordinates are provided. It is essential to correctly interpret and analyse the genomic intervals contained in genomic track files. Here, we demonstrate that this crucial metadatum (or rather datum) is often isolated from the genomic track files in public repositories and journal articles, which could be a major time thief. We propose best practices to ensure that genome build version is always carried along with genomic track files. Although not a substitute to the best practices, we also provide a tool to predict the genome build version of genomic track files.

bioinformatics

Hi-C-constrained physical models of human chromosomes recover functionally-related properties of genome organization

Combining genome-wide structural models with phenomenological data is at the forefront of efforts to understand the organizational principles regulating the human genome. Here, we use chromosome-chromosome contact data as knowledge- based constraints for large-scale three-dimensional models of the human diploid genome. The resulting models remain minimally entangled and acquire several functional features that are observed in vivo and that were never used as input for the model. We find, for instance, that gene-rich, active regions are drawn towards the nuclear center, while gene poor and lamina-associated domains are pushed to the periphery. These and other properties persist upon adding local contact constraints, suggesting their compatibility with non-local constraints for the genome organization. The results show that suitable combinations of data analysis and physical modelling can expose the unexpectedly rich functionally-related properties implicit in chromosome-chromosome contact data. Specific directions are suggested for further developments based on combining experimental data analysis and genomic structural modelling.

biophysics