Search bioRxivSearch

Biology subjects

Beyer, A.

Publications and source records attributed to Beyer, A..

3 recordsLinked to original sources

Detection of epistatic interactions with Random Forest

In order to elucidate the influence of genetic factors on phenotype variation, non-additive genetic interactions (i.e., epistasis) have to be taken into account. However, there is a lack of methods that can reliably detect such interactions, especially for quantitative traits. Random Forest was previously recognized as a powerful tool to identify the genetic variants that regulate trait variation, mainly due to its ability to take epistasis into account. However, although it can account for interactions, it does not specifically detect them. Therefore, we propose three approaches that extract interactions from a Random Forest by testing for specific signatures that arise from interactions, which we termed paired selection frequency, split asymmetry, and selection asymmetry. Since they complement each other for different epistasis types, an ensemble method that combines the three approaches was also created. We evaluated our approaches on multiple simulated scenarios and two different real datasets from different Saccharomyces cerevisiae crosses. We compared them to the commonly used exhaustive pair-wise linear model approach, as well as several two-stage approaches, where loci are pre-selected prior to interaction testing. The Random Forest-based methods presented here generally outperformed the other methods at identifying meaningful genetic interactions both in simulated and real data. Further examination of the results for the simulated and real datasets established how interactions are extracted from the Random Forest, and explained the performance differences between the methods. Thus, the approaches presented here extend the applicability of Random Forest for the genetic mapping of biological traits.\n\nAuthor summaryThe genetic mechanisms underlying biological traits are often complex, involving the effects of multiple genetic variants. Interactions between these variants, also called epistasis, are also common. The machine learning algorithm Random Forest can be used to study genotype-phenotype relationships, by using genetic variants to predict the phenotype. One of Random Forests strengths is its ability to implicitly model interactions. However, Random Forest does not give any information about which predictors specifically interact, i.e. which variants are in epistasis.\n\nHere, we developed three approaches that identify interactions in a Random Forest. We demonstrated their ability to detect genetic interactions using simulations and real data from Saccharomyces cerevisiae. Our Random Forest-based methods generally outperformed several other commonly used approaches at detecting epistasis.\n\nThis study contributes to the long-standing problem of extracting information about the underlying model from a Random Forest. Since Random Forest has many applications outside of genetic association, this work represents a valuable contribution to not only genotype-phenotype mapping research, but also other scientific applications where interactions between predictors in a Random Forest might be of interest.

bioinformatics

Multi-region proteome analysis quantifies spatial heterogeneity of prostate tissue biomarkers

Many tumors are characterized by large genomic heterogeneity and it remains unclear to what extent this impacts on protein biomarker discovery. Here, we quantified proteome intra-tissue heterogeneity (ITH) based on a multi-region analysis of 30 biopsy-scale prostate tissues using pressure cycling technology and SWATH mass spectrometry. We quantified 8,248 proteins and analyzed the ITH of 3,700 proteins. The level of ITH varied significantly depending on proteins and tissue types. Benign tissues exhibited generally more complex ITH patterns than malignant tissues. Spatial variability of ten prostate biomarkers was further validated by immunohistochemistry in an independent cohort (n=83) using tissue microarrays. PSA was preferentially variable in benign prostatic hyperplasia, while GDF15 substantially varied in prostate adenocarcinomas. Further, we found that DNA repair pathways exhibited a high degree of variability in tumorous tissues, which may contribute to the genetic heterogeneity of tumors. This study conceptually adds a new perspective to protein biomarker discovery by quantifying spatial proteome variation and it demonstrates the feasibility by exploiting recent technological progress.

systems biology

Individual nephron proteomes connect morphology and function in proteinuric kidney disease

In diseases of many parenchymatous organs, heterogenous detoriation of individual functional units determines the clinical prognosis. However, the molecular characterization of these subunits remains a technological challenge that needs to be addressed in order to better understand pathological mechanisms. Sclerotic and proteinuric glomerular kidney disease is a frequent and heterogeneous disease which affects a fraction of nephrons, glomeruli and draining tubules, to variable extents, and for which no treatment exists. Here, we developed and applied an antibody-independent methodology to investigate heterogeneity of individual nephron segment proteomes from mice with proteinuric kidney disease. This \"one-segment-one-proteome-approach\" defines mechanistic connections between upstream (glomerular) and downstream (tubular) nephron segment populations. In single glomeruli from two different mouse models of sclerotic glomerular disease, we identified a coherent protein expression module consisting of extracellular matrix protein deposition (reflecting glomerular sclerosis), glomerular albumin (reflecting proteinuria) and LAMP1, a lysosomal protein. This module was associated with a loss of podocyte marker proteins. In an attempt to target this protein co-expression module, genetic ablation of LAMP1-correlated lysosomal proteases in mice could ameliorate glomerular damage. Furthermore, individual glomeruli from patients with genetic sclerotic and non-sclerotic proteinuric diseases demonstrated increased abundance of lysosomal proteins, in combination with a decreased abundance of the mutated gene products. Therefore, increased glomerular lysosomal load is a conserved key mechanism in proteinuric kidney diseases, and the technology applied here can be implemented to address heterogeneous pathophysiology in a variety of diseases at a sub-biopsy scale

physiology