Search bioRxivSearch

Biology subjects

Stegle, O.

Publications and source records attributed to Stegle, O..

23 records · Page 2Linked to original sources

Joint genetic analysis using variant sets reveals polygenic gene-context interactions

Joint genetic models for multiple traits have helped to enhance association analyses. Most existing multi-trait models have been designed to increase power for detecting associations, whereas the analysis of interactions has received considerably less attention. Here, we propose iSet, a method based on linear mixed models to test for interactions between sets of variants and environmental states or other contexts. Our model generalizes previous interaction tests and in particular provides a test for local differences in the genetic architecture between contexts. We first use simulations to validate iSet before applying the model to the analysis of genotype-environment interactions in an eQTL study. Our model retrieves a larger number of interactions than alternative methods and reveals that up to 20% of cases show context-specific configurations of causal variants. Finally, we apply iSet to test for sub-group specific genetic effects in human lipid levels in a large human cohort, where we identify a gene-sex interaction for C-reactive protein that is missed by alternative methods.\n\nAuthor summaryGenetic effects on phenotypes can depend on external contexts, including environment. Statistical tests for identifying such interactions are important to understand how individual genetic variants may act in different contexts. Interaction effects can either be studied using measurements of a given phenotype in different contexts, under the same genetic backgrounds, or by stratifying a population into subgroups. Here, we derive a method based on linear mixed models that can be applied to both of these designs. iSet enables testing for interactions between context and sets of variants, and accounts for polygenic effects. We validate our model using simulations, before applying it to the genetic analysis of gene expression studies and genome-wide association studies of human blood lipid levels. We find that modeling interactions with variant sets offers increased power, thereby uncovering interactions that cannot be detected by alternative methods.

genetics

Genomic determinants of protein abundance variation in colorectal cancer cells

Assessing the extent to which genomic alterations compromise the integrity of the proteome is fundamental in identifying the mechanisms that shape cancer heterogeneity. We have used isobaric labelling and tribrid mass spectrometry to characterize the proteomic landscapes of 50 colorectal cancer cell lines and to decipher the relationships between genomic and proteomic variation. The robust quantification of 12,000 proteins and 27,000 phosphopeptides revealed how protein symbiosis translates to a co-variome which is subjected to a hierarchical order and exposes the collateral effects of somatic mutations on protein complexes. Targeted depletion of key chromatin modifiers confirmed the transmission of variation and the directionality as characteristics of protein interactions. Protein level variation was leveraged to build drug response predictive models towards a better understanding of pharmacoproteomic interactions in colorectal cancer. Overall, we provide a deep integrative view of the molecular structure underlying the variation of colorectal cancer cells.\n\nHighlightsO_LIThe cancer cell functional \"co-variome\" is a strong attribute of the proteome.\nC_LIO_LIMutations can have a direct impact on protein levels of chromatin modifiers.\nC_LIO_LITransmission of genomic variation is a characteristic of protein interactions.\nC_LIO_LIPharmacoproteomic models are strong predictors of response to DNA damaging agents.\nC_LI\n\nAbbreviations

systems biology

Scalable latent-factor models applied to single-cell RNA-seq data separate biological drivers from confounding effects

Single-cell RNA-sequencing (scRNA-seq) allows heterogeneity in gene expression levels to be studied in large populations of cells. Such heterogeneity can arise from both technical and biological factors, thus making decomposing sources of variation extremely difficult. We here describe a computationally efficient model that uses prior pathway annotation to guide inference of the biological drivers underpinning the heterogeneity. Moreover, we jointly update and improve gene set annotation and infer factors explaining variability that fall outside the existing annotation. We validate our method using simulations, which demonstrate both its accuracy and its ability to scale to large datasets with up to 100,000 cells. Moreover, through applications to real data we show that our model can robustly decompose scRNA-seq datasets into interpretable components and facilitate the identification of novel sub-populations.

bioinformatics

Genomic Rearrangements Considered as Quantitative Traits

To understand the population genetics of structural variants (SVs), and their effects on phenotypes, we developed an approach to mapping SVs, particularly transpositions, segregating in a sequenced population, and which avoids calling SVs directly. The evidence for a potential SV at a locus is indicated by variation in the counts of short-reads that map anomalously to the locus. These SV traits are treated as quantitative traits and mapped genetically, analogously to a gene expression study. Association between an SV trait at one locus and genotypes at a distant locus indicate the origin and target of a transposition. Using ultra-low-coverage (0.3x) population sequence data from 488 recombinant inbred Arabidopsis genomes, we identified 6,502 segregating SVs. Remarkably, 25% of these were transpositions. Whilst many SVs cannot be delineated precisely, PCR validated 83% of 44 predicted transposition breakpoints. We show that specific SVs may be causative for quantitative trait loci for germination, fungal disease resistance and other phenotypes. Further we show that the phenotypic heritability attributable to sequence anomalies differs from, and in the case of time to germination and bolting, exceeds that due to standard genetic variation. Gene expression within SVs is also more likely to be silenced or dysregulated. This approach is generally applicable to large populations sequenced at low-coverage, and complements the prevalent strategy of SV discovery in fewer individuals sequenced at high coverage.

genetics

Genome-wide Analysis of Differential Transcriptional and Epigenetic Variability Across Human Immune Cell Types

BackgroundA healthy immune system requires immune cells that adapt rapidly to environmental challenges. This phenotypic plasticity can be mediated by transcriptional and epigenetic variability.\n\nResultsWe applied a novel analytical approach to measure and compare transcriptional and epigenetic variability genome-wide across CD14+CD16- monocytes, CD66b+CD16+ neutrophils, and CD4+CD45RA+ naive T cells, from the same 125 healthy individuals. We discovered substantially increased variability in neutrophils compared to monocytes and T cells. In neutrophils, genes with hypervariable expression were found to be implicated in key immune pathways and to associate with cellular properties and environmental exposure. We also observed increased sex-specific gene expression differences in neutrophils. Neutrophil-specific DNA methylation hypervariable sites were enriched at dynamic chromatin regions and active enhancers.\n\nConclusionsOur data highlight the importance of transcriptional and epigenetic variability for the neutrophils key role as the first responders to inflammatory stimuli. We provide a resource to enable further functional studies into the plasticity of immune cells, which can be accessed from: http://blueprint-dev.bioinfo.cnio.es/WP10/hypervariability.

molecular biology