Search bioRxiv⌕ Search

Biology subjects

Lasky-Su, J. A.

Publications and source records attributed to Lasky-Su, J. A..

8 recordsLinked to original sources

Quantification of individual dataset contributions to prediction accuracy in cooperative learning

We consider cooperative learning, a recently proposed technique to leverage the predictive power of several datasets (also called data views) to improve the prediction accuracy of an outcome of interest. Cooperative learning uses a Lasso-type penalty to fit several datasets to a given outcome, while a so-called agreement penalty enforces that the predictions made by the individual datasets agree. We are interested in the question of whether the predictive power of each individual dataset can be quantified using cooperative learning. In this work, we demonstrate that certain trace plots, analogously to the ones for the classic Lasso, allow one to quantify the predictive power of each individual dataset. Importantly, this allows one to detect datasets which do not carry any predictive power on the outcome. In an experimental study, we quantify the predictive power of three real datasets in the context of the Childhood Asthma Management Program (CAMP), with the three datasets containing information on epidemiological variables, metabolites, and clinical data, respectively.

bioinformatics↗

Finding Salient Multi-Omic Interactomes Relevant to Multiple Biomedical Outcomes using Graph Ensemble Neural Networks

Although multi-omics integration relevant to patient outcome is typically characterized by an analyte interactome, current multi-omic integration methods either (1) model outcome without directly including associations between analytes, (2) model the interactome without directly evaluating the saliency of the model in the context of outcome, or (3) model outcome in a high-dimensional parameter space not suitable for small sample sizes (which are common in multi-omics studies). We introduce Graph Ensemble Neural Network (GENN), a methodology that learns the interactome most predictive of outcome in a low-dimensional parameter space built on complementary attributes for all possible analyte associations (metafeatures). We show that GENN is robust to noise in measurements using a theoretical model, outperforms the predictive performance of existing methods when evaluated on Tegafur drug response in NCI-60 cancer cell line data, and uncovers potentially novel multi-omic mechanisms driving total serum IgE levels in pediatric asthma and patient survival in glioblastomas.

bioinformatics↗

Genetic Architecture and Analysis Practices of Circulating Metabolites in the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program

Circulating metabolite levels partly reflect the state of human health and diseases and can be impacted by genetic determinants. Hundreds of loci associated with circulating metabolites have been identified; however, most findings focus on predominantly European ancestry or single-study analyses. Leveraging the rich metabolomics resources generated by the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program, we harmonized and accessibly cataloged 1,729 circulating metabolites among 25,058 ancestrally diverse samples. We provided a set of reasonable strategies for outlier and imputation handling to process metabolite data. Following the practical analysis framework, we further performed a genome-wide association analysis on 1,135 selected metabolites using whole genome sequencing data from 16,359 individuals passing the quality control filters, and discovered 1,778 independent loci associated with 667 metabolites. Among 108 novel locus-metabolite pairs, we detected not only novel loci within previously implicated metabolite associated genes but also novel genes (such as GAB3 and VSIG4 located in the X chromosome) that have putative roles in metabolic regulation. In the sex-stratified analysis, we revealed 85 independent locus-metabolite pairs with evidence of sexual dimorphism, including well-known metabolic genes such as FADS2, D2HGDH, SUGP1, UTG2B17, strongly supporting the importance of exploring sex difference in the human metabolome. Taken together, our study depicted the genetic contribution to circulating metabolite levels, providing additional insight into the understanding of human health.

bioinformatics↗

Network Analysis Reveals Protein Modules Associated with Childhood Respiratory Diseases

BackgroundThe first year of life is a period of rapid immune development that can impact health trajectories and the risk of developing respiratory-related diseases, such as asthma, recurrent infections, and eczema. However, the biology underlying subsequent disease development remains unknown. MethodsUsing weighted gene correlation network analysis (WGCNA), we derived modules of highly correlated immune-related proteins in plasma samples from children at age 1 year (N=294) from the Vitamin D Antenatal Asthma Reduction Trial (VDAART). We applied regression analyses to assess relationships between protein modules and development of childhood respiratory diseases up to age 6 years. We then characterized genomic, environmental, and metabolomic factors associated with modules. ResultsWGCNA identified four protein modules at age 1 year associated with incidence of childhood asthma and/or recurrent wheeze (Padj range: 0.02-0.03), respiratory infections (Padj range: 6.3x10-9-2.9x10-6), and eczema (Padj=0.01) by age 6 years; three modules were associated with at least one environmental exposure (Padj range: 2.8x10-10-0.03) and disrupted metabolomic pathway(s) (Padj range: 2.8x10-6-0.04). No genome-wide SNPs were identified as significant genetic risk factors for any protein module. Relationships between protein modules with clinical, environmental, and omic factors were temporally sensitive and could not be recapitulated in protein profiles at age 6 years. ConclusionThese findings suggested protein profiles as early as age 1 year predicted development of respiratory-related diseases through age 6 and were associated with changes in pathways related to amino acid and energy metabolism. These may inform new strategies to identify vulnerable individuals based on immune protein profiling.

bioinformatics↗

Integrative epigenetics and transcriptomics identify aging genes in human blood

Recent epigenome-wide studies have identified a large number of genomic regions that consistently exhibit changes in their methylation status with aging across diverse populations, but the functional consequences of these changes are largely unknown. On the other hand, transcriptomic changes are more easily interpreted than epigenetic alterations, but previously identified age-related gene expression changes have shown limited replicability across populations. Here, we develop an approach that leverages high-resolution multi-omic data for an integrative analysis of epigenetic and transcriptomic age-related changes and identify genomic regions associated with both epigenetic and transcriptomic age-dependent changes in blood. Our results show that these "multi-omic aging genes" in blood are enriched for adaptive immune functions, replicate more robustly across diverse populations and are more strongly associated with aging-related outcomes compared to the genes identified using epigenetic or transcriptomic data alone. These multi-omic aging genes may serve as targets for epigenetic editing to facilitate cellular rejuvenation.

systems biology↗

A genome-wide association study of mass spectrometry proteomics using the Seer Proteograph platform

Genome-wide association studies (GWAS) with proteomics are essential tools for drug discovery. To date, most studies have used affinity proteomics platforms, which have limited discovery to protein panels covered by the available affinity binders. Furthermore, it is not clear to which extent protein epitope changing variants interfere with the detection of protein quantitative trait loci (pQTLs). Mass spectrometry-based (MS) proteomics can overcome some of these limitations. Here we report a GWAS using the MS-based Seer ProteographTM platform with blood samples from a discovery cohort of 1,260 American participants and a replication in 325 individuals from Asia, with diverse ethnic backgrounds. We analysed 1,980 proteins quantified in at least 80% of the samples, out of 5,753 proteins quantified across the discovery cohort. We identified 252 and replicated 90 pQTLs, where 30 of the replicated pQTLs have not been reported before. We further investigated 200 of the strongest associated cis-pQTLs previously identified using the SOMAscan and the Olink platforms and found that up to one third of the affinity proteomics pQTLs may be affected by epitope effects, while another third were confirmed by MS proteomics to be consistent with the hypothesis that genetic variants induce changes in protein expression. The present study demonstrates the complementarity of the different proteomics approaches and reports pQTLs not accessible to affinity proteomics, suggesting that many more pQTLs remain to be discovered using MS-based platforms. Graphical AbstractSummarizing the approach taken to identify potential epitope effects. O_FIG O_LINKSMALLFIG WIDTH=166 HEIGHT=200 SRC="FIGDIR/small/596028v1_ufig1.gif" ALT="Figure 1"> View larger version (58K): org.highwire.dtl.DTLVardef@34d11aorg.highwire.dtl.DTLVardef@18c0211org.highwire.dtl.DTLVardef@dbb003org.highwire.dtl.DTLVardef@1009fa7_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗

OMICmAge: An integrative multi-omics approach to quantify biological age with electronic medical records

Biological aging is a multifactorial process involving complex interactions of cellular and biochemical processes that is reflected in omic profiles. Using common clinical laboratory measures in ~30,000 individuals from the MGB-Biobank, we developed a robust, predictive biological aging phenotype, EMRAge, that balances clinical biomarkers with overall mortality risk and can be broadly recapitulated across EMRs. We then applied elastic-net regression to model EMRAge with DNA-methylation (DNAm) and multiple omics, generating DNAmEMRAge and OMICmAge, respectively. Both biomarkers demonstrated strong associations with chronic diseases and mortality that outperform current biomarkers across our discovery (MGB-ABC, n=3,451) and validation (TruDiagnostic, n=12,666) cohorts. Through the use of epigenetic biomarker proxies, OMICmAge has the unique advantage of expanding the predictive search space to include epigenomic, proteomic, metabolomic, and clinical data while distilling this in a measure with DNAm alone, providing opportunities to identify clinically-relevant interconnections central to the aging process.

bioinformatics↗

A meta-analysis of immune cell fractions at high resolution reveals novel associations with common phenotypes and health outcomes

AbstractO_ST_ABSBackgroundC_ST_ABSChanges in cell-type composition of complex tissues are associated with a wide range of diseases, environmental risk factors and may be causally implicated in disease development and progression. However, these shifts in cell-type fractions are often of a low magnitude, or involve similar cell-subtypes, making their reliable identification challenging. DNA methylation profiling in a tissue like blood is a promising approach to discover shifts in cell-type abundance, yet studies have only been performed at a relatively low cellular resolution and in isolation, limiting their power to detect these shifts in tissue composition. MethodsHere we derive a DNA methylation reference matrix for 12 immune cell-types in human blood and extensively validate it with flow-cytometric count data and in whole-genome bisulfite sequencing data of sorted cells. Using this reference matrix and Stouffers method, we perform a meta-analysis encompassing 25,629 blood samples from 22 different cohorts, to comprehensively map associations between the 12 immune-cell fractions and common phenotypes, including health outcomes. ResultsOur meta-analysis reveals many associations with age, sex, smoking and obesity, many of which we validate with single-cell RNA-sequencing. We discover that T-regulatory and naive T-cell subsets are higher in women compared to men, whilst the reverse is true for monocyte, natural killer, basophil and eosinophil fractions. In a large subset encompassing 5000 individuals we find associations with stress, exercise, sleep and health outcomes, revealing that naive T-cell and B-cell fractions are associated with a reduced risk of all-cause mortality independently of age, sex, race, smoking, obesity and alcohol consumption. We find that decreased natural killer cell counts are associated with smoking, obesity and stress levels, whilst an increased count correlates with exercise, sleep and a reduced risk of all-cause mortality. ConclusionsThis work derives and extensively validates a high resolution DNAm reference matrix for blood, and uses it to generate a comprehensive map of associations between immune cell fractions and common phenotypes, including health outcomes. AvailabilityThe 12 immune cell-type DNAm reference matrices for Illumina 850k and 450k beadarrays alongside tools for cell-type fraction estimation are freely available from our EpiDISH Bioconductor R-package http://www.bioconductor.org/packages/devel/bioc/html/EpiDISH.html

bioinformatics↗