Search bioRxiv⌕ Search

Biology subjects

Gentry, A. E.

Publications and source records attributed to Gentry, A. E..

3 recordsLinked to original sources

Case-only rare variant analysis of severe alcohol dependence (AD) using a multivariate hierarchical gene clustering approach.

BackgroundVariation in genes involved in ethanol metabolism has been shown to influence risk for alcohol dependence (AD) including protective loss of function alleles in ethanol metabolizing genes. We therefore hypothesized that people with severe AD would exhibit different patterns of rare functional variation in genes with strong prior evidence for influencing ethanol metabolism and response when compared to genes not meeting these criteria. ObjectiveLeverage a novel case only design and Whole Exome Sequencing (WES) of severe AD cases from the island of Ireland to quantify differences in functional variation between genes associated with ethanol metabolism and/or response and their matched control genes. MethodsFirst, three sets of ethanol related genes were identified including those a) involved in alcohol metabolism in humans b) showing altered expression in mouse brain after alcohol exposure, and altering ethanol behavioral responses in invertebrate models. These genes of interest (GOI) sets were matched to control gene sets using multivariate hierarchical clustering of gene-level summary features from gnomAD. Using WES data from 190 individuals with severe AD, GOI were compared to matched control genes using logistic regression to detect aggregate differences in abundance of loss of function, missense, and synonymous variants, respectively. ResultsThree non-independent sets of 10, 117, and 359 genes were queried against control gene sets of 139, 1522, and 3360 matched genes, respectively. Significant differences were not detected in the number of functional variants in the primary set of ethanol-metabolizing genes. In both the mouse expression and invertebrate sets, we observed an increased number of synonymous variants in GOI over matched control genes. Post-hoc simulations showed the estimated effects sizes observed are unlikely to be under-estimated. ConclusionThe proposed method demonstrates a computationally viable and statistically appropriate approach for genetic analysis of case-only data for hypothesized gene sets supported by empirical evidence.

genetics↗

Determining the stability of genome-wide factors in BMI between ages 40 to 69 years.

Genome-wide association studies (GWAS) have successfully identified common variants associated with BMI. However, the stability of genetic variation influencing BMI from midlife and beyond is unknown. By analyzing BMI data collected from 165,717 men and 193,073 women from the UKBiobank, we performed BMI GWAS on six independent five-year age intervals between 40 and 73 years. We then applied genomic structural equation modeling (gSEM) to test competing hypotheses regarding the stability of genetic effects for BMI. LDSR genetic correlations between BMI assessed between ages 40 to 73 were all very high and ranged 0.89 to 1.00. Genomic structural equation modeling revealed that genetic variance in BMI at each age interval could not be explained by the accumulation of any age-specific genetic influences or autoregressive processes. Instead, a common set of stable genetic influences appears to underpin variation in BMI from middle to early old age in men and women alike.

genetics↗

Missingness Adapted Group Informed Clustered (MAGIC)-LASSO: A novel paradigm for prediction in data with widespread non-random missingness

The availability of large-scale biobanks linking rich phenotypes and biological measures is a powerful opportunity for scientific discovery. However, real-world collections frequently have extensive non-random missingness. While missing data prediction is possible, performance is significantly impaired by block-wise missingness inherent to many biobanks. To address this, we developed Missingness Adapted Group-wise Informed Clustered (MAGIC)-LASSO which performs hierarchical clustering of variables based on missingness followed by sequential Group LASSO within clusters. Variables are pre-filtered for missingness and balance between training and target sets with final models built using stepwise inclusion of features ranked by completeness. This research has been conducted using the UK Biobank (n>500k) to predict unmeasured Alcohol Use Disorders Identification Test (AUDIT) scores. The phenotypic correlation between measured and predicted total score was 0.67 while genetic correlations between independent subjects was high >0.86, demonstrating the method has significant accuracy and utility.

genetics↗