Search bioRxiv⌕ Search

bioRxiv · 10.1101/2022.09.03.506492

Statistical learning of large-scale genetic data: How to run a genome-wide association study of gene-expression data using the 1000 GenomesProject data

Abstract

Teaching statistics through engaging applications to contemporary large-scale datasets is essential to attracting students to the field. To this end, we developed a hands-on, week-long workshop for senior high-school or junior undergraduate students, without prior knowledge in statistical genetics but with some basic knowledge in data science, to conduct their own genome-wide association studies (GWAS). The GWAS was performed for open source gene expression data, using publicly-available human genetics data. Assisted by a detailed instruction manual, students were able to obtain [~]1.4 million p-values from a real scientific study, within several days. This early motivation kept students engaged in learning the theories that support their results, including regression, data visualization, results interpretation, and large-scale multiple hypothesis testing. To further their learning motivation by emphasizing the personal connection to this type of data analysis, students were encouraged to make short presentations about how GWAS has provided insights into the genetic basis of diseases that are present in their friends and/or families. The appended open source, step-by-step instruction manual includes descriptions of the datasets used, the software needed, and results from the workshop. Additionally, scripts used in the workshop are archived on Zenodo to further enhance reproducible research and training.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sugolov, A., Emmenegger, E., Sun, L., Paterson, A. D.. 2022-09-06. Statistical learning of large-scale genetic data: How to run a genome-wide association study of gene-expression data using the 1000 GenomesProject data. https://doi.org/10.1101/2022.09.03.506492

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

Where are they now? Academic and career trajectories of national laboratory STEM internship alumni from community colleges, compared to those from universities

It is well-established that participation in technical or research experiences can support university student retention in science, technology, engineering, and mathematics (STEM). However, there is very little published research on the availability and impact of opportunities for community college students, particularly those provided by Department of Energy (DOE) national laboratories. To address this gap, we collected data from students who participated in two DOE sponsored programs spanning from 2009 to 2016: the Community College Internship (CCI) and Science Undergraduate Laboratory Internship (SULI) at Lawrence Berkeley National Laboratory (LBNL). Among CCI alumni, 90% earned a STEM bachelors degree and 88% are on a STEM career pathway. For SULI alumni, 91% earned a STEM bachelors degree and 71% are on a STEM career pathway. Overall, 80% of CCI alumni and 56% of SULI alumni have entered the STEM workforce, 5% of CCI alumni and 11% of SULI alumni are in the health workforce, and 6% of CCI alumni and 13% of SULI alumni are in the non-STEM workforce. Our findings indicate that community college students who participate in STEM professional development activities (such as the CCI program) are likely to complete their academic degrees and pursue STEM careers at rates comparable to those of university students. This investment in providing internships at LBNL for community college students has effectively supported their entry into STEM careers and their desire in pursuing work within the DOE complex in the years following their participation in these programs.

scientific communication and education↗

A Large Language Model-Powered Map of Metabolomics Research

We present a comprehensive map of the metabolomics research landscape, synthesizing insights from over 80,000 publications. Using PubMedBERT, we transformed abstracts into 768-dimensional embeddings that capture the nuanced thematic structure of the field. Dimensionality reduction with t-SNE revealed distinct clusters corresponding to key domains such as analytical chemistry, plant biology, pharmacology, and clinical diagnostics. In addition, a neural topic modeling pipeline refined with GPT-4o mini reclassified the corpus into 20 distinct topics--ranging from "Plant Stress Response Mechanisms" and "NMR Spectroscopy Innovations" to "COVID-19 Metabolomic and Immune Responses." Temporal analyses further highlight trends including the rise of deep learning methods post-2015 and a continued focus on biomarker discovery. Integration of metadata such as publication statistics and sample sizes provide additional context to these evolving research dynamics. An interactive web application (https://metascape.streamlit.app/) enables dynamic exploration of these insights. Overall, this study offers a robust framework for literature synthesis that empowers researchers, clinicians, and policymakers to identify emerging research trajectories and address critical challenges in metabolomics, while also sharing our perspectives on key trends shaping the field.

scientific communication and education↗