bioRxiv · 10.64898/2026.09.24.754051
Large language model-based bibliometric evaluation of population descriptors in human genetics
Abstract
As the use of population descriptors such as race, ethnicity, and ancestry have become increasingly common in modern genetics research, there have been growing calls to critically examine their use. Most notably, in 2023, the National Academies of Science, Engineering, and Medicine (NASEM) published a report titled Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field, which included eight specific and actionable recommendations for researchers to implement the ethical and accurate use of population descriptors in genetic research. Here, we use the 2023 NASEM report as a benchmark to analyze the use of population descriptors in genome-wide association studies (GWAS). We develop a general toolkit for large language model-based bibliometrics, operationalize the report's recommendations into an evaluation framework, and apply this framework to evaluate all 4,007 papers from the GWAS Catalog published between 2007 and 2025 with full text available on PubMedCentral. We find significant improvements in adherence to NASEM report recommendations over time. However, most improvements predate the publication of the NASEM report itself, suggesting the report functioned primarily as a synthesis of existing best practices rather than a catalyst for change. We conclude by highlighting opportunities for growth in the field of human genetics.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gogarten, S. M., Martschenko, D. O., Patel, R., Sinnott-Armstrong, N.. 2026-09-28. Large language model-based bibliometric evaluation of population descriptors in human genetics. https://doi.org/10.64898/2026.09.24.754051
Cite the original work for its findings. Save a collection to share your selection of sources.