Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.04.14.718420

Community needs for FAIR pathogen data

Abstract

BackgroundDatasets related to infectious diseases are essential for public health decision-making, yet their reuse remains limited by persistent barriers to data sharing and integration. Achieving data that are Findable, Accessible, Interoperable, and Reusable (FAIR) is widely recognized as essential for accelerating scientific discovery and enabling coordinated responses to emerging threats, but the needs of the global pathogen data community have not been systematically characterized. AimThis study, conducted by the Pathogen Data Network (PDN), aims to identify infrastructural and educational priorities among stakeholders working with infectious disease-related data in order to guide community-responsive support for data sharing and interoperability. MethodsA cross-sectional stakeholder survey was disseminated to a well-defined expert population within PDN networks and via open professional channels. A total of 136 responses from researchers, healthcare professionals, bioinformaticians, and educators were analyzed descriptively to identify prioritized barriers, training needs, and preferred support mechanisms. ResultsRespondents consistently identified structural constraints as the primary impediments to effective data use, including limited funding (74%), data-aggregation challenges (68%), and a shortage of skilled personnel (52%). Respondents identified bioinformatics for infectious disease research (68%) as the highest priority for training, followed by guidance on using the integrated pathogen data and tools portal provided by the PDN, the Pathogens Portal (51%). The Pathogens Portal was also ranked as the most essential PDN resource (72%). Preferred training formats included virtual short courses (68%) and webinars (66%). Notably, while researchers emphasized technical subjects like machine learning, educators prioritized foundational case studies. ConclusionThese findings provide an evidence-based diagnostic of community needs and suggest that barriers to FAIR pathogen data are predominantly systemic rather than purely technological. The survey framework and openly available dataset offer a reusable template for assessing needs in other communities and regions. By aligning training, infrastructure development, and outreach with empirically identified priorities, organizations supporting infectious disease research can strengthen the interoperability and reuse of data and establish a benchmark for future community-driven improvements.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

van Geest, G., Thomas-Lopez, D., Feitzinger, A. A., Weissgold, L. A., Halabi, S., Cuesta, I., Hjerde, E., Gurwitz, K. T., Arora, N., Neves, A., Palagi, P. M., Williams, J. J.. 2026-04-15. Community needs for FAIR pathogen data. https://doi.org/10.64898/2026.04.14.718420

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Perceived Risk and Barriers to Open and Responsible Research Across Fifteen UK Universities

Open research practices are increasingly promoted to improve research transparency, reproducibility, accessibility, and societal impact. Despite growing support from research funders, institutions, and policy initiatives, adoption remains uneven across disciplines and research communities. This study examined perceived risks and barriers associated with 14 FORRT Guideline areas using qualitative responses from the UK Reproducibility Network Open and Transparent Research Practices Survey (N = 2,567), conducted across 15 UK higher education institutions. Free-text responses describing risks and barriers were analysed using inductive thematic analysis. A total of 3,951 relevant comments generated 3,433 coded references to barriers and risks. Three interconnected clusters emerged. Individual concerns included lack of motivation, fear of losing intellectual credit, and concerns about exposing mistakes and criticism. Systemic and institutional barriers included lack of time and resources, inadequate infrastructure, insufficient training and support, lack of incentives and recognition, unclear guidance, disciplinary and methodological challenges, and tensions between open practices and intellectual property requirements. Ethical and quality-related concerns included risks to participant confidentiality and privacy, challenges associated with sensitive data, concerns about inappropriate application of open research practices across different research traditions, perceived impacts on research quality and innovation, and potential effects on public trust in research. Lack of time and resources was the most frequently reported barrier. Across responses, barriers were commonly described as interdependent, with shortcomings in funding, infrastructure, institutional support, incentives, and training reinforcing one another. Respondents largely supported the principles underpinning open research but highlighted substantial practical, professional, and ethical challenges to implementation. These findings suggest that increasing open research adoption requires more than policy mandates or awareness-raising activities. Sustainable uptake will depend on aligned incentives, adequate infrastructure and support, recognition of methodological diversity, and approaches that enable openness to be implemented responsibly across different research contexts.

scientific communication and education↗

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

Undergraduate Biophysical Chemistry Series: Teaching through a Combination of a Purpose-built Textbook, Research-derived Biomolecular Samples and Computer Labs

Biophysics is a rapidly advancing field with an incredible breadth of topics. Thus, undergraduate biophysics instructors have to strategize and decide what topics they will cover in their courses. Educational institutions utilize a variety of biophysics textbooks. A common deficiency of each of the existing texts is that it serves well a given set of topics (theory, illustrations, practice problems) and leaves out other areas. A typical example includes good theory and problems for thermodynamics and kinetics while presenting molecular dynamics and various spectroscopic methods in a lacking or outdated way. The authors of this manuscript teach a capstone Biophysical Chemistry three-quarter series (Western Washington University/WWU, Bellingham, WA) which ideally should resonate with the general and major-specific courses the students take within their major at WWU. To achieve this goal and to enrich the traditional lecture-based delivery, the instructors have developed and brought together key pedagogical elements: purpose-built online textbook with a uniform structure of the academic content and practice problems, a study sample (oligopeptide) of biophysical significance with a growing set of experimental and computational data and student-centric in-class activities including computer labs. Our Biophysical series emphasizes concepts and methods of computational structural biology (Molecular Dynamics) and spectroscopic approaches (IR, UV and NMR). Here we describe the details of our integrative approach, summarize key outcomes and chart ways to advance the biophysical chemistry series further. Our textbook can be found through LibreText.

scientific communication and education↗