Search bioRxivSearch

bioRxiv · 10.1101/530691

Assessing metadata and curation quality: a case study from the development of a third-party curation service at Springer Nature

Abstract

Since 2017, the publisher Springer Nature has provided an optional Research Data Support service to help researchers deposit and curate data that support their peer-reviewed publications. This service builds on a Research Data Helpdesk, which since 2016 has provided support to authors and editors who need advice on the options available for sharing their research data. In this paper we describe a short project which aimed to facilitate an objective assessment of metadata quality, undertaken during the development of a third-party curation service for researchers (Research Data Support). We provide details on the single-blind user-testing which was undertaken, and the results gathered during this experiment. We also briefly describe the curation services which have been developed and introduced following an initial period of testing and piloting. This paper will be presented at the International Digital Curation Conference 2019, and has been submitted to the International Journal of Digital curation.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Grant, R., Smith, G., Hrynaszkiewicz, I.. 2019-01-27. Assessing metadata and curation quality: a case study from the development of a third-party curation service at Springer Nature. https://doi.org/10.1101/530691

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

A decade of empirical research on research integrity: what have we (not) looked at?

In the past decades, increasing visibility of research misconduct scandals created momentum for discourses on research integrity to such an extent that the topic became a field of research itself. Yet, a comprehensive overview of research in the field is still missing. Here we describe methods, trends, publishing patterns, and impact of a decade of research on research integrity.\n\nTo give a comprehensive overview of research on research integrity, we first systematically searched SCOPUS, Web of Science, and PubMed for relevant articles published in English between 2005 and 2015. We then classified each relevant article according to its topic, several methodological characteristics, its general focus and findings, and its citation impact.\n\nWe included 986 articles in our analysis. We found that the body of literature on research integrity is growing in importance, and that the field is still largely dominated by non-empirical publications. Within the bulk of empirical records (N=342), researchers and students are most often studied, but other actors and the social context in which they interact, seem to be overlooked. The few empirical articles that examined determinants of misconduct found that problems from the research system (e.g., pressure, competition) were most likely to cause inadequate research practices. Paradoxically, the majority of empirical articles proposing approaches to foster integrity focused on techniques to build researchers awareness and compliance rather than techniques to change the research system.\n\nOur review highlights the areas, methods, and actors favoured in research on research integrity, and reveals a few blindspots. Involving non-researchers and reconnecting what is known to the approaches investigated may be the first step to generate executable knowledge that will allow us to increase the success of future approaches.\n\nA word from the authorsWe find important to mention that this manuscript underwent peer review and was rejected from the following journals:\n\nO_LIPLOS ONE\nO_LISubmitted 19th December 2017\nC_LIO_LIPeer-review and rejection received 26th June 2018.\nC_LI\nC_LIO_LIJournal of Science and Engineering Ethics (JSEE)\nO_LISubmitted 18th August 2018\nC_LIO_LIPeer-review response with major revision request received 9th September 2019\nC_LIO_LIRevision submitted 24th October 218\nC_LIO_LIRejection received 27th December 2018\nC_LI\nC_LI\n\nWe regret not having submitted this preprint before our first submission. Nonetheless, now after over one year in submission processes, we thought that we should make this manuscript and its data available as a pre-print before undergoing further submissions.\n\nIn order to promote transparency however, we asked both journal whether anonymous reviews could be added alongside this pre-print to ensure that readers are informed of the issues that disqualified our manuscript.\n\nPLOS ONE agreed for us to share the anonymous reviews which are now available -- together with our itemized changes and responses -- in the Online Resource 5 - Peer Review Report. We thank the editors of the Journal of Science and Engineering Ethics and the integrity team of Springer Nature for thoroughly discussing our request, but unfortunately, given the closed peer review policy at Springer Nature, we were unable to provide information about the peer review from the Journal of Science and Engineering Ethics.\n\nWe advise our readers to look at the peer-review and be aware of the challenges and limitations attached with our work. Of course, we welcome comments and contributions to make our work better.\n\nSincerely,\n\nNoemie Aubert Bonn and Wim Pinxten

scientific communication and education

Hypothesis, analysis and synthesis: it's all Greek to me!

The linguistic foundations of science and technology have relied on a range of terms many of which are borrowed from ancient languages, a known but little researched fact from a statistical perspective. Precise definitions and novel concepts are often crafted with those -- frequently used -- terms, yet their etymology from Greek or Latin might not always be fully appreciated. Herein, we demonstrate that frequently used terms span almost the entire PubMed(R) database, while a handful of terms of Greek origin retrieve 80% of all entries. We argue that the etymology of those critical terms needs to be fully grasped to ensure correct use, in conjunction with other concepts. We further propose a number of terms for genomics, using prepositions that can accurately define subtle sub-disciplines of this ever-expanding field. Finally, we invite commentary by both the science community and the humanities, for possible adoption of suggested terms, not least to avoid inaccurate usage or inappropriate notions that may compromise clarity of meaning.

scientific communication and education