Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.09.16.613258

Updated science-wide author databases of standardized citation indicators including retraction data

Abstract

Citation metrics are widely used in research appraisal, but they provide incomplete views of scientists impact and research track record. Other indicators of research practices should be linked to citation data. We have updated a Scopus-based database of highly-cited scientists (top-2% in each scientific subfield according to a composite citation indicator) to incorporate retraction data. Using data from the Retraction Watch database (RWDB), retraction records were linked to Scopus citation data. Of 55,237 items in RWDB as of August 15, 2024, we excluded non-retractions, retractions clearly not due to any author error, retractions where the paper had been republished, and items not linkable to Scopus records. Eventually 39,468 eligible retractions were linked to Scopus. Among 217,097 top-cited scientists in career-long impact and 223,152 in single recent year (2023) impact, 7,083 (3.3%) and 8,747 (4.0%), respectively, had at least one retraction. Scientists with retracted publications had younger publication age, higher self-citation rates, and larger publication volume than those without any retracted publications. Retractions were more common in the life sciences and rare or nonexistent in several other disciplines. In several developing countries, very high proportions of top-cited scientists had retractions (highest in Senegal (66.7%), Ecuador (28.6%) and Pakistan (27.8%) in career-long citation impact lists). Variability in retraction rates across fields and countries suggests differences in research practices, scrutiny, and ease of retraction. Addition of retraction data enhances the granularity of top-cited scientists profiles, aiding in responsible research evaluation. However, caution is needed when interpreting retractions, as they do not always signify misconduct; further analysis on a case-by-case basis is essential. The database should hopefully provide a resource for meta-research and deeper insights into scientific practices.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ioannidis, J., Pezzullo, A. M., Cristiano, A., Boccia, S., Baas, J.. 2024-09-17. Updated science-wide author databases of standardized citation indicators including retraction data. https://doi.org/10.1101/2024.09.16.613258

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

The currency of research access: How undergraduates leverage social capital to gain research experience

BackgroundScience students who engage in undergraduate research experiences (UREs) benefit in numerous ways, including persisting in science at a higher rate compared to students who do not participate in UREs. However, UREs are a limited commodity and competition for access to these opportunities necessitates further investigation into why certain students succeed in accessing UREs, while others do not. Social capital, or the resources that students extract from their relationships with others, may play a key role in determining who engages in UREs. To begin to address this knowledge gap, we conducted semi-structured interviews with students who had recently started UREs (n=21). Informed by Lins conceptual definition of social capital, we qualitatively analyzed these interviews to characterize the social capital that science undergraduates found useful for accessing research. ResultsStudents described leveraging ten unique forms of social capital when accessing UREs that aligned with Lins conceptualization. Specifically, students described utilizing social capital to garner information about available UREs and how to best navigate the path to accessing them, to reinforce their confidence in pursuing UREs, to influence the views of opportunity holders, and to serve as social credentials lending credibility to their aptitude for research. Although students detailed how faculty served as sources of all of these forms of capital, they also described advisors, employers, peers, and family as influential sources. Furthermore, their institutions served as a source of capital, and students themselves engaged in a variety of proactive behaviors to access UREs. ConclusionsHere we describe the forms and sources of social capital that science undergraduates use to access research, which can be used as an operational definition of the construct. Such a definition is necessary for future research aimed at measuring social capital for undergraduate research and identifying its antecedents, correlates, and consequences. We also describe how institutions can serve as sources of capital, and how students proactive behaviors play a role in their pursuit of UREs. This work provides an important starting point for determining the influence of social capital in accessing UREs and further broadening access to research in the sciences.

scientific communication and education↗

Chromatin profiling for everyone: FFPE-CUTAC for the theory and practice of modern molecular biology

In 2025, together with the Fred Hutch Summer Undergraduate Research and Summer High School Internship programs, we developed and implemented a laboratory genomics research experience to introduce students to modern molecular biology techniques and bioinformatics. The course centered around using a new method we had developed in 2023 that uses readily available fixed tissue sections on glass slides. Students performed a series of steps to tagment genomic locations of RNA Polymerase II and then used PCR to enrich libraries for next-generation sequencing in a core facility. Students then visualized their data in genomic browser tracks and assessed the results. At the end of the summer, students prepared and presented their work and experiences in seminar format to their cohorts. Overall, the technical simplicity of on-slide chromatin profiling introduced the students to laboratory practice and current techniques in genomics, bioinformatics, and medical sciences.

scientific communication and education↗