Search bioRxivSearch

bioRxiv · 10.1101/2019.12.20.884031

Tracking self-citations in academic publishing

Abstract

Citation metrics have value because they aim to make scientific assessment a level playing field, but urgent transparency-based adjustments are necessary to ensure that measurements yield the most accurate picture of impact and excellence. One problematic area is the handling of self-citations, which are either excluded or inappropriately accounted for when using bibliometric indicators for research evaluation. Here, in favor of openly tracking self-citations we report on self-referencing behavior among various academic disciplines as captured by the curated Clarivate Analytics Web of Science database. Specifically, we examined the behavior of 385,616 authors grouped into 15 subject areas like Biology, Chemistry, Science & Technology, Engineering, and Physics. These authors have published 3,240,973 papers that have accumulated 90,806,462 citations, roughly five percent of which are self-citations. Up until now, very little is known about the buildup of self-citations at the author-level and in field-specific contexts. Our view is that hiding self-citation data is indefensible and needlessly confuses any attempts to understand the bibliometric impact of ones work. Instead we urge academics to embrace visibility of citation data in a community of peers, which relies on nuance and openness rather than curated scorekeeping.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kacem, A., Flatt, J. W., Mayr, P.. 2019-12-23. Tracking self-citations in academic publishing. https://doi.org/10.1101/2019.12.20.884031

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

Scientometric correlates of high-quality reference lists in ecological papers

It is said that the quality of a scientific publication is as good as the science it cites, but the properties of high-quality reference lists have never been numerically quantified. We examined seven numerical characteristics of reference lists of 50,878 primary research articles published in 17 ecological journals between 1997 and 2017. Over this 20-years period, there have been significant changes in reference lists’ properties. On average, more recent ecological papers have longer reference lists, cite more high Impact Factor papers, and fewer non-journal publications. Furthermore, we show that highly cited papers across the ecology literature have longer reference lists, cite more recent and impactful papers, and account for more self-citations. Conversely, the proportion of ‘classic’ papers and non-journal publications cited, as well as the temporal range of the reference list, have no significant influence on articles’ citations. From this analysis, we distill a recipe for crafting impactful reference lists.View Full Text

scientific communication and education

Inflated citations and metrics of journals discontinued from Scopus for publication concerns: the GhoS(t)copus Project

Background Scopus is considered a leading bibliometric database. It contains the largest number of abstracts and articles cited in peer reviewed publications. The journals included in Scopus are periodically re-evaluated to ensure they meet indexing criteria. Afterwards, some journals might be discontinued for publication concerns. Despite their discontinuation, previously published articles remain indexed and continue to be cited. The metrics and characteristics of journals discontinued for publication concerns have yet to be studied. This study aimed (1) to evaluate the main features and citation metrics of journals discontinued from Scopus for publication concerns, before and after discontinuation, and (2) to determine the extent of predatory journals among the discontinued journals.Materials and Methods Eight authors surveyed a list of discontinued journals from Scopus (version July 2019). Data regarding metrics, citations and indexing were extracted from Scopus or other scientific databases, for the journals discontinued for publication concerns.Results A total of 317 journals were evaluated. The mean number of citations per year after discontinuation was significantly higher than before (median of difference 64 citations, p<0.0001), and so was the number of citations per document (median of difference 0.4 citations, p<0.0001). The total number of citations after discontinuation was 607,261, with a median of 713 citations (IQR 254-2056, range 0-19468) per journal. The open access publishing model was declared for 93% (294/317) of the journals, but only nine of these are currently indexed in the Directory of Open Access Journals (DOAJ). Twenty-two percent (72/317) of the journals were included in the Cabell blacklist.Conclusions The citation count of journals discontinued for publication concerns, increases despite discontinuation. Countermeasures should be taken to ensure the validity and reliability of Scopus metrics both at journal- and author-level for the purpose of scientific assessment of publishing.View Full Text

scientific communication and education