Search bioRxivSearch

bioRxiv · 10.1101/779702

Transparent and Reproducible Research Practices in the Surgical Literature

Abstract

Previous studies have established a baseline of minimal reproducibility in the social science and biomedical literature. Clinical research is especially deficient in factors of reproducibility. Surgical journals contain fewer clinical trials than non-surgical ones, suggesting that it should be easier to reproduce the outcomes of surgical literature. In this study, we evaluated a broad range of indicators related to transparency and reproducibility in a random sample of 300 articles published in surgery-related journals between 2014 and 2018. A minority of our sample made available their materials (2/186, 95% C.I. 0-2.2%), protocols (1/196, 0-1.3%), data (19/196, 6.3-13%), or analysis scripts (0/196, 0-1.9%). Only one study was adequately pre-registered. No studies were explicit replications of previous literature. Most studies (162/292 50-61%) declined to provide a funding statement, and few declared conflicts of interest (22/292, 4.8-11%). Most have not been cited by systematic reviews (183/216, 81-89%) or meta-analyses (188/216, 83-91%), and most were behind a paywall (187/292, 58-70%). The transparency of surgical literature could improve with adherence to baseline standards of reproducibility.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hughes, T., Niemann, A., Tritz, D., Boyer, K., Robbins, H., Vassar, M.. 2019-10-07. Transparent and Reproducible Research Practices in the Surgical Literature. https://doi.org/10.1101/779702

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

Quantifying and contextualizing the impact of bioRxiv preprints through social media audience segmentation

Engagement with scientific manuscripts is frequently facilitated by Twitter and other social media platforms. As such, the demographics of a papers social media audience provide a wealth of information about how scholarly research is transmitted, consumed, and interpreted by online communities. By paying attention to public perceptions of their publications, scientists can learn whether their research is stimulating positive scholarly and public thought. They can also become aware of potentially negative patterns of interest from groups that misinterpret their work in harmful ways, either willfully or unintentionally, and devise strategies for altering their messaging to mitigate these impacts. In this study, we collected 331,696 Twitter posts referencing 1,800 highly tweeted bioRxiv preprints and leveraged topic modeling to infer the characteristics of various communities engaging with each preprint on Twitter. We agnostically learned the characteristics of these audience sectors from keywords each users followers provide in their Twitter biographies. We estimate that 96% of the preprints analyzed are dominated by academic audiences on Twitter, suggesting that social media attention does not always correspond to greater public exposure. We further demonstrate how our audience segmentation method can quantify the level of interest from non-specialist audience sectors such as mental health advocates, dog lovers, video game developers, vegans, bitcoin investors, conspiracy theorists, journalists, religious groups, and political constituencies. Surprisingly, we also found that 10% of the highly tweeted preprints analyzed have sizable (>5%) audience sectors that are associated with right-wing white nationalist communities. Although none of these preprints intentionally espouse any right-wing extremist messages, cases exist where extremist appropriation comprises more than 50% of the tweets referencing a given preprint. These results present unique opportunities for improving and contextualizing research evaluation as well as shedding light on the unavoidable challenges of scientific discourse afforded by social media.

scientific communication and education

No Relationship Between Perceived Health Anomalies and Perceived Experimental Success in Retired Breeder Male Hartley Albino Guinea Pigs

Guinea pigs used in our laboratory for cardiac research sometimes exhibit physical abnormalities. These issues may abate or intensify during the time they are housed in our facility. After using a guinea pig for research, experimentalists note the apparent health of an animal based on visible features and/or abnormal electrophysiology of the heart. There was an existing anecdotal observation that the health of the Guinea Pigs, and subsequently the experimental success rate, had a seasonal variation; therefore we sought to determine if there is a time of year in which our guinea pigs are more likely to be perceived as unhealthy, and whether any determined monthly pattern correlates with an experimentalists ability to complete an experimental protocol. An electronic log was created to record the perceived health of the animal and the ability to complete the experiment successfully. Irregular symptoms included, but were not limited to, severe weight or hair loss and irregularities with the heart found post thoracotomy or during baseline electrophysiological recordings of whole-heart preparations. Animals that did not exhibit significant weight or hair loss, or other ailments were considered "healthy". Overall, our results indicate that there are no monthly variations in perceived Hartley Albino guinea pig health or correlations with experimental completion rates, suggesting mild hair or weight loss that is common when shipping animals may not significantly affect the ability to conduct ex vivo whole-heart electrophysiological studies.

scientific communication and education