Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.04.02.645026

Estimating the replicability of Brazilian biomedical science

Abstract

Concerns over the replicability and reproducibility of published research have grown in many research fields, but empirical data to inform policies are still scarce. Biomedical research in Brazil expanded rapidly over the last three decades, with no systematic assessment of the replicability of its findings. With this in mind, we set up the Brazilian Reproducibility Initiative, a multicenter replication of published experiments from Brazilian science using three common experimental methods: the MTT assay, the reverse transcription polymerase chain reaction (RT-PCR) and the elevated plus maze (EPM). A total of 56 laboratories performed 143 replications of 56 experiments; of these, 90 replications of 45 experiments were considered valid by an independent committee. Replication rates for these experiments varied between 20 and 44% according to five predefined criteria. In median terms, ratios between group means were 58% lower in replications than in original experiments, while coefficients of variation were 82% higher. Effect size decrease was smaller for MTT experiments, original results with less variability and those considered more challenging to replicate, while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original articles last author. Deviations from preregistered protocols were very common in replications, most frequently due to reasons inherent to the experimental model or related to infrastructure and logistics. Our results highlight factors that limit the replicability of results published by researchers in Brazil and suggest ways by which this scenario can be improved.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Amaral, O. B., Carneiro, C. F. D., Neves, K., Sampaio, A. P. W., Gomes, B. V., Abreu, M. B. d., Tan, P. B., Mota, G. P. S., Goulart, R. N., Fernandes, N. R. d. S., Linhares, J. H., Ibelli, A. M. G., Sebollela, A., Andricopulo, A. D., Carvalho, A. A. V. d., Silva, A. P. e., Souza, A. S. O., Lima, A. C. C. d., Sousa, A. M. d., Birbrair, A., Borbely, A. U., Machado, A. M., Costa, A. d. C., Vasconcelos, A. P. d., Silva, A. H. B. d. L., Souza, A. d., Herrmann, A. P., Bastos, A. P. A., Nunes, A. C. C., Castrucci, A. M. d. L., Paula, A. B. R., Waltrick, A. P. F., Macedo, A. G. F. d., Mecawi, A. S., R. 2025-04-03. Estimating the replicability of Brazilian biomedical science. https://doi.org/10.1101/2025.04.02.645026

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

An All-In-One Software Solution for Automated Processing of LA-ICP-TOF-MS datasets

LA-ICP-TOF-MS provides rapid, high resolution elemental analysis of biological and non-biological samples. However, accurate real-time data analysis frequently requires the user to account for several instrumental and experimental variables that can change during data acquisition. AutoSpect is a novel software tool designed to automate the processing and fitting of LA-ICP-TOF-MS data, addressing key challenges such as time-dependent spectral drift, instrument sensitivity drift calibration inaccuracies, and peak deconvolution, enabling researchers to rapidly and accurately process complex datasets. The tool is optimized to be robustly applicable across scientific fields (e.g., geochemistry, biology, and materials science), providing a streamlined solution for end users seeking to maximize the potential of LA-ICP-TOF-MS for high-resolution elemental mapping and isotopic analysis. Significance to JAASAnalysis of fast transient signals using laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) has become mainstream for elemental mapping. Advancements in LA-ICP-TOF-MS technology continue to accelerate the collective understanding of the role inorganic chemistry plays in dynamic processes. To ensure accurate quantitative results, the vast amount of complex spectral data generated requires elegant solutions to perform a variety of functions including data partitioning, peak fitting, drift correction, mass-to-charge calibration, peak profiling, and spectral fitting. AutoSpect is an all-in-one software solution that provides high level automation with a user-friendly graphical interface to perform complex data analyses for ICP-TOF-MS datasets.

scientific communication and education↗

Could instructor talk drive CURE effectiveness? A comparative study of instructor talk in introductory lab courses

Course-based undergraduate research experiences (CUREs) are thought to enhance students motivation to continue in college, in science, and in research. Yet, how CUREs enhance student motivation is largely undefined. Theories of instructor immediacy, self-efficacy, and task values suggest that CURE instructors may talk in ways that influence students motivational beliefs. We characterized the non-content related talk of instructors teaching 48 introductory biology lab courses, half CUREs and half non-CUREs. We identified 14 types of instructor talk that fit these theoretical perspectives: fostering students closeness with their instructor (i.e., immediacy talk), building students confidence in their scientific abilities (i.e., self-efficacy talk), and promoting students sense of worth in their work (i.e., task value talk). Course type had a medium effect on talk type, with CURE instructors utilizing more immediacy, self-efficacy, and task values talk than non-CURE instructors but also showing more variation in these types of talk. Our results suggest that motivation-related instructor talk is more prevalent in CUREs than non-CUREs, but wide variation in CURE instructor talk indicates additional investigation is needed before non-content talk can be considered a mechanism for the motivational influences of CUREs. HIGHLIGHTThis study compares non-content instructor talk in CURE and non-CURE lab courses using immediacy, self-efficacy, and task value theories. CURE instructors use more talk than non-CURE instructors, but variation in CURE instructor talk leaves open the question of whether talk is a causal factor in the motivational influence of CURE instruction.

scientific communication and education↗