bioRxiv · 10.64898/2026.01.09.697335
systematic evaluation and benchmarking of text summarization methods for biomedical literature: From word-frequency methods to language models
Abstract
The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (LLMs), on 1,000 biomedical abstracts from 20 journals across ScienceDirect and Cell Press, using author-written highlights as reference summaries. Models were evaluated with a composite suite of lexical, semantic, and factual metrics, including ROUGE, BLEU, METEOR, embedding-based similarity, and factuality scores. General-purpose models (e.g., Mistral, GPT, Llama) achieved the highest overall performance across lexical and semantic dimensions, outperforming reasoning-oriented (e.g., DeepSeek, Magistral) and domain-specific (e.g., BioGPT, BioMistral) models. Notably, medium-sized models outperformed large-scale models, suggesting an optimal balance between model capacity and efficiency, while classical extractive methods lagged behind neural approaches. These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Baumgärtel, F., Bono, E., Galou, L., Keska-Izworska, K., Walter, S., Andorfer, P., Kratochwill, K., Perco, P., Ley, M.. 2026-01-13. systematic evaluation and benchmarking of text summarization methods for biomedical literature: From word-frequency methods to language models. https://doi.org/10.64898/2026.01.09.697335
Cite the original work for its findings. Save a collection to share your selection of sources.