bioRxiv · 10.1101/2022.10.11.511776
Enriching Biomedical Knowledge for Low-resource Language Through Translation
Abstract
Biomedical data and benchmarks are highly valuable yet very limited in low-resource languages other than English such as Vietnamese. In this paper, we make use of a state-of-theart translation model in English-Vietnamese to translate and produce both pretrained as well as supervised data in the biomedical domains. Thanks to such large-scale translation, we introduce ViPubmedT5, a pretrained Encoder-Decoder Transformer model trained on 20 million translated abstracts from the high-quality public PubMed corpus. ViPubMedT5 demonstrates state-of-the-art results on two different biomedical benchmarks in summarization and acronym disambiguation. Further, we release ViMedNLI a new NLP task in Vietnamese translated from MedNLI using the recently public En-vi translation model and carefully refined by human experts, with evaluations of existing methods against ViPubmedT5.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Phan, L., Dang, T., Tran, H., Phan, V., Chau, L. D., Trinh, T. H.. 2022-10-14. Enriching Biomedical Knowledge for Low-resource Language Through Translation. https://doi.org/10.1101/2022.10.11.511776
Cite the original work for its findings. Save a collection to share your selection of sources.