Parallel Corpora for the Biomedical Domain (L18-1)

Copied to clipboard

Challenge: Existing corpora of parallel corporata are being used in the biomedical domain . MT is known to support readers' access to textual documents in a language other than their native language .
Approach: They propose to leverage parallel corpora to implement cross-lingual information retrieval or machine translation tools.
Outcome: The proposed corpus is being used in the biomedical task at the conference on machine translation (WMT'16 and WMT'17) it can be leveraged to provide access to health information in languages other than English.

Similar Papers

MEDLINE as a Parallel Corpus: a Survey to Gain Insight on French-, Spanish- and Portuguese-speaking Authors’ Abstract Writing Practice (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpora are used to train and evaluate machine translation systems, but little information is available about the methods used for producing the corpus, including translation direction.
Approach: They used PubMed and publisher websites to obtain contact information for MEDLINE authors and asked about their abstract writing practices.
Outcome: The authors of MEDLINE articles included in the English/Spanish, English/FR, and English/Portuguese (EN/PT) WMT 2019 test sets reported a response rate of over 20% .
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
Leveraging High-Resource English Corpora for Cross-lingual Domain Adaptation in Low-Resource Japanese Medicine via Continued Pre-training (2025.findings-emnlp)

Copied to clipboard

Challenge: low-resource language corpora in professional domains like medicine hinder cross-lingual domain adaptation of pre-trained large language models.
Approach: They examine how linguistic features affect performance on a Japanese–English medical knowledge benchmark.
Outcome: The proposed model can leverage English-language resources in medical domains while ensuring sufficient coverage of language-specific expressions in a target language.
SciPar: A Collection of Parallel Corpora from Scientific Abstracts (2022.lrec-1)

Copied to clipboard

Challenge: SciPar is a collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries and national archives.
Approach: They propose to harvest and process openly available metadata from repositories to extract bilingual titles and abstracts from scientific publications.
Outcome: The proposed corpora could be useful for cross-lingual plagiarism detection or adapting Machine Translation systems for translation of scientific texts and academic writing in general.
BioMegatron: Larger Biomedical Domain Language Model (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies on domain language models do not study the factors affecting performance on domain languages.
Approach: They empirically evaluate factors that can affect performance on domain language applications . sub-word vocabulary set, model size, pre-training corpus, and domain transfer are important .
Outcome: The results show language models trained on biomedical text perform better on biomedicine benchmarks than those trained on general domain text corpora.
A Large Parallel Corpus of Full-Text Scientific Articles (L18-1)

Copied to clipboard

Challenge: Scielo database contains articles from several research domains.
Approach: They propose to build a parallel corpus from Scielo in three languages: English, Portuguese, and Spanish.
Outcome: The proposed system outperforms other systems on scientific articles in English, Portuguese, and Spanish.
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models.
Approach: They investigate the impact of parallel corpora quality and quantity, training objectives, and model size on performance of multilingual large language models enhanced with parallel corporeal.
Outcome: The proposed approach improves performance in bilingual and general-purpose tasks.
MultiMSD: A Corpus for Multilingual Medical Text Simplification from Online Medical References (2025.findings-acl)

Copied to clipboard

Challenge: Medical texts contain technical terms, and non-experts often cannot use information effectively.
Approach: They propose a method for training medical text simplification models to actively paraphrase medical terms.
Outcome: The proposed method improves the performance of medical text simplification in nine languages.
MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks .
Approach: They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain.
Outcome: The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English.
A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature (P18-1)

Copied to clipboard

Challenge: In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day.
Approach: They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured.
Outcome: The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations