| Challenge: | Existing models for low-resource languages often focus on creating the largest possible dataset for generic translation. |
| Approach: | They develop a dataset for the specific domain of health for a low-resource English to Irish language pair and compare it to other similar datasets. |
| Outcome: | The proposed model improved BLEU score by 22.2 points compared with top performing models from the LoResMT2021 Shared Task. |
Similar Papers
GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation (2023.emnlp-main)
Copied to clipboard
| Challenge: | a new study explores the effectiveness of bilingual lexica in machine translation models . cross-lingual vocabulary alignment is still highly imperfect in these models, despite the success of supervised and self-supervised training. |
| Approach: | They use a resource to improve translation performance on 200-language models . they show that lexica is more reliable than human-translated data . |
| Outcome: | The proposed approach improves on 200-language translation models with lexical data augmentation . the proposed approach is open-source and has 168 tail languages . |
ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent large-scale language models show remarkable achievements in key NLP tasks such as Question Answering and Text Summarization. |
| Approach: | They propose a domain-specific pre-trained Vietnamese language model that outperforms the general domain language models. |
| Outcome: | The proposed model outperforms the general domain language models in Vietnamese datasets while outperforming the general-domain language models. |
MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain (2024.lrec-main)
Copied to clipboard
Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johana Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, Andrea Zaninello
| Challenge: | Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks . |
| Approach: | They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. |
| Outcome: | The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English. |
Leveraging High-Resource English Corpora for Cross-lingual Domain Adaptation in Low-Resource Japanese Medicine via Continued Pre-training (2025.findings-emnlp)
Copied to clipboard
Kazuma Kobayashi, Zhen Wan, Fei Cheng, Yuma Tsuta, Xin Zhao, Junfeng Jiang, Jiahao Huang, Zhiyi Huang, Yusuke Oda, Rio Yokota, Yuki Arase, Daisuke Kawahara, Akiko Aizawa, Sadao Kurohashi
| Challenge: | low-resource language corpora in professional domains like medicine hinder cross-lingual domain adaptation of pre-trained large language models. |
| Approach: | They examine how linguistic features affect performance on a Japanese–English medical knowledge benchmark. |
| Outcome: | The proposed model can leverage English-language resources in medical domains while ensuring sufficient coverage of language-specific expressions in a target language. |
Lessons from Natural Language Inference in the Clinical Domain (D18-1)
Copied to clipboard
| Challenge: | State of the art models with deep neural networks lack generalization capabilities in specialized domains where training data is limited. |
| Approach: | They propose a dataset annotated by doctors performing a natural language inference task grounded in the medical history of patients. |
| Outcome: | The proposed model outperforms existing models in the clinical domain by incorporating domain knowledge from external data and lexical sources. |
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models (2024.findings-eacl)
Copied to clipboard
| Challenge: | Pretrained large language models (LLMs) can bridge the performance gap for under-resourced languages by substantial margins, as measured by both automatic and human evaluations. |
| Approach: | They propose to use pretrained large language models to bridge this gap by automating and evaluating data-to-text generation in under-resourced languages. |
| Outcome: | The proposed model can set the state of the art for under-resourced languages by substantial margins, as measured by both automatic and human evaluations. |
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains (2024.findings-acl)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, Richard Dufour
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine. |
| Approach: | They propose an open-source LLM tailored for the biomedical domain that utilizes Mistral as its foundation model and pre-trained on PubMed Central. |
| Outcome: | The proposed model outperforms existing models on a benchmark comprising 10 established medical question-answering tasks in English and is competitive with proprietary models. |
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | In this study, we explore massively multilingual low-resource neural machine translation. |
| Approach: | They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages. |
| Outcome: | The proposed approach is highly language-specific and can be tailored to the source language and its typology. |
A Balanced Data Approach for Evaluating Cross-Lingual Transfer: Mapping the Linguistic Blood Bank (2022.naacl-main)
Copied to clipboard
| Challenge: | Pretraining languages improve cross-lingual transfer for BERT-based models . Interestingly, PLMs exhibit zero-shot cross-linguistic abilities on downstream examples in languages seen only during pretraining. |
| Approach: | They develop a quadratic time complexity method to estimate pretraining languages' relations between linguistic features and two downstream tasks. |
| Outcome: | The proposed method is effective on a diverse set of languages spanning different linguistic features and two downstream tasks. |
A Persona-Based Corpus in the Diabetes Self-Care Domain - Applying a Human-Centered Approach to a Low-Resource Context (2024.lrec-main)
Copied to clipboard
| Challenge: | Human-centered design (HCD) is a new approach to natural language processing that uses personas, user profiles and other tools to build corpus. |
| Approach: | They propose to use personas to model interpersonal interaction in a healthcare domain to follow an HCD approach. |
| Outcome: | The proposed model improves the quality of human-centered design in a healthcare domain and overcomes the lack of in-depth human-centricity in the field. |