gaHealth: An English–Irish Bilingual Corpus of Health Data (2022.lrec-1)

Copied to clipboard

Challenge: Existing models for low-resource languages often focus on creating the largest possible dataset for generic translation.
Approach: They develop a dataset for the specific domain of health for a low-resource English to Irish language pair and compare it to other similar datasets.
Outcome: The proposed model improved BLEU score by 22.2 points compared with top performing models from the LoResMT2021 Shared Task.

Similar Papers

GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: a new study explores the effectiveness of bilingual lexica in machine translation models . cross-lingual vocabulary alignment is still highly imperfect in these models, despite the success of supervised and self-supervised training.
Approach: They use a resource to improve translation performance on 200-language models . they show that lexica is more reliable than human-translated data .
Outcome: The proposed approach improves on 200-language translation models with lexical data augmentation . the proposed approach is open-source and has 168 tail languages .
ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining (2022.lrec-1)

Copied to clipboard

Challenge: Recent large-scale language models show remarkable achievements in key NLP tasks such as Question Answering and Text Summarization.
Approach: They propose a domain-specific pre-trained Vietnamese language model that outperforms the general domain language models.
Outcome: The proposed model outperforms the general domain language models in Vietnamese datasets while outperforming the general-domain language models.
MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks .
Approach: They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain.
Outcome: The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English.
Leveraging High-Resource English Corpora for Cross-lingual Domain Adaptation in Low-Resource Japanese Medicine via Continued Pre-training (2025.findings-emnlp)

Copied to clipboard

Challenge: low-resource language corpora in professional domains like medicine hinder cross-lingual domain adaptation of pre-trained large language models.
Approach: They examine how linguistic features affect performance on a Japanese–English medical knowledge benchmark.
Outcome: The proposed model can leverage English-language resources in medical domains while ensuring sufficient coverage of language-specific expressions in a target language.
Lessons from Natural Language Inference in the Clinical Domain (D18-1)

Copied to clipboard

Challenge: State of the art models with deep neural networks lack generalization capabilities in specialized domains where training data is limited.
Approach: They propose a dataset annotated by doctors performing a natural language inference task grounded in the medical history of patients.
Outcome: The proposed model outperforms existing models in the clinical domain by incorporating domain knowledge from external data and lexical sources.
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Pretrained large language models (LLMs) can bridge the performance gap for under-resourced languages by substantial margins, as measured by both automatic and human evaluations.
Approach: They propose to use pretrained large language models to bridge this gap by automating and evaluating data-to-text generation in under-resourced languages.
Outcome: The proposed model can set the state of the art for under-resourced languages by substantial margins, as measured by both automatic and human evaluations.
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine.
Approach: They propose an open-source LLM tailored for the biomedical domain that utilizes Mistral as its foundation model and pre-trained on PubMed Central.
Outcome: The proposed model outperforms existing models on a benchmark comprising 10 established medical question-answering tasks in English and is competitive with proprietary models.
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: In this study, we explore massively multilingual low-resource neural machine translation.
Approach: They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages.
Outcome: The proposed approach is highly language-specific and can be tailored to the source language and its typology.
A Balanced Data Approach for Evaluating Cross-Lingual Transfer: Mapping the Linguistic Blood Bank (2022.naacl-main)

Copied to clipboard

Challenge: Pretraining languages improve cross-lingual transfer for BERT-based models . Interestingly, PLMs exhibit zero-shot cross-linguistic abilities on downstream examples in languages seen only during pretraining.
Approach: They develop a quadratic time complexity method to estimate pretraining languages' relations between linguistic features and two downstream tasks.
Outcome: The proposed method is effective on a diverse set of languages spanning different linguistic features and two downstream tasks.
A Persona-Based Corpus in the Diabetes Self-Care Domain - Applying a Human-Centered Approach to a Low-Resource Context (2024.lrec-main)

Copied to clipboard

Challenge: Human-centered design (HCD) is a new approach to natural language processing that uses personas, user profiles and other tools to build corpus.
Approach: They propose to use personas to model interpersonal interaction in a healthcare domain to follow an HCD approach.
Outcome: The proposed model improves the quality of human-centered design in a healthcare domain and overcomes the lack of in-depth human-centricity in the field.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations