Challenge: Existing corpora in African languages reveal significant spelling inconsistencies, contributing to poor-quality textual data when encountered in written form.
Approach: They propose a supervised method to identify loanwords in Portuguese . they employ traditional machine learning algorithms incorporating handcrafted features .
Outcome: The proposed method achieves the F1-score of 93% in Emakhuwa, borrowed from Portuguese.

Similar Papers

Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual Context (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study addresses the challenge of adapting loanwords during the translation process in low-resource languages.
Approach: They propose a method that augments source sentences with loanword constraints . they then integrate loanwords as external linguistic knowledge into machine translation systems .
Outcome: The proposed approach improves translation quality and handling loanword adaptation correctly in target languages.
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords (2025.acl-long)

Copied to clipboard

Challenge: Lexical borrowing is a ubiquitous linguistic phenomenon influenced by geopolitical, societal, and technological factors.
Approach: They propose a novel contrastive dataset comprising sentences with and without loanwords across 10 languages to examine how machine translation and language models process loanword .
Outcome: The proposed dataset shows that state-of-the-art models prefer loanwords over native terms and exhibit varying performance across languages.
Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks (2024.emnlp-main)

Copied to clipboard

Challenge: Emakhuwa is the most widely spoken language in Mozambique but has received limited attention in NLP research.
Approach: They propose a comprehensive collection of NLP resources for Emakhuwa, Mozambique's most widely spoken language.
Outcome: The proposed models show good performance in news topic classification and promising results in machine translation.
A Generalized Method for Automated Multilingual Loanword Detection (2022.coling-1)

Copied to clipboard

Challenge: Loanwords are words incorporated from one language into another without translation . authors present a method to automatically detect loanwords across language pairs .
Approach: They propose a method to automatically detect loanwords across language pairs . they incorporate edit distance, semantic similarity measures, phonetic alignment .
Outcome: The proposed method outperforms existing methods on single-pair loanword detection tasks and can generalize to unseen language pairs with sufficient data.
Massive vs. Curated Embeddings for Low-Resourced Languages: the Case of Yorùbá and Twi (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that word embeddings can be useful for training downstream natural language processing tasks.
Approach: They compare word embeddings obtained by word embeds from curated corpora with a language-dependent processing.
Outcome: The proposed model compares word embeddings with word embeds from curated corpora and a language-dependent processing on two African languages.
Evaluating Word Embeddings in Extremely Under-Resourced Languages: A Case Study in Bribri (2022.coling-1)

Copied to clipboard

Challenge: a case study of word embeddings in an under-resourced context is presented . Embeddings are critical for many NLP tasks, but their evaluation in actual under-sourced settings needs further investigation.
Approach: They adapt word embeddings to an under-resourced indigenous language in Bribri . they find that the best models find the appropriate semantic target 60% of the time .
Outcome: The proposed models were adapted to an under-resourced English language . the best models found the appropriate semantic target 60% of the time .
Data Collection Pipeline for Low-Resource Languages: A Case Study on Constructing a Tetun Text Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Labadain Crawler is a data collection pipeline designed to automate and optimize the process of constructing textual corpora from the web, with a specific target to low-resource languages.
Approach: They propose a data collection pipeline built on top of Nutch, an open-source web crawler and data extraction framework, and a tokenizer and identifier for Tetun.
Outcome: The proposed pipeline is based on Nutch, an open-source web crawler and data extraction framework, and is tested with Tetun, one of Timor-Leste’s official languages.
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
Toward Better Loanword Identification in Uyghur Using Cross-lingual Word Embeddings (C18-1)

Copied to clipboard

Challenge: Almost every natural language processing task suffers from data sparseness.
Approach: They propose a method which identify loanwords in monolingual corpora by using cross-lingual word embeddings as core feature and a log-linear model which combines several shallow features to predict the final results.
Outcome: The proposed method outperforms baseline models significantly on loanword identification and translation in four languages and eight translation directions.
Hypernymy Detection for Low-Resource Languages via Meta Learning (2020.acl-main)

Copied to clipboard

Challenge: Existing studies focus on monolingual hypernymy detection on high-resource languages, but few investigate low-resourced scenarios.
Approach: They propose to combine high-resource languages to solve low-resourced hypernymy detection problem . they extensively compare three joint training paradigms and propose meta learning .
Outcome: The proposed method significantly improves performance of extremely low-resource languages by preventing over-fitting on small datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations