Papers with low-

6 papers
DUMB: A Benchmark for Smart Evaluation of Dutch Models (2023.emnlp-main)

Copied to clipboard

Challenge: Current Dutch monolingual models under perform and suggest training larger models with other architectures and pre-training objectives.
Approach: They propose a Dutch Model Benchmark that compares performance of language models to a strong baseline that can be referred to in the future even when assessing different sets of language model.
Outcome: The proposed benchmark compares the performance of 14 pre-trained language models to a strong baseline . the results suggest training larger models with other architectures and pre-training objectives .
FLOR: On the Effectiveness of Language Adaptation (2024.lrec-main)

Copied to clipboard

Challenge: Large language models have amply proven their capabilities, but low- and mid-resource languages do not have access to the necessary means to train such models from scratch.
Approach: They use a 26B tokens corpus to further pre-train BLOOM, giving rise to FLOR models.
Outcome: The proposed model achieves consistent gains across Catalan and Spanish tasks.
Language Identification for Austronesian Languages (2022.lrec-1)

Copied to clipboard

Challenge: This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages.
Approach: They compare a classifier based on skip-gram embeddings with other methods . they then increase the number of non-Austronesian languages to 800 to evaluate their performance .
Outcome: The proposed model improves on the previous methods for low- and under-resourced languages in the Pacific region.
Viewing Knowledge Transfer in Multilingual Machine Translation Through a Representational Lens (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that translation quality alone is not sufficient for measuring knowledge transfer in multilingual neural machine translation.
Approach: They propose a method that measures representational similarities between languages to measure knowledge transfer.
Outcome: The proposed method improves translation quality for low- and mid-resource languages across multiple datasets and models.
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model (2025.emnlp-main)

Copied to clipboard

Challenge: Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content.
Approach: They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set .
Outcome: The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu.
PLAES: Prompt-generalized and Level-aware Learning Framework for Cross-prompt Automated Essay Scoring (2024.lrec-main)

Copied to clipboard

Challenge: Existing cross-prompt automatic essay scoring systems focus on obtaining shared knowledge specific to the target prompt, but this may not be feasible in practical situations because the target essay may not exist as training data.
Approach: They propose a novel learning framework for cross-prompt automatic essay scoring to capture more general knowledge across different prompts and improve the model’s capacity to distinguish between writing levels.
Outcome: The proposed learning framework captures more general knowledge across prompts and improves its capacity to distinguish between writing levels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations