Challenge: Vietnamese is a low-resource language, but each province has its own distinct pronunciation variations.
Approach: They propose a dataset that captures the rich diversity of 63 provincial dialects spoken in Vietnam.
Outcome: The proposed dataset captures the rich diversity of 63 provincial dialects spoken across Vietnam.

Similar Papers

ViGLUE: A Vietnamese General Language Understanding Benchmark and Analysis of Vietnamese Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing benchmarks for natural language understanding have been suggested, but there is a lack of such a benchmark in Vietnamese due to the difficulty in accessing datasets or the scarcity of task-specific datasets.
Approach: They propose to use a benchmark to evaluate Vietnamese language models in a variety of tasks and areas to explore the relationship between specific tasks and the number of shots.
Outcome: The proposed benchmark contains twelve tasks and encompasses over ten areas and subjects, enabling it to evaluate models comprehensively over a broad spectrum of aspects.
A Parallel Corpus for Vietnamese Central-Northern Dialect Text Transfer (2023.findings-emnlp)

Copied to clipboard

Challenge: Among these, the northern dialect is often treated as the standard i.e. the defacto text style of the language.
Approach: They propose a parallel corpus for Vietnamese central-northern dialect text transfer to facilitate research on this domain.
Outcome: The proposed model improves existing models on the central dialect domain with dedicated results in translation and text-image retrieval tasks.
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension (2024.eacl-long)

Copied to clipboard

Challenge: Existing datasets for machine reading comprehension tasks in Vietnamese focus on written documents, such as Wikipedia articles, online newspapers, or textbooks.
Approach: They propose to capture Vietnamese spoken language in natural settings and use it to create a machine-learning corpus for machine reading comprehension tasks.
Outcome: The proposed corpus consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube .
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Currently, there are no publicly available speech recognition datasets in the medical domain due to privacy restrictions.
Approach: They present a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical and 1200h of general-domain speech.
Outcome: The proposed model outperforms state-of-the-art models from 51.8% to 29.6% WER on test set.
Vietnamese Automatic Speech Recognition: A Revisit (2026.findings-eacl)

Copied to clipboard

Challenge: Existing datasets with low quality and inconsistent annotations are insufficient for high-quality models.
Approach: They propose a pipeline for aggregating and preprocessing high-quality ASR datasets from diverse, potentially noisy, open-source sources.
Outcome: The proposed pipeline provides a foundation for training and evaluating state-of-the-art Vietnamese ASR systems.
A Vietnamese Dataset for Evaluating Machine Reading Comprehension (2020.coling-main)

Copied to clipboard

Challenge: despite the lack of benchmark datasets for Vietnamese, there are few studies on machine reading comprehension (MRC) . MRC is an essential core for a range of natural language processing applications such as search engines and intelligent agents.
Approach: They propose to use Vietnamese Question Answering Dataset to evaluate machine reading comprehension in Vietnamese . they use over 23,000 human-generated question-answer pairs based on 5,109 Vietnamese articles .
Outcome: The proposed dataset includes over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.
VIMQA: A Vietnamese Dataset for Advanced Reasoning and Explainable Multi-hop Question Answering (2022.lrec-1)

Copied to clipboard

Challenge: Existing Vietnamese Question Answering (QA) datasets do not explore the model’s ability to perform advanced reasoning and provide evidence to explain the answer.
Approach: They propose to use Vietnamese as a question-answer dataset with 10,000 Wikipedia-based multi-hop question-and-answ pairs to test model's ability to reason and explain the answer.
Outcome: The proposed dataset is in Vietnamese, a low-resource language.
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Historically, studies investigating minority variants of languages have been limited to a select few languages.
Approach: They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors.
Outcome: The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap.
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: despite recent advances in speech processing, the majority of world languages and dialects remain uncovered.
Approach: They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset .
Outcome: The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni.
Revealing Weaknesses of Vietnamese Language Models Through Unanswerable Questions in Machine Reading Comprehension (2023.eacl-srw)

Copied to clipboard

Challenge: Existing problems in Vietnamese Machine Reading Comprehension systems are limited due to multilinguality, which limits the ability of multilingual models to develop state-of-the-art systems.
Approach: They propose to modify the process of annotating unanswerable questions to improve the quality of unanswered questions to a higher level of difficulty for Machine Reading Comprehension systems to solve.
Outcome: The proposed modification improves the quality of unanswerable questions to a higher level of difficulty for Machine Reading Comprehension systems to solve.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations