| Challenge: | Syntactic parsing improves the quality of natural language processing tasks. |
| Approach: | They evaluated Vietnamese Treebank model to find most suitable parsing method . they found that Vietnamese parsers produced limited training data and POS errors . |
| Outcome: | The proposed method improves the parsing quality in Vietnamese . the results highlight three possible sources of parser errors . |
Similar Papers
Revealing Weaknesses of Vietnamese Language Models Through Unanswerable Questions in Machine Reading Comprehension (2023.eacl-srw)
Copied to clipboard
| Challenge: | Existing problems in Vietnamese Machine Reading Comprehension systems are limited due to multilinguality, which limits the ability of multilingual models to develop state-of-the-art systems. |
| Approach: | They propose to modify the process of annotating unanswerable questions to improve the quality of unanswered questions to a higher level of difficulty for Machine Reading Comprehension systems to solve. |
| Outcome: | The proposed modification improves the quality of unanswerable questions to a higher level of difficulty for Machine Reading Comprehension systems to solve. |
BKTreebank: Building a Vietnamese Dependency Treebank (L18-1)
Copied to clipboard
| Challenge: | In this paper, we present the building of a dependency treebank for Vietnamese . |
| Approach: | They propose to build a Vietnamese dependency treebank using automatic taggers and automatic tagging. |
| Outcome: | The proposed treebank is a useful resource for Vietnamese language processing. |
An Empirical Comparison of Unsupervised Constituency Parsing Methods (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for unsupervised constituency parsing are inconsistent due to data preprocessing, lexicalization, and evaluation metrics. |
| Approach: | They propose to standardize experimental settings for better comparability between methods . they compare existing methods with those proposed by decade-old models . |
| Outcome: | The proposed methods perform better than decade-old models on English and Japanese, respectively, compared with decade- old models. |
A Pilot Study of Text-to-SQL Semantic Parsing for Vietnamese (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Semantic parsing is an important NLP task, but Vietnamese is a low-resource language. |
| Approach: | They extend EditSQL and IRNet semantic parsing baselines on Vietnamese datasets . they find automatic Vietnamese word segmentation improves parser results . |
| Outcome: | The proposed dataset improves on two strong parsing baselines for Vietnamese . the monolingual language model PhoBERT improves over the best multilingual language models. |
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases (2025.findings-naacl)
Copied to clipboard
| Challenge: | 1.2M original–paraphrase pairs were generated using a hybrid approach to generate high-quality paraphrases. |
| Approach: | They present a high-quality Vietnamese dataset for sentence paraphrasing . they used automatic paraphrase generation and manual evaluation to ensure high quality . |
| Outcome: | The proposed dataset is the first large-scale study on Vietnamese paraphrasing . it combines automatic paraphrase generation with manual evaluation to ensure high quality . |
Building a TOCFL Learner Corpus for Chinese Grammatical Error Diagnosis (L18-1)
Copied to clipboard
| Challenge: | Annotated learner corpus is valuable for research in second language acquisition, foreign language teaching, and contrastive interlanguage analysis. |
| Approach: | They construct a TOCFL learner corpus and use it for Chinese grammatical error diagnosis. |
| Outcome: | The constructed corpus is available to the public and will be used for shared tasks on Chinese grammatical error diagnosis. |
ViGLUE: A Vietnamese General Language Understanding Benchmark and Analysis of Vietnamese Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing benchmarks for natural language understanding have been suggested, but there is a lack of such a benchmark in Vietnamese due to the difficulty in accessing datasets or the scarcity of task-specific datasets. |
| Approach: | They propose to use a benchmark to evaluate Vietnamese language models in a variety of tasks and areas to explore the relationship between specific tasks and the number of shots. |
| Outcome: | The proposed benchmark contains twelve tasks and encompasses over ten areas and subjects, enabling it to evaluate models comprehensively over a broad spectrum of aspects. |
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges (2024.emnlp-main)
Copied to clipboard
| Challenge: | Vietnamese is a low-resource language, but each province has its own distinct pronunciation variations. |
| Approach: | They propose a dataset that captures the rich diversity of 63 provincial dialects spoken in Vietnam. |
| Outcome: | The proposed dataset captures the rich diversity of 63 provincial dialects spoken across Vietnam. |
On Parsing as Tagging (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to reduce constituency parsing to tagging are based on linearization, learning, and decoding . linearization of the derivation tree is the most critical factor in achieving accurate parsers as taggers . |
| Approach: | They propose a pipeline with three steps for reducing constituency parsing to tagging . they find that linearization and learning are critical factors for accurate parsers . |
| Outcome: | The proposed pipelines are linearized, learning, and decoded, and have three steps to achieve accurate parsing as taggers. |
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing open-source LLMs exhibit limited effectiveness in processing Vietnamese . lack of systematic benchmark datasets and metrics tailored for Vietnamese LLM evaluation exacerbates these issues. |
| Approach: | They propose to fine tune LLMs specifically for Vietnamese and develop a framework for evaluation . they find that larger models introduce more biases and uncalibrated outputs . |
| Outcome: | The proposed framework finetunes LLMs specifically for Vietnamese and provides a framework for evaluation . |