Introducing Two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness (N18-2)
Copied to clipboard
| Challenge: | Existing datasets for low-resource language Vietnamese assess semantic similarity . a dataset for word pairs with similarity levels is needed to evaluate these models . |
| Approach: | They present two new datasets for the low-resource language Vietnamese to assess models of semantic similarity. |
| Outcome: | The two datasets are comparable to the English datasets. |
Similar Papers
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)
Copied to clipboard
Nedjma Ousidhoum, Shamsuddeen Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Ahmad, Sanchit Ahuja, Alham Aji, Vladimir Araujo, Abinew Ayele, Pavan Baswani, Meriem Beloucif, Chris Biemann, Sofia Bourhim, Christine Kock, Genet Dekebo, Oumaima Hourrane, Gopichand Kanumolu, Lokesh Madasu, Samuel Rutunda, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Hailegnaw Tilaye, Krishnapriya Vishnubhotla, Genta Winata, Seid Yimam, Saif Mohammad
| Challenge: | SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text . |
| Approach: | They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu. |
| Outcome: | The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages. |
A Pilot Study of Text-to-SQL Semantic Parsing for Vietnamese (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Semantic parsing is an important NLP task, but Vietnamese is a low-resource language. |
| Approach: | They extend EditSQL and IRNet semantic parsing baselines on Vietnamese datasets . they find automatic Vietnamese word segmentation improves parser results . |
| Outcome: | The proposed dataset improves on two strong parsing baselines for Vietnamese . the monolingual language model PhoBERT improves over the best multilingual language models. |
SemR-11: A Multi-Lingual Gold-Standard for Semantic Similarity and Relatedness for Eleven Languages (L18-1)
Copied to clipboard
| Challenge: | SemR-11 is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages. |
| Approach: | This paper describes a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages. |
| Outcome: | The dataset is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages. |
Which Works Best for Vietnamese? A Practical Study of Information Retrieval Methods across Domains (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies on Large Language Models (LLMs) are limited to single domains or curated datasets. |
| Approach: | They propose a domain-normalized, multi-domain benchmark for Vietnamese IR . they evaluate lexical, neural-sparse, late-interaction, dense, and hybrid paradigms . |
| Outcome: | The proposed benchmarks cover six domains and ten datasets across education, legal, healthcare, customer support, lifestyle reviews, and open-domain knowledge. |
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)
Copied to clipboard
Maria Lymperaiou, George Manoliadis, Orfeas Menis Mastromichalakis, Edmund G. Dervakos, Giorgos Stamou
| Challenge: | Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies. |
| Approach: | They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances . |
| Outcome: | The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations. |
Similarity Measures for the Detection of Clinical Conditions with Verbal Fluency Tasks (N18-2)
Copied to clipboard
| Challenge: | Semantic Verbal Fluency tests have been used in the diagnosis of certain clinical conditions, like Dementia. |
| Approach: | They investigate three similarity measures for automatically identifying switches in semantic chains: semantic similarity from a manually constructed resource, word association strength and semantic relatedness, both calculated from corpora. |
| Outcome: | The proposed classifiers outperform those that use a gold standard taxonomy for clinical conditions. |
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges (2024.emnlp-main)
Copied to clipboard
| Challenge: | Vietnamese is a low-resource language, but each province has its own distinct pronunciation variations. |
| Approach: | They propose a dataset that captures the rich diversity of 63 provincial dialects spoken in Vietnam. |
| Outcome: | The proposed dataset captures the rich diversity of 63 provincial dialects spoken across Vietnam. |
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing open-source LLMs exhibit limited effectiveness in processing Vietnamese . lack of systematic benchmark datasets and metrics tailored for Vietnamese LLM evaluation exacerbates these issues. |
| Approach: | They propose to fine tune LLMs specifically for Vietnamese and develop a framework for evaluation . they find that larger models introduce more biases and uncalibrated outputs . |
| Outcome: | The proposed framework finetunes LLMs specifically for Vietnamese and provides a framework for evaluation . |
FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction (2025.findings-emnlp)
Copied to clipboard
| Challenge: | evaluating the usefulness of language models for literary-domain tasks remains challenging due to the cost of fine-grained annotation for long-form texts and data contamination concerns inherent in using public-domain literature. |
| Approach: | They use a dataset of long-form, recently written fiction to evaluate embedding models . they prioritize author agency and rely on continual, informed author consent . |
| Outcome: | The proposed dataset of long-form, recently written fiction is compared with existing models on this task. |
Representing Verbs with Visual Argument Vectors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes. |
| Approach: | They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities. |
| Outcome: | The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models. |