KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark (2024.lrec-main)
Copied to clipboard
| Challenge: | KoDialogBench is a benchmark designed to assess language models’ conversational capabilities in low-resource languages such as Korean. |
| Approach: | They propose a benchmark to assess language models’ conversational capabilities in Korean by collecting native Korean dialogues from public sources and translating them into diverse test datasets. |
| Outcome: | The proposed benchmark measures the conversational capabilities of language models in Korean, and shows that they can improve on previous training techniques. |
Similar Papers
xDial-Eval: A Multilingual Open-Domain Dialogue Evaluation Benchmark (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Currently, human evaluation is the most reliable way to holistically judge the quality of the dialogue. |
| Approach: | They propose to use English dialogue evaluation metrics to generalize them to other languages. |
| Outcome: | The proposed metrics outperform OpenAI’s ChatGPT in terms of average Pearson correlations over all datasets and languages. |
Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding (2024.emnlp-main)
Copied to clipboard
| Challenge: | despite its importance, there has been limited research on conversational grounding in recent years . pre-trained language models have been costly and time-consuming to evaluate . |
| Approach: | They evaluate the performance of large language models in various aspects of conversational grounding . they propose ways to enhance the capabilities of the models that lag in this aspect . |
| Outcome: | The proposed model performance is based on pre-trained language models and a large pre-training dataset. |
InterroLang: Exploring NLP Models and Datasets through Dialogue-based Explanations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on NLP explainability methods lacks a dialogue-based interpretability framework that can convey faithful explanations in human-understandable terms. |
| Approach: | They adapt the conversational explanation framework TalkToModel to the NLP domain and add new NLP-specific operations such as free-text rationalization to illustrate its generalizability. |
| Outcome: | The proposed framework can be used to explain models on three NLP tasks and is generalizable to different datasets, use cases and models. |
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models (2024.lrec-main)
Copied to clipboard
Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim, Jae cheol Lee, Je Won Yeom, Jihyu Jung, Jung woo Kim, Songseong Kim
| Challenge: | Existing evaluation tools rely on translations of English datasets or translation-specific benchmarks such as WMT 21 to assess large language models. |
| Approach: | They propose a dataset curated to challenge models lacking Korean cultural and contextual depth. |
| Outcome: | The HAE-RAE Bench challenges models lacking Korean cultural and contextual depth by highlighting their aptitude for recalling Korean-specific knowledge and cultural contexts. |
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Negation is a fundamental operation in natural language that reverses the meaning of an expression into its opposite. |
| Approach: | They propose a sentence-level negation understanding benchmark that measures negation performance in Korean. |
| Outcome: | The proposed benchmark improves negation understanding and broader comprehension in Korean. |
Language Model as an Annotator: Exploring DialoGPT for Dialogue Summarization (2021.acl-long)
Copied to clipboard
| Challenge: | Existing dialogue summarization systems encode text with a number of general semantic features, but these are often not available in open-domain tools. |
| Approach: | They propose to use DialoGPT to label three types of features on two datasets . they propose to employ pre-trained and non-pre-tried models as dialogue annotators . |
| Outcome: | The proposed method improves on two dialogue summarization datasets and achieves state-of-the-art performance. |
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset (P19-1)
Copied to clipboard
| Challenge: | EmpatheticDialogues dataset provides a benchmark for empathetic dialogue generation . human evaluators perceive dialogue models as more epathetic . |
| Approach: | They propose a benchmark for empathetic dialogue generation from a dataset of 25k conversations grounded in emotional situations. |
| Outcome: | The proposed benchmarks show that existing models are perceived to be more empathetic by human evaluators compared to models trained on large-scale Internet conversations. |
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI (2024.findings-eacl)
Copied to clipboard
Jianguo Zhang, Kun Qian, Zhiwei Liu, Shelby Heinecke, Rui Meng, Ye Liu, Zhou Yu, Huan Wang, Silvio Savarese, Caiming Xiong
| Challenge: | DialogStudio is the largest and most diverse collection of dialogue datasets . existing datasets lack diversity and comprehensiveness, authors say . |
| Approach: | They introduce DialogStudio: the largest and most diverse collection of dialogue datasets . DialogStuio aggregates more than 80 diverse dialogue dataset . |
| Outcome: | a new dataset is created to improve the quality and diversity of dialogue datasets . DialogStudio is the largest and most diverse collection of dialogue data . |
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Current evaluation practices of open domain dialogue systems are still highly dependent on human evaluation. |
| Approach: | They propose to use an annotated dataset to evaluate chatbots using large language models. |
| Outcome: | The proposed model improves over few-shot inferences on a GPT-3.5 generated dialogue dataset. |
A Dog Is Passing Over The Jet? A Text-Generation Dataset for Korean Commonsense Reasoning and Evaluation (2022.findings-naacl)
Copied to clipboard
Jaehyung Seo, Seounghoon Lee, Chanjun Park, Yoonna Jang, Hyeonseok Moon, Sugyeong Eo, Seonmin Koo, Heuiseok Lim
| Challenge: | Korean pretrained language models struggle to generate short sentences with a given condition based on compositionality and commonsense reasoning. |
| Approach: | They propose a Korean text-generation dataset for Korean generative commonsense reasoning and language model evaluation using a semi-automatic dataset construction approach. |
| Outcome: | The proposed dataset is available at http://aihub.or.kr/opendata/korea-university. |