Consistency Rating of Semantic Transparency: an Evaluation Method for Metaphor Competence in Idiom Understanding Tasks (2025.coling-main)
Copied to clipboard
| Challenge: | Idioms condense complex semantics into fixed phrases, making idiom comprehension a test of metaphor competence. |
| Approach: | They propose a method to evaluate the metaphor competence of LLMs for the idiom understanding task: the Consistency Rating of Semantic Transparency (CR-ST). |
| Outcome: | The proposed method assesses the difficulty of understanding idioms through two dimensions: overall semantic transparency and constituent semantic transparency, aiming to gauge LLMs’ mastery of metaphor competence. |
Similar Papers
Beyond Understanding: Evaluating the Pragmatic Gap in LLMs’ Cultural Processing of Figurative Language (2026.eacl-long)
Copied to clipboard
| Challenge: | Using figurative language as a proxy for cultural nuance and local knowledge, large language models struggle with connotative meaning. |
| Approach: | They evaluate large language models' ability to process culturally grounded language . they use figurative language as a proxy for cultural nuance and local knowledge . |
| Outcome: | The proposed models can understand and use figurative expressions that encode local knowledge and social nuance. |
On the Consistency of Commonsense in Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of commonsense for large language models focus on downstream knowledge tasks, failing to probe whether LLMs truly understand and utilize knowledge or merely memorize it. |
| Approach: | They propose to automatically construct a large benchmark named CoCo which measures LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks. |
| Outcome: | The proposed benchmark systematically assesses LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks. |
CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and Use (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks focus on narrow tasks such as multiple-choice cloze tests, isolated translation, or simple paraphrasing. |
| Approach: | They propose a benchmark to measure Chinese idioms' cultural and contextual nuances . they evaluate 2,937 human-verified examples covering 1,765 common idiomes . |
| Outcome: | The proposed benchmarks achieve 95% accuracy on Evaluative Connotation, but only 85% on Appropriateness and 40% top-1 accuracy in Open Cloze. |
Refining Idioms Semantics Comprehension via Contrastive Learning and Cross-Attention (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods based on deep learning struggle to grasp idiom semantics due to the figurative meanings of many idiomas deviating from their literal interpretations. |
| Approach: | They propose a Chinese idiom cloze test to capture comprehensive idiomatics and a semantic sense contrastive learning module to enhance the representation of idiomics. |
| Outcome: | The proposed model outperforms state-of-the-art models on the Chinese idiom cloze test and on other benchmark datasets. |
Metaphor Understanding Challenge Dataset for LLMs (2024.acl-long)
Copied to clipboard
| Challenge: | Metaphor understanding is an essential task for large language models (LLMs). |
| Approach: | They propose to evaluate the metaphor understanding capabilities of large language models (LLMs) the metaphor understanding challenge dataset provides over 10k paraphrases and 1.5k instances of inapt paraphrase. |
| Outcome: | The metaphor understanding challenge dataset evaluates the performance of large language models on a range of NLU tasks. |
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions. |
| Approach: | They propose to use a large-scale dataset of idioms in six languages to evaluate LLMs' idiomatic processing ability. |
| Outcome: | The proposed model integrates contextual cues and reasoning to improve idiom understanding in LLMs, suggesting that their performance is influenced by memorization and reasoning. |
Establishing Trustworthiness: Rethinking Tasks and Model Evaluation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing community has striven to model computationally for decades. |
| Approach: | They propose to rethink what constitutes tasks and model evaluation in NLP and pursue a more holistic view on language, placing trustworthiness at the center. |
| Outcome: | The proposed models are based on generative models and are being deployed in more real-world scenarios, including previously unforeseen zero-shot setups. |
Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Semantic phrases (SP) are lexical combinations whose meanings or usages may not be fully derived from their individual components. |
| Approach: | They propose to consolidate existing multiword expression resources into a unified testbed to assess language models in semantic phrase processing tasks. |
| Outcome: | The evaluation suite covers idiomatic expressions, noun compounds, and verbal constructions. |
MetaPro 2.0: Computational Metaphor Processing on the Effectiveness of Anomalous Language Modeling (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for metaphor interpretation are slow due to lack of annotated datasets and effective pre-trained language models. |
| Approach: | They propose a large annotated dataset and a PLM for the metaphor interpretation task. |
| Outcome: | The proposed method improves on metaphor identification and interpretation with comparable baselines on the new dataset. |
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context (2025.acl-long)
Copied to clipboard
| Challenge: | Existing models fail to resolve idiomaticity when it depends on contextual understanding . idiom frequency influences performance but does not guarantee accurate interpretation. |
| Approach: | They propose a novel contrastive dataset to assess whether large language models can effectively leverage context to disambiguate idiomatic meanings. |
| Outcome: | The proposed model performs better on sentences deemed more likely by the model . collocational frequency and sentence probability influence performance but not accuracy . |