Challenge: Existing implementations of the Rational Speech Act (RSA) framework do not account for figurative expressions or require modeling the implicit motivations behind using figurativ language in a setting-specific way.
Approach: They propose a framework which models figurative language use by considering a speaker's employed rhetorical strategy and a computational model which incorporates rhetorical strategies to support non-literal interpretation.
Outcome: The proposed framework enables human-compatible interpretations of non-literal utterances without modeling speaker's motivations for being non-lative.

Similar Papers

Testing the Ability of Language Models to Interpret Figurative Language (2022.naacl-main)

Copied to clipboard

Challenge: Existing work on figurative language has not been done on literal language models.
Approach: They propose a Winograd-style task to evaluate figurative phrases with divergent meanings by interpreting paired figurativ phrases with a human input.
Outcome: The proposed task outperforms state-of-the-art models on a nonliteral language understanding task in zero-shot settings.
Investigating Robustness of Dialog Models to Popular Figurative Language Constructs (2021.emnlp-main)

Copied to clipboard

Challenge: Existing dialog models are unable to handle popular figurative language constructs like metaphor and simile when faced with dialog contexts containing figurativ language.
Approach: They propose lightweight solutions to help existing dialog models become more robust to figurative language by simply using an external resource to translate figurativ language to literal (non-figurative) forms while preserving the meaning to the best extent possible.
Outcome: The proposed models show large drops in performance when faced with dialog contexts consisting of figurative language compared to contexts without figurativ language .
Beyond Understanding: Evaluating the Pragmatic Gap in LLMs’ Cultural Processing of Figurative Language (2026.eacl-long)

Copied to clipboard

Challenge: Using figurative language as a proxy for cultural nuance and local knowledge, large language models struggle with connotative meaning.
Approach: They evaluate large language models' ability to process culturally grounded language . they use figurative language as a proxy for cultural nuance and local knowledge .
Outcome: The proposed models can understand and use figurative expressions that encode local knowledge and social nuance.
FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and Korean (2025.emnlp-main)

Copied to clipboard

Challenge: Figurative language is a core component of everyday communication . existing benchmarks focus on sentence-level classification or inference tasks .
Approach: They propose a multilingual benchmark that evaluates figurative usage in dialogue . they use a sentence-level diagnostic task to embed figurativ choices into multi-turn contexts .
Outcome: The benchmark evaluates large language models' ability to use figurative expressions coherently in conversation.
Decision Biases and Intent-Irony Decoupling in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit impressive linguistic fluency, but it remains unclear whether they possess human-like Theory of Mind (ToM) or rely on statistical heuristics . a recent study examined the performance of LLMs against 300 human participants .
Approach: a study establishes a framework for large language models that modulates contextual contrast, linguistic cues, and cognitive mechanisms.
Outcome: a new evaluation framework compares ten state-of-the-art LLMs against 300 human participants . the framework systematically modulates contextual contrast, linguistic cues, and cognitive mechanisms .
IRFL: Image Recognition of Figurative Language (2023.findings-emnlp)

Copied to clipboard

Challenge: Figures of speech are ubiquitous in many forms of discourse, allowing people to convey complex, abstract ideas and evoke emotion.
Approach: They develop a dataset for multimodal figurative language understanding using human annotation and an automatic pipeline to generate a multimodal dataset.
Outcome: The proposed dataset performs better than human vision and language models compared with a human dataset .
It’s not Rocket Science: Interpreting Figurative Language in Narratives (2022.tacl-1)

Copied to clipboard

Challenge: Existing text representations by design rely on compositionality, while figurative language is often non-compositional.
Approach: They propose to use a pre-trained language model to interpret figurative language types to adopt human strategies for interpreting figurativ language types: inferring meaning from context and relying on constituent words’ literal meanings.
Outcome: The proposed models perform significantly worse than humans on discriminative and generative tasks, bridging the gap from human performance.
They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse (2025.findings-acl)

Copied to clipboard

Challenge: a recent study shows that large language models lack the pragmatic capabilities needed to interpret highly implicit content.
Approach: They propose to use transcribed italian political speeches to test their ability to interpret implicit content.
Outcome: The proposed model provides a fully correct explanation in only one-fourth of cases in the open-ended generation setup.
Identifying Exaggerated Language (2020.emnlp-main)

Copied to clipboard

Challenge: Recent studies on metaphor and metonymy have focused on hyperbole, but it is a relatively understudied phenomenon in the figurative language processing community.
Approach: They propose to use hyperbole detection to determine whether a sentence is hyperbolic . they also perform statistical and manual analyses of the corpus and address the automatic hyperbola detection task.
Outcome: The proposed dataset consists of 709 hyperbolic sentences with a non-hyperbolic version created by paraphrasing its hyperbolical counterpart.
Sheep’s Skin, Wolf’s Deeds: Are LLMs Ready for Metaphorical Implicit Hate Speech? (2025.acl-long)

Copied to clipboard

Challenge: specialized models fail to detect implicit hate speech due to its indirectly expressed hateful intent . advanced LLMs often misinterpret metaphorical implicit hate content, resulting in its propagation .
Approach: They propose a Jailbreaking strategy and Energy-based Constrained Decoding techniques to detect implicit hate speech in large language models.
Outcome: The proposed model can generate metaphorical implicit hate speech, but it fails to detect it effectively.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations