Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets (2026.eacl-long)
Copied to clipboard
| Challenge: | Novel metaphor comprehension involves complex semantic processes and linguistic creativity. |
| Approach: | They propose a cloze-style surprisal method that conditions on full-sentence context. |
| Outcome: | The proposed method shows that LM surprisal yields moderate correlations with scores/labels of metaphor novelty. |
Similar Papers
Towards a Similarity-adjusted Surprisal Theory (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that surprisal theory ignores the possibility of similarity between words and treats them as distinct entities. |
| Approach: | They propose a new measure of comprehension effort called information value that accounts for communicative equivalences between possible continuations. |
| Outcome: | The proposed measure of comprehension effort is based on the diversity index of the diversity of communicative units. |
A Corpus of Metaphor Novelty Scores for Syntactically-Related Word Pairs (L18-1)
Copied to clipboard
| Challenge: | Existing data on metaphor novelty are limited, making it difficult to perform research on this topic. |
| Approach: | They propose to release a corpus of metaphor novelty scores for syntactically related word pairs . they establish a performance benchmark to which future researchers can compare . |
| Outcome: | The proposed corpus of metaphor novelty scores is compared to other datasets . it performs better than chance or nave strategies, the authors show . |
Word Surprisal Correlates with Sentential Contradiction in LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing models are primarily optimized for task-specific performance, lacking well-defined objectives or linguistic grounding. |
| Approach: | They propose a token-to-word decoding algorithm that extends theoretically grounded probability estimation to open-vocabulary settings. |
| Outcome: | The proposed algorithm can localize sentence-level inconsistency at the word level, establishing a quantitative link between lexical uncertainty and sentential semantics. |
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal (2025.acl-long)
Copied to clipboard
| Challenge: | Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens. |
| Approach: | They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions. |
| Outcome: | The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000. |
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage (2025.findings-acl)
Copied to clipboard
| Challenge: | Language models (LMs) have been shown to flexibly capture many linguistic regularities from raw text, but the source stimuli of reading time datasets are often naturalistic text that are available online. |
| Approach: | They propose to replicate the negative relationship between language model size and the fit of surprisal to reading times using models trained on ‘leakage-free’ data that overlaps only minimally with the reading time corpora. |
| Outcome: | The proposed models show that language models trained on 'leakage-free' data are not driven by data leakage. |
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly (2026.eacl-short)
Copied to clipboard
| Challenge: | Recent work has observed an inverse scaling relationship between Transformers’ per-word estimated probability and the predictive power of their surprisal estimates on reading times. |
| Approach: | They conducted a more comprehensive evaluation using surprisal estimates from 17 pre-trained LMs on two functional magnetic resonance imaging datasets. |
| Outcome: | Recent work shows that surprisal from larger Transformer-based models is less predictive of reading times, resolving the inconclusive results and indicating that this trend is not specific to latency-based measures. |
MetaPro 2.0: Computational Metaphor Processing on the Effectiveness of Anomalous Language Modeling (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for metaphor interpretation are slow due to lack of annotated datasets and effective pre-trained language models. |
| Approach: | They propose a large annotated dataset and a PLM for the metaphor interpretation task. |
| Outcome: | The proposed method improves on metaphor identification and interpretation with comparable baselines on the new dataset. |
Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on metaphor processing have focused on single datasets and specific task settings, often using artificially constructed data through lexical replacement. |
| Approach: | They propose to evaluate the capabilities of Large Language Models (LLMs) in metaphor interpretation across multiple datasets, tasks, and prompt configurations. |
| Outcome: | The proposed frameworks are more realistic and efficient than current models and are more efficient than existing models. |
Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages (2022.acl-long)
Copied to clipboard
| Challenge: | Existing studies on pre-trained language models assume they encode metaphorical knowledge useful for NLP systems. |
| Approach: | They propose to probing metaphoricity information in PLMs and measure their generalization . they find that contextual representations in PMLs encode metaphorical knowledge . |
| Outcome: | The proposed model can encode metaphorical knowledge across languages and datasets . the model can be used to train and test NLP systems . |
Verifying Claims About Metaphors with Large-Scale Automatic Metaphor Identification (2024.naacl-short)
Copied to clipboard
| Challenge: | Existing studies on metaphors have focused on a small number of examples, whereas few studies verify claims with large corpus. |
| Approach: | They propose to use a large corpus to verify existing claims about verb metaphors . they apply metaphor detection to sentences extracted from Common Crawl . |
| Outcome: | The proposed method identifies verb metaphors with lower concreteness, imageability, familiarity and more emotional and subjective sentences. |