Characterizing Idioms: Conventionality and Contingency (2022.acl-long)

Copied to clipboard

Challenge: idioms have non-canonical meanings, but non-conventional meanings are contingent on other words . a recent study shows that idiomatic expressions are not homogeneous among idiomas .
Approach: They propose to use a contingency relationship between words in an idiom and non-canonical meanings of words in the idiome.
Outcome: a new study shows that idioms fall at the expected intersection of the two dimensions, but that the dimensions themselves are not correlated.

Similar Papers

Beyond Multiword Expressions: Processing Idioms and Metaphors (P18-5)

Copied to clipboard

Challenge: idioms and metaphors processing is a rapidly growing area in NLP, says dr. s. robertson . idiomatic idiomas are characteristic to all areas of human activity and to all types of discourse.
Approach: This tutorial will provide attendees with a clear notion of idioms and metaphors . it will provide them with computational models of linguistic characteristics and methods .
Outcome: This tutorial aims to provide attendees with a clear notion of the linguistic characteristics of idioms and metaphors . it outlines how to model idiomatic idiomes and their processing and what resources are available to support their use .
Rethinking the Idiomaticity Decomposability Hypothesis: Evidence from Distributional Learning (2026.acl-long)

Copied to clipboard

Challenge: Decomposability is thought to predict syntactic flexibility, but is not attributed to distributional experience.
Approach: They propose a model-internal measure of decomposability and relate it to human ratings, syntactic flexibility, and predictability while tracking idiom learning during pretraining.
Outcome: The proposed model-internal measure correlates weakly with human judgments and shows a small but consistent negative relationship with syntactic flexibility.
Dedicated Language Resources for Interdisciplinary Research on Multiword Expressions: Best Thing since Sliced Bread (2020.lrec-1)

Copied to clipboard

Challenge: Multiword expressions are challenging for disciplines like NLP, psycholinguistics and second language acquisition due to their more or less fixed character.
Approach: They propose to develop tools and language resources that are crucial for multifaceted research.
Outcome: The proposed tools and language resources are crucial for this kind of multifaceted research.
Paraphrases do not explain word analogies (2021.eacl-main)

Copied to clipboard

Challenge: Several attempts have been made to explain distributional word embeddings as linguistic regularities as directions.
Approach: They propose to use an analogy to explain why linguistic regularities should hold in distributional word embeddings.
Outcome: The proposed explanation does not hold empirically.
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions.
Approach: They propose to use a large-scale dataset of idioms in six languages to evaluate LLMs' idiomatic processing ability.
Outcome: The proposed model integrates contextual cues and reasoning to improve idiom understanding in LLMs, suggesting that their performance is influenced by memorization and reasoning.
The Word Analogy Testing Caveat (N18-2)

Copied to clipboard

Challenge: a number of word analogy tests are used to evaluate word embeddings . word embeds are used as a proxy for semantics and syntax à la Harris .
Approach: They propose to use word embeddings as a proxy for distributional similarity . they propose to apply a transfer learning approach to word embeds to improve performance .
Outcome: The proposed method improves performance across a wide range of NLP tasks.
Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss Weighting (2023.emnlp-main)

Copied to clipboard

Challenge: idioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts.
Approach: They propose to use retrieval-augmented models to increase the accuracy of a strong pretrained machine translation model on idiomatic sentences by up to 13%.
Outcome: The proposed techniques improve the accuracy of a strong pretrained model on idiomatic sentences by up to 13% in absolute accuracy, and holds potential benefits for non-idiomatic phrases.
ID10M: Idiom Identification in 10 Languages (2022.findings-naacl)

Copied to clipboard

Challenge: Identifying and understanding idioms in context is a key goal and challenge in Natural Language Understanding tasks.
Approach: They propose a multilingual Transformer-based system for the identification of idioms and a manually-curated evaluation benchmark.
Outcome: The proposed system performs well in 10 languages and is released on github.
Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are powerful representations that form the foundation of many natural language processing architectures.
Approach: They explore word embedding stability in a wide range of languages to gain insight into their stability.
Outcome: The proposed results provide insights into word embedding stability in English and other languages.
Transactions of the Association for Computational Linguistics, Volume 8 (2020.tacl-1)

Copied to clipboard

Challenge: null
Approach: null
Outcome: null

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations