| Challenge: | Existing corpora cover less than 5,000 instances of less than 100 different idiom types . large corpus allows for better evaluation of assumptions about idiomatic expressions . |
| Approach: | They propose to build the largest-to-date corpus of idioms for English using crowdsourcing methods. |
| Outcome: | The proposed corpus is larger than existing resources and contains rich metadata and is made publicly available. |
Similar Papers
Potential Idiomatic Expression (PIE)-English: Corpus for Classes of Idioms (2022.lrec-1)
Copied to clipboard
Tosin Adewumi, Roshanak Vadoodi, Aparajita Tripathy, Konstantina Nikolaido, Foteini Liwicki, Marcus Liwicki
| Challenge: | Potential Idiomatic Expression (PIE) dataset for NLP in English contains over 20,100 samples with almost 1,200 cases of idioms from 10 classes (or senses). |
| Approach: | They present a large Potential Idiomatic Expression (PIE) dataset for Natural Language Processing (NLP) in English. |
| Outcome: | The proposed dataset contains over 20,100 samples with almost 1,200 cases of idioms (with their meanings) from 10 classes (or senses). |
Designing a Russian Idiom-Annotated Corpus (L18-1)
Copied to clipboard
| Challenge: | a pilot experiment using the idiom-annotated corpus of Russian is described . corpora that could be used for training idiomatic classifiers are scarce, especially if one turns to other languages. |
| Approach: | They describe the development of an idiom-annotated corpus of Russian . the corpus is compiled from freely available online resources . |
| Outcome: | The proposed corpus is based on an online corpus of Russian texts . it is available for research purposes and can be used for linguistic studies and pedagogy . |
SLIDE - a Sentiment Lexicon of Common Idioms (L18-1)
Copied to clipboard
| Challenge: | Compositional solutions for phrase sentiment are not able to handle idioms because their sentiment is not derived from the sentiment of the individual words. |
| Approach: | They propose a crowdsourcing approach for collecting sentiment annotations of idiomatic expressions using crowdsourcing. |
| Outcome: | The proposed approach is able to capture sentiment strength and ambiguity in idiomatic expressions using crowdsourcing. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Beyond Multiword Expressions: Processing Idioms and Metaphors (P18-5)
Copied to clipboard
| Challenge: | idioms and metaphors processing is a rapidly growing area in NLP, says dr. s. robertson . idiomatic idiomas are characteristic to all areas of human activity and to all types of discourse. |
| Approach: | This tutorial will provide attendees with a clear notion of idioms and metaphors . it will provide them with computational models of linguistic characteristics and methods . |
| Outcome: | This tutorial aims to provide attendees with a clear notion of the linguistic characteristics of idioms and metaphors . it outlines how to model idiomatic idiomes and their processing and what resources are available to support their use . |
Dedicated Language Resources for Interdisciplinary Research on Multiword Expressions: Best Thing since Sliced Bread (2020.lrec-1)
Copied to clipboard
| Challenge: | Multiword expressions are challenging for disciplines like NLP, psycholinguistics and second language acquisition due to their more or less fixed character. |
| Approach: | They propose to develop tools and language resources that are crucial for multifaceted research. |
| Outcome: | The proposed tools and language resources are crucial for this kind of multifaceted research. |
AStitchInLanguageModels: Dataset and Methods for the Exploration of Idiomaticity in Pre-Trained Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing datasets are limited to providing the degree of idiomaticity of expressions along with the literal and, where applicable, (a single) non-literal interpretation of MWEs. |
| Approach: | They propose to use a dataset to test the effectiveness of a language model in generating representations of sentences containing idioms. |
| Outcome: | The proposed model performs reasonably well on the one-shot and few-shot scenarios, but there is scope for improvement in the zero-shot scenario. |
Examining the Tip of the Iceberg: A Data Set for Idiom Translation (L18-1)
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) has been widely used in recent years with significant improvements for many language pairs. |
| Approach: | They propose to use a large-scale data set to evaluate idiom translation in GermanEnglish. |
| Outcome: | The proposed dataset is used to perform preliminary NMT experiments on idiom translation in GermanEnglish. |
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages (2026.acl-long)
Copied to clipboard
Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto
| Challenge: | idioms are a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation. |
| Approach: | They propose a multilingual idiom dataset that provides idiomatic expressions in both sentence-level and conversational contexts. |
| Outcome: | The proposed model performs well with low-resource idioms, but lacks contextual inference. |
The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks (2026.findings-eacl)
Copied to clipboard
| Challenge: | a small-scale human evaluation confirms that the segments are highly parallel, making the dataset suitable for NLP applications. |
| Approach: | They present a first parallel corpus of Romansh idioms from 291 schoolbooks . they use automatic alignment methods to extract 207k multi-parallel segments from the books . |
| Outcome: | The proposed corpus is based on 291 schoolbook volumes, which are comparable in content for the five idioms. |