More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages (2024.emnlp-main)
Copied to clipboard
Dominik Schlechtweg, Pierluigi Cassotti, Bill Noble, David Alfter, Sabine Schulte Im Walde, Nina Tahmasebi
| Challenge: | Word Usage Graphs (WUGs) represent word sense clusters from simple pairwise word use judgments. |
| Approach: | They propose to use a weighted graph to represent human semantic proximity judgments for pairs of word uses to infer word sense clusters from simple pairwise word use judgments. |
| Outcome: | The proposed approach can be applied in a Word Sense Induction (WSI) setting or for Word sense disambiguation (WSD) it is the first and to date largest manually annotated, diachronic WUG dataset. |
Similar Papers
DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for graded contextual word meaning annotation have not been implemented yet. |
| Approach: | They propose a multi-round incremental annotation process and a clustering algorithm to group usages into senses to create a large-scale dataset. |
| Outcome: | The proposed method is the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments. |
Enriching Word Usage Graphs with Cluster Definitions (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing word usage graphs lack human interpretability of senses. |
| Approach: | They propose to enrich existing word usage graphs with cluster labels functioning as sense definitions. |
| Outcome: | The proposed dataset matches the definitions chosen from WordNet by two baseline systems. |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
The DURel Annotation Tool: Human and Computational Measurement of Semantic Proximity, Sense Clusters and Semantic Change (2024.eacl-demo)
Copied to clipboard
Dominik Schlechtweg, Shafqat Mumtaz Virk, Pauline Sander, Emma Sköldberg, Lukas Theuer Linke, Tuo Zhang, Nina Tahmasebi, Jonas Kuhn, Sabine Schulte Im Walde
| Challenge: | DURel is an open source tool for semantic proximity between word uses. |
| Approach: | They present an open-source tool for the annotation of semantic proximity between word uses. |
| Outcome: | The proposed tool supports standardized human annotation and computational annotation, building on recent advances with Word-in-Context models. |
DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing approaches to dialectology have been limited and rarely address language variation regarding lexical meaning. |
| Approach: | They propose to use existing framework DURel and framework-embedded Word Usage Graphs to distinguish, visualize and interpret diatopic lexical semantic variation of contextualized words in Spanish from these perspectives. |
| Outcome: | The proposed dataset exploits existing frameworks for annotating word senses in context and framework-embedded Word Usage Graphs (WUGs) . it distinguishes, visualizes and interprets lexical semantic variation of contextualized words in Spanish from these two perspectives, i.e., semasiological and onomasiology. |
Huge Automatically Extracted Training-Sets for Multilingual Word SenseDisambiguation (L18-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation is a crucial task in Natural Language Processing . supervised systems need to be trained on word-by-word basis, a problem that is beyond reach for resource-rich languages like English. |
| Approach: | They release six large-scale sense-annotated datasets in multiple languages to pave the way for supervised multilingual Word Sense Disambiguation. |
| Outcome: | The results show that large-scale sense annotations can be used as training sets for supervised systems. |
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)
Copied to clipboard
| Challenge: | The first workshop on crowdsourcing for NLP is open to all . |
| Approach: | The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks. |
| Outcome: | The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data . |
Proceedings of the Thirteenth Workshop on Graph-Based Methods for Natural Language Processing (TextGraphs-13) (D19-53)
Copied to clipboard
| Challenge: | TextGraphs is a workshop on graph-based methods for natural language processing . the workshop is being organized in conjunction with the 9th International Joint Conference on Natural Language Processing . |
| Approach: | TextGraphs is the 13th edition of the Workshop on Graph-Based Methods for Natural Language Processing . the workshop promotes synergy between GT and natural language processing . |
| Outcome: | the 2013 edition of TextGraphs is being held in conjunction with the 9th International Joint Conference on Natural Language Processing in Hong Kong. |
WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations (N19-1)
Copied to clipboard
| Challenge: | Existing word embeddings cannot model the dynamic nature of words’ semantics, i.e., the property of words to correspond to potentially different meanings. |
| Approach: | They propose a large-scale Word in Context dataset, called WiC, which is curated by experts and can be used to evaluate context-sensitive representations. |
| Outcome: | The proposed models outperform the standard evaluation dataset for the purpose and highlight their shortcomings. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |