DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing approaches to dialectology have been limited and rarely address language variation regarding lexical meaning. |
| Approach: | They propose to use existing framework DURel and framework-embedded Word Usage Graphs to distinguish, visualize and interpret diatopic lexical semantic variation of contextualized words in Spanish from these perspectives. |
| Outcome: | The proposed dataset exploits existing frameworks for annotating word senses in context and framework-embedded Word Usage Graphs (WUGs) . it distinguishes, visualizes and interprets lexical semantic variation of contextualized words in Spanish from these two perspectives, i.e., semasiological and onomasiology. |
Similar Papers
DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for graded contextual word meaning annotation have not been implemented yet. |
| Approach: | They propose a multi-round incremental annotation process and a clustering algorithm to group usages into senses to create a large-scale dataset. |
| Outcome: | The proposed method is the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments. |
Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)
Copied to clipboard
| Challenge: | Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power . |
| Approach: | They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus. |
| Outcome: | The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences. |
Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and Synonymy (2021.acl-long)
Copied to clipboard
| Challenge: | Existing models that represent different senses of words in context are not accurate for polysemous words. |
| Approach: | They propose a multilingual dataset that evaluates the ability of models to accurately represent different lexical-semantic relations such as homonymy and synonymy. |
| Outcome: | The proposed models can disambiguate homonyms in context, but fail to represent words with different senses when occurring in similar sentences. |
Syn2Vec: Synset Colexification Graphs for Lexical Semantic Similarity (2022.naacl-main)
Copied to clipboard
| Challenge: | In this paper we examine patterns of colexification as an aspect of lexical-semantic organization, and compare several approaches to build large scale graphs across 499 world languages. |
| Approach: | They propose to use patterns of colexification as an aspect of lexical-semantic organization to build large scale synset graphs across a typologically diverse set of 499 world languages. |
| Outcome: | The proposed models are evaluated against human judgments on a semantic similarity task for nine languages. |
Splits! Flexible Sociocultural Linguistic Investigation at Scale (2026.acl-long)
Copied to clipboard
| Challenge: | Variation in language use offers a rich lens into cultural perspectives, values, and opinions. |
| Approach: | They propose to construct a "sandbox" for systematic and flexible sociolinguistic research by splitting a reddit dataset into demographically/topically split SLPs. |
| Outcome: | The proposed method analyzes a demographically/topically split Reddit dataset validated by self-identification and replicating several known SLPs from existing literature. |
More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages (2024.emnlp-main)
Copied to clipboard
Dominik Schlechtweg, Pierluigi Cassotti, Bill Noble, David Alfter, Sabine Schulte Im Walde, Nina Tahmasebi
| Challenge: | Word Usage Graphs (WUGs) represent word sense clusters from simple pairwise word use judgments. |
| Approach: | They propose to use a weighted graph to represent human semantic proximity judgments for pairs of word uses to infer word sense clusters from simple pairwise word use judgments. |
| Outcome: | The proposed approach can be applied in a Word Sense Induction (WSI) setting or for Word sense disambiguation (WSD) it is the first and to date largest manually annotated, diachronic WUG dataset. |
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)
Copied to clipboard
| Challenge: | a new method is proposed to acquire typological evidence from "gold" treebanks for different languages. |
| Approach: | They propose a method for acquiring typological evidence from "gold" treebanks for different languages. |
| Outcome: | The proposed method can shed light on key issues of the linguistic typological literature. |
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)
Copied to clipboard
Dagmar Gromann, Hugo Goncalo Oliveira, Lucia Pitarch, Elena-Simona Apostol, Jordi Bernad, Eliot Bytyçi, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabik, Jorge Gracia, Letizia Granata, Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia Di Buono, Ana Ostroški Anić, Sigita Rackevičienė, Ricardo Rodrigues, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stanković, Ciprian-Octavian Truică, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
| Challenge: | Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions. |
| Approach: | They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge. |
| Outcome: | The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian. |
A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains (P19-1)
Copied to clipboard
| Challenge: | Existing models for diachronic and synchronic detection of lexical semantic divergences are superficial and lack of comparison. |
| Approach: | They propose to extend benchmark models on a common state-of-the-art evaluation task . they also demonstrate that the same evaluation task and modelling approaches can be utilised for synchronic detection of domain-specific sense divergences in the field of term extraction. |
| Outcome: | The proposed model can be utilised for the detection of domain-specific sense divergences in the field of term extraction. |
Representing Multiword Term Variation in a Terminological Knowledge Base: a Corpus-Based Study (2020.lrec-1)
Copied to clipboard
| Challenge: | Multiword terms are the most frequent type of lexical units in scientific and technical communication. rendering them in another language is not easy due to their cognitive complexity, proliferation of different forms, and their unsystematic representation in terminographic resources. |
| Approach: | They evaluated Spanish translation variants of multiword terms in three parallel corpora, two comparable corporales and two terminological resources. |
| Outcome: | The results show that multiword terms exhibit a significant degree of term variation . the proposed model is based on a set of criteria for determining which variants should be selected . |