Papers by Diana McCarthy
Acquiring Verb Classes Through Bottom-Up Semantic Verb Clustering (L18-1)
Copied to clipboard
| Challenge: | Existing methods for creating verbal classifications are limited or non-existent in most languages . a range of automatic verb classification approaches have been proposed, but high-quality resources are needed . |
| Approach: | They propose to use top-up semantic clustering to extract syntactic and semantic information from verbs in English, Polish and Croatian. |
| Outcome: | The proposed classifications in English, Polish and Croatian are compared with other languages. |
Towards Better Context-aware Lexical Semantics:Adjusting Contextualized Representations through Static Anchors (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent research has shown that contextualized models generate dynamic embeddings for words in context, but static embedds are often overlooked in this trend towards contextualized modeling. |
| Approach: | They propose a method that learns a transformation through static anchors and requires only another pre-trained model. |
| Outcome: | The proposed method improves a range of benchmark tasks that test contextual variations of meaning across different usages of a word and across different words as they are used in context. |
AM2iCo: Evaluating Word Meaning in Context across Low-Resource Languages with Adversarial Examples (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing multilingual evaluation datasets that evaluate lexical semantics "in-context" have various limitations, including limited coverage of high-resource languages and superficial cues. |
| Approach: | They propose to use a set of pretrained language models to evaluate lexical semantics in context. |
| Outcome: | The proposed set shows that current models lag behind human performance in interpreting word meaning in cross-lingual contexts. |
Spatial Multi-Arrangement for Clustering and Multi-way Similarity Dataset Construction (2020.lrec-1)
Copied to clipboard
Olga Majewska, Diana McCarthy, Jasper van den Bosch, Nikolaus Kriegeskorte, Ivan Vulić, Anna Korhonen
| Challenge: | Existing methods for creating large-scale semantic similarity resources are slow and expensive . a large verb similarity dataset is available for a number of verbs, but not for English. |
| Approach: | They propose a method for fast bottom-up creation of large-scale semantic similarity resources . they leverage semantic intuitions of native speakers and adapt a spatial multi-arrangement approach to lexical stimuli. |
| Outcome: | The proposed approach produces a large-scale verb similarity dataset containing similarity scores for 29,721 unique verb pairs and 825 target verbs. |
Manual Clustering and Spatial Arrangement of Verbs for Multilingual Evaluation and Typology Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods to learn general language representations from large volumes of unlabeled text have been used to improve multilingual NLP. |
| Approach: | They propose to use a spatial arrangement method to generate large-scale evaluation datasets that balance cross-lingual alignment with language specificity. |
| Outcome: | The proposed method produces semantic verb classes and fine-grained similarity scores for nearly 130 thousand verb pairs. |
Measuring Context-Word Biases in Lexical Semantic Datasets (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing pretrained contextualized models have been used to evaluate word-in-context representations in many lexical semantic tasks. |
| Approach: | They propose to quantify the degree of context or word biases in existing datasets by probing masked input. |
| Outcome: | The proposed model performs better when both word and context are available than with masked input. |