| Challenge: | Existing word embedding models resemble semantic similarity solely by distribution, but there seems to be a need for future judgments to measure similarity in full context and along more than a single spectrum. |
| Approach: | They propose a model-agnostic similarity goal standard for evaluating Danish word embeddings based on human judgments made by 42 native speakers of Danish. |
| Outcome: | The goal standard is applied to evaluate Danish word embeddings on 42 native speakers of Danish. |
Similar Papers
Just Rank: Rethinking Evaluation with Word and Sentence Similarities (2022.acl-long)
Copied to clipboard
| Challenge: | Word and sentence similarity tasks are the de facto evaluation method for embeddings. |
| Approach: | They propose a new intrinsic evaluation method called EvalRank which shows a much stronger correlation with downstream tasks. |
| Outcome: | The proposed method shows a much stronger correlation with downstream tasks and is released for future benchmarking purposes. |
On the Correlation of Word Embedding Evaluation Metrics (2020.lrec-1)
Copied to clipboard
| Challenge: | Word embeddings are geometrical representations of word paradigmatics and syntagmatics. |
| Approach: | They propose to investigate evaluation metrics on various datasets to find correlations . they propose a fast solution to select the best word embeddings among many others . |
| Outcome: | The proposed method could be used to select the best word embeddings among many others. |
A Rank-Based Similarity Metric for Word Embeddings (P18-2)
Copied to clipboard
| Challenge: | Word Embeddings have become a standard for word representations, with vector cosine being the only similarity metric. |
| Approach: | They propose to use rank-based similarity estimation metrics to measure word similarity . they find WE outperforms vector cosine in the recent outlier detection task . |
| Outcome: | The proposed rank-based measure outperforms vector cosine in the recent outlier detection task. |
SimLex-999 for Dutch (2024.lrec-main)
Copied to clipboard
| Challenge: | Word embeddings have revolutionised natural language processing by effectively representing words as dense vectors. |
| Approach: | They developed a Dutch variant of the SimLex-999 word similarity dataset by gathering similarity judgements from 235 native Dutch speakers. |
| Outcome: | The proposed model outperforms Bertje and RobBERT in terms of human similarity ratings and better represents semantic similarities between words. |
SemR-11: A Multi-Lingual Gold-Standard for Semantic Similarity and Relatedness for Eleven Languages (L18-1)
Copied to clipboard
| Challenge: | SemR-11 is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages. |
| Approach: | This paper describes a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages. |
| Outcome: | The dataset is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages. |
Word Embedding Evaluation in Downstream Tasks and Semantic Analogies (2020.lrec-1)
Copied to clipboard
| Challenge: | Language Models (LMs) are an oft studied area of natural language processing . Word Embeddings (WE) are vector space representations of a vocabulary . |
| Approach: | They evaluate Word Embeddings (WE) models for the Portuguese langauage . results show that a diverse corpus can often outperform a larger, less textually diverse corp. |
| Outcome: | The proposed models outperform a larger, less textually diverse corpus in two tasks . the evaluation shows that a diverse and comprehensive corpus outperformed a smaller, less diverse corp. |
Towards a Danish Semantic Reasoning Benchmark - Compiled from Lexical-Semantic Resources for Assessing Selected Language Understanding Capabilities of Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | a semantic reasoning benchmark for Danish is compiled from human-curated lexical-semantic resources. |
| Approach: | They present a semantic reasoning benchmark for Danish compiled semi-automatically from a number of human-curated lexical-semantic resources. |
| Outcome: | The proposed datasets are compiled semi-automatically from human-curated lexical-semantic resources. |
What’s in Your Embedding, And How It Predicts Task Performance (C18-1)
Copied to clipboard
| Challenge: | Attempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful. |
| Approach: | They propose a method that quantifies interpretable characteristics of word vector neighborhoods and shows how they correlate with performance on 14 extrinsic and intrinsic task datasets. |
| Outcome: | The proposed approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations. |
IceBATS: An Icelandic Adaptation of the Bigger Analogy Test Set (2022.lrec-1)
Copied to clipboard
| Challenge: | a new test set that measures word embeddings' ability to recognize linguistic regularities is presented in a paper in elijsson, iran . the test sets are a good quality estimator for extrinsic evaluation of word embedded models . |
| Approach: | They propose a test set that measures language models' ability to recognize linguistic regularities in a balanced way. |
| Outcome: | The proposed set is apt at measuring the capabilities of word embedding models. |
ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over Centuries (2021.emnlp-main)
Copied to clipboard
| Challenge: | Word embeddings learn implicit biases from word co-occurrence statistics . valNorm is a new intrinsic evaluation task and method to quantify affect in word embedded word sets . |
| Approach: | They propose a method to quantify valence dimension of affect in human-rated word sets . they apply ValNorm to embeddings from seven languages and 200 years of text . |
| Outcome: | The proposed method achieves a high accuracy in quantifying the valence of non-discriminatory, non-social group word sets. |