Challenge: a dataset for semantic model evaluation for Turkish is not available for the language . a similarity and relatedness evaluation resource is needed for higher level tasks .
Approach: They propose a semantic model evaluation dataset for Turkish that evaluates word similarity and word relatedness tasks while discriminating those two relations from each other.
Outcome: The proposed dataset is designed to evaluate word similarity and word relatedness tasks in Turkish.

Similar Papers

SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)

Copied to clipboard

Challenge: SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text .
Approach: They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu.
Outcome: The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages.
CoSimLex: A Resource for Evaluating Graded Word Similarity in Context (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to evaluate word embeddings ignore context and treat words in isolation.
Approach: They propose to build a new word embeddings-based dataset that provides context-dependent similarity measures.
Outcome: The proposed dataset provides context-dependent similarity measures and covers a well-resourced language (English) but a number of less-resource languages.
Indra: A Word Embedding and Semantic Relatedness Server (L18-1)

Copied to clipboard

Challenge: Word embedding/distributional semantic models are a fundamental component in many natural language processing (NLP) architectures.
Approach: They propose a multi-lingual word embedding/distributional semantics framework which supports creation, use and evaluation of word embedded models.
Outcome: The proposed tool supports the creation, use and evaluation of word embedding models.
SemR-11: A Multi-Lingual Gold-Standard for Semantic Similarity and Relatedness for Eleven Languages (L18-1)

Copied to clipboard

Challenge: SemR-11 is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Approach: This paper describes a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Outcome: The dataset is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
TurkishDelightNLP: A Neural Turkish NLP Toolkit (2022.naacl-demo)

Copied to clipboard

Challenge: a neural Turkish NLP toolkit performs computational linguistic analyses from morphological level to semantic level.
Approach: They propose a neural Turkish NLP toolkit that performs computational linguistic analyses from morphological level to semantic level.
Outcome: The proposed toolkit performs computational linguistic analyses from morphological level to semantic level in Turkish.
TR-MTEB: A Comprehensive Benchmark and Embedding Model Suite for Turkish Sentence Representations (2025.findings-emnlp)

Copied to clipboard

Challenge: TR-MTEB is the first large-scale, task-diverse benchmark for sentence embedding models for Turkish.
Approach: a new benchmark evaluates sentence embedding models for Turkish . TR-MTEB covers six core tasks and 26 high-quality datasets .
Outcome: The TR-MTEB benchmark covers six core tasks and includes 26 high-quality datasets . the models achieve competitive performance across most tasks and significantly improve on baseline models.
Towards a Gold Standard for Evaluating Danish Word Embeddings (2020.lrec-1)

Copied to clipboard

Challenge: Existing word embedding models resemble semantic similarity solely by distribution, but there seems to be a need for future judgments to measure similarity in full context and along more than a single spectrum.
Approach: They propose a model-agnostic similarity goal standard for evaluating Danish word embeddings based on human judgments made by 42 native speakers of Danish.
Outcome: The goal standard is applied to evaluate Danish word embeddings on 42 native speakers of Danish.
A Diverse Set of Freely Available Linguistic Resources for Turkish (2023.acl-long)

Copied to clipboard

Challenge: despite the abundance of Turkish speakers, linguistic resources for natural language processing remain scarce.
Approach: They propose a set of freely available linguistic resources for Turkish natural language processing . they provide corpora and pretrained models to help practitioners build their own applications .
Outcome: The proposed linguistic resources are first of their kind and easy to use in a broad range of implementations.
Data and Representation for Turkish Natural Language Inference (2020.emnlp-main)

Copied to clipboard

Challenge: Large annotated datasets in NLP are overwhelmingly in English . obtaining new annotation resources for each task in each language would be prohibitively expensive .
Approach: They propose to use machine translation to translate large annotated datasets into Turkish . they find that in-language embeddings are essential and morphological parsing can be avoided .
Outcome: The proposed model trains on human-translated evaluation sets.
Enhanced Word Representations for Bridging Anaphora Resolution (N18-2)

Copied to clipboard

Challenge: Existing word representations do not capture semantic similarity for bridging anaphora resolution.
Approach: They propose to use word embeddings to capture semantic similarity by exploring syntactic structure of noun phrases.
Outcome: The proposed model achieves 30% of accuracy for bridging anaphora resolution on ISNotes corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations