Challenge: Existing datasets for scoring text pairs in terms of semantic similarity contain instances whose resolution differs according to the degree of difficulty.
Approach: They propose to use lexical overlap to distinguish obvious from non-obvious text pairs by focusing on item difficulty and ground-truth labels to characterise existing datasets.
Outcome: The proposed models are based on lexical overlap and ground-truth labels and focus on cases of similarity which require more complex inference.

Similar Papers

A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
When Shallow is Good Enough: Automatic Assessment of Conceptual Text Complexity using Shallow Semantic Features (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches to automatic assessment of text complexity focus on syntactic and lexical complexity.
Approach: They propose to use graph-based deep semantic features to automatically assess conceptual text complexity by using DBpedia as a proxy to human knowledge.
Outcome: The proposed features outperform the state-of-the-art features on pairwise comparison of two versions of the same text and five-level classification task.
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies.
Approach: They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances .
Outcome: The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations.
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging.
Approach: They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned .
Outcome: The proposed methods are compared with existing models and compare them with existing ones.
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons (2024.lrec-main)

Copied to clipboard

Challenge: In this paper, we examine the ability of large language models (LLMs) to identify different meanings in sentences that are superficially similar.
Approach: They propose a challenge dataset for NLP with large lexical overlap which minimises the possibility of models discerning entailment solely based on token distinctions.
Outcome: The proposed model fails to distinguish between constructions with three classes of adjectives which cannot be distinguished by surface features.
FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction (2025.findings-emnlp)

Copied to clipboard

Challenge: evaluating the usefulness of language models for literary-domain tasks remains challenging due to the cost of fine-grained annotation for long-form texts and data contamination concerns inherent in using public-domain literature.
Approach: They use a dataset of long-form, recently written fiction to evaluate embedding models . they prioritize author agency and rely on continual, informed author consent .
Outcome: The proposed dataset of long-form, recently written fiction is compared with existing models on this task.
An Instance Level Approach for Shallow Semantic Parsing in Scientific Procedural Text (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to parse scientific text using grammatically similar labeled sentences are limited and expensive to create.
Approach: They propose a method where semantic labels from structurally similar sentences are copied to test sentences.
Outcome: The proposed approach outperforms baseline and prior methods by 0.75 to 3 F1 absolute in the wet lab protocol corpus and 1 F1 absolut in the materials science procedural text corpus.
Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status (2022.lrec-1)

Copied to clipboard

Challenge: comparing NLP systems by performance has become an essential question . comparing systems by performing performance criterion is criticized for allowing chance to determine superiority .
Approach: They propose to use bootstrap confidence intervals instead of state-of-the-art status and statistical significance testing to compare NLP system performance.
Outcome: The bootstrap confidence intervals are used to compare NLP system performance . the bootstrap test is more accurate than state-of-the-art status and statistical significance testing .
Meaning Representations for Natural Languages: Design, Models and Applications (2022.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation models.
Approach: This tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation models.
Outcome: This tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation models . it also reviews the applications of meaning representation in downstream NLP tasks and real-world applications .
Evaluating Historical Text Normalization Systems: How Well Do They Generalize? (N18-2)

Copied to clipboard

Challenge: Historical text normalization systems aim to convert historical wordforms to their modern equivalents . many of these systems have been developed and tested on a single language .
Approach: They propose to use a nave baseline system to evaluate historical text normalization systems . they show that the models generalize well to unseen words in tests on five languages .
Outcome: The proposed models generalize well to unseen words on five languages, but provide no clear benefit over the nave baseline.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations