| Challenge: | NegPar is the first parallel corpus annotated for negation in the narrative domain. |
| Approach: | They present NegPar, a parallel corpus annotated for negation in the narrative domain . they follow the annotation guidelines in the CONANDOYLE-NEG corpus . |
| Outcome: | The proposed corpus is based on the CONANDOYLE-NEG corpus and is reannotated to ensure more consistent and interpretable representations. |
Similar Papers
Negation Scope Conversion: Towards a Unified Negation-Annotated Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Negation scope resolution models that use pre-trained language models perform worse when fine-tuned on a combined dataset. |
| Approach: | They propose to automatically convert the negation scopes of BioScope and SFU to those of Sherlock and merge them into a unified dataset. |
| Outcome: | The proposed method improves on the unified dataset compared to the simply combined dataset. |
A review of Spanish corpora annotated with negation (C18-1)
Copied to clipboard
| Challenge: | Existing corpora annotated with negation information are small and not always compatible . negation is a linguistic phenomenon that is not addressed in English . |
| Approach: | They review existing corpora annotated with negation in Spanish and analyze compatibility . they propose to develop a supervised negation processing system for Spanish . |
| Outcome: | The proposed system will not be able to merge the small corpora in Spanish due to lack of compatibility in annotations. |
Cross-lingual Annotation Projection in Legal Texts (2020.coling-main)
Copied to clipboard
| Challenge: | a new study examines annotation projection in text classification problems where source documents are published in multiple languages. |
| Approach: | They propose to use word embeddings and dynamic time warping to create an annotation corpus for text classification problems where source documents are published in multiple languages. |
| Outcome: | The proposed method is based on word embeddings and dynamic time warping . the aim is to train linguistic tools for the target language without experts . |
Building an English-Chinese Parallel Corpus Annotated with Sub-sentential Translation Techniques (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study shows that human translators often resort to different non-literal translation techniques besides literal translation . however, they receive less attention in developing natural language processing (NLP) applications. |
| Approach: | They propose to have a better semantic control of extracting paraphrases from bilingual parallel corpora. |
| Outcome: | The proposed method can automatically recognize different non-literal translation techniques . the results confirm the hypothesis of the proposed method . |
This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have grammatical knowledge but fail to interpret negation . a recent study shows that LLMs struggle with negative sentences . |
| Approach: | They propose to use a dataset to grasp LLMs' generalization and inference capability . they also fine-tuned models to assess whether the understanding of negation can be trained . |
| Outcome: | The proposed model is able to generalize and infer negation in 400,000 sentences . but it is suboptimal when it comes to negation, a key step in natural language processing . |
An Analysis of Negation in Natural Language Understanding Corpora (2022.acl-short)
Copied to clipboard
| Challenge: | Using annotator-generated examples, one can evaluate systems with synthetic language that is not representative of language in the wild. |
| Approach: | They analyze negation in eight popular corpora spanning six natural language understanding tasks. |
| Outcome: | The proposed corpora have few negations compared to general-purpose English and are often unimportant . state-of-the-art transformers obtain significantly worse results with instances that contain negation, especially if the negations are important. |
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Negation is a fundamental operation in natural language that reverses the meaning of an expression into its opposite. |
| Approach: | They propose a sentence-level negation understanding benchmark that measures negation performance in Korean. |
| Outcome: | The proposed benchmark improves negation understanding and broader comprehension in Korean. |
Towards the Roots of the Negation Problem: A Multilingual NLI Dataset and Model Scaling Analysis (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Negations are key to determining sentence meaning, making them essential for logical reasoning. |
| Approach: | They construct and publish two new textual entailment datasets in four languages with paired examples differing in negation. |
| Outcome: | The results show that increasing the model size may improve the models’ ability to handle negations. |
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |
Parallel Text Alignment and Monolingual Parallel Corpus Creation from Philosophical Texts for Text Simplification (2021.naacl-srw)
Copied to clipboard
| Challenge: | Existing methods for text simplification require a lot of annotated data, however there are few suitable tools for this task. |
| Approach: | They propose an unsupervised method for aligning text based on Doc2Vec embeddings and an alignment algorithm capable of aligning texts at different levels. |
| Outcome: | The proposed method can be used to create a monolingual parallel corpus composed of the works of early modern philosophers and their corresponding simplified versions. |