NegPar: A parallel corpus annotated for negation (L18-1)

Copied to clipboard

Challenge: NegPar is the first parallel corpus annotated for negation in the narrative domain.
Approach: They present NegPar, a parallel corpus annotated for negation in the narrative domain . they follow the annotation guidelines in the CONANDOYLE-NEG corpus .
Outcome: The proposed corpus is based on the CONANDOYLE-NEG corpus and is reannotated to ensure more consistent and interpretable representations.

Similar Papers

Negation Scope Conversion: Towards a Unified Negation-Annotated Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Negation scope resolution models that use pre-trained language models perform worse when fine-tuned on a combined dataset.
Approach: They propose to automatically convert the negation scopes of BioScope and SFU to those of Sherlock and merge them into a unified dataset.
Outcome: The proposed method improves on the unified dataset compared to the simply combined dataset.
A review of Spanish corpora annotated with negation (C18-1)

Copied to clipboard

Challenge: Existing corpora annotated with negation information are small and not always compatible . negation is a linguistic phenomenon that is not addressed in English .
Approach: They review existing corpora annotated with negation in Spanish and analyze compatibility . they propose to develop a supervised negation processing system for Spanish .
Outcome: The proposed system will not be able to merge the small corpora in Spanish due to lack of compatibility in annotations.
Cross-lingual Annotation Projection in Legal Texts (2020.coling-main)

Copied to clipboard

Challenge: a new study examines annotation projection in text classification problems where source documents are published in multiple languages.
Approach: They propose to use word embeddings and dynamic time warping to create an annotation corpus for text classification problems where source documents are published in multiple languages.
Outcome: The proposed method is based on word embeddings and dynamic time warping . the aim is to train linguistic tools for the target language without experts .
Building an English-Chinese Parallel Corpus Annotated with Sub-sentential Translation Techniques (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that human translators often resort to different non-literal translation techniques besides literal translation . however, they receive less attention in developing natural language processing (NLP) applications.
Approach: They propose to have a better semantic control of extracting paraphrases from bilingual parallel corpora.
Outcome: The proposed method can automatically recognize different non-literal translation techniques . the results confirm the hypothesis of the proposed method .
This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have grammatical knowledge but fail to interpret negation . a recent study shows that LLMs struggle with negative sentences .
Approach: They propose to use a dataset to grasp LLMs' generalization and inference capability . they also fine-tuned models to assess whether the understanding of negation can be trained .
Outcome: The proposed model is able to generalize and infer negation in 400,000 sentences . but it is suboptimal when it comes to negation, a key step in natural language processing .
An Analysis of Negation in Natural Language Understanding Corpora (2022.acl-short)

Copied to clipboard

Challenge: Using annotator-generated examples, one can evaluate systems with synthetic language that is not representative of language in the wild.
Approach: They analyze negation in eight popular corpora spanning six natural language understanding tasks.
Outcome: The proposed corpora have few negations compared to general-purpose English and are often unimportant . state-of-the-art transformers obtain significantly worse results with instances that contain negation, especially if the negations are important.
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Negation is a fundamental operation in natural language that reverses the meaning of an expression into its opposite.
Approach: They propose a sentence-level negation understanding benchmark that measures negation performance in Korean.
Outcome: The proposed benchmark improves negation understanding and broader comprehension in Korean.
Towards the Roots of the Negation Problem: A Multilingual NLI Dataset and Model Scaling Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: Negations are key to determining sentence meaning, making them essential for logical reasoning.
Approach: They construct and publish two new textual entailment datasets in four languages with paired examples differing in negation.
Outcome: The results show that increasing the model size may improve the models’ ability to handle negations.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Parallel Text Alignment and Monolingual Parallel Corpus Creation from Philosophical Texts for Text Simplification (2021.naacl-srw)

Copied to clipboard

Challenge: Existing methods for text simplification require a lot of annotated data, however there are few suitable tools for this task.
Approach: They propose an unsupervised method for aligning text based on Doc2Vec embeddings and an alignment algorithm capable of aligning texts at different levels.
Outcome: The proposed method can be used to create a monolingual parallel corpus composed of the works of early modern philosophers and their corresponding simplified versions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations