Challenge: Most corpora for textual coherence evaluation are composed of randomly shuffled sentences that focus on sentence ordering.
Approach: They propose to use a variety of corruption strategies to build a corpus of incoherent pairs of sentences by swapping their discourse connective or a discourse argument.
Outcome: The proposed corpus is constructed from discourse argument pairs from the Penn Discourse Tree Bank and is compared with existing corpus models.

Similar Papers

How coherent are neural models of coherence? (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to model coherence are limited to small newswire corpora . evaluators need to be trained on lexical and document levels to perform evaluations .
Approach: They propose four generic evaluation tasks that capture coherence-specific properties . they aim at capturing correct use of discourse connectives and lexical cohesion .
Outcome: The proposed tasks capture coherence-specific properties, including correct use of discourse connectives, lexical cohesion, temporal consistency among events and participants in a story.
Is Incoherence Surprising? Targeted Evaluation of Coherence Prediction from Language Models (2021.naacl-main)

Copied to clipboard

Challenge: a common approach to coherence evaluation is shuffling the sentence order of a text, creating incoherent text samples that need to be discriminated from the original.
Approach: They propose an extendable set of test suites addressing different aspects of discourse and dialogue coherence.
Outcome: The proposed evaluation paradigm is suited to evaluate linguistic qualities that contribute to the notion of coherence.
Jigsaw Pieces of Meaning: Modeling Discourse Coherence with Informed Negative Sample Synthesis (2024.findings-eacl)

Copied to clipboard

Challenge: Existing work on creating “informed” incoherent samples for coherence modeling has focused on permutations of a coherent document .
Approach: They propose to use Constituency trees, Part-of-speech, semantic overlap to create “informed” negative samples that better represent or mimic incoherence.
Outcome: The proposed methods improve the quality of the negative sample.
DDisCo: A Discourse Coherence Dataset for Danish (2022.lrec-1)

Copied to clipboard

Challenge: Discourse coherence models have been developed using randomly shuffled texts instead of highly edited and coherent data.
Approach: They propose to annotate Danish Wikipedia and Reddit for discourse coherence using real-world text instead of artificially incoherent text for training and testing models.
Outcome: The proposed model performs well on annotated texts from the Danish Wikipedia and Reddit dataset.
Joint Modeling of Entities and Discourse Relations for Coherence Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work on coherence modeling focuses on entity features or discourse relation features, with little attention given to combining the two.
Approach: They propose two methods for jointly modeling entities and discourse relations for coherence assessment.
Outcome: The proposed methods significantly improve the performance of coherence models on three benchmark datasets.
Coherent or Not? Stressing a Neural Language Model for Discourse Coherence in Multiple Languages (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on coherence assessment using NLMs focuses on properties acquired from stand-alone sentences, but their ability to model discourse and pragmatic phenomena is still unclear.
Approach: They propose to use a Neural Language Model to assess coherence in multiple languages to compare models' performance and to examine their performance in a cross-language scenario.
Outcome: The proposed model can model coherent and incoherent text in multiple languages and in-domain settings.
COHESENTIA: A Novel Benchmark of Incremental versus Holistic Assessment of Coherence in Generated Texts (2023.emnlp-main)

Copied to clipboard

Challenge: linguistics has been used to assess the coherence of generated texts . a benchmark of coherency scores is developed to measure the quality of generated text .
Approach: They propose a benchmark to assess coherence of automatically generated texts . they use global and incremental methods to score sentences for coherency .
Outcome: The proposed benchmark measures human-perceived coherence of automatically generated texts . it uses global and incremental scoring, and shows that the models are unsatisfactory .
How to Find Strong Summary Coherence Measures? A Toolbox and a Comparative Study for Summary Coherence Measure Evaluation (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to evaluate summary coherence are often evaluated using disparate datasets and metrics.
Approach: They propose to use automatic evaluation to evaluate coherence of summaries by selecting high-scoring candidates.
Outcome: The proposed methods show that they can perform better on an even playing field.
DiscoScore: Evaluating Text Generation with BERT and Discourse Coherence (2023.eacl-main)

Copied to clipboard

Challenge: DiscoScore is a parametrized discourse metric that uses BERT to model discourse coherence . it is weak when operated at system level, and is therefore not reliable in a way to spot improvements .
Approach: They propose a parametrized discourse metric which uses BERT to model discourse coherence from different perspectives.
Outcome: The proposed model outperforms existing models on document-level machine translation and summarization.
A Novel Computational Modeling Foundation for Automatic Coherence Assessment (2025.naacl-long)

Copied to clipboard

Challenge: Existing models for text coherence assessment rely on a proxy task . however, this approach does not capture the full range of factors contributing to coherency.
Approach: They propose a formal linguistic definition of what makes a discourse coherent and formalize these conditions as respective computational tasks that are jointly trained.
Outcome: The proposed model improves on two human-rated coherence benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations