Challenge: Existing methods for detecting text reuse focus on recontextualization . current approaches focus on text reuse across a diachronic corpus .
Approach: They propose a framework that relies on topic relatedness for evaluating the diachronic change of context in which text is reused.
Outcome: The proposed framework evaluates biblical text reuse human-annotated with topic relatedness . it exhibits greater sensitivity to textual similarity than topic relatedity, the authors show .

Similar Papers

Diachronic Usage Relatedness (DURel): A Framework for the Annotation of Lexical Semantic Change (N18-2)

Copied to clipboard

Challenge: Existing frameworks for evaluating lexical semantic change are limited . evaluation of lexicals is a major obstacle in the field of semantic change detection .
Approach: They propose a framework that extends synchronic polysemy annotation to diachronic changes in lexical meaning to counteract lack of resources for evaluating computational models of lexiconal semantic change.
Outcome: The proposed framework exploits an intuitive notion of semantic relatedness and distinguishes between innovative and reductive meaning changes with high inter-annotator agreement.
Lexical and Semantic Features for Cross-lingual Text Reuse Classification: an Experiment in English and Latin Paraphrases (L18-1)

Copied to clipboard

Challenge: Analyzing historical languages is challenging because they lack primary material for certain time periods . under-resourced languages such as Ancient Greek and Latin lack advanced natural-language processing (NLP) techniques .
Approach: They propose to use machine learning to detect and classify paraphrastic text reuse in historical texts.
Outcome: The proposed method improves the accuracy of paraphrastic text reuse detection in historical languages.
Substitution-based Semantic Change Detection using Contextual Embeddings (2023.acl-short)

Copied to clipboard

Challenge: a simplified approach to measuring semantic change using contextual embeddings is proposed . the static word vectors used for measuring semantic changes are difficult to interpret .
Approach: They propose a simplified approach to measuring semantic change using contextual embeddings . they use the Jensen-Shannon Divergence between the distributions of most probable replacements for masked words in different time periods to measure semantic change.
Outcome: The proposed approach is interpretable, efficient and much more efficient than static embeddings.
Don’t Take This Out of Context!: On the Need for Contextual Models and Evaluations for Stylistic Rewriting (2023.emnlp-main)

Copied to clipboard

Challenge: Existing stylistic text rewriting methods ignore the context of the text, causing generic, incoherent, and generic outputs.
Approach: They propose a contextual evaluation metric that integrates preceding context into stylistic text rewriting.
Outcome: The proposed metric integrates the preceding textual context into rewriting and evaluation stages . human preferences are better reflected by the proposed criterio and other metrics .
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing relation extraction methods rely on exact matching with human-annotated reference relations, while GRE methods produce diverse and semantically accurate relations.
Approach: They propose a multi-dimensional assessment of relation extraction methods using human-annotated reference relations.
Outcome: The proposed method is consistent with human preferences for RE quality.
The DURel Annotation Tool: Human and Computational Measurement of Semantic Proximity, Sense Clusters and Semantic Change (2024.eacl-demo)

Copied to clipboard

Challenge: DURel is an open source tool for semantic proximity between word uses.
Approach: They present an open-source tool for the annotation of semantic proximity between word uses.
Outcome: The proposed tool supports standardized human annotation and computational annotation, building on recent advances with Word-in-Context models.
A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains (P19-1)

Copied to clipboard

Challenge: Existing models for diachronic and synchronic detection of lexical semantic divergences are superficial and lack of comparison.
Approach: They propose to extend benchmark models on a common state-of-the-art evaluation task . they also demonstrate that the same evaluation task and modelling approaches can be utilised for synchronic detection of domain-specific sense divergences in the field of term extraction.
Outcome: The proposed model can be utilised for the detection of domain-specific sense divergences in the field of term extraction.
Contextualized Topic Coherence Metrics (2024.findings-eacl)

Copied to clipboard

Challenge: Existing topic models that estimate the interpretability of topics are difficult to compare due to their nature as unsupervised models.
Approach: They propose to use contextualized topic coherence metrics to simulate human-centered coherency evaluation while maintaining the efficiency of other automated methods.
Outcome: The proposed metrics better reflect human judgment on topics extracted from short text collections by avoiding highly scored topics that are meaningless to humans.
Neural RST-based Evaluation of Discourse Coherence (2020.aacl-main)

Copied to clipboard

Challenge: Existing discourse parsers cannot predict coherent texts without using silver-standard features.
Approach: They propose a tree-recursive neural model which takes advantage of the text’s RST features produced by a state of the art RST parser and compares it to the current state of art.
Outcome: The proposed model achieves state-of-the-art accuracy on the Grammarly Corpus for Discourse Coherence (GCDC) and has 62% fewer parameters than existing models.
Automated Evaluation of Out-of-Context Errors (L18-1)

Copied to clipboard

Challenge: Existing methods to modify text understanding systems use only one sentence at a time . however, considering a larger context can improve performance for text understanding tasks.
Approach: They propose to modify existing text data to insert out-of-context errors . they use a 2016 TEDTalk corpus to evaluate computational models for text understanding .
Outcome: The proposed method targets real-world problems of transcription and translation systems by inserting authentic out-of-context errors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations