Semantic Pleonasm Detection (N18-2)

Copied to clipboard

Challenge: Pleonasms are words that are redundant.
Approach: They propose an annotated corpus of semantic pleonasms and compare it against other corpus resources.
Outcome: The proposed corpus is validated with interannotator agreement analyses and compares it with other corpus resources.

Similar Papers

Semantic Shift Stability: Efficient Way to Detect Performance Degradation of Word Embeddings and Pre-trained Language Models (2022.aacl-main)

Copied to clipboard

Challenge: Existing methods to detect time-series performance degradation of word embeddings and pre-trained language models are not efficient.
Approach: They propose a way to detect time-series performance degradation by calculating the degree of semantic shift.
Outcome: The proposed method detects time-series performance degradation in Japanese and English datasets.
Interannotator Agreement for Lexico-Semantic Annotation of a Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a method for lexico-semantic annotation of the Basic Corpus of Polish Metaphors is described . the procedure is composed of three steps: deciding whether a particular occurrence of a word is asemantics or strictly grammatical.
Approach: They propose a procedure for lexico-semantic annotation of the Basic Corpus of Polish Metaphor . procedure corrects morphosyntactic annotation of part of corpus that is automatically annotated .
Outcome: The proposed procedure corrects the morphosyntactic annotation of part of the corpus . it is composed of three steps: deciding whether a word is asemantic or strictly grammatical . preliminary results show that the procedure is adequate for the task .
Adversarial Semantic Collisions (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to generate semantic collisions for NLP tasks are vulnerable to adversarial examples.
Approach: They propose gradient-based approaches for generating semantic collisions given white-box access to a model and deploy them against several NLP tasks.
Outcome: The proposed approaches evade perplexity-based filtering and discuss other potential mitigations.
A Survey on Detection of LLMs-Generated Content (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance .
Approach: They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy.
Outcome: The proposed model will be able to detect human-written content in real time.
Oddballness: universal anomaly detection with language models (2025.coling-main)

Copied to clipboard

Challenge: a new method to detect anomalies in texts uses a metric called oddballness . the method considers probabilities generated by a language model but not low-likelihood tokens .
Approach: They propose a method to detect anomalies in texts using unsupervised language models . they define oddballness as a function that measures how strange a given token is .
Outcome: The proposed method is better than state-of-the-art models for grammatical error detection tasks.
Close or Cloze? Assessing the Robustness of Large Language Models to Adversarial Perturbations via Word Recovery (2025.coling-main)

Copied to clipboard

Challenge: Existing models implicitly recover the original text, but it is unclear when they rely on context and when they implicitly do so.
Approach: They propose to use a dictionary to recover adversarial words by using a phonetic, typo, and visual attack to study word recovery performance.
Outcome: The proposed model outperforms open-source models on hateful, offensive, and toxic classification tasks.
Semantic-pragmatic Annotations in the Prague Dependency Treebank (2026.findings-acl)

Copied to clipboard

Challenge: morphology and syntax work on sentence level, but semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context.
Approach: They present semantic-pragmatic specification and annotations in the Prague Dependency Treebank - Consolidated 2.0 release2 by annotating the entire corpus.
Outcome: The proposed annotations are based on the Prague Dependency Treebank -Consolidated 2.0 (PDT-C 2.0) the dataset contains more than 3 million tokens (of Czech) manually annotated from morphology to surface and deep syntax including several types of semantic-pragmatic annotations.
Mind Your Bias: A Critical Review of Bias Detection Methods for Contextual Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for detection of biases in contextual language models are inconsistent and inconclusive.
Approach: They propose to use word embedding association test to detect biases in contextual language models to compare them with other methods.
Outcome: The proposed methods are inconsistent and inconclusive for language models with word embeddings.
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations (D19-3)

Copied to clipboard

Challenge: Proceedings of the system demonstrations session were presented at the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) EMNMP-IjCNLP 2019 has a Best Demo Award for the first time .
Approach: Proceedings of the system demonstrations session are available online . they were presented at the conference on empirical methods in natural language processing .
Outcome: The system demonstrations session received 110 submissions, 22 of which were either invalid or withdrawn by the authors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations