Papers with detecting
Copyright Detective: A Forensic System to Evidence LLMs Flickering Copyright Leakage Risks (2026.acl-demo)
Copied to clipboard
Guangwei Zhang, Jianing Zhu, Cheng Qian, Neil Zhenqiang Gong, Rada Mihalcea, Zhaozhuo Xu, Jingrui He, Jiaqi W. Ma, Chaowei Xiao, Bo Li, Ahmed Abbasi, Dongwon Lee, Heng Ji, Denghui Zhang
| Challenge: | **Copyright Detective** is the first interactive forensic system for detecting, analyzing, and visualizing potential copyright risks in LLM outputs. |
| Approach: | They propose a system that detects copyright infringements and visualizes them . they use content recall testing, paraphrase-level similarity analysis and persuasive jailbreak probing . |
| Outcome: | The proposed system detects, analyzes, and visualizes potential copyright risks in LLM outputs. |
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification (2024.eacl-long)
Copied to clipboard
| Challenge: | Few-shot text classification systems are infeasible to deploy and use reliably due to their dependence on prompting and billion-parameter language models. |
| Approach: | They propose a modification to SetFit that fine-tunes a Sentence Transformer under a contrastive learning paradigm and achieves similar results to more unwieldy systems. |
| Outcome: | The proposed model fine-tunes a Sentence Transformer under a contrastive learning paradigm and achieves similar results to more unwieldy systems. |
Implicitly Abusive Comparisons – A New Dataset and Linguistic Analysis (2021.eacl-main)
Copied to clipboard
| Challenge: | Using crowdsourcing, we can detect implicitly abusive comparisons . Abusive language is defined as hurtful, derogatory or obscene utterances made by one person to another . |
| Approach: | They propose to use crowdsourcing to generate a dataset for detecting implicitly abusive comparisons . they also use a range of linguistic features to better understand abusive comparison mechanisms . |
| Outcome: | The proposed dataset includes measures to obtain representative and unbiased comparisons. |
BigTokDetect: A Clinically-Informed Vision–Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok (2026.eacl-long)
Copied to clipboard
Minh Duc Chu, Kshitij Pawar, Zihao He, Roxanna Sharifi, Ross M. Sonnenblick, Magdalayna Curry, Laura DAdamo, Lindsay Young, Stuart Murray, Kristina Lerman
| Challenge: | Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). |
| Approach: | They propose a framework for detecting pro-bigorexia content on TikTok using an expert-annotated multimodal benchmark dataset of over 2,200 Tiktok videos labeled by clinical psychiatrists. |
| Outcome: | The proposed framework improves on fine-grained subcategories while commercial models achieve the highest accuracy on primary categories. |
KIMERA: Injecting Domain Knowledge into Vacant Transformer Heads (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies show that transformer models lack specific domain knowledge and are under-performing in broad domains like the medical domain. |
| Approach: | They propose a method for retraining and instilling attention heads with structured domain knowledge by masking redundant attention heads. |
| Outcome: | The proposed method improves on seven datasets in the medical domain in information retrieval and clinical outcome prediction settings. |
Detecting, Generating, and Evaluating in the Writing Style of Different Authors (2025.naacl-srw)
Copied to clipboard
| Challenge: | In recent years, stylometry has been investigated in many different fields. |
| Approach: | They propose to use sentences from different books to generate and evaluate stylistic texts according to the authors' writing styles. |
| Outcome: | The proposed model can detect, generate, and evaluate documents according to the authors' writing styles with unpaired samples. |
NLP-ADBench: NLP Anomaly Detection Benchmark (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Anomaly detection (AD) is an important machine learning task, but its effectiveness in detecting harmful content, phishing attempts, and spam reviews is limited. |
| Approach: | They introduce NLP-ADBench, the most comprehensive NLP anomaly detection benchmark to date . it includes eight curated datasets and 19 state-of-the-art algorithms . |
| Outcome: | The NLP-ADBench benchmark includes 19 state-of-the-art methods and 8 curated datasets . no single model dominates across all datasets, indicating need for automated model selection . |
Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Strong attribute control can distort meaning, while prioritizing semantic preservation may weaken attribute alignment. |
| Approach: | They propose a method that restricts accepted samples to text meeting a minimum BERTScore threshold and applies gradient-assisted proposal generation to improve attribute alignment. |
| Outcome: | a new method for counterfactual text generation improves attribute alignment and semantic preservation . the proposed method achieved the best macro F1-score in two of three test sets . |
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing algorithms for audio deepfake detection are based on layer-wise analysis of self-supervised learning (SSL) models. |
| Approach: | They conduct a layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts. |
| Outcome: | The proposed models achieve competitive equal error rate (EER) scores even when employing a reduced number of layers. |
Recent Advances in Online Hate Speech Moderation: Multimodality and the Role of Large Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | HS is any communication demeaning a person or a group based on social or ethnic characteristics that undermines social harmony and individual safety . the recent Israel-Hamas conflict has escalated both anti-Muslim and anti-Semitic sentiments worldwide . |
| Approach: | They examine the role of large language models and large multimodal models in HS moderation . they examine how text, images, and audio interact to spread hate speech . |
| Outcome: | The findings highlight the need for solutions in low-resource settings and highlight the gaps in existing methods. |
How is BERT surprised? Layerwise detection of linguistic anomalies (2021.acl-long)
Copied to clipboard
| Challenge: | a number of studies have shown that transformer-based language models detect when a word is anomalous in context, but likelihood scores do not tell the cause of the anomaly. |
| Approach: | They propose to use Gaussian models for density estimation at intermediate layers of three language models to evaluate grammaticality. |
| Outcome: | The proposed method on BLiMP shows that language models employ different mechanisms to detect different types of linguistic anomalies. |
Can We Trust AI Doctors? A Survey of Medical Hallucination in Large Language and Large Vision-Language Models (2025.findings-acl)
Copied to clipboard
Zhihong Zhu, Yunyan Zhang, Xianwei Zhuang, Fan Zhang, Zhongwei Wan, Yuyan Chen, QingqingLong QingqingLong, Yefeng Zheng, Xian Wu
| Challenge: | Hallucination is a critical challenge for large language models and large vision-language models (LVLMs) however, dedicated research on medical hallucinations remains unexplored. |
| Approach: | They provide a unified perspective on medical hallucination for both LLMs and LVLMs, and delve into its causes. |
| Outcome: | The proposed models have demonstrated impressive performance on a variety of medical benchmarks. |
Robust AI-Generated Text Detection by Restricted Embeddings (2024.findings-emnlp)
Copied to clipboard
Kristian Kuznetsov, Eduard Tulchinskii, Laida Kushnareva, German Magai, Serguei Barannikov, Sergey Nikolenko, Irina Piontkovskaya
| Challenge: | Existing approaches for artificial text detection are score-based and classifier-based . however, score-driven methods often rely on a score-derived score. |
| Approach: | They investigate the ability of classifier-based detectors to transfer to unseen generators or semantic domains. |
| Outcome: | The proposed methods improve the out-of-distribution classification score by up to 9% and 14%. |