Challenge: GLTR is a tool to detect generated text that can be used by non-experts.
Approach: They propose a tool to detect generated text using a set of statistical methods that can be used by non-experts.
Outcome: The proposed method improves detection rate of fake text from 54% to 72% without training.

Similar Papers

Automatic Detection of Machine Generated Text: A Critical Survey (2020.coling-main)

Copied to clipboard

Challenge: Current text generative models excel in producing text that matches the style of human language reasonably well.
Approach: They conduct an in-depth error analysis of the state-of-the-art detector and discuss research directions to guide future work in this exciting area.
Outcome: The proposed detectors can distinguish between human and text generated by the model and can be used to generate fake news and fake product reviews.
Automatic Detection of Generated Text is Easiest when Humans are Fooled (2020.acl-main)

Copied to clipboard

Challenge: Recent advances in neural language modelling make it possible to rapidly generate vast amounts of human-sounding text.
Approach: They compare decoding methods with popular sampling-based decoding strategies . they show that multi-sentence excerpts can fool expert human raters over 30% of the time .
Outcome: The proposed methods improve with longer excerpt length, but multi-sentence excerpts fool human raters over 30% of the time.
Detecting Machine-Generated Text: Techniques and Challenges (2024.acl-tutorials)

Copied to clipboard

Challenge: This tutorial focuses on machine-generated text and deepfakes.
Approach: This tutorial aims to provide a comprehensive overview of text detection techniques . it will focus on machine-generated text and deepfakes .
Outcome: This tutorial focuses on machine-generated text and deepfakes.
Detecting Bot-Generated Text by Characterizing Linguistic Accommodation in Human-Bot Interactions (2021.findings-acl)

Copied to clipboard

Challenge: Language generation models' democratization makes it easier to generate human-like text at-scale for nefarious activities, from spreading misinformation to targeting specific groups with hate speech.
Approach: They propose to use linguistic alignment to detect bot-generated text rather than using it directly.
Outcome: The proposed methods are more robust across datasets and models if they use information about how people respond to it rather than using the bot's text directly.
IMGTB: A Framework for Machine-Generated Text Detection Benchmarking (2024.acl-demos)

Copied to clipboard

Challenge: MGTD methods are needed in many areas, such as prevention of disinformation spreading, plagiarism, impersonation and identity theft.
Approach: They propose a framework for machine-generated text detection that integrates custom methods and evaluation datasets into existing frameworks.
Outcome: The proposed framework simplifies the benchmarking of machine-generated text detection methods by easy integration of custom (new) methods and evaluation datasets.
Exploring the Limitations of Detecting Machine-Generated Text (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in the quality of the generation of text by large language models have spurred research into identifying machine-generated text.
Approach: They audit classification performance for detecting machine-generated text by evaluating on texts with varying writing styles.
Outcome: The proposed methods are highly sensitive to stylistic changes and complexity, and in some cases degrade entirely to random classifiers.
RoFT: A Tool for Evaluating Human Detection of Machine-Generated Text (2020.emnlp-demos)

Copied to clipboard

Challenge: Existing studies on how humans perceive machine-generated text are limited due to the prohibitive cost of running human evaluation studies.
Approach: They propose a task to detect the boundary at which a text passage starts off human-written transitions to being machine-generated.
Outcome: The proposed system evaluates machine-generated news articles on a wide range of domains.
Adversarial Text Generation via Sequence Contrast Discrimination (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to generate human-like texts are auto-regressive, but they suffer from exposure bias due to the dependence on the previous sampled output during the inferring phase.
Approach: They propose a sequence contrast loss driven text generation framework which learns the difference between real texts and generated texts and uses that difference.
Outcome: The proposed framework improves training stability and quality of generated texts and avoids the time-consuming sampling process.
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack (2024.lrec-main)

Copied to clipboard

Challenge: Despite the development of large language models, there are still significant challenges in detecting whether text is generated by a machine.
Approach: They propose a framework for a broader class of adversarial attacks to perform minor perturbations in machine-generated content to evade detection.
Outcome: The proposed framework can be compromised in as little as 10 seconds, and improves over iterative adversarial learning.
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text (2025.acl-long)

Copied to clipboard

Challenge: Qualitative analysis of experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity).
Approach: They hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions.
Outcome: The annotators who frequently use LLMs for writing tasks outperform commercial and open-source detectors even without evasion tactics like paraphrasing and humanization.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations