A Girl Has A Name: Detecting Authorship Obfuscation (2020.acl-main)

Copied to clipboard

Challenge: Existing authorship attribution methods are not stealthy as they degrade text smoothness in detectable manner.
Approach: They evaluate the stealthiness of authorship attribution methods under an adversarial threat model and show that they are not stealthy .
Outcome: The proposed methods can be identified with an average F1 score of 0.87 .

Similar Papers

Adversarial Authorship Attribution for Deobfuscation (2022.acl-long)

Copied to clipboard

Challenge: Existing authorship attribution approaches do not consider adversarial threat model . authors show adversarially trained authorship attributors can degrade effectiveness of existing obfuscators from 20-30% to 5-10% .
Approach: They propose to use rule-based and learning-based text obfuscation approaches to counter authorship attribution.
Outcome: The proposed approaches do not consider the adversarial threat model . authors show that adversarially trained attributors can degrade effectiveness of existing obfuscators from 20-30% to 5-10% .
Heuristic Authorship Obfuscation (P19-1)

Copied to clipboard

Challenge: Existing methods for authorship verification are insufficient to control the authorial style of a text.
Approach: They propose a novel method that models writing style difference as the Jensen-Shannon distance between character n-gram distributions of texts and manipulates an author’s subconsciously encoded writing style using heuristic search.
Outcome: The proposed approach defeats state-of-the-art verification approaches while keeping text changes at a minimum.
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to protect the identity and privacy of online authorship are lacking supervision data for diverse authorship and domains.
Approach: They propose an unsupervised inference-time approach to authorship obfuscation that uses a user-controlled, inference time algorithm to oblige the authorship.
Outcome: The proposed method outperforms state-of-the-art methods while performing competitively against a propriety model two orders of magnitudes larger.
Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verification (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models have been driven by large-scale training corpora drawn from diverse sources such as websites, news articles, and books.
Approach: They propose a framework for analyzing dynamic relationships among LLM-enabled AO, AM, and AV in the context of authorship privacy.
Outcome: The proposed framework analyzes the dynamic relationships among LLM-enabled AO, AM, and AV in the context of authorship privacy.
A Multifaceted Framework to Evaluate Evasion, Content Preservation, and Misattribution in Authorship Obfuscation Techniques (2022.emnlp-main)

Copied to clipboard

Challenge: Authorship obfuscation techniques are often evaluated based on their ability to hide the author’s identity (evasion) while preserving the content of the original text.
Approach: They propose to evaluate authorship obfuscation techniques on detection evasion and content preservation using competitive identification techniques in real-life scenarios.
Outcome: The proposed method reveals key weaknesses in state-of-the-art obfuscation techniques and surprisingly competitive effectiveness from a back-translation baseline in all evaluation aspects.
UPTON: Preventing Authorship Leakage from Public Text Release via Data Poisoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent authorship attribution models can reveal the true authorship of unseen texts with high accuracies, with some cases up to 95% accuracy.
Approach: They propose a solution that weakens authorship features in training samples and makes released texts unlearnable by exploiting black-box data poisoning methods.
Outcome: The proposed model weakens authorship features in training samples and makes released texts unlearnable.
Keep it Private: Unsupervised Privatization of Online Text (2024.naacl-long)

Copied to clipboard

Challenge: Authorship obfuscation has been evaluated in narrow settings in the NLP literature . superficial edit operations can lead to unnatural outputs, authors say .
Approach: They propose an automatic text privatization framework that fine-tunes a large language model via reinforcement learning to produce rewrites that balance soundness, sense, and privacy.
Outcome: The proposed method maintains high text quality according to automated metrics and human evaluation, and successfully evades several automated authorship attacks.
Authorship Obfuscation in Multilingual Machine-Generated Text Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Language Modeling have birthed Large Language Models (LLMs), which exhibit significant improvements, including the ability to generate texts easily misconstrued as humanwritten.
Approach: They compare authorship obfuscation methods against machine-generated text (MGT) in 11 languages and analyze their performance against 37 well-known AO methods.
Outcome: The proposed methods can cause evasion of detection in all languages, with homoglyph attacks particularly successful.
StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements (2024.emnlp-main)

Copied to clipboard

Challenge: Authorship obfuscation methods that ignore author-specific stylistic features are often too rigid and lead to degradation of fluency and grammaticality.
Approach: They propose an adaptive obfuscation method that perturbs stylistic elements of text . authors release a large set of 30K high-quality, long-form texts from a diverse set of 14 authors .
Outcome: The proposed method outperforms state-of-the-art methods on an array of domains on automatic and human evaluation.
Open-World Authorship Attribution (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for large language models do not evaluate their performance in academic research . authors aim to identify authors from anonymous text without additional information .
Approach: They propose a benchmark to quantitatively assess LLMs' ability to infer author from text . they propose 'open-world' authorship attribute' to be a two-stage framework .
Outcome: The proposed approach achieves 60.7% accuracy and 44.3% accuracy in two stages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations