A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts (2024.acl-long)
Copied to clipboard
Nafis Irtiza Tripto, Saranya Venkatraman, Dominik Macko, Robert Moro, Ivan Srba, Adaku Uchendu, Thai Le, Dongwon Lee
| Challenge: | Using a computational approach, we discover that diminishing performance in text classification models is closely associated with the extent of deviation from the original author’s style. |
| Approach: | They propose to use large language models to determine whether a text retains original authorship when it undergoes numerous paraphrasing iterations. |
| Outcome: | The results suggest that authorship should be task-dependent . |
Similar Papers
Can Large Language Models Identify Authorship? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated exceptional capacity for reasoning and problem-solving, but their potential in authorship analysis remains under-explored. |
| Approach: | They propose to integrate explicit linguistic features into LLMs to provide explanations into their reasoning processes. |
| Outcome: | The proposed models demonstrate their ability to perform zero-shot, end-to-end authorship verification effectively and provide explainability through explicit linguistic features. |
Large Language Models Threaten Language’s Epistemic and Communicative Foundations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models are reshaping the norms of human communication, sometimes decouping words from genuine human thought. |
| Approach: | They propose to model humans, LLMs, and texts in a provenance network . they propose to use epistemic doppelgängers to generate texts that are indis- tinguishable from human-authored texts . |
| Outcome: | The proposed models induce semantic drift, erode account-ability, and obfuscate intent and authorship. |
All That Glitters is Not Novel: Plagiarism in AI Generated Research (2025.acl-long)
Copied to clipboard
| Challenge: | Recent studies claim autonomous research agents can generate novel research ideas. |
| Approach: | They ask experts to evaluate whether existing work is similar to new ones . they find 24% of the 50 evaluated documents to be either paraphrased or significantly borrowed . |
| Outcome: | The authors find that 24% of the 50 evaluated research documents are either paraphrased, or significantly borrowed from existing work. |
What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | a new study examines the evolution of language models as a time-specific collection of models of interest. |
| Approach: | They investigate the problem of "Ship of Language Models" where scientific evolution takes form of continuous retrofits of key *existing* terms. |
| Outcome: | The proposed model is based on recent NLP publications and is quantitatively analyzed. |
The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for detecting texts generated by large language models are disputed . authors argue that there are limitations in the current technology . |
| Approach: | They propose to make LLM detectors robust against domain shifts and build benchmarks . they argue that the limitations lie elsewhere, and open the realm of authorship analysis technology . |
| Outcome: | The proposed method systematically analyzes the benchmarks and validates it using state-of-the-art detectors. |
Authorship Attribution for Neural Text Generation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in deep learning have enabled the generation of realistic artifacts . however, the qualities of texts generated by these models are better, often confusing classifiers if they are not real. |
| Approach: | They propose to use neural network-based language models to generate realistic texts . they investigate the authorship attribution problem in three versions of a text . |
| Outcome: | The proposed models generate texts that are difficult to distinguish from human-written ones . the results show that most generators still generate texts significantly different from human ones compared to other models . |
A Survey on LLMs for Story Generation (2025.findings-emnlp)
Copied to clipboard
Maria Teleki, Vedangi Bengali, Xiangjue Dong, Sai Tejas Janjur, Haoran Liu, Tian Liu, Cong Wang, Ting Liu, Yin Zhang, Frank Shipman, James Caverlee
| Challenge: | Methods for story generation with Large Language Models (LLMs) have come into the spotlight recently. |
| Approach: | They propose a novel taxonomy of LLMs for story generation consisting of two major paradigms: independent story generation by an LLM, and author-assistance for story creation . |
| Outcome: | The proposed taxonomy compares existing work on the topic with those of novel author-assistance models. |
Authorship Attribution in Multilingual Machine-Generated Texts (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have reached human-like fluency and coherence, but distinguishing machine-generated text from human-written content becomes increasingly difficult. |
| Approach: | They propose a problem of multilingual authorship attribution (AA) that involves attributing texts to human or multiple LLM generators across diverse languages. |
| Outcome: | The proposed method can be adapted to multilingual settings, but still has significant limitations and challenges. |
Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verification (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large language models have been driven by large-scale training corpora drawn from diverse sources such as websites, news articles, and books. |
| Approach: | They propose a framework for analyzing dynamic relationships among LLM-enabled AO, AM, and AV in the context of authorship privacy. |
| Outcome: | The proposed framework analyzes the dynamic relationships among LLM-enabled AO, AM, and AV in the context of authorship privacy. |
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution (2024.emnlp-main)
Copied to clipboard
| Challenge: | Authorship attribution relies on manual features and fails to capture long-range correlations, limiting their effectiveness. |
| Approach: | They propose to use Bayesian methods to calculate the probability that a text entails previous writings of an author. |
| Outcome: | The proposed model can achieve 85% accuracy on the IMDb and blog datasets. |