Challenge: a novel authorship attribution method is developed for short texts . delta measures are well-established, but N-gram tracing is not robust enough .
Approach: They propose to use delta measures and N-gram tracing to compare short texts . they find they are highly sensitive to the choice of authors and texts in the corpus .
Outcome: The proposed methods are highly sensitive to the selection of authors and texts in the comparison corpus.

Similar Papers

Topic or Style? Exploring the Most Useful Features for Authorship Attribution (C18-1)

Copied to clipboard

Challenge: Existing approaches to authorship attribution rely on individual's writing style and/or preferred topics.
Approach: They analyse four widely used datasets to explore how different types of features affect authorship attribution accuracy under varying conditions.
Outcome: The proposed model outperforms the state-of-the-art on two out of the four datasets used.
Authorship Attribution for Neural Text Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in deep learning have enabled the generation of realistic artifacts . however, the qualities of texts generated by these models are better, often confusing classifiers if they are not real.
Approach: They propose to use neural network-based language models to generate realistic texts . they investigate the authorship attribution problem in three versions of a text .
Outcome: The proposed models generate texts that are difficult to distinguish from human-written ones . the results show that most generators still generate texts significantly different from human ones compared to other models .
The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for detecting texts generated by large language models are disputed . authors argue that there are limitations in the current technology .
Approach: They propose to make LLM detectors robust against domain shifts and build benchmarks . they argue that the limitations lie elsewhere, and open the realm of authorship analysis technology .
Outcome: The proposed method systematically analyzes the benchmarks and validates it using state-of-the-art detectors.
AttributionBench: How Hard is Automatic Attribution Evaluation? (2024.findings-acl)

Copied to clipboard

Challenge: generative search engines enhance the reliability of large language model responses by providing cited evidence.
Approach: They propose to use a benchmark to evaluate whether a large language model supports the generated responses or not .
Outcome: The proposed benchmark shows that even a fine-tuned GPT-3.5 only achieves around 80% macro-F1 under a binary classification formulation.
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution (2024.emnlp-main)

Copied to clipboard

Challenge: Authorship attribution relies on manual features and fails to capture long-range correlations, limiting their effectiveness.
Approach: They propose to use Bayesian methods to calculate the probability that a text entails previous writings of an author.
Outcome: The proposed model can achieve 85% accuracy on the IMDb and blog datasets.
Whodunit? Learning to Contrast for Authorship Attribution (2022.aacl-main)

Copied to clipboard

Challenge: Existing approaches to authorship attribution are dataset-dependent and yield inconsistent performance across corpora.
Approach: They propose to fine-tune pre-trained generic language representations with a contrastive objective to learn author-specific representations by identifying clusters of authors.
Outcome: The proposed approach improves on multiple human and machine authorship attribution benchmarks, but at the cost of sacrificing performance for some authors.
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)

Copied to clipboard

Challenge: n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear.
Approach: They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics.
Outcome: The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand.
What represents “style” in authorship attribution? (C18-1)

Copied to clipboard

Challenge: Authorship attribution uses all information representing content and style whereas stylometry is robust in cross-domain settings.
Approach: They analyze the role of syntax and lexical words in representing style . they show that syntax may be helpful for cross-genre attribution .
Outcome: The proposed model may not be effective alone and needs to be combined with other robust models.
Evaluating and Modeling Attribution for Cross-Lingual Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Open-retrieval question answering systems are lacking in attribution for cross-lingual question answering . open-research questions are available in 20 languages, but their raw generation often falls short in factuality .
Approach: They are the first to study attribution for cross-lingual question answering . they collect data in 5 languages to assess the attribution level of a state-of-the-art QA system .
Outcome: The proposed approach improves the attribution level of a state-of-the-art cross-lingual QA system.
CCTAA: A Reproducible Corpus for Chinese Authorship Attribution Research (2022.lrec-1)

Copied to clipboard

Challenge: a lack of standard, reproducible testbeds for authorship attribution in Chinese language documents impedes progress.
Approach: They propose a Chinese Cross-Topic Authorship Attribution corpus for Chinese prose . it is the first standard testbed for authorship attribution on contemporary Chinese pros.
Outcome: The proposed testbed is the first standard testbed for authorship attribution on Chinese prose.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations