Challenge: Existing approaches to authorship attribution rely on individual's writing style and/or preferred topics.
Approach: They analyse four widely used datasets to explore how different types of features affect authorship attribution accuracy under varying conditions.
Outcome: The proposed model outperforms the state-of-the-art on two out of the four datasets used.

Similar Papers

What represents “style” in authorship attribution? (C18-1)

Copied to clipboard

Challenge: Authorship attribution uses all information representing content and style whereas stylometry is robust in cross-domain settings.
Approach: They analyze the role of syntax and lexical words in representing style . they show that syntax may be helpful for cross-genre attribution .
Outcome: The proposed model may not be effective alone and needs to be combined with other robust models.
Mode Effects’ Challenge to Authorship Attribution (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies on authorship attribution have shown that authorial style changes with respect to sentence length, word use, readability, and certain part-of-speech ratios.
Approach: They propose to measure the effect of writing mode on authorial style in a corpus of documents composed online and offline using a traditional word processor.
Outcome: The authors show that online writing differs from offline writing in terms of sentence length, word use, readability, and certain part-of-speech ratios.
The Topic Confusion Task: A Novel Evaluation Scenario for Authorship Attribution (2021.findings-emnlp)

Copied to clipboard

Challenge: Autorship attribution is the problem of identifying the most plausible author of an anonymous text from a set of candidate authors.
Approach: They propose a topic confusion task where they switch the author-topic configuration between training and testing sets and propose attribution errors that are caused by the topic shift and by the features’ inability to capture the writing styles.
Outcome: The proposed task combines author-topic configuration with other features to lower topic confusion and higher attribution accuracy.
Can Authorship Representation Learning Capture Stylistic Features? (2023.tacl-1)

Copied to clipboard

Challenge: Existing methods to disentangle an author's style from the content of their writing are limited by the reliance on human labels and the narrow focus of stylistic distinctions.
Approach: They propose to use a surrogate task to learn authorship representations that are sensitive to writing style and to validate their hypothesis .
Outcome: The proposed representations are sensitive to writing style and may be robust to topic drift over time.
Who’s the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attribution (2025.findings-emnlp)

Copied to clipboard

Challenge: despite growing interest in explainable NLP, it remains unclear how explanation strategies shape user behavior in tasks like authorship identification.
Approach: They propose two explanation types to support their analysis of user behavior . they use example-based style rewrites and feature-based rationales to generate explanations .
Outcome: The proposed explanations support appropriate reliance, whereas explanations increase AI overreliance, the study finds .
Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layers (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to authorship attribution model only learn from the output layer of pre-trained transformers, ignoring representations learned at other layers.
Approach: They propose a model that leverages the various linguistic representations learned at different layers of pre-trained transformer-based models to model the authorship attribution task more effectively.
Outcome: The proposed model performs better on out-of-domain and in-domain scenarios, while ignoring representations learned at other layers.
Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution (2025.coling-main)

Copied to clipboard

Challenge: Recent authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications.
Approach: They propose a method for interpreting latent authorship representations by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point.
Outcome: The proposed method outperforms baseline methods on the authorship attribution task by +20% on average when aided with explanations from the method.
The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for detecting texts generated by large language models are disputed . authors argue that there are limitations in the current technology .
Approach: They propose to make LLM detectors robust against domain shifts and build benchmarks . they argue that the limitations lie elsewhere, and open the realm of authorship analysis technology .
Outcome: The proposed method systematically analyzes the benchmarks and validates it using state-of-the-art detectors.
Experiments with Convolutional Neural Networks for Multi-Label Authorship Attribution (L18-1)

Copied to clipboard

Challenge: Existing methods for authorship attribution tasks are difficult, but they are effective.
Approach: They propose a CNN that averaging author probability distributions at sentence level for longer documents and treating smaller documents as sentences adapts to single-label datasets and various document sizes.
Outcome: The proposed method outperforms state-of-the-art models on a single-label AA benchmark dataset.
Rethinking the Authorship Verification Experimental Setups (2022.emnlp-main)

Copied to clipboard

Challenge: Identifying the author of a text is one of the most versatile NLP tasks, with applications ranging from plagiarism detection to forensics and monitoring the activity of cyber-criminals.
Approach: They propose five new public splits over the PAN dataset to isolate and identify biases related to the text topic and to the author’s writing style.
Outcome: The proposed models are competitive with state-of-the-art methods and generalize better on dark reddit datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations