| Challenge: | In traditional linguistics, there exists a famous saying that one should know a word by the company it keeps. |
| Approach: | They propose a method which uses word embeddings to identify pairwise relational features in the context of authorship attribution. |
| Outcome: | The proposed method is based on three literary corpora and shows that word similarity is a key feature in the authorship attribution task. |
Similar Papers
What represents “style” in authorship attribution? (C18-1)
Copied to clipboard
| Challenge: | Authorship attribution uses all information representing content and style whereas stylometry is robust in cross-domain settings. |
| Approach: | They analyze the role of syntax and lexical words in representing style . they show that syntax may be helpful for cross-genre attribution . |
| Outcome: | The proposed model may not be effective alone and needs to be combined with other robust models. |
What company do words keep? Revisiting the distributional semantics of J.R. Firth & Zellig Harris (2022.naacl-main)
Copied to clipboard
| Challenge: | linguists J.R. Firth and Zellig Harris are often credited with the invention of "distributional semantics" a close reading of their work uncovers two distinct and in many ways divergent theories of meaning . |
| Approach: | They propose to compare two different theories of meaning that focus on internal workings of linguistic forms with a broader cultural and situational context. |
| Outcome: | The authors examine the differences between their theories of meaning and the internal workings of linguistic forms . they find that Firth could guide the field towards a more culturally grounded notion of semantics . |
Topic or Style? Exploring the Most Useful Features for Authorship Attribution (C18-1)
Copied to clipboard
| Challenge: | Existing approaches to authorship attribution rely on individual's writing style and/or preferred topics. |
| Approach: | They analyse four widely used datasets to explore how different types of features affect authorship attribution accuracy under varying conditions. |
| Outcome: | The proposed model outperforms the state-of-the-art on two out of the four datasets used. |
Exploring the Value of Personalized Word Embeddings (2020.coling-main)
Copied to clipboard
| Challenge: | a subset of words belonging to specific psycholinguistic categories vary more in their representations across users . combining generic and personalized word embeddings yields the best performance . |
| Approach: | They propose personalized word embeddings and compare their performance to generic ones . they show that personalized word representations can be leveraged for improved performance . |
| Outcome: | The proposed model can be used for authorship attribution. |
Who’s the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attribution (2025.findings-emnlp)
Copied to clipboard
| Challenge: | despite growing interest in explainable NLP, it remains unclear how explanation strategies shape user behavior in tasks like authorship identification. |
| Approach: | They propose two explanation types to support their analysis of user behavior . they use example-based style rewrites and feature-based rationales to generate explanations . |
| Outcome: | The proposed explanations support appropriate reliance, whereas explanations increase AI overreliance, the study finds . |
Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution (2025.coling-main)
Copied to clipboard
| Challenge: | Recent authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. |
| Approach: | They propose a method for interpreting latent authorship representations by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point. |
| Outcome: | The proposed method outperforms baseline methods on the authorship attribution task by +20% on average when aided with explanations from the method. |
Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | Word embeddings are powerful representations that form the foundation of many natural language processing architectures. |
| Approach: | They explore word embedding stability in a wide range of languages to gain insight into their stability. |
| Outcome: | The proposed results provide insights into word embedding stability in English and other languages. |
Can Authorship Representation Learning Capture Stylistic Features? (2023.tacl-1)
Copied to clipboard
Andrew Wang, Cristina Aggazzotti, Rebecca Kotula, Rafael Rivera Soto, Marcus Bishop, Nicholas Andrews
| Challenge: | Existing methods to disentangle an author's style from the content of their writing are limited by the reliance on human labels and the narrow focus of stylistic distinctions. |
| Approach: | They propose to use a surrogate task to learn authorship representations that are sensitive to writing style and to validate their hypothesis . |
| Outcome: | The proposed representations are sensitive to writing style and may be robust to topic drift over time. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Mode Effects’ Challenge to Authorship Attribution (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing studies on authorship attribution have shown that authorial style changes with respect to sentence length, word use, readability, and certain part-of-speech ratios. |
| Approach: | They propose to measure the effect of writing mode on authorial style in a corpus of documents composed online and offline using a traditional word processor. |
| Outcome: | The authors show that online writing differs from offline writing in terms of sentence length, word use, readability, and certain part-of-speech ratios. |