Papers by Matti Wiegmann

7 papers
The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for detecting texts generated by large language models are disputed . authors argue that there are limitations in the current technology .
Approach: They propose to make LLM detectors robust against domain shifts and build benchmarks . they argue that the limitations lie elsewhere, and open the realm of authorship analysis technology .
Outcome: The proposed method systematically analyzes the benchmarks and validates it using state-of-the-art detectors.
Argumentation and Domain Discourse in Scholarly Articles on the Theory of International Relations (2025.coling-main)

Copied to clipboard

Challenge: SKILL project aims to provide students with AI tools to facilitate analysis of argumentation in scholarly articles on international relations.
Approach: They propose to use AI to analyze argumentation in scholarly articles on international relations . they use a dataset, discourse analysis, and baseline experiments to examine argumentation and domain content types .
Outcome: The proposed method enables educationally-relevant insight into scholarly IR discourse . it requires domain-specific training and fine-tuning on relation and content type prediction tasks.
Trigger Warning Assignment as a Multi-Label Document Classification Problem (2023.acl-long)

Copied to clipboard

Challenge: a trigger warning is used to warn people about potentially disturbing content . a webis dataset of 1 million fanfiction works contains up to 36 different warnings per document .
Approach: They introduce a multi-label task to assign a trigger warning to fanfiction . they map 41 million free-form tags assigned by authors into a single taxonomy of trigger warnings .
Outcome: The proposed model achieves micro-F1 scores of about 0.5, which reveals the difficulty of the task.
Celebrity Profiling (P19-1)

Copied to clipboard

Challenge: Using a corpus of 71,706 verified accounts, we construct a profile of a wide cross-section of local and global celebrities.
Approach: They propose to use Twitter feeds of 71,706 verified accounts to build a corpus of celebrity profiles using Wikidata crawling.
Outcome: The proposed corpus contains an average of 29,968 words per profile and up to 239 pieces of personal information.
Trigger Warnings: Bootstrapping a Violence Detector for Fan Fiction (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing guidelines for proactively alerting readers of potentially disturbing content have been proposed.
Approach: They propose to use a labeled corpus of narrative fiction from a popular fan fiction site to determine whether to assign a trigger warning to an English story.
Outcome: The proposed task achieves F1 scores between 0.8 and 0.9 on three datasets . the authors show that assigning trigger warnings for violence is feasible .
Crowdsourcing a Large Corpus of Clickbait on Twitter (C18-1)

Copied to clipboard

Challenge: Clickbait is a nuisance on social media.
Approach: a corpus of 38,517 annotated Twitter tweets was constructed to detect clickbait . the corpus was annotating tweets on 4-point scale by five annotators at Amazon's Mechanical Turk .
Outcome: The corpus of 38,517 annotated Twitter tweets was used to evaluate 12 clickbait detectors submitted to the Clickbait Challenge 2017 .
Analyzing Persuasion Strategies of Debaters on Social Media (2022.coling-1)

Copied to clipboard

Challenge: Existing studies on the analysis of persuasion in online discussions focus on the effectiveness of comments in individual discussions and ignore the effectiveness analysis of debaters over multiple discussions.
Approach: They propose to quantify debaters effectiveness in the online discussion platform "ChangeMyView" they aim to explore diverse insights into their persuasion strategies .
Outcome: The proposed analysis of debater effectiveness in the ChangeMyView subreddit reveals that debaters have different levels of effectiveness, behavioral characteristics and text stylistic features .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations