Papers by Tolga Buz
LLMs Cannot (Yet) Match the Specificity and Simplicity of Online Communities in Long Form Question Answering (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent years have positioned Large Language Models (LLMs) as powerful question answering (QA) tools, shifting users away from interacting in communities towards discourse with AI-driven conversational interfaces. |
| Approach: | They propose to use a QA preference dataset to fine-tune and align Large Language Models (LLMs) from more than 7.4 million submissions and 82 million comments from 2008 to 2022 in Reddit’s 15 largest finance communities. |
| Outcome: | The proposed framework improves on the social quality of the data, and the proposed framework is more accurate and more specific. |
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations (2026.eacl-long)
Copied to clipboard
Yindong Wang, Martin Preiß, Margarita Bugueño, Jan Vincent Hoffbauer, Abdullatif Ghajar, Tolga Buz, Gerard de Melo
| Challenge: | Evaluating 9 state-of-the-art LLMs reveals two critical limitations: 61% of incorrect span predictions are semantically unrelated to actual errors. |
| Approach: | They propose a benchmark of 1,001 expert-annotated question-answer pairs with span-level error annotations derived from Reddit's r/AskScience. |
| Outcome: | Evaluating 9 state-of-the-art LLMs, we find that comparative judgment is paradoxically harder than independent detection when comparing answers side-by-side. |
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit’s Showerthoughts (2024.starsem-1)
Copied to clipboard
| Challenge: | Recent Large Language Models (LLMs) have shown the ability to generate content that is difficult or impossible to distinguish from human writing. |
| Approach: | They compare GPT-2 and GPT-Neo fine-tuned on Reddit data and GTP-3.5 invoked in a zero-shot manner, against human-authored texts. |
| Outcome: | The proposed model can generate short, creative texts that are difficult to distinguish from human writing, but human evaluators rate them worse than the model. |