Papers by Ivan Kartáč

1 papers
Reasoning Gets Harder for LLMs Inside A Dialogue (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in task-oriented dialogue (TOD).
Approach: They propose to use a dynamic benchmark to examine how framing reasoning tasks within task-oriented dialogue (TOD) affect LLM performance.
Outcome: The proposed model performs well on isolated tasks and in task-oriented dialogues, but performance is inconsistent between them.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations