Challenge: a grammar-based dialogue dataset, GRICE, is designed to bring implicature into pragmatic reasoning in conversations . implicature recovery is a key component of open-ended dialogue reasoning .
Approach: They propose a grammar-based dialogue dataset to bring implicature into pragmatic reasoning . they use a hierarchical grammar model to generate the entire dataset .
Outcome: The proposed model shows a significant performance gap between baseline methods and human models . the model shows an overall performance boost in conversational reasoning .

Similar Papers

Understanding Conversational Implicatures in Humans and LLMs (2026.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) interpret conversational implicatures using humans as a baseline . et al.: do LLMs exhibit a human-like sensitivity to pragmatic inference?
Approach: They adopt a surprisal-based and response-based metric to measure the accuracy of implicatures . they find that LLMs employing the response- based meter exhibit human-like behavior .
Outcome: The proposed model performs better in the literal condition than in the implied condition . the model differs from humans in its understanding of conversational implicatures .
PragmatiCQA: A Dataset for Pragmatic Question Answering in Conversations (2023.findings-acl)

Copied to clipboard

Challenge: Mars? - PragmatiCQA
Approach: Mars? - The Paper .
Outcome: The proposed dataset features 6873 QA pairs that explores pragmatic reasoning in conversations over a diverse set of topics.
Data Collection and End-to-End Learning for Conversational AI (D19-2)

Copied to clipboard

Challenge: tutorial aims to familiarise research community with recent advances in statistical dialogue systems . focus of tutorial is on learning end-to-end from data and their relation to more common modular systems.
Approach: This tutorial aims to familiarise the research community with the latest advances in statistical dialogue systems . the focus of the tutorial is on recently introduced end-to-end learning for dialogue systems and their relation to more common modular systems.
Outcome: This tutorial aims to familiarise the research community with the recent advances in statistical dialogue systems for open-domain and task-based dialogue paradigms.
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .
SIGA: A Naturalistic NLI Dataset of English Scalar Implicatures with Gradable Adjectives (2024.lrec-main)

Copied to clipboard

Challenge: scalar implicatures are a phenomenon by which a speaker conveys the negation of a more informative utterance by producing a less informative .
Approach: They propose to use a dataset to investigate the ability of language models to interpret utterances with scalar implicatures.
Outcome: The proposed models perform significantly worse on in-domain and out-of-domain examples than other types of NLI examples.
DIRECT: Direct and Indirect Responses in Conversational Text Corpus (2021.findings-emnlp)

Copied to clipboard

Challenge: Neural conversation models have been able to generate fluent responses through training on a dialogue corpus, but they lack the ability to reveal the implied intentions of users.
Approach: They propose to train neural conversation models on a dialogue corpus that provides pragmatic paraphrases to advance techniques for natural language understanding in dialogue systems.
Outcome: The proposed corpus provides 71,498 pairs of indirect–direct utterance pairs accompanied by a multi-turn dialogue history extracted from the MultiWoZ dataset.
Stephanie: Step-by-Step Dialogues for Mimicking Human Interactions in Social Conversations (2025.findings-naacl)

Copied to clipboard

Challenge: a new paradigm for dialogue systems is being developed to mimic human interactions . the current single-step dialogue paradigm lacks the depth and fluidity of human interactions.
Approach: They propose a step-by-step dialogue paradigm that mimics human interactions . they use a dataset to fine-tune existing language models .
Outcome: The proposed system mimics the dynamic nature of human conversations . it is compared with existing paradigms and will be released later this year .
InterroLang: Exploring NLP Models and Datasets through Dialogue-based Explanations (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work on NLP explainability methods lacks a dialogue-based interpretability framework that can convey faithful explanations in human-understandable terms.
Approach: They adapt the conversational explanation framework TalkToModel to the NLP domain and add new NLP-specific operations such as free-text rationalization to illustrate its generalizability.
Outcome: The proposed framework can be used to explain models on three NLP tasks and is generalizable to different datasets, use cases and models.
DRInQ: Evaluating Conversational Implicature with Controlled Context Variation (2026.acl-long)

Copied to clipboard

Challenge: Recent large language models exhibit strong conversational fluency but are unreliable when interpretation depends on reasoning that integrates social and contextual cues.
Approach: They propose a semi-automated pipeline that produces question-context-interpretation instances with systematic variation to isolate pragmatic variation while holding each question’s surface form fixed.
Outcome: The proposed framework isolates pragmatic variation while holding each question’s surface form fixed.
CICERO: A Dataset for Contextualized Commonsense Inference in Dialogues (2022.acl-long)

Copied to clipboard

Challenge: Fig. 1a shows an example where commonsense knowledge is crucial in sifting relevant information from the context.
Approach: They curate a dataset of dyadic conversations with five types of utterance-level reasoning-based inferences: cause, subsequent event, prerequisite, motivation, and emotional reaction.
Outcome: The dataset contains 53,105 of such inferences from 5,672 dialogues.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations