Challenge: Pragmatics understanding is not well studied in LLMs, but their understanding of pragmatics is lacking.
Approach: They propose to use a dataset to measure LLMs' understanding of pragmatics to evaluate their models.
Outcome: The proposed dataset includes 14 tasks in four pragmatics phenomena, namely; Implicature, Presupposition, Reference, and Deixis.

Similar Papers

Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that large language models are robust in commonsense reasoning . however, some variations in questions can lead to incorrect responses .
Approach: They propose a large-scale bilingual benchmark consisting of 11,200 cases . they conduct extensive experiments on 41 representative LLMs .
Outcome: The proposed benchmark systematically evaluates the robustness of large language models in commonsense reasoning.
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being used to automate programming tasks.
Approach: They propose a benchmark to evaluate LLMs' reasoning abilities on program semantics.
Outcome: The proposed benchmark shows that LLMs perform well with simple control flows but struggle with more complex structures, especially loops, even with advanced prompting.
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark (2023.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks of social language are lacking for large language models.
Approach: They propose a new benchmark that measures how well large language models understand social language by grouping 58 tasks into five categories: humor & sarcasm, offensiveness, sentiment & emotion, and trustworthiness.
Outcome: The proposed model performs well at 58 tasks that are divided into five categories: humor & sarcasm, offensiveness, sentiment & emotion, and trustworthiness.
Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to assess social-pragmatic inference in large language models are inadequacy, and preferential tuning is the best approach.
Approach: They propose to use free-form models' responses as a measure to assess social-pragmatic reasoning and advocate for preference optimization over supervised finetuning (SFT).
Outcome: The proposed model outperforms supervised finetuning (SFT) and offers a near-free launch in pragmatic abilities without compromising general capabilities.
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent speech-LLMs have shown impressive performance in tasks like transcription and translation, yet they remain limited in understanding the paralinguistic aspects of speech crucial for social and emotional intelligence.
Approach: They propose a benchmark for evaluating speech-LLMs on contextual paralinguistic reasoning . the benchmark includes curated question answering datasets requiring both linguistic and empathetic understanding .
Outcome: The proposed benchmark reveals a key gap in existing evaluations and offers insights into building more context-aware and emotionally intelligent LLMs.
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities in mathematical reasoning, but their effectiveness is limited to specific mathematical topics.
Approach: They propose to use the MaTT benchmark to assess large language models' accuracy in multiple-choice scenarios.
Outcome: The proposed model achieved 54% accuracy in a multiple-choice scenario, while the Chain-of-Thought prompting did not improve.
CUTE: Measuring LLMs’ Understanding of Their Tokens (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) perform well on a wide variety of tasks, authors say . they lack direct access to characters, which can be difficult to generalize to new languages .
Approach: They propose a benchmark to test the orthographic knowledge of Large Language Models . they find that most LLMs seem to know the spelling of their tokens - yet fail to manipulate text .
Outcome: The proposed benchmark tests the orthographic knowledge of large language models . it finds that most LLMs seem to know the spelling of their tokens, but fail to manipulate text .
From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs? (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmark datasets focus on low-level cognitive tasks while providing limited coverage of higher-level reasoning skills.
Approach: They analyze the cognitive depth of popular LLM benchmarks using Bloom’s Taxonomy to evaluate both the cognitive and knowledge dimensions.
Outcome: The results show that incorporating higher-level cognitive instructions into the current instruction fine-tuning process improves model performance.
Towards a Danish Semantic Reasoning Benchmark - Compiled from Lexical-Semantic Resources for Assessing Selected Language Understanding Capabilities of Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: a semantic reasoning benchmark for Danish is compiled from human-curated lexical-semantic resources.
Approach: They present a semantic reasoning benchmark for Danish compiled semi-automatically from a number of human-curated lexical-semantic resources.
Outcome: The proposed datasets are compiled semi-automatically from human-curated lexical-semantic resources.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations