Papers with *reasoning

2 papers
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks for large language models for domain specific tasks are coarse and do not provide a multidimensional evaluation of a model's ability to interpret domain specific data.
Approach: They propose a diagnostic benchmark grounded in national qualification exams that exposes critical gaps across four dimensions: expert visual reasoning of charts, logical validity via expert-verified rationales, Korean-specific geo-cultural comprehension, and fine-grained domain analysis.
Outcome: The proposed model outperforms global models in local contexts, demonstrating that parameter scaling alone cannot resolve cultural dependencies.
Argument-Based Consistency in Toxicity Explanations of LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods to evaluate free-form toxicity explanations are overly relying on input text perturbations.
Approach: They propose a multi-dimensional criterion to evaluate LLMs' reasoning about toxicity . they conduct experiments on three Llama models and an 8B Ministral model .
Outcome: The proposed criterion measures the extent to which LLMs’ free-form toxicity explanations reflect an ideal and logical argumentation process.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations