Papers with *reasoning
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation frameworks for large language models for domain specific tasks are coarse and do not provide a multidimensional evaluation of a model's ability to interpret domain specific data. |
| Approach: | They propose a diagnostic benchmark grounded in national qualification exams that exposes critical gaps across four dimensions: expert visual reasoning of charts, logical validity via expert-verified rationales, Korean-specific geo-cultural comprehension, and fine-grained domain analysis. |
| Outcome: | The proposed model outperforms global models in local contexts, demonstrating that parameter scaling alone cannot resolve cultural dependencies. |
Argument-Based Consistency in Toxicity Explanations of LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods to evaluate free-form toxicity explanations are overly relying on input text perturbations. |
| Approach: | They propose a multi-dimensional criterion to evaluate LLMs' reasoning about toxicity . they conduct experiments on three Llama models and an 8B Ministral model . |
| Outcome: | The proposed criterion measures the extent to which LLMs’ free-form toxicity explanations reflect an ideal and logical argumentation process. |