Papers with Mathematics
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains (2025.findings-naacl)
Copied to clipboard
Guoli Yin, Haoping Bai, Shuang Ma, Feng Nan, Yanchao Sun, Zhaoyang Xu, Shen Ma, Jiarui Lu, Xiang Kong, Aonan Zhang, Dian Ang Yap, Yizhe Zhang, Karsten Ahnert, Vik Kamath, Mathias Berglund, Dominic Walsh, Tobias Gindele, Juergen Wiest, Zhengfeng Lai, Xiaoming Simon Wang, Jiulong Shan, Meng Cao, Ruoming Pang, Zirui Wang
| Challenge: | Existing benchmarks focus on specific application scenarios, emphasizing task completion but failing to dissect the underlying skills that drive these outcomes. |
| Approach: | They propose a Massive Multitask Agent Understanding benchmark that evaluates LLMs across five domains and offline tasks. |
| Outcome: | The Massive Multitask Agent Understanding (MMAU) benchmark evaluates models across five domains including Tool-use, Directed Acyclic Graph (DAG) QA, Data Science and Machine Learning coding, Contest-level programming and Mathematics. |
Mathematical Entities: Corpora and Benchmarks (2024.lrec-main)
Copied to clipboard
| Challenge: | a limited amount of annotated data is available for mathematical language processing . mathematics is a highly specialized domain with its own unique set of challenges . |
| Approach: | They provide annotated corpora that can be used to study the language of mathematics . they provide part-of-speech tags, lemmas, and dependency trees . |
| Outcome: | The proposed corpora provide part-of-speech tags, lemmas, and dependency trees . the learning assistant grants access to the content of the corporata in a context-sensitive manner . |
Assessing the Reasoning Capabilities of LLMs in the context of Evidence-based Claim Verification (2025.findings-acl)
Copied to clipboard
John Dougrez-Lewis, Mahmud Elahi Akhter, Federico Ruggeri, Sebastian Löbbers, Yulan He, Maria Liakata
| Challenge: | Large Language Models (LLMs) have shown remarkable proficiency in complex tasks where reasoning capabilities are paramount. |
| Approach: | They propose a framework to break down claims into atomic reasoning types needed for verification. |
| Outcome: | The proposed framework breaks down claims into atomic reasoning types needed for verification. |