Papers with Mathematics

3 papers
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains (2025.findings-naacl)

Copied to clipboard

Challenge: Existing benchmarks focus on specific application scenarios, emphasizing task completion but failing to dissect the underlying skills that drive these outcomes.
Approach: They propose a Massive Multitask Agent Understanding benchmark that evaluates LLMs across five domains and offline tasks.
Outcome: The Massive Multitask Agent Understanding (MMAU) benchmark evaluates models across five domains including Tool-use, Directed Acyclic Graph (DAG) QA, Data Science and Machine Learning coding, Contest-level programming and Mathematics.
Mathematical Entities: Corpora and Benchmarks (2024.lrec-main)

Copied to clipboard

Challenge: a limited amount of annotated data is available for mathematical language processing . mathematics is a highly specialized domain with its own unique set of challenges .
Approach: They provide annotated corpora that can be used to study the language of mathematics . they provide part-of-speech tags, lemmas, and dependency trees .
Outcome: The proposed corpora provide part-of-speech tags, lemmas, and dependency trees . the learning assistant grants access to the content of the corporata in a context-sensitive manner .
Assessing the Reasoning Capabilities of LLMs in the context of Evidence-based Claim Verification (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable proficiency in complex tasks where reasoning capabilities are paramount.
Approach: They propose a framework to break down claims into atomic reasoning types needed for verification.
Outcome: The proposed framework breaks down claims into atomic reasoning types needed for verification.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations