Papers with FERMAT

2 papers
Can Vision-Language Models Evaluate Handwritten Math? (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in Vision-Language Models (VLMs) have significantly enhanced the ability to interpret both textual and visual data.
Approach: They propose a benchmark to assess VLMs’ ability to detect, localize and correct errors in handwritten mathematical content.
Outcome: The proposed benchmark covers over 2,200 handwritten math solutions from 609 manually curated problems from grades 7-12 with intentionally introduced perturbations.
FERMAT: An Alternative to Accuracy for Numerical Reasoning (2023.acl-long)

Copied to clipboard

Challenge: Existing numerical reasoning models are too weak for downstream tasks like fact-checking . FERMAT evaluates models on number understanding, mathematical operations, and training dependency .
Approach: They propose a multi-view evaluation set for numerical reasoning in English that evaluates models on key numerical reasoning aspects instead of reporting a single score on a whole dataset.
Outcome: FERMAT evaluates models on number understanding, mathematical operations, and training dependency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations