Papers with FERMAT
Can Vision-Language Models Evaluate Handwritten Math? (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in Vision-Language Models (VLMs) have significantly enhanced the ability to interpret both textual and visual data. |
| Approach: | They propose a benchmark to assess VLMs’ ability to detect, localize and correct errors in handwritten mathematical content. |
| Outcome: | The proposed benchmark covers over 2,200 handwritten math solutions from 609 manually curated problems from grades 7-12 with intentionally introduced perturbations. |
FERMAT: An Alternative to Accuracy for Numerical Reasoning (2023.acl-long)
Copied to clipboard
| Challenge: | Existing numerical reasoning models are too weak for downstream tasks like fact-checking . FERMAT evaluates models on number understanding, mathematical operations, and training dependency . |
| Approach: | They propose a multi-view evaluation set for numerical reasoning in English that evaluates models on key numerical reasoning aspects instead of reporting a single score on a whole dataset. |
| Outcome: | FERMAT evaluates models on number understanding, mathematical operations, and training dependency. |