Papers by Michael Krumdick
BizBench: A Quantitative Reasoning Benchmark for Business and Finance (2024.acl-long)
Copied to clipboard
| Challenge: | Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge. |
| Approach: | They propose a benchmark for evaluating models’ ability to reason about realistic financial problems by focusing on question-answering over financial data via program synthesis. |
| Outcome: | The proposed benchmark evaluates models' financial background knowledge, ability to parse financial documents, and capacity to solve complex problems with code. |
DocFinQA: A Long-Context Financial Reasoning Dataset (2024.acl-short)
Copied to clipboard
| Challenge: | Existing work on automating financial numerical reasoning focuses on unrealistically specific document snippets, failing to reflect the broader and more realistic scenarios faced by analysts. |
| Approach: | They propose a long-document financial QA task that augments 7,437 questions from existing FinQA dataset with full-document context, extending the average context length from under 700 words in FinQA to 123k words in DocFinQA. |
| Outcome: | The proposed task extends the average context length from under 700 words in FinQA to 123k words in DocFinQA. |
Language Model Probabilities are Not Calibrated in Numeric Contexts (2025.acl-long)
Copied to clipboard
Charles Lovering, Michael Krumdick, Viet Dac Lai, Varshini Reddy, Seth Ebner, Nilesh Kumar, Rik Koncel-Kedziorski, Chris Tanner
| Challenge: | Using language model outputs, we find that even in simple settings, the best LMs (1) are poorly calibrated and (2) have systematic biases. |
| Approach: | They argue that language model outputs should capture natural distributions over multiple options within their textual contexts. |
| Outcome: | The proposed model outputs are calibrated to the numeric content of their contexts. |
An Analysis of Multilingual FActScore (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in LLMs have demonstrated significant capabilities in many applications. |
| Approach: | They propose a dataset for FActScore on texts generated by strong multilingual LLMs and evaluate their performance in other languages. |
| Outcome: | The proposed dataset shows that LLMs exhibit distinct behaviors in fact extraction and fact scoring tasks. |
On Finding Inconsistencies in Documents (2026.findings-acl)
Copied to clipboard
Charles Lovering, Seth Ebner, Brandon Smock, Michael Krumdick, Muhammad Saad Rabbani, Ahmed Muhammad, Varshini Reddy, Chris Tanner
| Challenge: | Language models can be used to quickly and easily detect inconsistencies in documents . |
| Approach: | They propose a benchmark to measure language models' ability to detect inconsistencies in documents . they use a document with an inconsistent inserted manually by a domain expert . |
| Outcome: | The best-performing model recovered 64% of the inserted inconsistencies on 50 arXiv papers and found that the original authors had already found inconsistent inconsistances. |