Papers by Michael Krumdick

5 papers
BizBench: A Quantitative Reasoning Benchmark for Business and Finance (2024.acl-long)

Copied to clipboard

Challenge: Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge.
Approach: They propose a benchmark for evaluating models’ ability to reason about realistic financial problems by focusing on question-answering over financial data via program synthesis.
Outcome: The proposed benchmark evaluates models' financial background knowledge, ability to parse financial documents, and capacity to solve complex problems with code.
DocFinQA: A Long-Context Financial Reasoning Dataset (2024.acl-short)

Copied to clipboard

Challenge: Existing work on automating financial numerical reasoning focuses on unrealistically specific document snippets, failing to reflect the broader and more realistic scenarios faced by analysts.
Approach: They propose a long-document financial QA task that augments 7,437 questions from existing FinQA dataset with full-document context, extending the average context length from under 700 words in FinQA to 123k words in DocFinQA.
Outcome: The proposed task extends the average context length from under 700 words in FinQA to 123k words in DocFinQA.
Language Model Probabilities are Not Calibrated in Numeric Contexts (2025.acl-long)

Copied to clipboard

Challenge: Using language model outputs, we find that even in simple settings, the best LMs (1) are poorly calibrated and (2) have systematic biases.
Approach: They argue that language model outputs should capture natural distributions over multiple options within their textual contexts.
Outcome: The proposed model outputs are calibrated to the numeric content of their contexts.
An Analysis of Multilingual FActScore (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in LLMs have demonstrated significant capabilities in many applications.
Approach: They propose a dataset for FActScore on texts generated by strong multilingual LLMs and evaluate their performance in other languages.
Outcome: The proposed dataset shows that LLMs exhibit distinct behaviors in fact extraction and fact scoring tasks.
On Finding Inconsistencies in Documents (2026.findings-acl)

Copied to clipboard

Challenge: Language models can be used to quickly and easily detect inconsistencies in documents .
Approach: They propose a benchmark to measure language models' ability to detect inconsistencies in documents . they use a document with an inconsistent inserted manually by a domain expert .
Outcome: The best-performing model recovered 64% of the inserted inconsistencies on 50 arXiv papers and found that the original authors had already found inconsistent inconsistances.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations