Papers by Jacob Pfau

1 papers
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback (2025.findings-acl)

Copied to clipboard

Challenge: Existing static benchmarks that measure task performance often rely on a simple input-output configuration.
Approach: They propose an evaluation pipeline that evaluates code models with different feedback types in an interactive setting.
Outcome: The proposed evaluation pipeline compares model-user collaboration with static benchmarks by obfuscating inputs to a simulated user.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations