Papers by Alexander Gill

2 papers
On Evaluating Explanation Utility for Human-AI Decision Making in NLP (2024.findings-emnlp)

Copied to clipboard

Challenge: a lack of evidence that explanations help people in situations they are introduced for is a problem in NLP . prior work on explainability has focused on overcoming technical challenges and used proxy evaluations.
Approach: They propose to use existing metrics to evaluate the effectiveness of explanations in NLP . they argue that providing AI predictions does not cause decision makers to speed up work .
Outcome: The proposed evaluations show that providing AI predictions does not cause decision makers to speed up their work without compromising performance.
What Has Been Lost with Synthetic Evaluation? (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study evaluated the validity and difficulty of large language models for evaluation benchmarks . large language model evaluation benchmarking is challenging and requires specific phenomena to be addressed .
Approach: They compare LLM-generated reasoning-over-text benchmarks to those generated through crowdsourcing . they find they are *less challenging for LLMs* than their human-authored counterparts .
Outcome: The results show that LLMs can generate variants that are valid according to annotation guidelines, but less challenging than human-authored counterparts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations