Papers by Jeffrey Cheng

2 papers
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of test-time scaling assume that a reasoning system should always give an answer to any question provided.
Approach: They propose to increase compute budget at inference time to increase confidence in correct responses by considering settings with non-zero levels of response risk.
Outcome: The proposed model can answer more questions correctly and have higher confidence in correct responses.
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers? (2025.findings-emnlp)

Copied to clipboard

Challenge: CLAIMCHECK is an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews from OpenReview.
Approach: They annotate NeurIPS 2023 and 2024 submissions and reviews for weaknesses and dispute them for fine-grained labels of validity, objectivity, and type of the identified weaknesses.
Outcome: The proposed dataset is richly annotated by ML experts for weaknesses statements in the reviews and the claims that they dispute, as well as fine-grained labels of validity, objectivity, and type of the identified weaknesses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations