Papers by Juexi Shao
NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for evaluating novelty have been proposed, but there is no systematic evaluation of their ability to generate novelty evaluations. |
| Approach: | They propose a benchmark to evaluate large language models’ ability to generate novelty evaluations in support of human peer review. |
| Outcome: | The proposed framework evaluates the quality of LLM-generated novelty evaluations under different prompting strategies. |