Papers by Dingyi Chang

1 papers
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation (2026.eacl-short)

Copied to clipboard

Challenge: Existing studies on LLM factuality evaluation have not investigated the reliability of static evaluation benchmarks.
Approach: They examine five popular factuality benchmarks and eight LLMs released over different years to assess their reliability.
Outcome: The proposed method compared five popular factuality benchmarks and eight LLMs released over different years.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations