Papers by Dingyi Chang
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation (2026.eacl-short)
Copied to clipboard
| Challenge: | Existing studies on LLM factuality evaluation have not investigated the reliability of static evaluation benchmarks. |
| Approach: | They examine five popular factuality benchmarks and eight LLMs released over different years to assess their reliability. |
| Outcome: | The proposed method compared five popular factuality benchmarks and eight LLMs released over different years. |