Papers by Yuecong Min
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks only evaluate models in clean settings due to hallucinations . |
| Approach: | They propose a diagnostic benchmark that evaluates models in four modes for faithfulness and factuality. |
| Outcome: | The proposed benchmark evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items. |