Papers by Haoqi Hu
LRBench and Judge-R1: Principled Evaluation and Training of LLM-Based Judges for Long-Context Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating large language models (LLMs) under long contexts are underexplored. |
| Approach: | They propose a large-scale benchmark for evaluating large language models (LLMs) that combines reinforcement learning with multi-turn search to enable grounded and principle-aware evaluation. |
| Outcome: | The proposed model outperforms single-turn baselines across domains and principles. |