Papers by Huizhi Liang
KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models are limited by static and narrow questions, leading to limited coverage and misleading evaluations. |
| Approach: | They propose a Knowledge Graph-based hallucination benchmark that assesses Large Language Models across the breadth and depth of their knowledge and provides a fairer and more comprehensive insight into LLM truthfulness. |
| Outcome: | The proposed framework assesses LLMs across breadth and depth of their knowledge, and provides a fairer and more comprehensive insight into LLM truthfulness. |
Tracing the Light of Thought: A Probabilistic Self- and Cross-Consistency Verification Mechanism Improving Mathematical Reasoning in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for evaluating reasoning paths are not efficient, but they are prone to errors. |
| Approach: | They propose a probabilistic self- and cross-consistency framework for mathematical reasoning that employs an accept-reject mechanism to encourage high-quality reasoning paths. |
| Outcome: | The proposed framework improves on 9 LLMs across 4 challenging benchmarks. |