Papers by Ruihan Yang
SELFGOAL: Your Language Agents Already Know How to Achieve High-level Goals (2025.naacl-long)
Copied to clipboard
Ruihan Yang, Jiangjie Chen, Yikai Zhang, Siyu Yuan, Aili Chen, Kyle Richardson, Yanghua Xiao, Deqing Yang
| Challenge: | Existing approaches to improve the performance of language agents without training are not available. |
| Approach: | They propose an automatic approach to break down high-level goals into tree structure of more practical subgoals during interaction with environments while identifying the most useful subgoal. |
| Outcome: | The proposed approach significantly improves the performance of language agents across various tasks, including competitive, cooperative, and deferred feedback environments. |
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing work lacks direct and fair evaluation of Large Language Models’ ability to express uncertainty effectively in long-form generation. |
| Approach: | They propose a benchmark to evaluate uncertainty expression in both long- and short-form question answering (QA) they propose prompt-based and training-based methods to improve models’ performance. |
| Outcome: | The proposed method mitigates this issue but a misalignment persists in uncertainty expression between long- and short-form generation. |
LoGU: Long-form Generation with Uncertainty Expressions (2025.acl-long)
Copied to clipboard
Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Sen Yang, Nigel Collier, Dong Yu, Deqing Yang
| Challenge: | Large Language Models (LLMs) generate factually incorrect content, i.e., hallucinations, despite impressive performance. |
| Approach: | They propose a framework to enable models to express uncertainty when unsure . they propose atomic claims to refine uncertainty and refine it using supervised fine-tuning and direct preference optimization to enhance uncertainty expression. |
| Outcome: | The proposed framework significantly improves accuracy, reduces hallucinations, and maintains comprehensiveness of responses. |
Confidence Estimation for LLMs in Multi-turn Interactions (2026.findings-acl)
Copied to clipboard
Caiqi Zhang, Ruihan Yang, Xiaochen Zhu, Chengzu Li, Tiancheng Hu, Yijiang River Dong, Deqing Yang, Nigel Collier
| Challenge: | Despite recent progress, most prior work studies confidence in single-turn question answering. |
| Approach: | They propose a logit-based probe that measures confidence in multi-turn dialogues . they propose 'infoECE' and a "hinter-guesser" paradigm for generating controlled evaluations based on data . |
| Outcome: | The proposed framework is grounded in calibration and monotonicity of confidence as more information becomes available. |
GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick (2024.acl-long)
Copied to clipboard
| Challenge: | Large language models generate human-like content, but they also pose a problem with generation diversity, negatively impacting generation diversity and user experience. |
| Approach: | They propose a Logits-Addition watermark and three variants that aim to enhance diversity to overcome generation diversity challenges. |
| Outcome: | The Logits-Addition watermark outperforms the Logits+Trick-based watermark in diversity tests and outperformed other decoding-based methods by 0.1 to 0.3. |
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents (2026.acl-long)
Copied to clipboard
Ruihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong, Xiaoliang Yang, Yuxuan Gu, Zekun Zhou, Yunfei Lu, Haoyu Ren, Kun Chen, Dandan Tu, Bing Qin
| Challenge: | Existing GUI benchmarks lack fine-grained diagnostics to identify which capabilities lead to task failures. |
| Approach: | They propose a multilingual P R GUI Benchmark to assess LVLMs' language capabilities . they propose XLI to align non-English hidden states with English ones during inference . |
| Outcome: | The proposed benchmark reveals consistent gaps between English and non-English settings . it reduces the cross-lingual gaps with an average gain of 6.5% in non- English settings compared to static benchmarks . |