Papers by Shuyue Zhu
Rhetorical Device-Aware Sarcasm Detection with Counterfactual Data Augmentation (2025.findings-acl)
Copied to clipboard
| Challenge: | Sarcasm is a complex form of sentiment expression widely used in human daily life. |
| Approach: | They propose a device-aware sarcasm dataset with counterfactually augmented data to capture its complexity. |
| Outcome: | The proposed dataset shows that it is more balanced than zero-shot models. |
CHROMIC: Chronological Reasoning Across Multi-Panel Comics (2026.eacl-long)
Copied to clipboard
Bingxuan Hou, Jiayi Lin, Chenyang Zhang, Dapeng Yin, Shuyue Zhu, Qingqing Hong, Mengna Gao, Junli Wang
| Challenge: | Large-scale vision–language models have achieved remarkable progress on various reasoning tasks, but most studies focus on natural photographic images and pay limited attention to multi-panel visual narratives such as comics. |
| Approach: | They propose a benchmark dataset for chronological reasoning in multi-panel comics that covers six types of reasoning questions and spans both Western and Japanese comic styles. |
| Outcome: | The proposed dataset covers six types of reasoning questions and spans both Western and Japanese comic styles. |
StruNRAG: Evaluation of OCR-Induced Structural Noise on RAG Robustness (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of RAG systems ignore structural noise, authors say . complex layouts can cause OCR failures and disrupt semantic flow of text . advanced LLMs demonstrate robustness against local noise, but struggle to maintain reasoning capabilities under severe structural disruption that fragments global context. |
| Approach: | They propose a benchmark to evaluate RAG robustness against OCR-induced structural perturbations. |
| Outcome: | The proposed benchmark systematically injects three categories of real-world structural noise into a bilingual dataset of 2,132 question-answer pairs . results show that advanced LLMs demonstrate robustness against local noise, but struggle to maintain reasoning capabilities under severe structural disruption . |
MdEval: Massively Multilingual Code Debugging (2026.findings-acl)
Copied to clipboard
Shukai Liu, Linzheng Chai, Jian Yang, Jiajun Shi, He Zhu, Liran Wang, Jin Ke, Wei Zhang, Hualei Zhu, Shuyue Guo, Tao Sun, Jiaheng Liu, Yunlong Duan, Yu Hao, Liqun Yang, Guanglin Niu, Ge Zhang, Zhoujun Li
| Challenge: | Existing benchmarks primarily focus on Python and are limited in terms of language diversity. |
| Approach: | They propose a multilingual debugging benchmark that includes 3.9K test samples of 20 programming languages and introduces the debug instruction corpora MdEval-Instruct by injecting bugs into the correct multilingual queries and solutions. |
| Outcome: | The proposed benchmark includes 3.9K test samples of 20 programming languages and covers the automated program repair task, bug localization task, and bug identification task. |
Reinforcement Learning for Large Language Models via Group Preference Reward Shaping (2025.emnlp-main)
Copied to clipboard
Huaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao, Zhimeng Guo, Shijie Zhou, Shuyue Hu, Vasant G. Honavar
| Challenge: | Existing methods for fine-tuning Large Language Models (LLMs) are expensive and sensitive to reward model quality. |
| Approach: | They propose a method that leverages preference-based comparisons rather than precise numerical rewards. |
| Outcome: | Experiments show that GPRS outperforms critic-model-free RL algorithms on RLHF and reasoning tasks. |
LIME: Less Is More for MLLM Evaluation (2025.findings-acl)
Copied to clipboard
King Zhu, Qianbo Zang, Shian Jia, Siwei Wu, Feiteng Fang, Yizhi Li, Shuyue Guo, Tianyu Zheng, Jiawei Guo, Bo Li, Haoning Wu, Xingwei Qu, Jian Yang, Ruibo Liu, Xiang Yue, Jiaheng Liu, Chenghua Lin, Hamid Alinejad-Rokny, Min Yang, Shiwen Ni, Wenhao Huang, Ge Zhang
| Challenge: | Existing MLLM benchmarks and unified evaluation frameworks cannot accurately and efficiently reflect the ability of MLMLs. |
| Approach: | They propose a semi-automated benchmark curated using a pipeline that filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding. |
| Outcome: | The proposed benchmark reduces the number of samples by 76% and evaluation time by 77% while it can more effectively distinguish different models’ abilities. |