Papers by Zhuohan Li
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning (2026.acl-long)
Copied to clipboard
Zhuohan Xie, Daniil Orel, Rushil Thareja, Dhruv Sahnan, Hachem Madmoun, Fan Zhang, Debopriyo Banerjee, Georgi Nenkov Georgiev, Xueqing Peng, Lingfei Qian, Jimin Huang, Jinyan Su, Aaryamonvikram Singh, Rui Xing, Rania Elbadry, Chen Xu, Haonan Li, Fajri Koto, Ivan Koychev, Tanmoy Chakraborty, Yuxia Wang, Salem Lahlou, Veselin Stoyanov, Sophia Ananiadou, Preslav Nakov
| Challenge: | Existing benchmarks emphasize final numerical answers while neglecting intermediate reasoning steps. |
| Approach: | They propose a symbolic benchmark for verifiable Chain-of-Thought evaluation in finance . FINCHAIN spans 58 topics across 12 financial domains and three difficulty levels . |
| Outcome: | The proposed benchmark aims to bridge symbolic reasoning and factual verification. |
DeltaScore: Fine-Grained Story Evaluation with Perturbations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation metrics for stories are limited in assessing intricate aspects of storytelling, such as fluency and interestingness. |
| Approach: | They propose a novel method that uses perturbation techniques to evaluate story aspects . they compare fluency, coherence, relatedness, logicality, interestingness and interestingness to existing metrics . |
| Outcome: | The proposed method shows that one specific perturbation is highly effective in capturing multiple aspects. |
Hint-Based Training for Non-Autoregressive Machine Translation (D19-1)
Copied to clipboard
| Challenge: | AutoRegressive Translation models have to generate tokens sequentially during decoding and thus suffer from high inference latency. |
| Approach: | They propose to use hidden states and word alignments to help train NART models. |
| Outcome: | The proposed model improves on the WMT14 En-De and De-En datasets but is faster in inference than the current models. |
VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration (2025.findings-acl)
Copied to clipboard
Jiahui Geng, Qing Li, Zongxiong Chen, Yuxia Wang, Derui Zhu, Zhuohan Xie, Chenyang Lyu, Xiuying Chen, Preslav Nakov, Fakhri Karray
| Challenge: | Existing safety calibration methods focus on model undersafety, where the model responds to hazardous queries, while neglecting oversafetiness, where models refuse to answer safe queries. |
| Approach: | They propose safety calibration which addresses both undersafety and oversafetiness by comparing model responses to a novel dataset of 3,600 image-text pairs. |
| Outcome: | The proposed methods have been used to evaluate safety calibration across image-centric and text-centric scenarios. |