Papers by Lei Lyu
ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge (2025.emnlp-main)
Copied to clipboard
Chaoyue He, Xin Zhou, Yi Wu, Xinjia Yu, Yan Zhang, Lei Zhang, Di Wang, Shengfei Lyu, Hong Xu, Wang Xiaoqiao, Wei Liu, Chunyan Miao
| Challenge: | ESGenius is a comprehensive benchmark for evaluating Large Language Models on ESG and sustainability knowledge. |
| Approach: | They introduce ESGenius, a benchmark for evaluating and enhancing ESG proficiency . they use a rigorous two-stage evaluation protocol and a repository of foundational frameworks . |
| Outcome: | ESGenius is a benchmark for evaluating and enhancing the proficiency of Large Language Models (LLMs) in ESG and sustainability-focused question answering. |
RESIN-11: Schema-guided Event Prediction for 11 Newsworthy Scenarios (2022.naacl-demo)
Copied to clipboard
Xinya Du, Zixuan Zhang, Sha Li, Pengfei Yu, Hongwei Wang, Tuan Lai, Xudong Lin, Ziqi Wang, Iris Liu, Ben Zhou, Haoyang Wen, Manling Li, Darryl Hannan, Jie Lei, Hyounghun Kim, Rotem Dror, Haoyu Wang, Michael Regan, Qi Zeng, Qing Lyu, Charles Yu, Carl Edwards, Xiaomeng Jin, Yizhu Jiao, Ghazaleh Kazeminejad, Zhenhailong Wang, Chris Callison-Burch, Mohit Bansal, Carl Vondrick, Jiawei Han, Dan Roth, Shih-Fu Chang, Martha Palmer, Heng Ji
| Challenge: | Existing methods for event prediction are incomplete and noisy. |
| Approach: | They propose to use news-related event schemas to extract newsworthy events . they build a demo website and include a video demonstrating the framework . |
| Outcome: | The proposed framework can be applied to a wide variety of newsworthy scenarios. |
ClinAlign: Scaling Healthcare Alignment from Clinician Preference (2026.findings-acl)
Copied to clipboard
Shiwei Lyu, Xidong Wang, Hao Zhu, Lei Liu, Chaohe Zhang, Jian Wang, Jinjie Gu, Benyou Wang, Yue Shen
| Challenge: | Existing methods for aligning open-ended outputs with fine-grained clinician preferences are weakly grounded in professional guidelines. |
| Approach: | They propose a framework to align large language models' outputs with fine-grained clinician preferences . they propose 119 broadly reusable, clinically grounded principles organized by clinical dimensions . |
| Outcome: | The proposed framework outperforms existing models on HealthBench-Hard and Deepseek-R1 and o3. |
Uni-Dubbing: Zero-Shot Speech Synthesis from Visual Articulation (2024.acl-long)
Copied to clipboard
Songju Lei, Xize Cheng, Mengjiao Lyu, Jianqiao Hu, Jintao Tan, Runlin Liu, Lingyu Xiong, Tao Jin, Xiandong Li, Zhou Zhao
| Challenge: | Multimodal speech synthesis is a key challenge due to the scarcity of datasets that pair audio with corresponding video. |
| Approach: | They propose a method that incorporates modality alignment during the pre-training phase on multimodal datasets and freezes the video modality extraction component and the encoder module within the pretrained weights. |
| Outcome: | The proposed method achieves a reduced word error rate (WER) of 31.73%, surpassing the previous best of 33.9% with single-modality audio. |
Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards (2026.acl-long)
Copied to clipboard
| Challenge: | Existing agentic training data are narrow in task variety and easily solved . real-world APIs lack diversity and are unstable for large-scale reinforcement learning rollout processes. |
| Approach: | They propose a framework that synthesizes diverse tool-use training data and simulates complete environments. |
| Outcome: | The proposed framework synthesizes diverse tool-use training data and simulates complete environments. |
SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory (2026.acl-long)
Copied to clipboard
| Challenge: | Existing defenses for Large Language Models suffer from a 'memory gap' parameter-modifying methods are computationally expensive and inference-time filters cannot retain or reuse defense knowledge across interactions. |
| Approach: | They propose a framework that secures Large Language Models through a dual-component safety memory system. |
| Outcome: | The proposed framework significantly reduces attack success rates while preserving interpretability and efficiency. |
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding (2024.acl-long)
Copied to clipboard
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li
| Challenge: | Large language models (LLMs) can only handle texts a few thousand tokens long, limiting their applications on longer sequence inputs, such as books, reports, and codebases. |
| Approach: | They propose a bilingual, multi-task benchmark for long context understanding that extends context windows and more sophisticated memory mechanisms to improve models' long context capabilities. |
| Outcome: | The proposed model outperforms open-source models but struggles on longer contexts. |
At Your Own PACE: A Causal Framework for Evaluating EQ in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Emotional Quotient (EQ) has emerged as a competency for seamless human-AI integration. |
| Approach: | They propose a framework for a closed-loop EQ evaluation using a PACE taxonomy to define four dimensions of LLM EQ. |
| Outcome: | The proposed framework achieves high alignment of 89.31% with human preferences while maintaining robust consistency of 83.6%. |
KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents (2025.findings-naacl)
Copied to clipboard
Yuqi Zhu, Shuofei Qiao, Yixin Ou, Shumin Deng, Shiwei Lyu, Yue Shen, Lei Liang, Jinjie Gu, Huajun Chen, Ningyu Zhang
| Challenge: | Large Language Models (LLMs) fail to effectively guide the planning trajectories during task solving and result in planning hallucinations. |
| Approach: | They propose a novel approach to enhance the planning capabilities of large language models by incorporating explicit action knowledge. |
| Outcome: | The proposed approach can achieve comparable or superior performance to existing baselines on HotpotQA and ALFWorld. |