Papers by Keqin Li
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning (2025.findings-acl)
Copied to clipboard
Xiaoyuan Li, Moxin Li, Rui Men, Yichang Zhang, Keqin Bao, Wenjie Wang, Fuli Feng, Dayiheng Liu, Junyang Lin
| Challenge: | Existing studies show that large language models are robust in commonsense reasoning . however, some variations in questions can lead to incorrect responses . |
| Approach: | They propose a large-scale bilingual benchmark consisting of 11,200 cases . they conduct extensive experiments on 41 representative LLMs . |
| Outcome: | The proposed benchmark systematically evaluates the robustness of large language models in commonsense reasoning. |
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation (2026.acl-long)
Copied to clipboard
Xiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu
| Challenge: | Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. |
| Approach: | They propose to use a multi-turn reasoning evaluation framework to cover multi-turn interactions with the environments of large language models. |
| Outcome: | The proposed framework covers diverse reasoning capabilities, fine-grained difficulty granularity, and necessitates multi-turn interactions with the environments. |
Dynamic Graph Neural ODE Network for Multi-modal Emotion Recognition in Conversation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing graph-based multimodal emotion recognition methods fail to capture dynamic changes in emotions. |
| Approach: | They propose a Dynamic Graph Neural Ordinary Differential Equation Network (DGODE) which combines dynamic changes of emotions to capture temporal dependencies of speakers’ emotions. |
| Outcome: | The proposed model can capture the temporal dependencies caused by dynamic changes in emotions and can improve on two publicly available multimodal emotion recognition datasets. |
Towards Making the Most of ChatGPT for Machine Translation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Prior studies have shown that ChatGPT achieves comparable results to commercial systems for high-resource languages, but lags behind in complex tasks, e.g., low-resourced and distant-language-pairs translation. |
| Approach: | They propose task-specific prompts and domain-specific prompts which are based on task information and domain information and a task-specific prompt. |
| Outcome: | The proposed prompts improve the performance of ChatGPT in complex tasks and generate hallucinations for non-English-centric tasks. |