Papers by Pengfei Zhu
Towards a Unified Multi-Dimensional Evaluator for Text Generation (2022.emnlp-main)
Copied to clipboard
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, Jiawei Han
| Challenge: | Existing evaluation frameworks for natural language generation are dominated by similarity-based metrics. |
| Approach: | They propose a multi-dimensional evaluator for natural language generation that integrates multiple dimensions into one evaluer. |
| Outcome: | The proposed evaluator improves on three typical NLG tasks and improves with external knowledge. |
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do (2026.acl-long)
Copied to clipboard
| Challenge: | Existing open-source models often yield only marginal overall improvements, possibly due to an overemphasis on mathematical reasoning at the expense of broader capabilities. |
| Approach: | They evaluate 12 multimodal tasks using 14 non-reasoning models and 8 reasoning models. |
| Outcome: | The proposed method is effective in multimodal reasoning tasks, the authors show . they show that it lacks the ability to maintain deep visual introspection throughout the reasoning process. |
Combining Curriculum Learning and Knowledge Distillation for Dialogue Generation (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have shown that curriculum learning facilitates dialogue generation tasks while knowledge distillation can yield significant performance boosts for student models. |
| Approach: | They propose a combination of curriculum learning and knowledge distillation for dialogue generation models . they cluster training cases according to their complexity and employ an adversarial training strategy . |
| Outcome: | The proposed model improves compared with baselines. |
FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation (D18-1)
Copied to clipboard
| Challenge: | Empirical results show that even the most competitive few-shot learning models struggle on this task, especially as compared with humans. |
| Approach: | They propose a Few-Shot Relation Classification Dataset consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. |
| Outcome: | The proposed methods perform well on the most competitive few-shot learning models, especially as compared with humans. |
Modeling Multi-turn Conversation with Deep Utterance Aggregation (C18-1)
Copied to clipboard
| Challenge: | Existing work on retrieval-based context modeling for multi-turn conversation ignores interactions among previous utterances. |
| Approach: | They propose retrieval-based response matching for multi-turn conversation . they propose to combine previous utterances into context using a deep utterrance aggregation model . |
| Outcome: | The proposed model outperforms state-of-the-art methods on three multi-turn conversation benchmarks including an e-commerce dialogue corpus. |
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models (2026.findings-acl)
Copied to clipboard
Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
| Challenge: | Existing medical benchmarks for diagnostic reasoning are limited in their ability to perform complex tasks. |
| Approach: | They propose to benchmark diagnostic capabilities of large language models to assess their accuracy and generalization bottlenecks. |
| Outcome: | The proposed model achieves 45.82%, 31.09%, and 17.79% accuracy, compared to current models, o3-mini, e1 and DeepSeek-R1 . |
M3TQA: Massively Multilingual Multitask Table Question Answering (2026.findings-acl)
Copied to clipboard
Daixin Shu, Jian Yang, Zhenhe Wu, Xianjie Wu, Xianfu Cheng, Guan Xiangyuan, Yanghai Wang, Pengfei Wu, Tingyang Yang, Hualei Zhu, Wei Zhang, Ge Zhang, Jiaheng Liu, Zhoujun Li
| Challenge: | Existing multilingual table benchmarks suffer from geolinguistic imbalance - overrepresenting certain languages and lacking sufficient scale for rigorous cross-lingual analysis. |
| Approach: | They propose a framework for massively multilingual table question answering that includes tables expanded to 97 languages from Chinese and English sources. |
| Outcome: | Experiments on state-of-the-art LLMs show that synthetically generated training data significantly boosts performance, especially for low-resource languages. |
Candidate Soups: Fusing Candidate Results Improves Translation Quality for Non-Autoregressive Translation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to improve NAT model's performance but do not fully utilize it. |
| Approach: | They propose a non-autoregressive translation method which can obtain high-quality translations while maintaining the inference speed of NAT models. |
| Outcome: | The proposed method outperforms the autoregressive translation model on three translation tasks with 7.6 speedup. |
Lingke: a Fine-grained Multi-turn Chatbot for Customer Service (C18-2)
Copied to clipboard
| Challenge: | e-commerce chatbots usually need a mass of human dialogue data to train, but for multi-turn conversations, the performance is poor. |
| Approach: | They propose an information retrieval augmented multi-turn chatbot which can answer questions based on unstructured documents and deal with multi-turned conversations. |
| Outcome: | The proposed solution outperforms all other models in multi-turn conversations and can learn from conversation records. |