Papers by Yaxin Zhu
DraDDP: A Multimodal Multi-Party Dialogue Discourse Parsing Dataset (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on multi-party dialogue discourse parsing focus on textual modality and two-party dialog . et al., 2016) focused on text-based discourse parses, ignoring the complexity and richness of multimodal interactions in real-world scenarios. |
| Approach: | They construct the first publicly available English multimodal dataset for multi-party dialogue discourse parsing based on American TV dramas. |
| Outcome: | The proposed dataset contains 495 dialogue segments with 6,374 utterances and 9.1 hours of parallel video content, covering rich multi-party interaction scenarios. |
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models (2024.emnlp-main)
Copied to clipboard
Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, Tetsuya Sakai, Tian Feng, Hayato Yamana
| Challenge: | Currently, tool-augmented large language models (LLMs) only achieve total scores of 45.3 and 37.0, respectively, on a scale of 100. |
| Approach: | They propose a multi-level diagnostic process to assess the LLM's hallucinations through two perspectives: depth and breadth. |
| Outcome: | The proposed diagnostic process assesses the hallucinations of large language models through two perspectives: depth and breadth. |
QAP: A Quantum-Inspired Adaptive-Priority-Learning Model for Multimodal Emotion Recognition (2023.findings-acl)
Copied to clipboard
| Challenge: | Experimental results show that multimodal emotion recognition is a state-of-the-art technique . textual, visual and acoustic modalities are involved in multimodal video emotion recognition . |
| Approach: | They propose a quantum-inspired adaptive-priority-learning model to address the challenges . they use quantum state to model modal features and Q-attention to integrate three modalities . |
| Outcome: | Experimental results show that QAP improves on previous models. |
Uncertainty-Aware Cross-Modal Alignment for Hate Speech Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for detecting hate speech ignore misalignment and uncertainty between modalities . social media platforms have become conduits for the rapid dissemination of hate speech . |
| Approach: | They propose an uncertainty-aware cross-modal alignment framework for hate speech detection that minimizes the misalignment of image and text in memes. |
| Outcome: | The proposed framework produces a competitive performance compared with existing methods. |
Improving Dialogue Discourse Parsing through Discourse-aware Utterance Clarification (2025.acl-long)
Copied to clipboard
| Challenge: | Extensive experiments on the STAC and Molweni datasets demonstrate that our approach effectively resolves ambiguities and significantly outperforms the state-of-the-art (SOTA) baselines. |
| Approach: | They propose a Discourse-aware Clarification Module (DCM) that generates clarifications for the parser through systematic clarification type reasoning and discourse goal reasoning. |
| Outcome: | Extensive experiments on the STAC and Molweni datasets demonstrate that the proposed module significantly outperforms the state-of-the-art (SOTA) framework. |
Enhancing Goal-oriented Proactive Dialogue Systems via Consistency Reflection and Correction (2025.acl-long)
Copied to clipboard
| Challenge: | Unlike traditional dialogue systems, goal-oriented proactive dialogue systems focus on achieving specific objectives by actively guiding and anticipating user needs. |
| Approach: | They propose a model-agnostic two-stage Consistency Reflection and Correction framework that allows the model to reflect on discrepancies between generated responses and dialogue contexts and suggest possible corrections. |
| Outcome: | The proposed framework significantly improves the consistency between generated responses and dialogue contexts on three datasets. |
Faithfully Explainable Recommendation via Neural Logic Reasoning (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing models for explainable recommendation have neglected faithfulness of KG reasoning . |
| Approach: | They propose to draw on interpretable logical rules to guide path-reasoning process for explanation generation. |
| Outcome: | The proposed method delivers high-quality recommendations and ascertains the faithfulness of the derived explanation. |
A Distance-Aware Multi-Task Framework for Conversational Discourse Parsing (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies have focused on graph-based and transition-based discourse parsing, but no study has investigated the advantages of both paradigms for conversational discourse paring. |
| Approach: | They propose a distance-aware multi-task framework that incorporates the strengths of transition-based paradigms to facilitate conversational discourse parsing. |
| Outcome: | The proposed framework improves the graph-based paradigm on long-distance dependency links. |
Incomplete Utterance Rewriting with Editing Operation Guidance and Utterance Augmentation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing generation methods on Incomplete Utterance Rewriting (IUR) can generate coherent utterances, but they often include irrelevant and redundant tokens in rewritten utteras . |
| Approach: | They propose a multi-task learning framework that uses editing operation labels to guide generation model to focus on critical tokens in dialogue context. |
| Outcome: | The proposed model outperforms state-of-the-art models on open-domain and task-oriented dialogues on three datasets. |
Improving Multi-party Dialogue Generation via Topic and Rhetorical Coherence (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on multi-party dialogue generation focus on the reply-to structure of dialogue histories, but they neglect the coherence between generated responses and target utterances. |
| Approach: | They propose a Reinforcement Learning approach emphasizing Topic and Rhetorical Coherence to enhance the model's perception of coherence with the target utterance. |
| Outcome: | The proposed approach significantly outperforms the state-of-the-art baselines on two popular datasets. |
TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making (2025.emnlp-main)
Copied to clipboard
Kechen Jiao, Zhirui Fang, Jiahao Liu, Bei Li, Qifan Wang, Xinyu Liu, Junhao Ruan, Zhongjian Qiao, Yifan Zhu, Yaxin Xu, Jingang Wang, Xiu Li
| Challenge: | Existing post-SFT methods for embodied AI are constrained by sparse rewards and action-only optimization, resulting in low sample efficiency, poor consistency, and model degradation. |
| Approach: | They propose to integrate Thought-Centric Preference Optimization (TCPO) into embodied decision-making by transforming sparse reward signals into richer step sample pairs. |
| Outcome: | The proposed approach achieves an average success rate of 26.67% in the ALFWorld environment, and a 6% improvement over RL4VLM. |
Faithful Inference Chains Extraction for Fact Verification over Multi-view Heterogeneous Graph with Causal Intervention (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for fact verification do not extract faithful inference chains due to the diversity of relation paths. |
| Approach: | They propose a multi-view heterogeneous Graph with causal intervention to extract evidence graphs from the knowledge graph. |
| Outcome: | The proposed model provides precise evidence graphs and achieves state-of-the-art performance on the public KG-based fact verification dataset FactKG. |
ICXML: An In-Context Learning Framework for Zero-Shot Extreme Multi-Label Classification (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing research has focused on fully supervised XMC, but real-world scenarios often lack supervision signals, highlighting the importance of zero-shot settings. |
| Approach: | They propose a framework that generates a set of candidate labels through in-context learning and then reranks them. |
| Outcome: | The proposed framework advances state-of-the-art on two diverse public benchmarks. |
Enhancing Goal-oriented Proactive Dialogue Systems via Dynamic Multi-dimensional Consistency Optimization (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on goal-oriented proactive dialogue systems failed to address the multi-dimensional consistency issue between generated responses and key contextual elements. |
| Approach: | They propose a Dynamic Multi-dimensional Consistency Reinforcement Learning framework which measures the impact of each consistency dimension on overall dialogue quality and provides feedback to improve response quality. |
| Outcome: | The proposed framework significantly improves the consistency of generated responses on two datasets. |
Predicting Prerequisite Relations for Unseen Concepts (2022.emnlp-main)
Copied to clipboard
| Challenge: | Concept prerequisite learning (CPL) is a task of building a concept graph by structuring open knowledge in prerequisite relations. |
| Approach: | They propose to use both content-based and graph-based models to build a concept graph by structuring open knowledge in prerequisite relations. |
| Outcome: | The proposed approach improves F1 scores by 10% on three public benchmarks. |
Simulating Dual-Process Thinking in Dialogue Topic Shift Detection (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for topic shift detection focus on shallow local reasoning, overlooking the importance of considering the global historical structure and local details to elucidate the underlying causes of topic shift. |
| Approach: | They propose a dual-process theory for dialogue topic shift detection that employs Large Language Models to extract and store the global topic structure of historical dialogue, while a reasoning module introduces a LLM to generate reasoning samples between the response and the most recent topic of historical dialog. |
| Outcome: | The proposed framework outperforms the state-of-the-art on three public datasets and is based on a dual-process theory. |
Document-level Relationship Extraction by Bidirectional Constraints of Beta Rules (2023.emnlp-main)
Copied to clipboard
| Challenge: | Document-level Relation Extraction (DocRE) aims to extract relations among entity pairs in documents. |
| Approach: | They propose a logic constraint framework that uses bidirectional constraints to model rules by beta contribtion and reconstruct rule consistency loss by bidirectional constraint. |
| Outcome: | The proposed framework outperforms existing models in relation extraction performance and logical consistency. |
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model (2025.emnlp-main)
Copied to clipboard
Zhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Yaxin Peng, Chaomin Shen, Feifei Feng, Yi Xu
| Challenge: | Recent advances in vision-language-action models prioritize robotic action mastery . however, models trained on visual-text pairs struggle to interpret multimodal data . |
| Approach: | They propose a framework that integrates multimodal data after initial control mastery and a Mixture-of-Experts architecture to minimize task interference. |
| Outcome: | The proposed framework surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks and achieves six times higher performance on visual question-answering datasets. |
Data Augmentation with Adversarial Training for Cross-Lingual NLI (2021.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to train cross-lingual models with labeled data are subpar, resulting in subpar results. |
| Approach: | They propose a data augmentation strategy that enriches data to reflect more diversity in a semantically faithful way and leverages adversarial training regimens to achieve greater robustness. |
| Outcome: | The proposed approach improves cross-lingual inference by leveraging the data to reflect more diversity in a semantically faithful way. |
Improving Dialogue Discourse Parsing via Reply-to Structures of Addressee Recognition (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to learn dialogue discourse parsing with related tasks require additional annotation, thus limiting their generality. |
| Approach: | They propose a multitasking framework that integrates dialogue discourse parsing with addressee recognition to reflect relation-based structure of dialogue. |
| Outcome: | The proposed framework outperforms baselines on the Molweni and STAC datasets. |
Uncertainty-Guided Modal Rebalance for Hateful Memes Detection (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for integrating hate information from different modalities ignore the modality uncertainty caused by the contribution degree of each modality to hate sentiment. |
| Approach: | They propose an Uncertainty-guided Modal Rebalance framework for hateful memes detection . they propose to combine cross-modal fusion features with unimodal features . |
| Outcome: | The proposed framework produces state-of-the-art performance on four widely-used datasets. |
Enhancing Multi-party Dialogue Discourse Parsing with Explanation Generation (2025.coling-main)
Copied to clipboard
| Challenge: | Multi-party dialogue discourse parsing is an important and challenging task in natural language processing. |
| Approach: | They propose a model to integrate external knowledge from Large Language Models to analyze dialogue discourse structures and semantic relations between utterances in multi-party conversations. |
| Outcome: | The proposed model outperforms the state-of-the-art (SOTA) models on two public datasets. |
Two-stage Incomplete Utterance Rewriting on Editing Operation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods to generate rewritten utterances based on dialogue context ignore coreference and ellipsis in dialogues. |
| Approach: | They propose a framework where the first stage generates editing operations and the second stage rewrites incomplete utterances utilizing the generated editing operations. |
| Outcome: | The proposed framework outperforms the existing models on three IUR datasets. |
Not Just Classification: Recognizing Implicit Discourse Relation on Joint Modeling of Classification and Generation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods of implicit discourse relation recognition (IDRR) focus on three aspects: enhancing discourse units representation, enhancing semantic interaction, and joint learning with other tasks. |
| Approach: | They propose a joint model to recognize the relation label and generate the target sentence containing the meaning of relations simultaneously. |
| Outcome: | The proposed model achieves the best performance against several state-of-the-art systems on Chinese and English datasets. |
Non-Emotion-Centric Empathetic Dialogue Generation (2025.coling-main)
Copied to clipboard
| Challenge: | Empathy is a social psychology theory that enables individuals to comprehend each other's experiences and emotions, thereby fostering more intimate interpersonal relationships. |
| Approach: | They propose a framework for empathetic dialogue generation based on contrastive learning and context-sensitive entity and social commonsense that punishes responses with incorrect emotions and improves the quality of emotions. |
| Outcome: | The proposed framework improves the quality of empathetic generation and generates more diverse responses in comparison with the state-of-the-art baselines. |