Papers by Xiuyi Chen
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Currently, vision-Language Models are optimized for direct visual question-answering tasks. |
| Approach: | They propose a visual-language-based VLM that prioritizes reasoning within the perception process. |
| Outcome: | The proposed model outperforms existing models and domain-specific open-source models in the chemical domain. |
DualGATs: Dual Graph Attention Networks for Emotion Recognition in Conversations (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies focus on speaker-aware context modeling, overlooking the discourse structure of the conversation. |
| Approach: | They propose Dual Graph ATtention networks to capture contextual dependencies in conversational contexts and integrate it into a speaker-aware GAT module. |
| Outcome: | The proposed model outperforms state-of-the-art models on four datasets and is highly efficient. |
Learning to Use Tools via Cooperative and Interactive Agents (2024.findings-emnlp)
Copied to clipboard
Zhengliang Shi, Shen Gao, Xiuyi Chen, Yue Feng, Lingyong Yan, Haibo Shi, Dawei Yin, Pengjie Ren, Suzan Verberne, Zhaochun Ren
| Challenge: | Existing methods for large language models (LLMs) use one agent to iterate and execute tools, but they suffer from performance degradation when addressing practical tasks. |
| Approach: | They propose a tool learning framework that coordinates three specialized agents for tool selection, tool execution, and action calibration separately. |
| Outcome: | The proposed framework outperforms baseline models on three datasets with 14% higher success rate. |
Unsupervised Knowledge Selection for Dialogue Generation (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing knowledge selection tasks require the preidentified knowledge to generate informative dialogues. |
| Approach: | They propose a novel method to supervise knowledge selection when the gold knowledge label is unknown by obtaining an oracle knowledge label via distant supervision and leverage knowledge distillation to alleviate the noisy labeling problem of distant supervision. |
| Outcome: | The proposed method outperforms strong supervised baselines on two knowledge-grounded dialogue datasets and generates more informative responses. |
Progressive LoRA for Multimodal Continual Instruction Tuning (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to MCIT address Catastrophic Forgetting and Knowledge Transfer (KT) but using a fixed number of shared LoRA blocks across tasks can lead to knowledge interference. |
| Approach: | They propose a framework that uses a fixed number of shared LoRA blocks to reduce knowledge interference. |
| Outcome: | The proposed framework outperforms existing approaches on the latest MCIT benchmark. |
Learning to Ground Visual Objects for Visual Dialog (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to ground visual objects are inadequate for visual dialog . a posterior distribution is inferred from context and questions, while posterior distributions are used to facilitate visual objects grounding. |
| Approach: | They propose a method to learn to ground visual objects for visual dialog using prior and posterior distributions over visual objects to facilitate visual objects grounding. |
| Outcome: | The proposed approach improves the existing models in generative and discriminative settings by a significant margin. |
A Working Memory Model for Task-oriented Dialog Response Generation (P19-1)
Copied to clipboard
| Challenge: | Existing models to integrate external Knowledge Base information, one form of world knowledge, confound dialog history with KB tuples and store them into one memory. |
| Approach: | They propose a working memory model that interacts with two long-term memories to generate dialog responses. |
| Outcome: | The proposed model outperforms the state-of-the-art models on two task-oriented dialog datasets. |
Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies focus on implicit exploration of multimodal coreference but neglect the importance of locating the objects explicitly in the visual content, which is associated with textual entities. |
| Approach: | They propose a multimodal incremental transformer with visual grounding which aims to explicitly locate related objects in the image guided by textual entities. |
| Outcome: | The proposed model achieves comparable performance on the VisDial v0.9 and v1.0 datasets. |
GoG: Relation-aware Graph-over-Graph Network for Visual Dialog (2021.findings-acl)
Copied to clipboard
| Challenge: | Experimental results show that our model outperforms the strong baseline in both generative and discriminative settings by a significant margin. |
| Approach: | They propose a relation-aware graph-over-graph network (GoG) for visual dialog . their model outperforms the strong baseline in both generative and discriminative settings . |
| Outcome: | The proposed model outperforms baseline models in both generative and discriminative settings by a significant margin. |
Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue Generation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogue models lack prior and posterior knowledge selection . prior selection module may not learn to select knowledge properly because of lack of posterior information . |
| Approach: | They propose a knowledge distillation-based training strategy to remove the exposure bias of knowledge selection. |
| Outcome: | The proposed model improves on two knowledge-grounded dialogue datasets. |
Knowledge Aware Emotion Recognition in Textual Conversations via Multi-Task Incremental Transformer (2020.coling-main)
Copied to clipboard
| Challenge: | Existing models for ERTC use a few non-neutral categories to identify the emotion of each utterance. |
| Approach: | They propose a novel Knowledge Aware Incremental Transformer with Multi-task Learning to address these challenges by leveraging commonsense knowledge to leverage context. |
| Outcome: | The proposed model outperforms state-of-the-art models across five benchmark datasets. |
Flexible Weight Tuning and Weight Fusion Strategies for Continual Named Entity Recognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for Named Entity Recognition (CNER) use knowledge distillation to retain old knowledge, but they are too expensive and fail to integrate with existing state-of-the-art models. |
| Approach: | They propose a weight tuning and weightfusion strategy to learn new entity types while mitigating catastrophic forgetting of old models. |
| Outcome: | The proposed strategies improve the performance of existing models and are model-agnostic. |
Continual Named Entity Recognition without Catastrophic Forgetting (2023.emnlp-main)
Copied to clipboard
| Challenge: | Named Entity Recognition (CNER) is a burgeoning area of research . a new paradigm has ushered NER into a non-entity type at the current step t . |
| Approach: | They propose a pooled feature distillation loss that skillfully navigates the trade-off between retaining knowledge of old entity types and acquiring new ones. |
| Outcome: | The proposed method outperforms state-of-the-art approaches on ten CNER settings using three datasets. |
TSAM: A Two-Stream Attention Model for Causal Emotion Entailment (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies on EAC focus on Emotion Recognition in Conversations (ERC), i.e., recognizing emotion labels of utterances. |
| Approach: | They propose a two-stream attention model to capture correlations between utterances in a global view and classify multiple utterrances synchronously to capture emotion and speaker information in parallel. |
| Outcome: | The proposed model outperforms baselines and achieves new State-Of-The-Art (SOTA) performance. |