Papers by Wei Gong
Beyond Chunks and Graphs: Retrieval-Augmented Generation through Triplet-Driven Thinking (2026.findings-acl)
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) is critical for reducing hallucinations and incorporating external knowledge into Large Language Models (LLMs). |
| Approach: | They propose a framework that leverages an LLM to decompose questions into searchable triplets with placeholders. |
| Outcome: | Empirical results show that T2RAG outperforms state-of-the-art multi-round and Graph RAG methods while reducing retrieval costs by up to 45%. |
VGA: Vision GUI Assistant - Minimizing Hallucinations through Image-Centric Fine-Tuning (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing Large Vision-Language Models (VLMs) often overly rely on internal text-based knowledge while neglecting visual inputs. |
| Approach: | They propose a model that balances attention image and text to enhance interpretation and reduce hallucinations by using a visual input. |
| Outcome: | The proposed model improves interpretation and reduces hallucinations by balancing attention image and text to enhance interpretation and reduction of hallucinosity. |
CPT-Agent: A Cognitive Process Theory-driven Framework for Student Simulation in Writing Development (2026.acl-long)
Copied to clipboard
| Challenge: | Existing LLMs model overly capable learners who over-apply feedback, resulting in pedagogically implausible behavior. |
| Approach: | They propose a framework that decouples cognitive ability from writing proficiency and models their interaction during writing and revision. |
| Outcome: | The proposed model produces distinguishable proficiency levels and is consistent with instructional theories. |
TableGPT: Few-shot Table-to-Text Generation with Table Structure Reconstruction and Content Matching (2020.coling-main)
Copied to clipboard
| Challenge: | Recent studies show that pre-trained language models can produce informative and fluent text with the help of large-scale datasets, but they suffer insufficient learning problem with limited training data. |
| Approach: | They propose to use table transformation module with template to rewrite structured table in natural language as input for GPT-2 and exploit multi-task learning with two auxiliary tasks to preserve table’s structural information. |
| Outcome: | The proposed model outperforms existing systems on most few-shot settings. |
IFlyEA: A Chinese Essay Assessment System with Automated Rating, Review Generation, and Recommendation (2021.acl-demo)
Copied to clipboard
| Challenge: | Automated Essay Assessment (AEA) aims to judge students’ writing proficiency in an automatic way. |
| Approach: | They propose to use Chinese AEA system IFlyEssayAssess to evaluate essays written by native Chinese students from primary and junior schools. |
| Outcome: | The proposed system provides application services for essay scoring, review generation, recommendation, and explainable analytical visualization. |
Outlier Suppression+: Accurate quantization of large language models by equivalent and effective shifting and scaling (2023.emnlp-main)
Copied to clipboard
| Challenge: | asymmetric outliers in transformer language models are a challenge for post-training quantization . we propose a framework for outlier suppression that can be seamlessly migrated into subsequent modules . |
| Approach: | They propose a framework for post-training quantization that includes the channel-wise shifting and scaling for concentration. |
| Outcome: | The proposed framework can be migrated into subsequent modules while maintaining equivalence. |
ELTLM: Evaluation of Longitudinal Temporal Large Multimodal Models in Clinical Scenarios (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation benchmarks focus on static evaluation of large multimodal models . existing evaluation paradigms neglect a critical aspect of clinical practice: longitudinal analysis . |
| Approach: | They propose a temporal perception and reasoning benchmark to assess models' temporal grounding and consistency. |
| Outcome: | ELTLM features a hierarchical task taxonomy comprising Temporal Perception QA and Temporal Reasoning QA. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |
FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing models generate explanations that appear coherent while containing unfaithful intermediate steps. |
| Approach: | They propose a causality-inspired framework for evaluating CoT quality using controlled perturbations as an instrumental signal to separate genuine step-to-step dependence from bias-driven artifacts. |
| Outcome: | Experiments on GSM8K, MATH, and CommonsenseQA show that FACT-E improves reasoning-trajectory selection and yields stronger in-context learning exemplars. |
Mitigating Posterior Salience Attenuation in Long-Context LLMs with Positional Contrastive Decoding (2025.acl-short)
Copied to clipboard
| Challenge: | Current solutions incur prohibitive training costs, leaving statistical behaviors and cost-effective approaches underexplored. |
| Approach: | They propose a positional contrast decoding technique that contrasts long-aware attention with designed local-awn attention. |
| Outcome: | The proposed model achieves state-of-the-art performance on long-context benchmarks. |
Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules (2024.emnlp-main)
Copied to clipboard
| Challenge: | Empirical results show that MoMs consistently outperform vanilla transformers . |
| Approach: | They propose an architecture that allows for a mixture-of-modules computation that uses a finite set of modules defined by multi-head attention and feed-forward networks. |
| Outcome: | The proposed architecture outperforms vanilla Transformers and their variants in multiple ways. |
Contextual Fine-to-Coarse Distillation for Coarse-grained Response Selection in Open-Domain Conversations (2022.acl-long)
Copied to clipboard
Wei Chen, Yeyun Gong, Can Xu, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan
| Challenge: | Existing studies focus on coarse-grained response selection in retrieval-based dialogue systems. |
| Approach: | They propose a Contextual Fine-to-Coarse (CFC) distilled model for coarse-grained response selection in open-domain conversations. |
| Outcome: | The proposed model improves over baseline methods on two datasets based on the Reddit comments dump and Twitter corpus compared with baseline methods. |
Enhancing the Transferability of Jailbreak Attacks on Large Language Models via Exploiting Reparameterization Invariance (2026.acl-long)
Copied to clipboard
| Challenge: | Existing token-level attacks have shown efficacy on open-source models but suffer from poor cross-model transferability. |
| Approach: | They propose a framework to improve cross-model transferability by modifying model parameters and generating update directions according to differences in output distributions rather than parameter-space distances. |
| Outcome: | The proposed framework improves cross-model transferability and success rates on open-source models. |
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models (2024.findings-acl)
Copied to clipboard
Xueliang Zhao, Xinting Huang, Tingchen Fu, Qintong Li, Shansan Gong, Lemao Liu, Wei Bi, Lingpeng Kong
| Challenge: | Multimodal reasoning is a key capability for large vision-language models . however, the vanilla Chain-of-Thought method fails to address critical steps in multi-step reasoning tasks. |
| Approach: | They propose a bi-modal Behavioral Alignment method to augment multimodal reasoning . they use domain-specific language to integrate multimodal information into a precise alternative form . |
| Outcome: | The proposed method significantly improves GPT-4V(ision) on geometry problem solving, chess positional advantage prediction and molecular property prediction. |
Sentence Ordering with a Coherence Verifier (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent sentence ordering studies can be classified into 2 categories: pair-wise ranking-based and sequence generation-based methods. |
| Approach: | They propose a sentence ordering method by plugging a coherence verifier into ranking-based and sequence generation-based methods. |
| Outcome: | The proposed method improves on topological sorting-based and pointer network-based methods with topological and point-based models. |
Mask Attention Networks: Rethinking and Strengthen Transformer (2021.naacl-main)
Copied to clipboard
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang
| Challenge: | Existing research explores to enhance the two sublayers separately to improve the capability of Transformer for text representation. |
| Approach: | They propose to combine SAN and Feed-Forward Networks to create a dynamic mask attention network with a learnable mask matrix which can model localness adaptively. |
| Outcome: | The proposed model outperforms the original Transformer on translation and text summarization tasks. |
Toward Safe and Human-Aligned Game Conversational Recommendation via Multi-Agent Decomposition (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing systems for conversational recommender systems (CRS) have strong results in movies, but games present distinct challenges . MATCHA framework provides specialized agents for intent parsing, tool-augmented retrieval, multi-LLM ranking, and stronger safety. |
| Approach: | They propose a framework for conversational recommender systems that assigns specialized agents for intent parsing, tool-augmented retrieval, multi-LLM ranking and risk control. |
| Outcome: | MATCHA outperforms baselines on real user request dataset, improves Hit@5 by 20%, reduces popularity bias by 24%, and achieves 97.9% adversarial defense. |
R.R.: Unveiling LLM Training Privacy through Recollection and Ranking (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing privacy attacks focus on membership inference or data extraction, but reconstructing specific personally identifiable information (PII) in training data remains challenging. |
| Approach: | They propose a two-step privacy stealing attack that enables attackers to reconstruct PII entities from scrubbed training data where the PI I entities have been masked. |
| Outcome: | The proposed attack can reconstruct PII entities from scrubbed training data where the PI I entities have been masked. |
AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora (2026.acl-long)
Copied to clipboard
Jiaxin Bai, Wei Fan, Qi Hu, Qing Zong, Chunyang Li, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Haoyu Huang, Xiao Zhou, Feng Qin, Tianshi Zheng, Xi Peng, Xin Yao, Huiwen Yang, Leijie Wu, JI Yi, Gong Zhang, Renhai Chen, Yangqiu Song
| Challenge: | Existing knowledge graph construction frameworks require predefined schemas, limiting their scalability and domain coverage. |
| Approach: | They propose a framework for fully autonomous knowledge graph construction that eliminates the need for predefined schemas. |
| Outcome: | The proposed framework outperforms state-of-the-art models on multi-hop QA tasks and enhances LLM factuality. |
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination (2024.emnlp-main)
Copied to clipboard
| Challenge: | Despite the success of Large Vision-Language Models, they suffer from hallucination. |
| Approach: | They propose a training-free strategy that "D**ive into" the attention of LVLMs to "R**educe" object hallucination by using classification tokens of ViT. |
| Outcome: | The proposed method reduces the impact of outlier tokens on LVLMs . the proposed method is based on LLaVA-1.5, LLvaVA-NeXT and InstructBLIP . |
LLM Inductive Reasoning Through Multi-Agent Enhanced Monte Carlo Tree Search (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for enhancing inductive reasoning of large language models often lack explicit optimization guidance and effective error correction. |
| Approach: | They propose a plug-and-play test-time framework that integrates multi-agent coordination with Monte Carlo Tree Search to improve inductive reasoning. |
| Outcome: | The proposed framework outperforms existing methods on four benchmarks and shows consistent improvements on QWQ-32B and Deepseek-V3 . |
Shorten After You’re Right: Lazy Length Penalties for Reasoning RL (2026.findings-acl)
Copied to clipboard
Danlong Yuan, Tian Xie, Shaohan Huang, Huishuai Zhang, Zhuocheng Gong, Chong Luo, Furu Wei, Dongyan Zhao
| Challenge: | Existing shortening methods for long reasoning models rely on additional supervision or multi-stage post-training. |
| Approach: | They propose a lazy length penalty that imposes length pressure on models without extra training stages. |
| Outcome: | The proposed method significantly reduces response length without extra training stages while maintaining or improving performance. |
XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation (2020.emnlp-main)
Copied to clipboard
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou
| Challenge: | XGLUE provides a benchmark dataset to train large-scale cross-lingual pre-trained models . XCLUE provides 11 diversified tasks that cover both understanding and generation scenarios . |
| Approach: | They introduce a new benchmark dataset to train large-scale cross-lingual pre-trained models using multilingual and bilingual corpora. |
| Outcome: | The proposed dataset is labeled in English and includes only natural language understanding tasks. |
An Enhanced Knowledge Injection Model for Commonsense Generation (2020.coling-main)
Copied to clipboard
Zhihao Fan, Yeyun Gong, Zhongyu Wei, Siyuan Wang, Yameng Huang, Jian Jiao, Xuanjing Huang, Nan Duan, Ruofei Zhang
| Challenge: | a recent study shows that digging the relationship of concepts from scratch is non-trivial for commonsense generation tasks. |
| Approach: | They use a retrieve-and-edit framework to retrieve a prototype with these concepts . they use qt and qq to generate commonsense questions at scale . |
| Outcome: | The proposed method significantly improves the performance on commonsense generation tasks. |
Fin-STAR: Structure-as-Semantics to Resolve Implicitness in Financial Retrieval (2026.findings-acl)
Copied to clipboard
Yu Zou, Yan Chen, Lida He, Qi Zhou, Xiaorui Zhou, Aixi Zhong, Yi Wang, Wei Li, Qingyu Wang, Jiatao Li, Wei Gong, Jialei Zeng, Jingmei Zhao, Ke Jiang, Qing Li
| Challenge: | Existing Retrieval-Augmented Generation systems treat structure as a physical navigational skeleton rather than intrinsic semantic knowledge. |
| Approach: | They propose a framework that redefining hierarchy as intrinsic semantics and uses snippets to enrich hierarchical lineage. |
| Outcome: | The proposed framework outperforms state-of-the-art hierarchical and graph-based benchmarks on FinTierQA Gold. |
LLM Factoscope: Uncovering LLMs’ Factual Discernment through Measuring Inner States (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) produce outputs that deviate from factual reality, especially in sensitive applications such as medical consultation and legal advice. |
| Approach: | They propose a Siamese network-based model that leverages LLMs’ inner states for factual detection. |
| Outcome: | The proposed model achieves over 96% accuracy on a custom-collected factual detection dataset. |
GeoSQA: A Benchmark for Scenario-based Question Answering in the Geography Domain at High School Level (D19-1)
Copied to clipboard
| Challenge: | SQA is an emerging application of NLP in the medical, geography, and legal domains. |
| Approach: | They propose a dataset of 1,981 scenarios and 4,110 multiple-choice questions in geography domain at high school level. |
| Outcome: | The proposed dataset consists of 1,981 scenarios and 4,110 multiple-choice questions in the geography domain at high school level. |
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)
Copied to clipboard
Tianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang
| Challenge: | Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level. |
| Approach: | They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context. |
| Outcome: | The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set. |
Be Cautious When Merging Unfamiliar LLMs: A Phishing Model Capable of Stealing Privacy (2025.findings-acl)
Copied to clipboard
| Challenge: | Model merging is a widespread technology in large language models that integrates multiple task-specific LLMs into a unified one. |
| Approach: | They propose a model merging approach that trains a phishing model capable of stealing privacy using a privacy phish instruction dataset. |
| Outcome: | The proposed model cloaking method mimics a specialized capability to conceal attack intent, luring users into merging the phishing model. |
PreQuant: A Task-agnostic Quantization Approach for Pre-trained Language Models (2023.findings-acl)
Copied to clipboard
Zhuocheng Gong, Jiahao Liu, Qifan Wang, Yang Yang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, Rui Yan
| Challenge: | Quantization is a viable solution for pre-trained language models, but most existing methods are task-specific and require customized training and quantization with a large number of trainable parameters. |
| Approach: | They propose a "quantize before fine-tuning" framework that allows for quantization with a large number of trainable parameters on each individual task. |
| Outcome: | The proposed framework is compatible with quantization-aware training and post-training quantization and corrects quantization errors. |
GUI0: Self-Evolving Foundational GUI Agents in Super App Ecosystems (2026.acl-long)
Copied to clipboard
Xinyi Wang, Wei Dai, Kyle Qiao, Ke Wang, Peng Chen, Gang Cao, null Kangqin, Zhongpu Wang, Xiaode Zhang, Yanming Liu, Jihao Gu, Jingtao Xu, Gong Zhi
| Challenge: | Automated interaction with graphical user interfaces (GUIs) is central to general artificial intelligence, but remains challenging within Super App ecosystems. |
| Approach: | They propose a framework synergizing autonomous data synthesis with dual-agent co-evolution . GUI0 establishes a domain-aware foundation model via synthesized corpora and employs curriculum-driven reinforcement learning . |
| Outcome: | The proposed framework outperforms Gemini-2.5-Pro and Claude-4-Sonnet in the SuperAPP benchmark and has universal efficacy across base models. |
Chinese Metaphorical Relation Extraction: Dataset and Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Metaphor identification is a core task in metaphor processing, which involves recognizing and analyzing metaphorical expressions in text. |
| Approach: | They propose a new formulation of metaphor identification as a relation extraction problem . they use a dataset to analyze metaphorical relations between two spans, a target and a source . |
| Outcome: | The proposed model can capture the properties of the target and source in Chinese sentences. |
IoTMigrator: LLM-driven Embedded IoT Code Migration across Different OSes for Cloud-device Integration (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Neither outline-based code generation nor common code translation techniques can adequately address this challenge, despite their prevalence in existing systems. |
| Approach: | They have developed an algorithm that employs a multi-agent pipeline to handle embedded code migration under the TSL paradigm. |
| Outcome: | The proposed algorithm outperforms the baseline by 50.5% for pass rate and 13.0% for completeness across all tasks in RIOT and Zephyr. |
Enhancing Content Planning for Table-to-Text Generation with Data Understanding and Verification (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Table-to-text models that select and order salient data and verbalize them fluently are lacking in content planning stage. |
| Approach: | They propose to enhance neural content planning by understanding data values with contextual numerical value representations that bring the sense of value comparison into content planning. |
| Outcome: | The proposed model outperforms existing systems with respect to content planning metrics on ROTOWIRE and MLB datasets. |
A Triple-View Framework for Fine-Grained Emotion Classification with Clustering-Guided Contrastive Learning (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have focused on dealing with only one of the two difficulties of coarse-grained emotion classification. |
| Approach: | They propose a triple-view framework that treats FEC as an instance-label joint embedding learning problem to tackle both difficulties concurrently by considering three complementary views. |
| Outcome: | The proposed framework achieves significant and consistent improvements on two widely-used benchmark datasets. |
DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation (2022.acl-long)
Copied to clipboard
Wei Chen, Yeyun Gong, Song Wang, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Yi Mao, Weizhu Chen, Biao Cheng, Nan Duan
| Challenge: | Existing pre-trained dialog models shed light on various downstream tasks in natural language processing (NLP). |
| Approach: | They propose a dialog pre-training framework that introduces latent variables into the enhanced encoder-decoder pre-train framework to increase relevance and diversity of responses. |
| Outcome: | The proposed model achieves state-of-the-art on personaChat, DailyDialog, and DSTC7-AVSD datasets. |
Optimizing Chinese Lexical Simplification Across Word Types: A Hybrid Approach (2024.emnlp-main)
Copied to clipboard
| Challenge: | Expensive large language models outperform small models in simplifying complex content words and Chinese idioms from the dictionary. |
| Approach: | They propose to use a retrieval-based interpretation augmentation strategy to refine small models to simplify complex content words and Chinese idioms. |
| Outcome: | The proposed framework outperforms large language models in Chinese Lexical Simplification (CLS) and improves on OOD models. |
CodeM: Less Data Yields More Versatility via Ability Matrix (2024.findings-acl)
Copied to clipboard
Daoguang Zan, Ailun Yu, Wei Liu, Bo Shen, Shaoxin Lin, Yongshun Gong, Yafen Yao, Yan Liu, Bei Guan, Weihua Luo, Yongji Wang, Qianxiang Wang, Lizhen Cui
| Challenge: | Recent efforts to train code large language models have been booming recently . however, this will incur significant costs in constructing data and training model considering the countless downstream scenarios. |
| Approach: | They propose a data construction strategy which decouples code LLMs’ abilities into two dimensions and constructs a lightweight training corpus that only covers a subset of target scenarios. |
| Outcome: | The proposed model can train a multilingual multitasking model using less data and training data. |