Papers by Yang Qiao
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? (2025.acl-long)
Copied to clipboard
Runqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma GongQue, Shanglin Lei, YiFan Zhang, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Yida Xu, Peiqing Yang, Zhimin Bao, Muxi Diao, Chen Li, Honggang Zhang
| Challenge: | Existing benchmarks focus more on end-to-end performance, but neglect the underlying principles of knowledge acquisition and generalization. |
| Approach: | They propose a benchmark specifically designed to explore the problem-solving principles by decomposing 6.5K visual math problems into 10.9K step-level questions for evaluation. |
| Outcome: | The proposed benchmark covers 6.5K visual math problems and 10.9K step-level questions spanning 5 layers of knowledge granularity and 67 hierarchical knowledge concepts. |
Towards Self-Evolving Agents: Enabling Autonomy through Interactive Experience Refinement (2026.findings-acl)
Copied to clipboard
Cheng Yang, Xuemeng Yang, Licheng Wen, Daocheng Fu, Jianbiao Mei, Rong Wu, Pinlong Cai, Yufan Shen, Nianchen Deng, Jia Xu, Botian Shi, Yu Qiao, Haifeng Li
| Challenge: | Large Language Models struggle with complex, multi-step operational tasks because they remain static during inference and cannot learn from past experience. |
| Approach: | They propose a framework that organizes cross-domain insights to facilitate orchestration of long-horizon workflows. |
| Outcome: | The proposed framework outperforms existing methods on the TAC productivity benchmark and shows strong cross-task transferability. |
FinToolSyn: A forward synthesis Framework for Financial Tool-Use Dialogue Data with Dynamic Tool Retrieval (2026.findings-acl)
Copied to clipboard
Caishuang Huang, Yang Qiao, Rongyu Zhang, Junjie Ye, Pu Lu, null Wuwenxi, Meng Zhou, Xiku Du, Qi Zhang, Tao Gui, Xuanjing Huang
| Challenge: | Existing data synthesis methods rely on static tools to generate queries . this approach fails to capture the implicit, event-driven nature of real-world needs . |
| Approach: | They propose a forward synthesis framework to generate high-quality financial dialogues . they construct a repository of 43,066 tools and synthesize over 148k dialogue instances . |
| Outcome: | Experiments show that models trained on FinToolSyn achieve a 21.06% improvement . the framework is designed to generate high-quality financial dialogues . |
Enabling Stroke-Level Structural Analysis of Hieroglyphic Scripts without Language-Specific Priors (2026.findings-acl)
Copied to clipboard
Fuwen Luo, Zihao Wan, Ziyue Wang, Yaluo Liu, Pau Tong Lin Xu, Xuanjia Qiao, Xiaolong Wang, Peng Li, Yang Liu
| Challenge: | Existing structural analysis methods for hieroglyphic scripts are script-specific and labor-intensive. |
| Approach: | They propose a hieroglyphic Stroke Analyzer framework that captures character-internal structures and semantics without handcrafted data. |
| Outcome: | The proposed framework captures character-internal structures and semantics without priors . it can be used to generalize hieroglyphic scripts across languages . |
Neural Models for Reasoning over Multiple Mentions Using Coreference (N18-2)
Copied to clipboard
| Challenge: | Existing Recurrent Neural Network (RNN) layers are biased towards short-term dependencies and hence not suited to such tasks. |
| Approach: | They propose a recurrent layer which is instead biased towards coreferent dependencies and uses coreference annotations extracted from an external system to connect entity mentions belonging to the same cluster. |
| Outcome: | The proposed layer improves performance on Wikihop, LAMBADA and the bAbi AI datasets with large gains when training data is scarce. |
CB-Whisper: Contextual Biasing Whisper Using Open-Vocabulary Keyword-Spotting (2024.lrec-main)
Copied to clipboard
Yuang Li, Yinglu Li, Min Zhang, Chang Su, Jiawei Yu, Mengyao Piao, Xiaosong Qiao, Miaomiao Ma, Yanqing Zhao, Hao Yang
| Challenge: | End-to-end automatic speech recognition systems struggle to recognize rare name entities such as personal names, organizations and terminologies that are not frequently encountered in the training data. |
| Approach: | They propose a convolutional neural network-based ASR system that performs open-vocabulary keyword-spotting before the decoder to match the features between the entities and the utterances. |
| Outcome: | The proposed system significantly improves mixed-error-rate (MER) and entity recall compared to the original Whisper model on three internal datasets and two publicly available datasets. |
SmartSpanNER: Making SpanNER Robust in Low Resource Scenarios (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is one of the most fundamental tasks in natural language processing. |
| Approach: | They propose a method which introduces a Named Entity Head (NEH) prediction task to SpanNER and performs multi-task learning together with task of span classification. |
| Outcome: | The proposed method improves the robustness of SpanNER in low resource scenarios on the CoNLL03, Few-NERD, GENIA and ACE05 benchmark datasets. |
Rumor Detection on Social Media with Crowd Intelligence and ChatGPT-Assisted Networks (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on rumor detection challenges the expressive power of text encoding sequences, and insufficient mining of semantic structural information. |
| Approach: | They propose a Crowd Intelligence-based semantic feature learning module to capture textual content’s sequential and hierarchical features and a knowledge-based structural mining module that leverages ChatGPT for knowledge enhancement. |
| Outcome: | The proposed system achieves performance improvement in rumor detection tasks validating the effectiveness and rationality of using large language models as auxiliary tools. |
Threshold-driven Pruning with Segmented Maximum Term Weights for Approximate Cluster-based Sparse Retrieval (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using rank score thresholding, sparse retrieval skips the index at cluster and document levels. |
| Approach: | They propose a pruning control scheme with a probabilistic guarantee on rank-safeness competitiveness. |
| Outcome: | The proposed pruning control scheme improves accuracy and safeness while delivering low latency on single-threaded CPU. |
When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to early exit reasoning often rely on handcrafted or empirical indicators that are unreliable and impractical. |
| Approach: | They propose a framework that allows LRMs to assess the sufficiency of its chain-of-thought and determine the optimal point for early exit. |
| Outcome: | The proposed framework reduces reasoning length by 28.9%–34.9% with minimal performance loss, effectively mitigating overthinking. |
Inference-Time Language Model Alignment via Integrated Value Guidance (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models are fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. |
| Approach: | They propose a method that uses implicit and explicit value functions to guide language model decoding at token and chunk-level respectively. |
| Outcome: | The proposed method outperforms traditional methods and circumvents the complexities of fine-tuning. |
Cross-Domain Audio Deepfake Detection: Dataset and Analysis (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing audio deepfake detection datasets are outdated and lack generalization capabilities. |
| Approach: | They construct a new cross-domain audio deepfake detection dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models. |
| Outcome: | The proposed models achieve 4.1% and 6.5% error rates in the cross-domain ADD dataset generated by five advanced zero-shot TTS models. |
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey (2024.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are now commonplace in conversation applications, but their misuse for generating harmful responses has raised serious societal concerns. |
| Approach: | They provide a comprehensive overview of recent studies covering attacks, defenses, and evaluations of Large Language Models (LLMs) . |
| Outcome: | The proposed review summarizes three aspects of LLM conversation safety: attacks, defenses, and evaluations. |
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent approaches to language model alignment assume homogeneous human preferences, but actual human preferences vary widely and are hard to satisfy with a single language model. |
| Approach: | They propose an RL-free extension of Direct Preference Optimization (DPO) that folds language modeling directly into reward modeling and trains language models as collective reward models that combine all objectives with specific weights. |
| Outcome: | The proposed method matches or outperforms existing methods in safety alignment and long-form question answering. |
From Mimesis to Metamorphosis: Evolving VLM Judges via In-Context Comparing and Knowledge Internalization (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to subjective assessment are inconsistent and inconsistent due to inconsistent scales and inherent preference biases. |
| Approach: | They propose a framework that operationalizes subjective assessment as comparative analysis and internalizes it via Language Buttons. |
| Outcome: | The proposed framework achieves state-of-the-art performance across multiple benchmarks and is scale-steerable. |
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agents (2026.acl-long)
Copied to clipboard
Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang, Qingyun Li, Yu Qiao, Zun Wang, Zichen Ding
| Challenge: | Vision-Language Models (VLMs) lack visual-aware tutorial retrieval and historical visual context curation and pruning. |
| Approach: | They propose a framework that integrates an orchestrator and a Reflection-Memory Agent for robust automation. |
| Outcome: | Experimental results show that OS-Symphony delivers substantial performance gains across model scales. |
Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Tool-integrated reasoning (TIR) enables LLM agents to solve tasks through planning, tool use, and iterative revision, but outcome-only reinforcement learning suffers from sparse, delayed rewards and weak step-level credit assignment. |
| Approach: | They propose a tool-integrated reasoning approach that localizes the first irrecoverable step and leverages it for fine-grained credit assignment. |
| Outcome: | The proposed algorithm outperforms strong Agentic RL benchmarks in math, science QA, and code execution with additional gains in Pass@K and Major@K scaling, rollout ranking quality, and tool-call efficiency. |
Aspect and Sentiment Aware Abstractive Review Summarization (C18-1)
Copied to clipboard
| Challenge: | Abstractive summarization is a task that generates short and concise summaries of user generated reviews. |
| Approach: | They propose an interactive attention mechanism to learn the representations of context and aspect words within reviews, acted as an encoder. |
| Outcome: | The proposed model achieves impressive results compared to other strong competitors on a real-life dataset. |
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs (2025.acl-long)
Copied to clipboard
Tao Zhang, ChengLIn Zhu, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Tao Zhang, Fan Yang, Mingan Lin, Yujing Qiao, Weipeng Chen, Bin Cui, Wentao Zhang, Zenan Zhou
| Challenge: | Existing evaluations of Large Language Models (LLMs) focus on fragmented constraints or narrow scenarios, but they overlook the comprehensiveness and authenticity of constraints from the user’s perspective. |
| Approach: | They propose a Chinese Comprehensive Constraints Following Benchmark for LLMs that compiles constraints from real-world instructions and constructs a systematic framework for constraint types. |
| Outcome: | The proposed framework integrates multi-dimensional assessment criteria with requirement prioritization, covering various perspectives of constraints, instructions, and requirement fulfillment. |
V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me (2025.acl-long)
Copied to clipboard
Runqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Jiapeng Wang, YiFan Zhang, Zhuoma GongQue, Chong Sun, Yida Xu, Yadong Xue, Ye Tian, Zhimin Bao, Lan Yang, Chen Li, Honggang Zhang
| Challenge: | Deciphering oracle bone scripts using AI technology is not an overnight task due to the evolution of written language over millennia. |
| Approach: | They propose a framework that utilizes Large Multi-modal Models (LMMs) for interpreting Oracle Bone Script (OBS). |
| Outcome: | The proposed framework provides quantitative analyses and superior deciphering capability. |
A Sentiment-Controllable Topic-to-Essay Generator with Topic Knowledge Graph (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Topic-to-essay generation is a promising task for natural language generation. |
| Approach: | They propose a Sentiment Controllable topic-to- essay generator with a Topic Knowledge Graph enhanced decoder to generate essays with only several given topic words. |
| Outcome: | The proposed model outperforms the state-of-the-art model on automatic and human evaluation. |
XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation (2020.emnlp-main)
Copied to clipboard
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou
| Challenge: | XGLUE provides a benchmark dataset to train large-scale cross-lingual pre-trained models . XCLUE provides 11 diversified tasks that cover both understanding and generation scenarios . |
| Approach: | They introduce a new benchmark dataset to train large-scale cross-lingual pre-trained models using multilingual and bilingual corpora. |
| Outcome: | The proposed dataset is labeled in English and includes only natural language understanding tasks. |
SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods focus on single-step reasoning, ignoring logical dependencies between steps. |
| Approach: | They propose a method that maximizes a structure-based return to facilitate structured reasoning and explanation. |
| Outcome: | The proposed method outperforms state-of-the-art methods on EntailmentBank and STREET benchmarks. |
Deputy: Accelerating Large Language Model Inference with Dynamic Low-Rank Substitution (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing dynamic schemes such as early-exit and layer-drop reduce FLOPs but break batch processing or introduce KV-cache inconsistency. |
| Approach: | They propose a dynamic low-rank substitution framework that employs a lightweight decision module at each layer to dynamically determine the execution branch for different tokens. |
| Outcome: | The proposed model reduces computation by approximately 40% compared to the original dense model while outperforming existing baseline methods. |
On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning (2026.acl-long)
Copied to clipboard
Changyu Liu, Yiyang Liu, Taowen Wang, Qiao Zhuang, James Chenhao Liang, Wenhao Yang, Renjing Xu, Qifan Wang, Dongfang Liu, Cheng Han
| Challenge: | Existing vision-language-action models are unsuitable for simulated or physical-world deployments . current methods fail when confronted with inherent real-world dynamic variability. |
| Approach: | They propose a test-time reinforcement learning framework that enables on-the-fly policy adaptation during inference. |
| Outcome: | Empirical results show that the proposed framework improves adaptability, stability and task success in dynamic, previously unseen scenarios. |
MiniConGTS: A Near Ultimate Minimalist Contrastive Grid Tagging Scheme for Aspect Sentiment Triplet Extraction (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches within the pretraining-finetuning paradigm tend to meticulously craft complex tagging schemes and classification heads, or incorporate external semantic enhancements to enhance performance. |
| Approach: | They propose to integrate a minimalist tagging scheme and a novel token-level contrastive learning strategy to improve pretrained representations. |
| Outcome: | The proposed framework achieves comparable or superior performance compared to state-of-the-art techniques while featuring a more compact design and reduced computational overhead. |
Exploring Pre-trained Language Models for Event Extraction and Generation (P19-1)
Copied to clipboard
| Challenge: | Existing methods to extract event data are laborious to create and limited in size. |
| Approach: | They propose an event extraction model to overcome the roles overlap problem by separating the argument prediction in terms of roles. |
| Outcome: | The proposed method surpasses existing methods on the ACE2005 dataset and improves on the previous methods. |
Compact Token Representations with Contextual Quantization for Efficient Document Re-ranking (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work has adopted a late interaction architecture with pre-computed contextual token representations at the cost of a large online storage. |
| Approach: | They propose to decouple document-specific and document-independent ranking contributions during codebook-based compression for effective online decompression and embedding composition for better search relevance. |
| Outcome: | The proposed model achieves high relevance and space efficiency with minimal computational cost and low computational complexity. |
Decoder-Only LLMs can be Masked Auto-Encoders (2025.acl-short)
Copied to clipboard
Dan Qiao, Yuan Gao, Zheming Yang, Di Yang, Ziheng Wu, Pengcheng Lu, Minghui Qiu, Juntao Li, Min Zhang
| Challenge: | Modern NLP workflows require different models for generation and embedding tasks. |
| Approach: | They propose a method that transforms an LLM into a Uni-Directional Masked Auto-Encoder. |
| Outcome: | The proposed method achieves state-of-the-art under unsupervised conditions with merely 100 training steps. |
OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use (2025.acl-long)
Copied to clipboard
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu
| Challenge: | a new generation of (M)LLMs is enabling the creation of superintelligent AI assistants . OS Agents can complete tasks autonomously and have the potential to significantly enhance the lives of billions of users worldwide. |
| Approach: | They propose to build OS Agents that operate within operating systems' GUIs and GUIs . they examine evaluation metrics and benchmarks to identify promising directions . |
| Outcome: | The proposed agents are based on operating systems (OS) and operating systems frameworks. |
MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense Disambiguation (2025.emnlp-main)
Copied to clipboard
Kaiyuan Zhang, Qian Liu, Luyang Zhang, Chaoqun Zheng, Shuaimin Li, Bing Xu, Muyun Yang, Xinxiao Qiao, Wenpeng Lu
| Challenge: | Word sense disambiguation (WSD) is a fundamental yet challenging task in natural language processing. |
| Approach: | a novel multi-agent Debate framework for adversarial word Sense disambiguation is proposed . the framework simulates a real-world debate environment where multiple agents engage in discussions about ambiguous words in the context of adversarials. |
| Outcome: | The proposed framework integrates with existing LLMs and improves models in Chinese language . it shows that it can be used to improve models in the Chinese language and improve performance . |
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms (2025.findings-emnlp)
Copied to clipboard
Xiaojun Bi, Shuo Li, Junyao Xing, Ziyue Wang, Fuwen Luo, Weizheng Qiao, Lu Han, Ziwei Sun, Peng Li, Yang Liu
| Challenge: | Dongba pictographic is the only pictograph script still in use in the world. |
| Approach: | DongbaMIE is the first dataset focusing on multimodal information extraction of Dongbe pictographs. |
| Outcome: | The dataset contains 23,530 sentence-level and 2,539 paragraph-level high-quality text-image pairs. |
Empower Large Language Model to Perform Better on Industrial Domain-Specific Question Answering (2023.emnlp-industry)
Copied to clipboard
Fangkai Yang, Pu Zhao, Zezhong Wang, Lu Wang, Bo Qiao, Jue Zhang, Mohit Garg, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
| Challenge: | Large Language Models (LLMs) have gained popularity but lack specific domain knowledge in domain-specific tasks. |
| Approach: | They propose a model interaction paradigm that empowers LLM to achieve better performance on domain-specific tasks where it is not proficient. |
| Outcome: | The proposed approach outperforms the commonly used LLM with retrieval methods in domain-specific tasks. |
IG-Pruning: Input-Guided Block Pruning for Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for efficient inference rely on fixed block masks, which can lead to suboptimal performance. |
| Approach: | They propose an input-aware block-wise pruning method that dynamically selects layer masks at inference time. |
| Outcome: | The proposed method outperforms state-of-the-art static depth pruning methods . it is particularly suitable for resource-constrained deployment scenarios . |
Partial Order-centered Hyperbolic Representation Learning for Few-shot Relation Extraction (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for few-shot relation extraction are limited to labeled instances and rely on data labeling. |
| Approach: | They propose a partial order-centered hyperbolic representation learning framework which imposes constraints on relations on instances by modeling partial order in hyperbolical space. |
| Outcome: | The proposed framework outperforms baseline methods on three benchmark datasets on 1-shot settings lacking relation descriptions. |
Multimodal Machine Translation with Text-Image In-depth Questioning (2025.findings-acl)
Copied to clipboard
| Challenge: | Multimodal machine translation (MMT) models focus on intermodal interactions, but focus on simple interactions between nouns and entities in image, overlooking global semantic alignment. |
| Approach: | They propose a Text-Image In-depth Questioning method to deepen interactions and optimize translations by utilizing visual data to capture global semantic alignment. |
| Outcome: | The proposed method achieves state-of-the-art results on five translation directions of Multi30K and AmbigCaps, with +2.35 BLEU on the challenging MSCOCO benchmark. |