Papers by Yutong Wang
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions (2024.findings-acl)
Copied to clipboard
Honglin Lin, Siyu Li, Guoshun Nan, Chaoyue Tang, Xueting Wang, Jingxin Xu, Rong Yankai, Zhouzhili Zhouzhili, Yutong Gao, Qimei Cui, Xiaofeng Tao
| Challenge: | Existing approaches to image retrieval from contextual descriptions (IRCD) lag behind human performance in IRCD. |
| Approach: | They propose a method that relies on a doubly contextual alignment scheme for challenging IRCD. |
| Outcome: | The proposed method can yield comparable results with GPT-4V, despite fewer parameters. |
AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing MAS initialization methods do not fully account for the collaborative needs of the generated agents in subsequent stages. |
| Approach: | They propose to use a Natural Language to Format mechanism to optimize the structure of agent teams and incorporate a natural language to format mechanism to ensure consistency and standardization. |
| Outcome: | The proposed method outperforms state-of-the-art initialization methods and pre-defined strategies across various frameworks and tasks while reducing token consumption. |
PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations (2026.acl-long)
Copied to clipboard
Yuhe Wu, Guangyu Wang, Yuran Chen, Jiatong Zhang, Yutong Zhang, Yujie Chen, Jiaming Shang, Guang Zhang, Zhuang Liu
| Challenge: | Existing benchmarks for hallucination evaluation rely on mixed queries and posterior evaluation, which quantifies hallucinosity severity but offers limited insight into where and why they occur. |
| Approach: | They propose a controlled benchmark that disentangles hallucinations into four dimensions: knowledge missing, knowledge errors, reasoning errors, and instruction-following errors. |
| Outcome: | The proposed model disentangles hallucinations into four dimensions: knowledge missing, knowledge errors, reasoning errors, and instruction-following errors. |
AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for MAS suffer from high token consumption and inefficiency due to frequent generation and communication among multiple agents. |
| Approach: | They propose a multi-agent system based on large language models that identifies redundant agents and communication across different communication rounds by optimizing the adjacency matrices of the communication graphs and eliminates them to enhance both token efficiency and task performance. |
| Outcome: | The proposed method reduces prompt token consumption and completion token consumption by 18.4% and improves task performance by 1.14. |
DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to idiom translation are limited by the constraints of static parametric memory and retrieval noise . idiomatic expressions are non-compositional units where figurative meanings diverge from literal interpretations . |
| Approach: | They propose a detect-retrieve-arbitrate framework that detects idiomatic spans by reasoning over semantic conflicts between literal and contextual meanings. |
| Outcome: | The proposed framework improves GPT-5-mini and Emerging Slang datasets on various model scales. |
Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language Models (2025.emnlp-main)
Copied to clipboard
Yi Feng, Jiaqi Wang, Wenxuan Zhang, Zhuang Chen, Shen Yutong, Xiyao Xiao, Minlie Huang, Liping Jing, Jian Yu
| Challenge: | Existing approaches to mental health support lack realism and capture therapeutic progression over time. |
| Approach: | They propose a framework that simulates expert narrative therapists by planning therapeutic stages, guiding reflection levels, and generating contextually appropriate responses through retrieval-augmentation. |
| Outcome: | The proposed framework outperforms standard methods in quality and depth on 260 simulated clients and 230 human participants. |
Enhancing Zero-shot and Few-shot Stance Detection with Commonsense Knowledge Graph (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for stance detection are not applicable to zero-shot and few-shot scenarios. |
| Approach: | They propose a model that integrates commonsense knowledge into a stance detection model. |
| Outcome: | The proposed model outperforms state-of-the-art methods on zero-shot and few-shot stance detection tasks. |
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time (2025.emnlp-main)
Copied to clipboard
Junyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, Huan Zhang
| Challenge: | Existing monotonic scaling methods for large reasoning models are not reliable. |
| Approach: | They propose a universal framework for modulating reasoning progress in large reasoning models at test time. |
| Outcome: | The proposed framework unifies and generalizes existing monotonic scaling methods and enables flexible and dense slow-to-fast reasoning modulation. |
Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio (2026.acl-long)
Copied to clipboard
Kaixiong Gong, Kaituo Feng, Bohao Li, Yibing Wang, Mofan Cheng, Shijia Yang, Jiaming Han, Benyou Wang, Yutong Bai, Zhuoran Yang, Xiangyu Yue
| Challenge: | Recent multimodal large language models lack robust audio-visual integration ability and performance on DeafTest is highly correlated with AV-Odyssey accuracy. |
| Approach: | They propose a benchmarking tool that integrates audio-visual reasoning with audio-video cues to infer solutions. |
| Outcome: | The proposed model performs well on DeafTest, but lacks audio perception in simple audio tasks. |
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines (2025.naacl-long)
Copied to clipboard
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo
| Challenge: | Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts. |
| Approach: | They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset. |
| Outcome: | The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages. |
Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and Challenge (2023.acl-long)
Copied to clipboard
| Challenge: | CR is the ability to understand and navigate the world using basic knowledge and understanding shared by most people. |
| Approach: | They propose to incorporate pretrained knowledge into NMT models and use them as robust testbeds for investigating CR in NMT. |
| Outcome: | The proposed method improves the training of NMT models with high CR abilities and provides accurate evaluation metrics. |
Aligning Language Models with Real-time Knowledge Editing (2026.acl-long)
Copied to clipboard
| Challenge: | Mainstream knowledge editing methods are static and fail to keep pace with the evolving real-world knowledge. |
| Approach: | They propose a new paradigm for knowledge editing that integrates edit augmentation and self-adaptive post-alignment inference into CRAFT to improve edit success. |
| Outcome: | The proposed method shows significant performance gain on CRAFT and traditional datasets compared to existing methods. |
AscendKernelGen: LLM-Driven Kernel Generation for NPUs (2026.findings-acl)
Copied to clipboard
Xinzi Cao, Jianyang Zhai, Pengfei Li, Zhiheng Hu, Cen Yan, null Mubingxu, Guanghuan Fang, Bin She, Jiayu Li, Yihan Su, Dongyang Tao, Feidiao Yang, Chang-Dong Wang, Yutong Lu, Weicheng Xue, Bin Zhou, Yonghong Tian
| Challenge: | Neural Processing Units (NPUs) are critical for AI infrastructure, but their development remains a bottleneck due to vendor-specific Domain-Specific Languages (DSLs). |
| Approach: | They propose a framework for NPU kernel development that bridges the gap in hardware-specific coding . compiler success on complex Level-2 kernels improves from 0% to 95.5%, they say . |
| Outcome: | The proposed framework bridges the gap in hardware-specific coding, showing a near-zero success rate on complex kernels. |
TLUE: A Tibetan Language Understanding Evaluation Benchmark (2025.emnlp-main)
Copied to clipboard
Fan Gao, Cheng Huang, Yutong Liu, Nyima Tashi, Xiangxiang Wang, Thupten Tsering, Ban Ma-bao, Renzeng Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Xiao Feng Cd, Yongbin Yu, Hao Wang
| Challenge: | Low-resource languages, like Tibetan, remain underrepresented in large language models' evaluations. |
| Approach: | They propose a Tibetan Language Understanding Evaluation Benchmark to assess LLMs' proficiency in Tibetan . they use a multi-task understanding benchmark and a safety benchmark to evaluate models . |
| Outcome: | The proposed benchmark shows that most large language models perform below the random baseline, especially in Tibetan language processing. |
TasTe: Teaching Large Language Models to Translate through Self-Reflection (2024.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to enhance LLMs' performance in machine translation are unable to fully exploit their instruction-following capabilities. |
| Approach: | They propose a framework for translating through self-reflection that involves two stages of inference . they propose to use the framework to refine LLMs' preliminary translations . |
| Outcome: | The proposed framework can produce translation outputs that match the quality of NMT systems. |
Interactive Plot Manipulation using Natural Language (2021.naacl-demos)
Copied to clipboard
| Challenge: | a new interactive plotting agent is available for programming with natural language . the interactive aspect allows users to manipulate plots using natural language instructions. |
| Approach: | They propose an interactive natural language interface for plotting that maps language to plot updates. |
| Outcome: | The proposed system maps language to plot updates within an interactive programming environment. |