Papers by Yutong Wang

16 papers
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to image retrieval from contextual descriptions (IRCD) lag behind human performance in IRCD.
Approach: They propose a method that relies on a doubly contextual alignment scheme for challenging IRCD.
Outcome: The proposed method can yield comparable results with GPT-4V, despite fewer parameters.
AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing MAS initialization methods do not fully account for the collaborative needs of the generated agents in subsequent stages.
Approach: They propose to use a Natural Language to Format mechanism to optimize the structure of agent teams and incorporate a natural language to format mechanism to ensure consistency and standardization.
Outcome: The proposed method outperforms state-of-the-art initialization methods and pre-defined strategies across various frameworks and tasks while reducing token consumption.
PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for hallucination evaluation rely on mixed queries and posterior evaluation, which quantifies hallucinosity severity but offers limited insight into where and why they occur.
Approach: They propose a controlled benchmark that disentangles hallucinations into four dimensions: knowledge missing, knowledge errors, reasoning errors, and instruction-following errors.
Outcome: The proposed model disentangles hallucinations into four dimensions: knowledge missing, knowledge errors, reasoning errors, and instruction-following errors.
AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for MAS suffer from high token consumption and inefficiency due to frequent generation and communication among multiple agents.
Approach: They propose a multi-agent system based on large language models that identifies redundant agents and communication across different communication rounds by optimizing the adjacency matrices of the communication graphs and eliminates them to enhance both token efficiency and task performance.
Outcome: The proposed method reduces prompt token consumption and completion token consumption by 18.4% and improves task performance by 1.14.
DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to idiom translation are limited by the constraints of static parametric memory and retrieval noise . idiomatic expressions are non-compositional units where figurative meanings diverge from literal interpretations .
Approach: They propose a detect-retrieve-arbitrate framework that detects idiomatic spans by reasoning over semantic conflicts between literal and contextual meanings.
Outcome: The proposed framework improves GPT-5-mini and Emerging Slang datasets on various model scales.
Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to mental health support lack realism and capture therapeutic progression over time.
Approach: They propose a framework that simulates expert narrative therapists by planning therapeutic stages, guiding reflection levels, and generating contextually appropriate responses through retrieval-augmentation.
Outcome: The proposed framework outperforms standard methods in quality and depth on 260 simulated clients and 230 human participants.
Enhancing Zero-shot and Few-shot Stance Detection with Commonsense Knowledge Graph (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for stance detection are not applicable to zero-shot and few-shot scenarios.
Approach: They propose a model that integrates commonsense knowledge into a stance detection model.
Outcome: The proposed model outperforms state-of-the-art methods on zero-shot and few-shot stance detection tasks.
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time (2025.emnlp-main)

Copied to clipboard

Challenge: Existing monotonic scaling methods for large reasoning models are not reliable.
Approach: They propose a universal framework for modulating reasoning progress in large reasoning models at test time.
Outcome: The proposed framework unifies and generalizes existing monotonic scaling methods and enables flexible and dense slow-to-fast reasoning modulation.
Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio (2026.acl-long)

Copied to clipboard

Challenge: Recent multimodal large language models lack robust audio-visual integration ability and performance on DeafTest is highly correlated with AV-Odyssey accuracy.
Approach: They propose a benchmarking tool that integrates audio-visual reasoning with audio-video cues to infer solutions.
Outcome: The proposed model performs well on DeafTest, but lacks audio perception in simple audio tasks.
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines (2025.naacl-long)

Copied to clipboard

Challenge: Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts.
Approach: They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset.
Outcome: The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages.
Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and Challenge (2023.acl-long)

Copied to clipboard

Challenge: CR is the ability to understand and navigate the world using basic knowledge and understanding shared by most people.
Approach: They propose to incorporate pretrained knowledge into NMT models and use them as robust testbeds for investigating CR in NMT.
Outcome: The proposed method improves the training of NMT models with high CR abilities and provides accurate evaluation metrics.
Aligning Language Models with Real-time Knowledge Editing (2026.acl-long)

Copied to clipboard

Challenge: Mainstream knowledge editing methods are static and fail to keep pace with the evolving real-world knowledge.
Approach: They propose a new paradigm for knowledge editing that integrates edit augmentation and self-adaptive post-alignment inference into CRAFT to improve edit success.
Outcome: The proposed method shows significant performance gain on CRAFT and traditional datasets compared to existing methods.
AscendKernelGen: LLM-Driven Kernel Generation for NPUs (2026.findings-acl)

Copied to clipboard

Challenge: Neural Processing Units (NPUs) are critical for AI infrastructure, but their development remains a bottleneck due to vendor-specific Domain-Specific Languages (DSLs).
Approach: They propose a framework for NPU kernel development that bridges the gap in hardware-specific coding . compiler success on complex Level-2 kernels improves from 0% to 95.5%, they say .
Outcome: The proposed framework bridges the gap in hardware-specific coding, showing a near-zero success rate on complex kernels.
TLUE: A Tibetan Language Understanding Evaluation Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Low-resource languages, like Tibetan, remain underrepresented in large language models' evaluations.
Approach: They propose a Tibetan Language Understanding Evaluation Benchmark to assess LLMs' proficiency in Tibetan . they use a multi-task understanding benchmark and a safety benchmark to evaluate models .
Outcome: The proposed benchmark shows that most large language models perform below the random baseline, especially in Tibetan language processing.
TasTe: Teaching Large Language Models to Translate through Self-Reflection (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to enhance LLMs' performance in machine translation are unable to fully exploit their instruction-following capabilities.
Approach: They propose a framework for translating through self-reflection that involves two stages of inference . they propose to use the framework to refine LLMs' preliminary translations .
Outcome: The proposed framework can produce translation outputs that match the quality of NMT systems.
Interactive Plot Manipulation using Natural Language (2021.naacl-demos)

Copied to clipboard

Challenge: a new interactive plotting agent is available for programming with natural language . the interactive aspect allows users to manipulate plots using natural language instructions.
Approach: They propose an interactive natural language interface for plotting that maps language to plot updates.
Outcome: The proposed system maps language to plot updates within an interactive programming environment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations