Papers by Xingyu Wang

37 papers
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination? (2024.naacl-long)

Copied to clipboard

Challenge: Existing large language models (LLMs) suffer from hallucinations and unfaithful reasoning due to keyword/entity biases.
Approach: They propose a new probing method and benchmark to quantify this phenomenon by using a keyword/entity biases-based probing technique called EUREQA.
Outcome: The proposed method achieves 62% accuracy on multi-hop and complex QA benchmarks.
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities but their application in open-ended, knowledge-intensive, complex reasoning scenarios is limited.
Approach: They propose a framework that integrates risk assessment of intermediate reasoning states with dynamic retrieval-augmented generation within a Monte Carlo tree search paradigm.
Outcome: The proposed framework outperforms the state-of-the-art KAR methods by up to 23.10% and the latest RAG-equipped large reasoning models by upto 25.37%.
Research Replication Prediction Using Weakly Supervised Learning (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to predict scientific claims’ replicability use only hand-extracted statistics features without utilizing research papers’ text information.
Approach: They propose two weakly supervised learning approaches that use automatically extracted text information of research papers to improve the prediction accuracy of research replication using both labeled and unlabeled datasets.
Outcome: The proposed methods achieve an accuracy of 75.76% over real-world datasets.
MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills (2026.eacl-long)

Copied to clipboard

Challenge: Current clinical LLM benchmarks fail to evaluate advanced clinical skills in AI and large language models (LLMs).
Approach: They propose a framework to evaluate large language models (LLMs) using two instruction-following tasks designed to reflect real clinical scenarios.
Outcome: The proposed framework evaluates LLMs through two instruction-following tasks designed to reflect real clinical scenarios.
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Conventional speculative decoding methods use a predefined length policy for proposing drafts, but the reality deviates from this assumption.
Approach: They propose a self-verification length policy that adaptively determines the lengths of draft sequences by referring to the draft entropy.
Outcome: The proposed method achieves 17% speedup on MT-Bench and 22% speedup in long-form reasoning.
JiraiBench: A Bilingual Benchmark for Evaluating Large Language Models’ Detection of Human risky health behavior Content in Jirai Community (2026.eacl-long)

Copied to clipboard

Challenge: a cross-lingual dataset captures a transnational cultural phenomenon . risky health behaviors (RHB) are often linked to complex mental health conditions .
Approach: They present the first cross-lingual dataset that captures a transnational cultural phenomenon . their dataset of more than 15,000 annotated social media posts forms the core of JiraiBench .
Outcome: The study shows that cultural context can be more influential than linguistic similarity . the study also shows that the Japanese prompts better handle Chinese content .
DocMMIR: A Framework for Document Multi-modal Information Retrieval (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multi-modal information retrieval models lack a comprehensive exploration of document-level retrieval . existing models suffer from the absence of cross-domain datasets at this granularity.
Approach: They propose a multi-modal document retrieval framework to unify diverse document formats and domains with a comprehensive retrieval scenario.
Outcome: The proposed framework improves document retrieval performance on a large multimodal dataset.
LaMP-Val: Large Language Models Empower Personalized Valuation in Auction (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, most research focuses on the bidding algorithms used within auction mechanisms.
Approach: They propose a personalized valuation framework that integrates Large Language Models to incorporate personalized semantic preference into users valuation process.
Outcome: The proposed framework incorporates Large Language Models to incorporate personalized semantic preference into users valuation process.
RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) have demonstrated impressive capabilities in code generation across various domains, but their ability to replicate complex, multi-panel visualizations remains largely unassessed.
Approach: They propose a large-scale benchmark to evaluate chart generation from large- scale raw data and assess iterative code refinement in a multi-turn conversational setting.
Outcome: The new benchmark evaluates 14 leading VLMs on real-world data and shows they struggle with complex plot structures and authentic data.
Robust Preference Optimization via Dynamic Target Margins (2025.findings-acl)

Copied to clipboard

Challenge: Direct Preference Optimization (DPO) is an efficient method for ensuring safety and reliability in practical applications.
Approach: They propose a dynamic target margin preference optimization algorithm that adjusts reward margins at the pairwise level.
Outcome: The proposed method achieves an average 4.4% improvement over baselines, setting new benchmarks for state-of-the-art performance.
Alignment for Efficient Tool Calling of Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in tool learning have enabled large language models to integrate external tools, enhancing their task performance by expanding their knowledge boundaries.
Approach: They propose a framework that combines probabilistic knowledge boundary estimation with dynamic decision-making to allow LLMs to better assess when to invoke tools based on their confidence.
Outcome: The proposed framework shows significant improvements in tool efficiency by reducing unnecessary tool usage.
Social Welfare Function Leaderboard: On the Emergence of LLM Agents as the Welfare Dictator (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare.
Approach: They evaluate 20 state-of-the-art Large language models (LLMs) and 20 LLM dictators to create a social welfare function benchmark.
Outcome: The proposed model creates dilemma between maximizing collective efficiency and ensuring distributive fairness.
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly tasked with creative generation, but their ability to portray non-prosocial, antagonistic personas remains largely unexamined.
Approach: They propose a moral alignment benchmark to test the safety of large language models . they find that models struggle with traits directly antithetical to safety principles .
Outcome: The proposed model fails to accurately portray morally ambiguous or villainous characters . the model fails most with traits directly antithetical to safety principles .
CLEAR: Can Language Models Really Understand Causal Graphs? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing language models lack a conceptual framework for understanding causal graphs, but there is still potential for improvement.
Approach: They develop a framework to define causal graph understanding by assessing language models’ behaviors through four practical criteria derived from diverse disciplines.
Outcome: The proposed framework defines three complexity levels and encompasses 20 causal graph-based tasks across 20 different levels.
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have evolved from statistical sequence predictors to sophisticated autonomous agents capable of reasoning, planning, and sustaining multi-turn conversa-tions.
Approach: They propose a system that instantiates a "Sentient Agent" that simulates human-like emotional changes and inner thoughts to provide a more realistic evaluation of the model in multi-turn conversations.
Outcome: The proposed framework measures the agent's higher-order social cognition in multi-turn conversations.
Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Model (LLM)-based agents extend the utility of LLMs by interacting with dynamic environments.
Approach: They propose a parameter fusion framework based on directional consensus evaluation that disentangles knowledge updates through a two-stage process.
Outcome: The proposed framework disentangles knowledge updates through a two-stage process with minimal computational overhead and parameter updates.
TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Industry Systems (2024.emnlp-industry)

Copied to clipboard

Challenge: Large language models have demonstrated proficiency in addressing tasks that necessitate a combination of task planning and the usage of external tools.
Approach: They propose a framework to enhance the task planning and tool usage abilities of LLMs in industrial systems.
Outcome: The proposed framework enhances the task planning and tool usage abilities of LLM-based agents in industrial systems.
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches for optimizing human annotation efforts are limited . et al., 2015) suggest that densely annotated image captions improve vision-language alignment .
Approach: They propose an AI-in-the-loop methodology to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints.
Outcome: The proposed method improves annotation speed and retrieval performance over the parallel method.
Crossroads of Optimization under Uncertainty: How to Choose the Optimal Model (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to Optimization under Uncertainty (OuU) have inherent limitations and advantages.
Approach: They propose a framework that automates the modeling and solving of six types of uncertainty models and generates mapping pairs to explore the potential relationship between optimization problems and optimal models.
Outcome: The proposed framework achieves superior performance even on specific model types, with correlation analysis showing that data scale and specific scenario significantly influence model selection.
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to interleaved reasoning are limited by the cost of re-encoding pixel-dense images.
Approach: They propose a framework that unifies dynamic state evolution with precise perceptual modeling.
Outcome: The proposed framework outperforms existing approaches on multimodal reasoning benchmarks.
Not All Citations Are Equal:Entropy-Guided Citation Selection for Noise-Resistant Medical LLM (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have demonstrated extensive potential in medical applications . however, their practical deployment in healthcare faces significant challenges .
Approach: They propose a training-free multi-turn reasoning framework and a post-training methodology that provides external knowledge support for large language models.
Outcome: The proposed framework elicits internal thought, external thought, and fusion thought, with an entropy-based reward that encourages selective citation of beneficial external knowledge while penalizing noisy citations.
Chain of Strategy Optimization Makes Large Language Models Better Emotional Supporter (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing supervised fine-tuning (SFT) fails to address these issues, as it trains models on single gold-standard responses without modeling nuanced strategy trade-offs.
Approach: They propose a two-stage framework that optimizes strategy selection preferences at each dialogue turn.
Outcome: The proposed framework improves strategy selection preferences at each dialogue turn.
ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA (2025.emnlp-main)

Copied to clipboard

Challenge: Multi-hop question answering (QA) is a central challenge in natural language processing . early mistakes can cause errors and undermine the final result, authors say .
Approach: They propose a reversible multi-agent reasoning framework that backtracks to earlier valid states when conflicts arise.
Outcome: Empirical evaluation shows that the framework improves on forward-only benchmarks by 6% . the approach enables agents to backtrack to valid states when conflicts arise .
CodeRipple: Wavelet-Based Detection of LLM-Generated Code (2026.acl-long)

Copied to clipboard

Challenge: Existing training-free detectors rely on global statistics of the Token Perplexity Sequence (TPS) and struggle with code.
Approach: They propose a training-free detection framework that characterizes TPS morphology across scales.
Outcome: The proposed framework outperforms existing training-free detectors on three challenging benchmarks spanning programming languages, multiple generating LLMs, and various evasion strategies.
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large vision-language models produce hallucinations that compromise output reliability.
Approach: They propose a dual-stage framework for mitigating hallucinations without performance degradation . they propose semantic-aware component disentanglement and interpretable parameter updates .
Outcome: The proposed model reduces hallucinations by 23.4% while maintaining 97.4% of general generative capability.
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via sequence-level likelihood (2026.acl-long)

Copied to clipboard

Challenge: Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs).
Approach: They propose a token-level framework that leverages sequence-level likelihood to link group-level rewards with individual tokens via token- level aggregation and introduces a KL-Divergence mask constraint that targets tokens with positive advantages and decreasing entropy to mitigate abrupt policy updates.
Outcome: Experiments show that TEPO achieves state-of-the-art performance on mathematical reasoning benchmarks and reduces convergence time by 50% compared with GRPO/DAPO.
Rethinking Word-Level Auto-Completion in Computer-Aided Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models for word-level auto-completion (WLAC) do not meet the criterion of good auto-completes.
Approach: They propose a measurable criterion to address the question: what kind of words are good auto-completions? they propose an approach to enhance WLAC performance by promoting adherence to the cri-terion.
Outcome: The proposed approach outperforms the top-performing system submitted to the WLAC shared tasks in WMT2022 while using significantly smaller model sizes.
PAR2-RAG: Planned Active Retrieval and Reasoning for Multi-Hop Question Answering (2026.acl-industry)

Copied to clipboard

Challenge: Multi-hop question answering is a practical bottleneck in industry applications . large language models (LLMs) fail frequently when evidence coverage is incomplete or reasoning trajectories drift .
Approach: They propose a training-free two-stage framework that separates coverage from commitment . it performs breadth-first anchoring to build a high-recall evidence frontier . compared with IRCoT, it achieves 23.5% higher answer accuracy .
Outcome: The proposed framework outperforms strong baselines in MHQA benchmarks and achieves 23.5% higher answer accuracy and 10.5% NDCG gains in retrieval quality.
Massive End-to-end Speech Recognition Models with Time Reduction (2024.naacl-long)

Copied to clipboard

Challenge: Using the neural architecture of Google’s universal speech model, we reduce the frame rate and speed up training and inference.
Approach: They propose to use the neural architecture of Google’s universal speech model with additional funnel pooling layers to significantly reduce the frame rate and speed up training and inference.
Outcome: The proposed methods work with both connectionist temporal classification (CTC) and RNN-Transducer (RNN-T) and over two domains.
HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Current RAG system retrieves evidence from knowledge graphs and text documents but has limitations in multi-hop reasoning, multi-entity questions, and source verification.
Approach: They propose a training-free framework that unifies graph topology, document semantics, and source reliability to support deep, faithful reasoning in large language models.
Outcome: The proposed framework outperforms the current hybrid model-based model-driven system by 20.3% and 30.1% on seven benchmark datasets.
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to train low-latency multilayer perceptrons (MLPs) on graph tasks are based on graph nodes and lack graph structural information.
Approach: They propose to distill graph structural information from Graph Neural Networks (GNNs) to low-latency multilayer perceptrons (MLPs) on graph tasks.
Outcome: The proposed method does not require graph edges (edge-free setting) yet learns structure-aware MLPs.
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution (2026.acl-industry)

Copied to clipboard

Challenge: Xue et al., 2025): deploying autonomous web agents in production remains difficult due to site heterogeneity and long-horizon instability.
Approach: They propose a knowledge-evolving agent that can be used to automate web workflows . they use human-in-the-loop knowledge adaptation and knowledge-aligned progressive summarization .
Outcome: Experiments on WebArena, WebChoreAren and industrial deployment show it outperforms baselines.
SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing large-scale large-context models suffer from performance degradation when processing long numerical sequences.
Approach: They propose a framework to mitigate attention dispersion by strategically inserting separator tokens into the model to recalibrat attention to local segments while preserving global context.
Outcome: The proposed framework improves accuracy and reduces inference token consumption by 16.4% on 9 widely-adopted LLMs.
BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches often fail to leverage the linguistic intelligence of Large Language Models (LLMs) Existing models lack the ability to follow text instructions for controllable Text-to-Speech (TTS).
Approach: They propose a framework where an LLM acts as a conductor, understanding user instructions and generating a textual plan - explicit vocal features.
Outcome: The proposed model outperforms open- and closed-source models in speech synthesis and achieves zero-shot cross-lingual generalization.
Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge (2023.findings-acl)

Copied to clipboard

Challenge: Open-ended Visual Question Answering (VQA) requires models to reason over visual and natural language inputs using world knowledge.
Approach: They propose a new VQA pipeline that deploys a generate-then-select strategy guided by world knowledge for the first time.
Outcome: The proposed pipeline expands the knowledge coverage from in-domain training data by 4.1% on OK-VQA, without additional computation cost.
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities on general text, but their proficiency in specialized scientific domains remains uncharacterized.
Approach: They evaluate the capabilities of large language models in metabolomics research using MetaBench . they found that models perform well on text generation tasks, but cross-database identifier grounding remains challenging .
Outcome: The evaluation of 25 open- and closed-source LLMs reveals distinct performance patterns across metabolomics tasks.
Logic-of-Thought: Injecting Logic into Contexts for Full Reasoning in Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks but their performance in complex logical reasoning tasks remains unsatisfactory.
Approach: They propose a propositional logic prompting method which generates expanded logical information descriptions and utilizes them as an additional augmentation to original contexts.
Outcome: Extensive experiments show that Logic-of-Thought boosts the performance of various prompting methods with a striking margin across five logical reasoning tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations