Papers by Yixiao Wang

13 papers
VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic Levels (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Video Large Language Models (Video-LLMs) enhance the ability of VAU models to describe and interpret anomalies.
Approach: They propose a benchmark that explicitly defines anomalies across five semantic levels and provides detailed temporal boundaries and detailed textual descriptions for each.
Outcome: The proposed benchmark defines anomalies across five semantic levels and provides detailed descriptions for each.
Personal Large Language Model Agents: A Case Study on Tailored Travel Planning (2024.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are becoming more autonomous and capable of handling real-world tasks through their access to tools, various planning strategies, and memory, referred to as LLM agents.
Approach: They introduce a personalized version of TravelPlanner and establish baselines for personal LLM agents by comparing generic and personal plans.
Outcome: The proposed model encapsulates user-related information, preferences, and personal concepts and provides baselines for personal LLM agents.
How Do LLMs "Trust" Unknown Knowledge? An Unknown Knowledge Based Jailbreak Attack (2026.findings-acl)

Copied to clipboard

Challenge: Existing research on how to effectively utilize unknown knowledge has focused on how it can be used to enhance LLMs' performance in specialized fields.
Approach: They propose a completely unrestricted and fully randomized jailbreak attack that embeds malicious queries within trust-enhanced unknown knowledge.
Outcome: The proposed method achieves 99% to 100% ASR on all tested LLMs, including the latest GPT-5.1, and becomes SOTA.
BotChat: Evaluating LLMs’ Capabilities of Having Multi-Turn Dialogues (2024.findings-naacl)

Copied to clipboard

Challenge: Modern Large Language Models (LLMs) facilitate high-quality, multi-turn dialogues with humans, but human-based evaluation of such a capability requires substantial manual effort.
Approach: They propose to evaluate LLMs' ability to emulate human-like, multi-turn conversations using an LLM-centric approach.
Outcome: The proposed model emulates human-like, multi-turn conversations using an LLM-centric approach.
Sentence Selection Strategies for Distilling Word Embeddings from BERT (2022.lrec-1)

Copied to clipboard

Challenge: Using language models to learn word embeddings is a key feature of transformer-based language models.
Approach: They propose to use language models to learn high-quality word vectors from as few as 5 to 10 sentences with a careful selection strategy.
Outcome: The proposed strategies can learn high-quality word vectors from as few as 5 to 10 sentences.
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots (2025.findings-naacl)

Copied to clipboard

Challenge: Multi-modal Large Language Models have shown remarkable progress in visual contexts, yet their ability to convert visual figures into executable code remains underexplored.
Approach: They propose to use a set of visual coding metrics to assess MLLMs' visual . pass rate, text-match ratio, and GPT-4V rating judgement to assess the quality of generated code and rendered images.
Outcome: The proposed benchmark includes 132 high-quality matplotlib plots across six plot types, as well as 150 and 86 plots from Python’s and R’s plotly libraries respectively, totaling 368 plots.
Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) can be effective at rewriting toxic content, but they often default to overly polite rewrites, distorting the emotional tone and communicative intent.
Approach: They evaluate 17 large language models with variant architectures to evaluate their ability to rewrite toxic content while preserving the speaker's original intent.
Outcome: The first Chinese detoxification dataset explicitly designed to preserve sentiment polarity is evaluated across five real-world scenarios.
LLaMA Pro: Progressive LLaMA with Block Expansion (2024.acl-long)

Copied to clipboard

Challenge: Existing studies have demonstrated that pre-trained LLMs are limited in certain domains, such as programming, mathematics, biomedical, or finance.
Approach: They propose a new post-pretraining method with an expansion of Transformer blocks to tune the expanded blocks using only new corpus, efficiently and effectively improving the model’s knowledge while mitigating forgetting.
Outcome: The proposed model outperforms existing models in programming and math and its instruction-following counterpart LLaMA Pro-8.3B in general tasks, programming, and mathematics.
kNN-LM Does Not Improve Open-ended Text Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Interpolation-based retrieval-augmented language models (LMs) are a subtype of retrieval augmented language model that computes the probability of the next token by interpolating between the softmax distribution of the original LM and a token distribution formed by retrieving over an external datastore.
Approach: They propose to interpolate the predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix.
Outcome: The proposed methods do not exhibit improvements in open-ended generation quality, as measured by automatic evaluation metrics and human evaluations.
Investigating the Personality Consistency in Quantized Role-Playing Dialogue Agents (2024.emnlp-industry)

Copied to clipboard

Challenge: Using the Big Five personality traits model, we evaluate how stable assigned personalities are for Quantized Role-Playing Dialog Agents (QRPDA) during multi-turn interactions.
Approach: They propose a non-parametric method to evaluate the stability of assigned personalities in quantized large language models (LLMs) for role-playing scenarios.
Outcome: The proposed method shows that it maintains consistent personality traits in QRPDA, and it is more reliable in real-world applications.
SaCa: A Highly Compatible Reinforcing Framework for Knowledge Graph Embedding via Structural Pattern Contrast (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing knowledge Graph Embedding approaches lack structural semantics of knowledge graphs . structure-aware calibration (SaCa) is a framework designed to calibrate KGEs based on global structural patterns.
Approach: a new framework is designed to calibrate knowledge graphs using global structural patterns.
Outcome: a new framework can calibrate KGE models using global structural patterns . the framework consistently boosts performance across ten models on link prediction and entity classification tasks .
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Recent reinforcement learning approaches have advanced reasoning in Large Language Models (LLMs), yet their adaptation to multimodal LLMs remains underexplored.
Approach: They propose a reinforcement learning framework that eliminates KL penalties and rewards consistency . they propose GRPO-CARE, which outperforms standard GR PO, with a base reward for accuracy and an adaptive bonus for consistency.
Outcome: The proposed framework outperforms standard GRPO on the most difficult evaluation level and reasoning consistency test benchmarks.
Evaluating and Mitigating Object Hallucination in Large Vision-Language Models: Can They Still See Removed Objects? (2025.naacl-long)

Copied to clipboard

Challenge: LVLMs often mistakenly determine objects as present in images where they do not exist . authors propose a new benchmark to evaluate object hallucinations by removing objects from images and asking the model whether it can still see the removed objects.
Approach: They propose a benchmark to evaluate object hallucinations by removing objects from images . they propose oDPO, a direct preference optimization objective based on visual objects .
Outcome: The proposed benchmark reduces the likelihood of object hallucinations by removing objects from images and asking the model whether it can still see the removed objects.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations