Papers by Siyu Zhao
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios (2025.emnlp-main)
Copied to clipboard
| Challenge: | a number of tools are used to perform complex tasks, but the tool utilization process can cause errors. |
| Approach: | They propose a critique evaluation benchmark for tool learning that analyzes function-calling errors on tool evaluation benchmarks. |
| Outcome: | The proposed critique evaluation benchmark holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios. |
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing RAG frameworks rely on Automatic Speech Recognition to process speech input, which discards crucial audio information and increases computational overhead. |
| Approach: | They propose a retrieval augmented generation framework with native, end-to-end audio support that integrates audio and text into a unified knowledge representation. |
| Outcome: | The proposed framework can perform 10x faster than current pipelines while delivering 10x acceleration. |
Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing Large Language Models (LLMs) mainly address isolated tasks such as emotion analysis or stance detection. |
| Approach: | They propose a large-scale model that combines large-level annotations with hyperbolic space to model human cognitive states. |
| Outcome: | The proposed model outperforms baseline models on cognitive dimensions on single dimension tasks while retaining strong hierarchical structure. |
Learning from Adjective-Noun Pairs: A Knowledge-enhanced Framework for Target-Oriented Multimodal Sentiment Classification (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods to determine sentiment polarity of opinion target are inconsistent and lack visual attention. |
| Approach: | They propose a framework which can exploit adjective-noun pairs extracted from images to improve visual attention and sentiment prediction capability of the TMSC task. |
| Outcome: | The proposed framework outperforms state-of-the-art on two public datasets. |
TARE: Lightweight Token-Aware Representation Editing for Fine-tuning Transformer-like Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing PEFT methods can be costly and underfit token-level contexts. |
| Approach: | They propose a PEFT method that performs fine-grained, token-specific edits with a small additional inference overhead and minimal tuning. |
| Outcome: | The proposed method outperforms state-of-the-art methods in 8 tasks and GLUE with a minimal tuning overhead and inference overhead. |
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies attributed verbosity to biased labels, but new research shows that DPO can be effective in mitigating verboses. |
| Approach: | They propose to use a method to reduce the amount of verbosity in LLMs by using a downsampling approach. |
| Outcome: | The proposed approach overcomes the problem of verbosity by reducing the length reliance of the proposed algorithm. |
T⋆: Progressive Block Scaling for Masked Diffusion Language Models Through Trajectory Aware Reinforcement Learning (2026.acl-short)
Copied to clipboard
| Challenge: | Autoregressive (AR) modeling via next-token prediction dominates scaling practice and deployed systems. |
| Approach: | They propose a TraceRL-based curriculum for progressive block-size scaling in masked diffusion language models. |
| Outcome: | The proposed curriculum outperforms direct large-block TraceRL on two SDAR scales and three benchmarks and retains block-size-specific non-monotone updates while improving accuracy. |