Papers by Yiqun Yao
SelfRACG: Enabling LLMs to Self-Express and Retrieve for Code Generation (2025.emnlp-main)
Copied to clipboard
Qian Dong, Jia Chen, Qingyao Ai, Hongning Wang, Haitao Li, null Yiwu, Yao Hu, Yiqun Liu, Shaoping Ma
| Challenge: | Existing retrieval-augmented code generation methods fail to accurately fetch the knowledge required for code generation for consecutive code fragments. |
| Approach: | They propose a paradigm that enables large language models to Self-express their information needs to enhance retrieval-augmented code generation methods. |
| Outcome: | Experiments show that SelfRACG can retrieve external knowledge that better aligns with the LLM’s own information needs, resulting in superior generation performance compared to vanilla RACG. |
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs (2026.acl-long)
Copied to clipboard
Siqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang, Kang Liu, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang
| Challenge: | Existing benchmarks for large language models (LLMs) fail to capture these dynamics, focusing on static, open-ended evaluations. |
| Approach: | They propose a benchmark to assess lifelong learning in large language models . they use two episodic datasets rich in narrative structure and character interactions . |
| Outcome: | Experiments on LLMs show that non-parametric methods outperform parametric ones in managing stateful learning. |
The World in My Mind: Visual Dialog with Adversarial Multi-modal Feature Encoding (N19-1)
Copied to clipboard
| Challenge: | Visual Dialog is a multi-modal task that requires a model to participate in a dialog grounded on an image and generate correct, human-like responses. |
| Approach: | They propose a framework for effective and robust auxiliary training of visual dialog systems using multi-modal encoding. |
| Outcome: | The proposed framework outperforms supervised learning baselines and fine-tuning methods on most metrics of VisDial v0.5/v0.9 generative tasks. |
MUSER: MUltimodal Stress detection using Emotion Recognition as an Auxiliary Task (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to detect stress have not explored the inter-dependence between emotion and stress. |
| Approach: | They propose a transformer-based model architecture and a novel multi-task learning algorithm with speed-based dynamic sampling strategy to improve stress detection. |
| Outcome: | The proposed model is effective with internal and external auxiliary tasks and achieves state-of-the-art results. |
Modality-specific Learning Rates for Effective Multimodal Additive Late-fusion (2022.findings-acl)
Copied to clipboard
| Challenge: | Multimodal machine learning uses additive late-fusion to combine feature representations from different modalities into a joint representation. |
| Approach: | They propose a Modality-Specific Learning Rate method to build late-fusion multimodal models from fine-tuned unimodal models. |
| Outcome: | The proposed method outperforms global learning rates on multiple tasks and settings and enables the models to effectively learn each modality. |
Cascaded Mutual Modulation for Visual Reasoning (D18-1)
Copied to clipboard
| Challenge: | Visual reasoning is a multi-step and compositional problem that requires intensive text-vision interactions. |
| Approach: | They propose a visual reasoning model that uses a feature-wise linear modulation technique to enable textual/visual pipelines to mutually control each other. |
| Outcome: | The proposed model outperforms existing models on visual reasoning benchmarks CLEVR and NLVR . it can generate a textual answer to a visual question answering problem with images . |
Prompt Refinement with Image Pivot for Text-to-Image Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-image generation have markedly expanded the boundaries of digital artistry, enabling the creation of visually compelling images with unprecedented ease. |
| Approach: | They propose to decompose the prompt refinement process into two tasks: inferring user-preferred images from user languages and translating them into system languages. |
| Outcome: | Experiments show that PRIP outperforms baselines and transfers to unseen systems in a zero-shot manner. |