Papers by Yong Cao
Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Introducing **MARK**, a framework for cultural value survey simulation . based on type dynamics theory, it improves accuracy and interpretation of models . |
| Approach: | They propose a framework that integrates psychological theory into cultural value survey simulations. |
| Outcome: | The proposed framework outperforms baseline models on the World Values Survey by 10% accuracy and reduces divergence between model predictions and human preferences. |
AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation (2026.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Approach: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Outcome: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge (2025.naacl-long)
Copied to clipboard
Li Zhou, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao, Wenyu Chen, Haizhou Li, Daniel Hershcovich
| Challenge: | Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet lack a robust methodology to dissect these phenomena comprehensively. |
| Approach: | They propose a multilingual dataset centered on food-related cultural facts and variations in food practices. |
| Outcome: | The proposed model incorporates cultural context significantly and improves its ability to access cultural knowledge. |
An Empirical Comparison of Unsupervised Constituency Parsing Methods (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for unsupervised constituency parsing are inconsistent due to data preprocessing, lexicalization, and evaluation metrics. |
| Approach: | They propose to standardize experimental settings for better comparability between methods . they compare existing methods with those proposed by decade-old models . |
| Outcome: | The proposed methods perform better than decade-old models on English and Japanese, respectively, compared with decade- old models. |
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation (2026.eacl-long)
Copied to clipboard
Bolei Ma, Yong Cao, Indira Sen, Anna-Carolina Haensch, Frauke Kreuter, Barbara Plank, Daniel Hershcovich
| Challenge: | Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena. |
| Approach: | They argue that open-endedness is essential for realistic social simulations . they argue that it captures expressiveness and individuality . |
| Outcome: | The proposed frameworks can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias. |
Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples! (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models have made remarkable strides in various tasks, but whether they are competitive few-shot solvers remains an open question. |
| Approach: | They propose an adaptive filter-then-rerank paradigm to combine the strengths of LLMs and SLMs. |
| Outcome: | The proposed system achieves promising improvements on various IE tasks with acceptable time and cost investment. |
Geo-Encoder: A Chunk-Argument Bi-Encoder Framework for Chinese Geographic Re-Ranking (2024.eacl-long)
Copied to clipboard
| Challenge: | Chinese geographic re-ranking task aims to find the most relevant addresses among retrieved candidates. |
| Approach: | They propose a framework to integrate Chinese geographic semantics into re-ranking pipelines. |
| Outcome: | The proposed framework improves on two Chinese benchmark datasets. |
Explore More Guidance: A Task-aware Instruction Network for Sign Language Translation Enhanced with Data Augmentation (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies focus on the recognition step, while paying less attention to sign language translation. |
| Approach: | They propose a task-aware instruction network, namely TIN-SLT, for sign language translation, by introducing the isntruction module and the learning-based feature fuse strategy into a Transformer network. |
| Outcome: | The proposed system outperforms existing solutions on two benchmark datasets, PHOENIX-2014-T and ASLG-PC12, and outperformed previous best solutions by 1.65 and 1.42 in terms of BLEU-4. |
Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for training large language models require additional annotations to adjust to shifted distributions. |
| Approach: | They propose an algorithm that allows LLMs and reward models to update alternatively via a min-max game to improve their alignment. |
| Outcome: | The proposed framework improves existing alignment baselines in terms of LLM helpfulness and harmlessness. |
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Autoregressive (AR) large audio language models are expensive in data and computation . prior work shows diffusion-based LALMs can improve audio understanding under matched settings . |
| Approach: | They propose a diffusion-based LALM that upgrades the speech encoder and employs dual semantic and acoustic adapters. |
| Outcome: | a new model improves over existing autoregressive large language models and is competitive to strong AR models . the proposed model can make use of limited training data and improve inference efficiency . a recent study shows that diffusion-based models can improve audio understanding . |
Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine Translation (2022.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to train multilingual models to learn the inductive bias of a shared vocabulary and set of parameters across languages. |
| Approach: | They propose to use a multilingual crossover encoder-decoder to fuse language pairs at an instance level to encourage sharing of input and output spaces. |
| Outcome: | The proposed approach improves quality on English-to-Many, Many-to English and zero-shot translation tasks from +0.5 BLEU up to +5.5 BLUE points. |
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents (2025.acl-long)
Copied to clipboard
Zhigen Li, Jianxiang Peng, Yanmeng Wang, Yong Cao, Tianhao Shen, Minghui Zhang, Linxi Su, Shang Wu, Yihang Wu, YuQian Wang, Ye Wang, Wei Hu, Jianfeng Li, Shaojun Wang, Jing Xiao, Deyi Xiong
| Challenge: | Existing models that use Large Language Models (LLMs) show superior performance in various tasks, but lack of controllability leads to unfocused conversations or task failure. |
| Approach: | They propose a standard operating procedure (SOP) framework to regulate dialogue flow by integrating Chain of Thought reasoning and supervised fine-tuning for SOP prediction. |
| Outcome: | The proposed method achieves a 27.95% improvement in action accuracy compared to baseline models based on GPT-3.5 and also shows notable gains for open-source models. |
Attribution-Based Analysis and Optimization of Modular Agentic Workflows (2026.findings-acl)
Copied to clipboard
Yingxuan Yang, Bo Huang, Siyuan Qi, Chao Feng, Haoyi Hu, Yuxuan Zhu, Jinbo Hu, Haoran Zhao, Ziyi He, Xiao Liu, ZongYu Wang, Muning Wen, Lin Qiu, Xuezhi Cao, Xunliang Cai, Yong Yu, Weinan Zhang
| Challenge: | Large Language Models (LLMs) have driven the rise of agentic workflows . yet, how can we attribute performance gains to individual upgrades and their interactions? |
| Approach: | They propose a game-theoretic framework that models component upgrades as players and evaluates component coalitions to compute Shapley values. |
| Outcome: | The proposed framework provides interaction-aware attribution and recommendation for model allocation under a fixed workflow structure. |
Pay More Attention to Relation Exploration for Knowledge Base Question Answering (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches focus on entity representation and final answer reasoning, which results in limited supervision for this task. |
| Approach: | They propose a framework that utilizes relations to enhance entity representation and introduce additional supervision. |
| Outcome: | The proposed framework improves the F1 score on two benchmark datasets by 5.8% . it improves by 6.7% on WebQSP, better than state-of-the-art methods . |
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations (2025.naacl-long)
Copied to clipboard
| Challenge: | Prior work has focused on using large language models to simulate human behaviors . but, LLMs are known to generate erroneous, stereotypical, or overconfident answers . |
| Approach: | They propose to specialize large language models for simulating survey response distributions by first-token probabilities. |
| Outcome: | The proposed model outperforms other methods and zero-shot classifiers on unseen questions, countries, and a completely unseened survey. |
A + B: A General Generator-Reader Framework for Optimizing LLMs to Unleash Synergy Potential (2024.findings-acl)
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) is an effective solution to supplement necessary knowledge to large language models. |
| Approach: | They propose a "generate-then-read" pipeline to replace retrieval stage with generation from the LLM itself. |
| Outcome: | The proposed framework outperforms single models in the base and chat versions and addresses safety and helpfulness post-adaptation challenges. |
RealVul: Can We Detect Vulnerabilities in Web Applications with LLM? (2024.emnlp-main)
Copied to clipboard
| Challenge: | a lack of research specifically focused on vulnerabilities in the PHP language hinders the model’s ability to effectively capture the characteristics of specific vulnerabilities. |
| Approach: | They propose a framework that can isolate potential vulnerability triggers while streamlining code and eliminating unnecessary semantic information. |
| Outcome: | The proposed framework can isolate potential vulnerability triggers while streamlining the code and eliminating unnecessary semantic information. |
Scholar Inbox: Personalized Paper Recommendations for Scientists (2025.acl-demo)
Copied to clipboard
Markus Flicke, Glenn Angrabeit, Madhav Iyengar, Vitalii Protsenko, Illia Shakun, Jovan Cicvaric, Bora Kargi, Haoyu He, Lukas Schuler, Lewin Scholz, Kavyanjali Agnihotri, Yong Cao, Andreas Geiger
| Challenge: | Scholar Inbox is an open-access platform designed to address the challenges researchers face in staying current with the rapidly expanding volume of scientific publications. |
| Approach: | They propose to provide personalized recommendations, continuous updates from open-access archives, visual paper summaries, semantic search, and a range of tools to streamline research workflows and promote open access publications. |
| Outcome: | The proposed platform is based on a dataset of 800k user ratings and an extensive user study. |
Dependency Parsing via Sequence Generation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for dependency parsing are transition-based, graph-based and sequence-to-sequence method. |
| Approach: | They propose to achieve dependency parsing (DP) via Sequence Generation (SG) by utilizing only the pre-trained language model without any auxiliary structures. |
| Outcome: | The proposed method performs well on DP benchmarks including PTB, UD2.2, SDP15 and SemEval16. |
Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys (2024.findings-eacl)
Copied to clipboard
| Challenge: | integrating cultural dimensions with dialogue encoding features can enhance the predictive accuracy and quality of dialogue agents. |
| Approach: | They propose to incorporate cultural dimensions into dialogue encoding features to enhance the predictive accuracy of dialogue agents. |
| Outcome: | The proposed model improves the accuracy and quality of dialogue predictions by incorporating cultural dimensions with dialogue encoding features. |
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks (2026.acl-long)
Copied to clipboard
Lingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan, Yaoming Zhu, Lin Qiu, ZongYu Wang, Xuezhi Cao, Xunliang Cai, Weiwen Liu, Weinan Zhang, Yong Yu
| Challenge: | Existing large language models for software engineering rely on coarse-grained pass rates obscuring specific cognitive bottlenecks. |
| Approach: | They propose a repository-level benchmark that dissects coding capabilities through atomized tasks. |
| Outcome: | The proposed framework achieves a 78.55% validity yield, surpassing the 31.7% retention rate of SWE-bench-Verified. |
EvoWiki: Evaluating LLMs on Evolving Knowledge (2025.acl-long)
Copied to clipboard
Wei Tang, Yixin Cao, Yang Deng, Jiahao Ying, Bo Wang, Yizhe Yang, Yuyue Zhao, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Yong Liao
| Challenge: | Existing knowledge evolution benchmarks are static and fail to capture the evolving nature of LLMs and knowledge. |
| Approach: | They propose an evolving dataset that categorizes information into stable, evolved, and uncharted states. |
| Outcome: | The proposed dataset is auto-updatable and enables evaluation of continuously changing knowledge and newly released LLMs. |