Papers by Yunze Xiao

13 papers
JiraiBench: A Bilingual Benchmark for Evaluating Large Language Models’ Detection of Human risky health behavior Content in Jirai Community (2026.eacl-long)

Copied to clipboard

Challenge: a cross-lingual dataset captures a transnational cultural phenomenon . risky health behaviors (RHB) are often linked to complex mental health conditions .
Approach: They present the first cross-lingual dataset that captures a transnational cultural phenomenon . their dataset of more than 15,000 annotated social media posts forms the core of JiraiBench .
Outcome: The study shows that cultural context can be more influential than linguistic similarity . the study also shows that the Japanese prompts better handle Chinese content .
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents (2026.acl-long)

Copied to clipboard

Challenge: a fundamental pillar of trustworthiness is calibration, which refers to an agent’s ability to express confidence that reliably reflects its actual performance.
Approach: They propose a reinforcement learning framework that jointly optimizes task accuracy and calibration, supported by a holistic benchmark of reward designs.
Outcome: The proposed framework improves calibration across tool types and shows that trained agents achieve superior calibration and exhibit robust generalization from local training environments to noisy web settings and to distinct domains such as mathematical reasoning.
Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens (2026.findings-eacl)

Copied to clipboard

Challenge: anthropological accounts of culture often focus on static facts or homogeneous values . large language models are being implemented in translation systems, educational tools and search engines .
Approach: They propose to categorize how benchmarks frame culture such as knowledge, preference, performance, or bias.
Outcome: The proposed framework categorizes how benchmarks frame culture, such as knowledge, preference, performance, or bias.
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs (2024.lrec-main)

Copied to clipboard

Challenge: Lexical-syntactic flexibility is a hallmark of English morphology . conversion involves placing a word with one part of speech in a non-prototypical context .
Approach: They propose to test lexical-syntactic flexibility in the form of conversion . conversion is a process where a word with one part of speech is placed in a non-prototypical context .
Outcome: The proposed task tests the ability of five language models to generalize over words with a non-prototypical part of speech.
Sentipolis: Emotion-Aware Agents for Social Simulations (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in reasoning and long-context memory are making large language models (LLMs) appear increasingly human-like, which has led researchers to adopt LLM agents as a substrate for social simulation.
Approach: They propose a framework for emotionally stateful agents that integrates continuous Pleasure-Arousal-Dominance representation, dual-speed emotion dynamics, and emotion–memory coupling.
Outcome: The proposed framework improves emotional grounded behavior, boosting communication, and emotional continuity across thousands of interactions over multiple base models and evaluators.
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics (2025.emnlp-main)

Copied to clipboard

Challenge: a study of multi-dimensional persona effects in AI-AI debates shows that personas influence moral stances and debate outcomes . political ideology and personality traits exert the strongest influence, according to our study .
Approach: They propose to use a 6-dimensional persona space to simulate structured debates . they find political ideology and personality traits exert the strongest influence .
Outcome: The study shows that personas affect moral stances and debate outcomes . political ideology and personality traits exert the strongest influence .
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing large language models struggle with systematically perturbed data designed to evade detection mechanisms.
Approach: They propose a large language model with homophonic substitutions and emoji transformations to test their models' robustness against cloaking perturbations.
Outcome: The proposed model underperforms in detecting offensive content when perturbations are applied to Chinese language datasets.
InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews (2024.acl-long)

Copied to clipboard

Challenge: Existing methods focus on knowledge and linguistic patterns of characters.
Approach: They propose to evaluate character fidelity of role-playing agents with psychological scales . they propose to use psychological scale to measure personality traits of RPAs based on personality traits.
Outcome: The proposed model reproduces character fidelity with psychological scales and shows that it is effective in measuring personality traits.
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language model evaluation benchmarks focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities.
Approach: They propose a comprehensive benchmark covering 29 languages, built on an English benchmark.
Outcome: The MMLU-ProX is a comprehensive benchmark covering 29 languages, built on an English benchmark.
TartanMaroon: Multi-Agent Academic Advising with Iterative Negotiation and Transparent Collaboration (2026.acl-demo)

Copied to clipboard

Challenge: Academic advising is a critical yet resourceintensive component of higher education . monolithic model must simultaneously maintain awareness of heterogeneous institutional constraints .
Approach: They propose a multi-agent academic advising system that handles the full complexity spectrum of student queries.
Outcome: The proposed system handles the full complexity spectrum of student queries . it also provides a real-time transparency interface streaming agent reasoning and negotiation rounds to users . the system is released open-source and has been rated highly by users based on their results .
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction (2026.findings-acl)

Copied to clipboard

Challenge: Accurate estimation of item (question or task) difficulty suffers from the cold start problem.
Approach: They propose to use large-scale empirical analysis to examine human-AI Difficulty Alignment . they find that models struggle to simulate the capability limitations of students .
Outcome: The proposed model size is not reliably helpful for human-AI alignment . high performance often impedes accurate difficulty estimation, the authors say .
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on single agentic capability, failing to capture long-horizon real-world scenarios.
Approach: They propose a benchmark that evaluates 6 agentic capabilities across 32 real-world scenarios.
Outcome: Experiments show that closed-source models outperform open-source model (48.4% vs 32.1%) integrating models with advanced scaffolds to form autonomous agents is a paradigm shift.
Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit anthropomorphism characteristics – human-like qualities portrayed across their outlook, language, behavior, and reasoning functions.
Approach: They propose that anthropomorphism should be treated as a design concept that can be intentionally tuned to support user goals.
Outcome: The proposed design should reflect interaction between artifact designers and interpreters, and should be based on cues embedded in the artifactor and the (cognitive) responses of interpreters to the cue.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations