Papers by Jianxun Lian
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds (2025.naacl-long)
Copied to clipboard
| Challenge: | Evaluating role-playing capabilities in large language models is challenging due to complex dynamics involved in role-playering. |
| Approach: | They propose a simulation sandbox that generates situational fine-grained character behavior trajectories to enhance LLM performance. |
| Outcome: | The proposed model generates situational fine-grained character behavior trajectories to enhance performance. |
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models lack information asymmetry with real-world situations. |
| Approach: | They propose a benchmark to evaluate the human-like motivational and behavioral reasoning ability of LLMs with detailed, realistic situations. |
| Outcome: | The proposed benchmark compared LLMs with real-world scenarios on seven model families and found that the most advanced models struggle with understanding "love & belonging" needs. |
TrendSim: Simulating Trending Topics in Social Media Under Poisoning Attacks with LLM-based Multi-agent System (2025.findings-naacl)
Copied to clipboard
Zeyu Zhang, Jianxun Lian, Chen Ma, Yaning Qu, Ye Luo, Lei Wang, Rui Li, Xu Chen, Yankai Lin, Le Wu, Xing Xie, Ji-Rong Wen
| Challenge: | Trending topics bring in a new channel for poisoning attacks, resulting in negative impacts on society. |
| Approach: | They propose an LLM-based multi-agent system to simulate trending topics in social media . they propose a time-aware interaction mechanism, centralized message dissemination, and an interactive system . |
| Outcome: | The proposed system simulates trending topics under poisoning attacks on social media platforms. |
Knowing-but-Doing: Diagnosing and Defending Role-Play-Driven LLMs Jailbreaks via Moral Disengagement (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used in role-play scenarios, but their safety implications remain under-characterized. |
| Approach: | They propose a diagnostic benchmark for role-play jailbreaks based on Bandura’s Moral Disengagement theory and propose 'MD-Trace' based defense that reduces attack success while maintaining Role Fidelity. |
| Outcome: | The proposed framework improves safety behavior for benign personas while increasing unsafe compliance for malicious ones. |
Eliminating Out-of-Domain Recommendations in LLM-based Recommender Systems: A Unified View (2026.findings-acl)
Copied to clipboard
Hao Liao, Jiwei Zhang, Jianxun Lian, Wensheng Lu, Mingqi Wu, null Shuowangg, Yong Zhang, Yitian Huang, Mingyang Zhou, Rui Mao
| Challenge: | Existing approaches to reduce OOD recommendations fall into three grounding paradigms: retrieval, constrained generation and discrete item tokenizer generation. |
| Approach: | They propose a framework that instantiates three grounding paradigms under a single architecture . embedding-based retrieval, constrained generation and discrete item-tokenizer methods are implemented . |
| Outcome: | The proposed framework eradicates OOD recommendations across all variants and achieves state-of-the-art accuracy compared to strong ID-based and LLM-based baselines. |
Pretraining Context Compressor for Large Language Models with Embedding-Based Memory (2025.acl-long)
Copied to clipboard
| Challenge: | Efficient processing of long contexts in large language models is essential for real-world applications such as retrieval-augmented generation and in-context learning. |
| Approach: | They propose a decoupled compressor-LLM framework that preserves contextual information within condensed embedding representations. |
| Outcome: | The proposed framework outperforms baseline models in three domains and across eight datasets while adapting to different downstream LLMs. |
SocialCC: Interactive Evaluation for Cultural Competence in Language Agents (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have evaluated cultural knowledge of large language models, but they fail to assess dynamic cultural competence. |
| Approach: | They propose a benchmark to assess cultural competence through intercultural scenarios that span 60 countries across six continents. |
| Outcome: | The proposed benchmark measures the ability of large language models to apply cultural knowledge effectively in cross-cultural interactions. |
PTUM: Pre-training User Model from Unlabeled User Behaviors via Self-supervision (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for user modeling cannot exploit useful information in unlabeled data . Existing models only model task-specific user information and do not exploit universal user information encoded in user behaviors. |
| Approach: | They propose to pre-train user models from large-scale unlabeled user behavior data. |
| Outcome: | The proposed method can model relatedness between historical and future behaviors on two real-world datasets. |
Can AI Revise Research Papers with Human Review Feedback? An Empirical Study and Benchmark (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are fundamentally reshaping the scientific landscape, transitioning the role of AI from passive tools to active partners within a new paradigm of Human-AI collaboration. |
| Approach: | They propose a benchmark to evaluate the ability of Large Language Models to improve papers with human feedback. |
| Outcome: | The proposed benchmark tests the skills of Large Language Models (LLMs) on paper interpretation, experimental implementation, and paper formulation, using authors’ camera-ready versions as natural human baselines. |
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing research focuses on character-level settings and static evaluation formats fail to capture the complexity of everyday social interactions. |
| Approach: | They propose a dynamic simulation framework for evaluating and improving persona-level role-playing in large language models (LLMs). |
| Outcome: | The proposed framework leverages user-generated social content to construct a nuanced persona bank and elicits multi-turn, context-rich interactions within simulated social environments. |
Reasoning-Aware AIGC Detection via Alignment and Reinforcement (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to AIGC detection have relied on statistical classifiers or black-box neural models, which exploit surface-level patterns and struggle to generalize as LLMs evolve. |
| Approach: | They propose a framework that generates interpretable reasoning chains before classification using supervised fine-tuning and reinforcement learning to improve accuracy. |
| Outcome: | The proposed framework achieves state-of-the-art performance across multiple benchmarks, offering a robust and transparent solution for AIGC detection. |
Hierarchical Reward Modeling for Fault Localization in Large Code Repositories (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have limited fault localization capabilities due to limited context length. |
| Approach: | They propose a hierarchical localization reward model to evaluate and select the most accurate fault localization candidates from the outputs of LLMs. |
| Outcome: | The proposed model improves the final line-level localization recall by 12% on the SWE-Bench-Lite dataset. |
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic nature of human-AI interactions. |
| Approach: | They propose a new paradigm of interactive ToM evaluation with both perspective and metric shifts. |
| Outcome: | The proposed approach improves the performance of four representative LLM enhancement techniques using real-world datasets and a user study. |
Towards Better Entity Linking with Multi-View Enhanced Distillation (2023.acl-long)
Copied to clipboard
Yi Liu, Yuan Tian, Jianxun Lian, Xinlong Wang, Yanan Cao, Fang Fang, Wen Zhang, Haizhen Huang, Weiwei Deng, Qi Zhang
| Challenge: | Entity linking is a fundamental task in Natural Language Processing (NLP), connecting mentions within unstructured contexts to their corresponding entities in a Knowledge Base (KB). |
| Approach: | They propose a dual-encoder framework that can efficiently match mentions to two-encoding frameworks by a global-view. |
| Outcome: | The proposed framework achieves state-of-the-art on several entity linking benchmarks. |
Aligning Large Language Models for Controllable Recommendations (2024.acl-long)
Copied to clipboard
| Challenge: | Existing literature focuses on integrating domain-specific knowledge into LLMs to enhance accuracy using a fixed task template. |
| Approach: | They propose a collection of supervised learning tasks augmented with labels derived from a conventional recommender model to improve LLMs’ proficiency in adhering to recommendation-specific instructions. |
| Outcome: | The proposed approach significantly improves the capability of LLMs to respond to instructions within recommender systems, reducing formatting errors while maintaining a high level of accuracy. |
MIND: A Large-scale Dataset for News Recommendation (2020.acl-main)
Copied to clipboard
Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, Ming Zhou
| Challenge: | Personalized news recommendation is an important technique for personalized news service. |
| Approach: | They propose to build a large-scale news recommendation dataset from Microsoft News . they demonstrate that news recommendation relies on the quality of news content understanding . |
| Outcome: | The proposed dataset contains 1 million users and more than 160k English news articles, each of which has rich textual content such as title, abstract and body. |
Explaining Length Bias in LLM-Based Preference Evaluations (2025.findings-emnlp)
Copied to clipboard
Zhengyu Hu, Linxin Song, Jieyu Zhang, Zheyuan Xiao, Tianfu Wang, Zhengyu Chen, Nicholas Jing Yuan, Jianxun Lian, Kaize Ding, Hui Xiong
| Challenge: | a preference evaluation metric is often biased towards longer responses, revealing a reliability problem . a decomposition of the preference evaluation into two components is needed to understand this bias. |
| Approach: | They propose to decompose the preference evaluation metric into two key components . the first component is length-dependent and related to trustworthiness . |
| Outcome: | The proposed evaluation metric is based on two components: desirability and information mass. |