Papers by Fangwei Zhong
Simple Role Assignment is Extraordinarily Effective for Safety Alignment (2026.findings-acl)
Copied to clipboard
Zhou Ziheng, Jiakun Ding, Zhaowei Zhang, Ruosen Gao, Ying Nian Wu, Demetri Terzopoulos, Yipeng Kang, Fangwei Zhong, Junqi Wang
| Challenge: | a new study proposes a role-conditioned pipeline for value alignment . principles alone are incomplete, and they provide little guidance on when and how a value applies in context. |
| Approach: | They propose a role-conditioned pipeline with role-based critics and a model-free approach that is based on role conditioning. |
| Outcome: | The proposed approach outperforms principle-based, Chain-of-Thought and other benchmarks. |
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective (2025.findings-acl)
Copied to clipboard
Yipeng Kang, Junqi Wang, Yexin Li, Mengmeng Wang, Wenming Tu, Quansen Wang, Hengli Li, Tingjun Wu, Xue Feng, Fangwei Zhong, Zilong Zheng
| Challenge: | Current approaches to value alignment focus on a few core values, such as helpfulness, harmlessness, and honesty. |
| Approach: | They propose to use latent causal value graphs to guide two lightweight value-steering methods . role-based prompting and sparse autoencoder (SAE) steering are also used . |
| Outcome: | Experiments on Gemma-2B-IT and Llama3-8B- IT show that the proposed methods are effective and controllable. |
Communication-Efficient Desire Alignment for Proactive Embodied Human–Agent Interaction (2026.acl-long)
Copied to clipboard
| Challenge: | Effective real-world human–agent interactions are long-term and repeated. |
| Approach: | They propose a simulation that uses a proxy user with value-driven preferences and natural language behavior to evaluate how agents adapt to users across interactions and satisfy their desires. |
| Outcome: | HA-Desire, a home assistance simulation, shows that agents can adapt to user needs and provide proactive assistance within limited communication. |
How do Role Models Shape Collective Morality? Exemplar-Driven Moral Learning in Multi-Agent Simulation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that role models influence morality, but they are not uniformly interpreted and appropriated in groups with heterogeneous motivations. |
| Approach: | They build a multi-agent simulation where agents with diverse intrinsic drives interact and adapt through a four-stage cognitive loop. |
| Outcome: | The proposed model can significantly reshape morality of agents with diverse intrinsic drives . the simulations show that identity-driven conformity can substantially reshaped initial dispositions . |
Why Are We Moral? An LLM-based Agent Simulation Approach to the Study of Moral Evolution (2026.acl-long)
Copied to clipboard
Zhou Ziheng, Huacong Tang, Mingjie Bi, Wanying He, Fang Sun, Yizhou Sun, Ying Nian Wu, Demetri Terzopoulos, Yipeng Kang, Fangwei Zhong
| Challenge: | Existing models of moral evolution must abstract away cognitive processes . et al. (2017): evolution of morality presents a puzzle: natural selection favors selfish . |
| Approach: | They propose an LLM-based agent simulation framework that manipulates cognitive factors to understand moral evolution. |
| Outcome: | The proposed model exploits cognitive realism to explore moral evolution in a hunter-gatherer society. |
CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language Games (2025.acl-long)
Copied to clipboard
| Challenge: | Metaphors are crucial for humans to express complex or subtle ideas by comparing one concept to another, often from a different domain. |
| Approach: | They propose a framework that enables LLMs to engage in metaphor processing by combining hypothesis-based metaphor reasoner and metaphor generator. |
| Outcome: | The proposed framework enhances agents' ability to interpret and apply metaphors in language games. |