Papers by Fangwei Zhong

6 papers
Simple Role Assignment is Extraordinarily Effective for Safety Alignment (2026.findings-acl)

Copied to clipboard

Challenge: a new study proposes a role-conditioned pipeline for value alignment . principles alone are incomplete, and they provide little guidance on when and how a value applies in context.
Approach: They propose a role-conditioned pipeline with role-based critics and a model-free approach that is based on role conditioning.
Outcome: The proposed approach outperforms principle-based, Chain-of-Thought and other benchmarks.
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective (2025.findings-acl)

Copied to clipboard

Challenge: Current approaches to value alignment focus on a few core values, such as helpfulness, harmlessness, and honesty.
Approach: They propose to use latent causal value graphs to guide two lightweight value-steering methods . role-based prompting and sparse autoencoder (SAE) steering are also used .
Outcome: Experiments on Gemma-2B-IT and Llama3-8B- IT show that the proposed methods are effective and controllable.
Communication-Efficient Desire Alignment for Proactive Embodied Human–Agent Interaction (2026.acl-long)

Copied to clipboard

Challenge: Effective real-world human–agent interactions are long-term and repeated.
Approach: They propose a simulation that uses a proxy user with value-driven preferences and natural language behavior to evaluate how agents adapt to users across interactions and satisfy their desires.
Outcome: HA-Desire, a home assistance simulation, shows that agents can adapt to user needs and provide proactive assistance within limited communication.
How do Role Models Shape Collective Morality? Exemplar-Driven Moral Learning in Multi-Agent Simulation (2026.acl-long)

Copied to clipboard

Challenge: Existing studies show that role models influence morality, but they are not uniformly interpreted and appropriated in groups with heterogeneous motivations.
Approach: They build a multi-agent simulation where agents with diverse intrinsic drives interact and adapt through a four-stage cognitive loop.
Outcome: The proposed model can significantly reshape morality of agents with diverse intrinsic drives . the simulations show that identity-driven conformity can substantially reshaped initial dispositions .
Why Are We Moral? An LLM-based Agent Simulation Approach to the Study of Moral Evolution (2026.acl-long)

Copied to clipboard

Challenge: Existing models of moral evolution must abstract away cognitive processes . et al. (2017): evolution of morality presents a puzzle: natural selection favors selfish .
Approach: They propose an LLM-based agent simulation framework that manipulates cognitive factors to understand moral evolution.
Outcome: The proposed model exploits cognitive realism to explore moral evolution in a hunter-gatherer society.
CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language Games (2025.acl-long)

Copied to clipboard

Challenge: Metaphors are crucial for humans to express complex or subtle ideas by comparing one concept to another, often from a different domain.
Approach: They propose a framework that enables LLMs to engage in metaphor processing by combining hypothesis-based metaphor reasoner and metaphor generator.
Outcome: The proposed framework enhances agents' ability to interpret and apply metaphors in language games.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations