Papers by Pengyuan Liu
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are being used in urban planning but there is concern that they reproduce or amplify such biases. |
| Approach: | They propose a framework to evaluate spatial gender bias in large language models . they use a taxonomy of 62 urban micro-spaces, a prompt library and three diagnostic layers . |
| Outcome: | The proposed framework identifies structured gender-space associations that go beyond the public-private divide, forming nuanced micro-level mappings. |
Investigating Value-Reasoning Reliability in Small Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | sLLMs have been widely deployed in practical applications, but little attention has been paid to their value-reasoning abilities, particularly in terms of reasoning reliability. |
| Approach: | They propose a systematic evaluation framework for assessing the Value-Reasoning Reliability of small Large Language models (sLLMs) . framework includes three core tasks: Repetition Consistency task, Interaction Stability task, and Open-ended Expression Consistencies task. |
| Outcome: | The proposed framework incorporates self-reported confidence scores to evaluate the model’s value reasoning reliability from two perspectives: the model's self awareness of its values, and its value-based decision-making. |
From Polarity to Intensity: Mining Morality from Semantic Space (2022.coling-1)
Copied to clipboard
| Challenge: | Existing approaches to compute moral intensity are limited to word-level measurement and heavily rely on human labelling. |
| Approach: | They propose a weakly-supervised framework that can automatically measure moral intensity from text. |
| Outcome: | The proposed framework can measure moral intensity from text with moral polarity labels, which are more robust and easier to acquire. |
MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization (2024.findings-acl)
Copied to clipboard
Zhiyu Yang, Zihan Zhou, Shuo Wang, Xin Cong, Xu Han, Yukun Yan, Zhenghao Liu, Zhixing Tan, Pengyuan Liu, Dong Yu, Zhiyuan Liu, Xiaodong Shi, Maosong Sun
| Challenge: | Scientific data visualization is an essential process in research, but its use of large language models remains unexplored. |
| Approach: | They propose a model-agnostic LLM agent framework to automate scientific data visualization tasks. |
| Outcome: | The proposed framework improves performance of commercial and open-source models. |
Beyond Detection: Evaluating Fallacy Awareness of LLMs in Interactive Scenarios (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models fail to recognize fallacious reasoning in real-world interactions despite strong performance on static fallacy detection tasks. |
| Approach: | They propose a Chinese benchmark to assess fallacy awareness without explicit cues . they propose 'fate' evaluation framework that assesses fallacy without explicit . |
| Outcome: | The proposed framework assesses fallacy awareness without explicit cues, combining natural dialogue responses and reasoning-based decisions. |
Attribution and Application of Multiple Neurons in Multimodal Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to identify multimodal neurons in MLLMs are insufficiently understood . previous studies focused on identifying neurons corresponding to single-tokens . |
| Approach: | They propose a method to identify multimodal neurons in Transformer-based MLLMs . they introduce fuzzy set theory to model the complex relationship between neurons and semantic concepts . |
| Outcome: | The proposed method improves performance on the Visual Question Answering task. |
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Proper moral beliefs are fundamental for language models, yet assessing these beliefs poses a significant challenge. |
| Approach: | They propose a framework to evaluate the moral beliefs of four large language models . they use a dataset containing 472 moral choice scenarios in Chinese . |
| Outcome: | The proposed framework evaluates the moral beliefs of four large language models. |
What’s the most important value? INVP: INvestigating the Value Priorities of LLMs through Decision-making in Social Scenarios (2025.coling-main)
Copied to clipboard
| Challenge: | Large scale language models (LLMs) have demonstrated impressive performance in various tasks and are increasingly integrated into the decision-making process. |
| Approach: | They propose a framework for INvestigating Value Priorities through decision-making in social scenarios and evaluate seven popular LLMs. |
| Outcome: | The proposed framework covers 1613 scenarios and 3226 decisions across 283 topics and focuses on Universalism and Benevolence, while Power and Hedonism are given lower priority. |
CLGC: A Corpus for Chinese Literary Grace Evaluation (2022.lrec-1)
Copied to clipboard
| Challenge: | Literature grace is a key element of the style and quality of articles in China. |
| Approach: | They propose to annotate a Chinese literary grace corpus with 10,000 texts and 1.85 million tokens and build a literary grace evaluation task to assess the literary grace level. |
| Outcome: | The proposed model achieves 79.71% on the weighted average F1-score. |