Papers by Yun-Shiuan Chuang
Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding (2025.emnlp-main)
Copied to clipboard
Yun-Shiuan Chuang, Sameer Narendran, Nikunj Harlalka, Alexander Cheung, Sizhe Gao, Siddharth Suresh, Junjie Hu, Timothy T. Rogers
| Challenge: | a common real-world skill of guesstimation is underexplored in large language model research . a recent study suggests that LLMs encode a world model that supports approximate reasoning . |
| Approach: | They propose to decode a guesstimation dataset using MARBLES, FUTURE, and ELECPRED . they replicate WOC effects in human participants and find similar benefits . |
| Outcome: | The proposed model improves accuracy over greedy, self-consistency, and mean decoding in human participants. |
Beyond Demographics: Aligning Role-playing LLM-based Agents Using Human Belief Networks (2024.findings-emnlp)
Copied to clipboard
Yun-Shiuan Chuang, Krirk Nirunwiroj, Zach Studdiford, Agam Goyal, Vincent Frigo, Sijia Yang, Dhavan Shah, Junjie Hu, Timothy Rogers
| Challenge: | Existing large language models can be prompted to role-play as individuals with particular demographic traits, but results are often human-like. |
| Approach: | They found that seeding LLM-based agents with a single belief improved alignment . they say that role-playing based on demographic information does not improve alignment a . |
| Outcome: | The proposed approach improves LLM alignment with human behavior . seeding agents with a single belief improves alignment for topics related to the belief network . |
Ada-RS: Adaptive Rejection Sampling for Selective Thinking (2026.acl-industry)
Copied to clipboard
Yirou Ge, Yixi Li, Alec M. Chiu, Shivani Shekhar, Zijie Pan, Avinash Thangali, Yun-Shiuan Chuang, Chaitanya Kulkarni, Uma Kona, Linsey Pang, Prakhar Mehrotra
| Challenge: | Large language models are increasingly being deployed in cost- and latency-sensitive settings . chain-of-thought improves reasoning, but it can waste tokens on simple requests . |
| Approach: | They introduce an algorithm-agnostic sample filtering framework for learning selective reasoning . they show that Ada-RS reduces average output tokens by 80% and reducing thinking rate by 5% . |
| Outcome: | The proposed framework reduces output tokens by 80% and thinking rate by 95% on a synthetic tool call-oriented e-commerce benchmark. |
Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents (2026.acl-industry)
Copied to clipboard
Yun-Shiuan Chuang, Chaitanya Kulkarni, Alec M. Chiu, Avinash Thangali, Zijie Pan, Shivani Shekhar, Yirou Ge, Yixi Li, Uma Kona, Linsey Pang, Prakhar Mehrotra
| Challenge: | Existing agentic benchmarks rely on deterministic backends and are costly to build and iterate. |
| Approach: | They propose a framework that preserves final state-based evaluation without a deterministic database. |
| Outcome: | The proposed framework produces stable, model-differentiating rankings across families and inference-time reasoning efforts. |
Simulating Opinion Dynamics with Networks of LLM-based Agents (2024.findings-naacl)
Copied to clipboard
Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, Timothy Rogers
| Challenge: | Existing approaches to simulating opinion dynamics often over-simplify human behavior . authors propose refining LLMs with real-world discourse to better simulate evolution of beliefs . |
| Approach: | They propose to use large language models to simulate opinion dynamics in groups of simulated agents . they found that LLM agents produce more accurate information than ABMs . |
| Outcome: | The proposed model can be used to better simulate opinion dynamics in real-world discourses. |