Papers by Yun-Shiuan Chuang

5 papers
Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding (2025.emnlp-main)

Copied to clipboard

Challenge: a common real-world skill of guesstimation is underexplored in large language model research . a recent study suggests that LLMs encode a world model that supports approximate reasoning .
Approach: They propose to decode a guesstimation dataset using MARBLES, FUTURE, and ELECPRED . they replicate WOC effects in human participants and find similar benefits .
Outcome: The proposed model improves accuracy over greedy, self-consistency, and mean decoding in human participants.
Beyond Demographics: Aligning Role-playing LLM-based Agents Using Human Belief Networks (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing large language models can be prompted to role-play as individuals with particular demographic traits, but results are often human-like.
Approach: They found that seeding LLM-based agents with a single belief improved alignment . they say that role-playing based on demographic information does not improve alignment a .
Outcome: The proposed approach improves LLM alignment with human behavior . seeding agents with a single belief improves alignment for topics related to the belief network .
Ada-RS: Adaptive Rejection Sampling for Selective Thinking (2026.acl-industry)

Copied to clipboard

Challenge: Large language models are increasingly being deployed in cost- and latency-sensitive settings . chain-of-thought improves reasoning, but it can waste tokens on simple requests .
Approach: They introduce an algorithm-agnostic sample filtering framework for learning selective reasoning . they show that Ada-RS reduces average output tokens by 80% and reducing thinking rate by 5% .
Outcome: The proposed framework reduces output tokens by 80% and thinking rate by 95% on a synthetic tool call-oriented e-commerce benchmark.
Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents (2026.acl-industry)

Copied to clipboard

Challenge: Existing agentic benchmarks rely on deterministic backends and are costly to build and iterate.
Approach: They propose a framework that preserves final state-based evaluation without a deterministic database.
Outcome: The proposed framework produces stable, model-differentiating rankings across families and inference-time reasoning efforts.
Simulating Opinion Dynamics with Networks of LLM-based Agents (2024.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to simulating opinion dynamics often over-simplify human behavior . authors propose refining LLMs with real-world discourse to better simulate evolution of beliefs .
Approach: They propose to use large language models to simulate opinion dynamics in groups of simulated agents . they found that LLM agents produce more accurate information than ABMs .
Outcome: The proposed model can be used to better simulate opinion dynamics in real-world discourses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations