Papers by Dasol Choi

2 papers
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: HAERAE-Vision benchmarks feature clear, explicit prompts but are often informal and underspecified . state-of-the-art models achieve under 50% on original queries, compared to GPT-5 and Gemini 2.5 Pro .
Approach: They propose a benchmark of 653 real-world visual questions from Korean online communities . they find that even state-of-the-art models achieve under 50% on original queries .
Outcome: HAERAE-Vision benchmarks from Korean online communities yield 1,306 query variants . state-of-the-art models achieve under 50% on original queries, compared with smaller models . authors show that query explicitation alone yields 8 to 22 point improvements .
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Large language models are being rapidly adopted across a wide range of domains, including healthcare, finance, and the public sector.
Approach: They propose a framework to evaluate whether large language models comply with policies . they apply COMPASS to eight diverse industry scenarios to validate models .
Outcome: The proposed framework evaluates whether LLMs comply with allowlist and denylist policies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations