Papers by Heeseung Yun

3 papers
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games (2025.emnlp-main)

Copied to clipboard

Challenge: Existing game benchmarks lack diversity and evaluate GUI agents on completing entire storylines.
Approach: They propose a benchmark of 34 Flash-based adventure games to test full story arc completion and tackle observation-behavior gap.
Outcome: The proposed benchmarks show GUI agents struggle with full story arc completion while others improve on observation-behavior gaps.
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations (2026.findings-acl)

Copied to clipboard

Challenge: Large audio-language models extend language understanding into the auditory domain, yet their ability to perform low-level listening, such as pitch and duration detection, remains underexplored.
Approach: They propose a global benchmark to evaluate low-level auditory perception and cognition using marine mammal vocalizations to better assess models’ low- level listening.
Outcome: The proposed models show performance far below human levels, indicating a need for stronger auditory grounding in LALMs.
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in multimodal systems have demonstrated remarkable capabilities in generating multimodal content from multimodal inputs.
Approach: They propose a benchmark that leverages large language models to generate deceptive text samples to exploit compositional vulnerabilities across different modalities.
Outcome: The proposed approach exploits compositional vulnerabilities across images, videos, and audios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations