Papers by Heeseung Yun
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing game benchmarks lack diversity and evaluate GUI agents on completing entire storylines. |
| Approach: | They propose a benchmark of 34 Flash-based adventure games to test full story arc completion and tackle observation-behavior gap. |
| Outcome: | The proposed benchmarks show GUI agents struggle with full story arc completion while others improve on observation-behavior gaps. |
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations (2026.findings-acl)
Copied to clipboard
| Challenge: | Large audio-language models extend language understanding into the auditory domain, yet their ability to perform low-level listening, such as pitch and duration detection, remains underexplored. |
| Approach: | They propose a global benchmark to evaluate low-level auditory perception and cognition using marine mammal vocalizations to better assess models’ low- level listening. |
| Outcome: | The proposed models show performance far below human levels, indicating a need for stronger auditory grounding in LALMs. |
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in multimodal systems have demonstrated remarkable capabilities in generating multimodal content from multimodal inputs. |
| Approach: | They propose a benchmark that leverages large language models to generate deceptive text samples to exploit compositional vulnerabilities across different modalities. |
| Outcome: | The proposed approach exploits compositional vulnerabilities across images, videos, and audios. |