Papers by Kunat Pipatanakul

4 papers
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation (2026.eacl-long)

Copied to clipboard

Challenge: Current speech evaluation systems rely on specialized systems for individual audio characteristics and poor correlation between automatic methods and human preferences.
Approach: They propose a unified evaluation framework for Large Audio Models as a Judge, AudioJudge . they propose specialized judges that can be prompted to perform audio characteristic detection tasks .
Outcome: The proposed method improves performance across audio characteristic detection and human preference simulation tasks.
Prior Prompt Engineering for Reinforcement Fine-Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on algorithms, reward shaping, and data curation, but prior prompt engineering is understudied.
Approach: They investigate prior prompt engineering (pPE) in reinforcement fine-tuning . they translate five representative iPE strategies into corresponding pPE approaches .
Outcome: The proposed approaches outperform iPE-prompted models on in-domain and out-of-domain benchmarks.
Mind the Gap: Static and Interactive Evaluations of Large Audio Models (2025.acl-long)

Copied to clipboard

Challenge: Recent work has focused on evaluating large audio models (LAMs) that directly accept audio inputs.
Approach: They propose an interactive approach to evaluate large audio models and collect 7,500 LAM interactions from 484 participants.
Outcome: The proposed model is based on a set of user-generated audio interfaces with 7,500 interactions from 484 participants.
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior work has introduced context-extension methods (e.g. YaRN) on unimodal LLMs, yet their application to LALMs remains unexplored.
Approach: They propose a training-free, modality-decoupled extension method that modifies only audio token positions, leaving text positions intact to preserve the base LLM’s text capabilities.
Outcome: The proposed method outperforms the original models across wide range of settings and provides significant performance improvement on long audio of unseen lengths.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations