Papers by Kunat Pipatanakul
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation (2026.eacl-long)
Copied to clipboard
Potsawee Manakul, Woody Haosheng Gan, Michael J Ryan, Ali Sartaz Khan, Warit Sirichotedumrong, Kunat Pipatanakul, William Barr Held, Diyi Yang
| Challenge: | Current speech evaluation systems rely on specialized systems for individual audio characteristics and poor correlation between automatic methods and human preferences. |
| Approach: | They propose a unified evaluation framework for Large Audio Models as a Judge, AudioJudge . they propose specialized judges that can be prompted to perform audio characteristic detection tasks . |
| Outcome: | The proposed method improves performance across audio characteristic detection and human preference simulation tasks. |
Prior Prompt Engineering for Reinforcement Fine-Tuning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have focused on algorithms, reward shaping, and data curation, but prior prompt engineering is understudied. |
| Approach: | They investigate prior prompt engineering (pPE) in reinforcement fine-tuning . they translate five representative iPE strategies into corresponding pPE approaches . |
| Outcome: | The proposed approaches outperform iPE-prompted models on in-domain and out-of-domain benchmarks. |
Mind the Gap: Static and Interactive Evaluations of Large Audio Models (2025.acl-long)
Copied to clipboard
Minzhi Li, William Barr Held, Michael J Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, Diyi Yang
| Challenge: | Recent work has focused on evaluating large audio models (LAMs) that directly accept audio inputs. |
| Approach: | They propose an interactive approach to evaluate large audio models and collect 7,500 LAM interactions from 484 participants. |
| Outcome: | The proposed model is based on a set of user-generated audio interfaces with 7,500 interactions from 484 participants. |
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models (2026.eacl-long)
Copied to clipboard
Yuatyong Chaichana, Pittawat Taveekitworachai, Warit Sirichotedumrong, Potsawee Manakul, Kunat Pipatanakul
| Challenge: | Prior work has introduced context-extension methods (e.g. YaRN) on unimodal LLMs, yet their application to LALMs remains unexplored. |
| Approach: | They propose a training-free, modality-decoupled extension method that modifies only audio token positions, leaving text positions intact to preserve the base LLM’s text capabilities. |
| Outcome: | The proposed method outperforms the original models across wide range of settings and provides significant performance improvement on long audio of unseen lengths. |