Papers by Yuatyong Chaichana
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models (2026.eacl-long)
Copied to clipboard
Yuatyong Chaichana, Pittawat Taveekitworachai, Warit Sirichotedumrong, Potsawee Manakul, Kunat Pipatanakul
| Challenge: | Prior work has introduced context-extension methods (e.g. YaRN) on unimodal LLMs, yet their application to LALMs remains unexplored. |
| Approach: | They propose a training-free, modality-decoupled extension method that modifies only audio token positions, leaving text positions intact to preserve the base LLM’s text capabilities. |
| Outcome: | The proposed method outperforms the original models across wide range of settings and provides significant performance improvement on long audio of unseen lengths. |