Papers by Liang-Hsuan Tseng
Introducing Semantics into Speech Encoders (2023.acl-long)
Copied to clipboard
Derek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim, Zhaojiang Lin, Bing Liu, Akshat Shrivastava, Shang-Wen Li, Liang-Hsuan Tseng, Guan-Ting Lin, Alexei Baevski, Hung-yi Lee, Yizhou Sun, Wei Wang
| Challenge: | Existing self-supervised speech encoders contain primarily acoustic rather than semantic information. |
| Approach: | They propose a task-agnostic unsupervised way to incorporate semantic information from large language model (LLM) systems into self-supervised speech encoders without labeled audio transcriptions. |
| Outcome: | The proposed approach improves spoken language understanding (SLU) performance by over 5% on intent classification (IC), with modest gains in named entity resolution (NER) and slot filling (SF), and spoken question answering (SQA) score by over 22%. |
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation (2026.findings-acl)
Copied to clipboard
Chan-Jan Hsu, Liang-Hsuan Tseng, Yi-Cheng Lin, Yen-Chun Kuo, Ju-Chieh Chou, Kai-Wei Chang, Hung-yi Lee, Carlos Busso
| Challenge: | Generative spoken language models are often evaluated using global token perplexity, which overlooks fundamental differences between speech and text modalities. |
| Approach: | They propose a variety of likelihood- and generative-based evaluation methods that serve in place of naive global token perplexity. |
| Outcome: | The proposed evaluations more faithfully reflect perceived generation quality, as evidenced by stronger correlations with human-rated mean opinion scores (MOS). |