Papers by Liang-Hsuan Tseng

2 papers
Introducing Semantics into Speech Encoders (2023.acl-long)

Copied to clipboard

Challenge: Existing self-supervised speech encoders contain primarily acoustic rather than semantic information.
Approach: They propose a task-agnostic unsupervised way to incorporate semantic information from large language model (LLM) systems into self-supervised speech encoders without labeled audio transcriptions.
Outcome: The proposed approach improves spoken language understanding (SLU) performance by over 5% on intent classification (IC), with modest gains in named entity resolution (NER) and slot filling (SF), and spoken question answering (SQA) score by over 22%.
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Generative spoken language models are often evaluated using global token perplexity, which overlooks fundamental differences between speech and text modalities.
Approach: They propose a variety of likelihood- and generative-based evaluation methods that serve in place of naive global token perplexity.
Outcome: The proposed evaluations more faithfully reflect perceived generation quality, as evidenced by stronger correlations with human-rated mean opinion scores (MOS).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations