Papers by Kwanghee Choi

6 papers
Wav2Gloss: Generating Interlinear Glossed Text from Speech (2024.acl-long)

Copied to clipboard

Challenge: Interlinear Glossed Text (IGT) is a form of linguistic annotation that can support documentation and resource creation for endangered languages.
Approach: They propose a task in which these four annotation components are extracted automatically from speech and introduce a dataset to lay the groundwork for future research on IGT generation from speech.
Outcome: The proposed dataset provides the first dataset to lay the groundwork for future research on IGT generation from speech, including end-to-end versus cascaded, monolingual versus multilingual, and single-task versus multiple-task approaches.
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model (2026.acl-long)

Copied to clipboard

Challenge: Phone-level modeling of speech is a common approach to speech recognition, but it relies on task-specific architectures and datasets.
Approach: They propose a phonetic framework capable of performing multiple phone-related tasks . they propose 'Phonetic Open Whisper-style Speech Model' that can perform these tasks together .
Outcome: The proposed model outperforms or matches specialized PR models of similar size while supporting G2P, P2G, and ASR.
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration (2026.findings-acl)

Copied to clipboard

Challenge: We show that script information is linearly encoded in the activation space of multilingual speech models . modifying activations at inference time induces script change even in unconventional pairings .
Approach: They propose to add script vectors to activations at test time to induce script change . they also show that script information is linearly encoded in the activation space of multilingual speech models .
Outcome: The proposed approach can induce script change even in unconventional language-script pairings.
[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on how self-supervised speech models encode rich phonetic information have not explored how they are structured.
Approach: They conduct a comprehensive analysis of the underlying structure of S3M representations with particular attention to phonological vectors.
Outcome: The proposed model encodes phonologically interpretable and compositional vectors, demonstrating phonology vector arithmetic.
PRiSM: Benchmarking Phone Realization in Speech Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of phone recognition systems only measure surface-level transcription accuracy.
Approach: They propose to standardize transcription-based evaluation and assess downstream utility in clinical, educational, and multilingual settings with transcription and representation probes.
Outcome: The proposed system outperforms LALMs in clinical, educational, and multilingual settings.
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment (2025.naacl-long)

Copied to clipboard

Challenge: Recent phoneme classifiers treat allophonic variation as a single phoneme . atypical pronunciation assessment requires distinguishing between a typical and asymmetric pronunciations .
Approach: They propose a new approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters.
Outcome: The proposed approach achieves state-of-the-art across dysarthric and non-native speech datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations