Papers by Tom Ko
MOSPC: MOS Prediction Based on Pairwise Comparison (2023.acl-short)
Copied to clipboard
| Challenge: | et al., 2016a) show that MOS prediction model can improve ranking accuracy of speech quality. |
| Approach: | They propose a general framework for MOS prediction based on pair comparison . they use C-Mixup algorithm to enhance generalization performance of MOSPC . |
| Outcome: | The proposed model outperforms baselines on most correlation coefficient metrics . it also surpasses the strong baseline in ranking accuracy on each fine-grained segment. |
RepCodec: A Speech Representation Codec for Speech Tokenization (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have led to discrete speech tokenization, but this discretization can be costly and impedes performance. |
| Approach: | They propose a new speech representation codec for semantic speech tokenization that reconstructs speech representations from speech encoders like HuBERT or data2vec. |
| Outcome: | The proposed method outperforms the widely used k-means clustering approach in speech understanding and generation. |
Parameter-Efficient Transfer Learning for End-to-end Speech Translation (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to improve end-to-end speech translation are limited by the availability of labeled data. |
| Approach: | They propose a method which utilizes two lightweight adaptation techniques to modulate Attention and the Feed-Forward Network while preserving the capabilities of pre-trained models. |
| Outcome: | The proposed method outperforms baseline models and significantly improves performance in low-resource settings. |
CTC-based Non-autoregressive Speech Translation (2023.acl-long)
Copied to clipboard
Chen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun, Yuhao Zhang, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao, Anxiang Ma, Jingbo Zhu
| Challenge: | End-to-end speech translation (E2E ST) and non-autoregressive (NAR) generation are promising in language and speech processing for their advantages of less error propagation and low latency. |
| Approach: | They develop a model that uses connectionist temporal classification to predict the source and target texts. |
| Outcome: | The proposed model achieves an average BLEU score of 29.5 with a speed-up of 5.67. |
Selective Prompting Tuning for Personalized Conversations with LLMs (2024.findings-acl)
Copied to clipboard
| Challenge: | Personalization in conversational AI requires persona profiles and contextual understanding to create meaningful conversations. |
| Approach: | They propose a method that softly prompts LLMs for personalized conversations in a selective way. |
| Outcome: | The proposed approach improves response diversity by up to 90% on the CONVAI2 dataset. |
Learning Retrieval Augmentation for Personalized Dialogue Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Personalized dialogue generation is a popular approach for conversational AI applications . however, persona profiles may not provide comprehensive descriptions of the persona . |
| Approach: | They propose a method that leverages persona profiles and dialogue context to generate personalized dialogues by leveraging personas and persona profile. |
| Outcome: | The proposed method outperforms baselines on the CONVAI2 dataset . it is expected to generate personalized dialogues based on persona profiles and dialogue context . |
DUB: Discrete Unit Back-translation for Speech Translation (2023.findings-acl)
Copied to clipboard
| Challenge: | Discrete unit back-translation (DUB) is a back-translated speech-to-text translation (ST) technique that can be applied to ST . a modality gap between speech and text makes it difficult to transfer these techniques to ST due to the modality of the speech-text model. |
| Approach: | They propose a method to represent speech with discrete units instead of continuous features in direct ST. |
| Outcome: | The proposed method achieves comparable performance to existing methods that rely on large-scale external data. |
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing (2022.acl-long)
Copied to clipboard
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei
| Challenge: | Existing work shows that pre-trained models can improve in various natural language processing tasks. |
| Approach: | They propose a unified-modal encoder-decoder framework that pre-trains speech-text representations using large-scale unlabeled speech and text data. |
| Outcome: | The proposed framework is superior to existing models on speech-to-text processing tasks. |