Papers by Yu-An Chung
Supervised and Unsupervised Transfer Learning for Question Answering (N18-1)
Copied to clipboard
| Challenge: | Several QA scenarios and datasets have been introduced over the past few years. |
| Approach: | They conduct extensive experiments to investigate the transferability of knowledge from a source QA dataset to a target dataset using two QA models. |
| Outcome: | The proposed model outperforms the previous best model on TOEFL listening comprehension test by 7% on target datasets. |
Speech-to-Speech Translation for a Real-world Unwritten Language (2023.findings-acl)
Copied to clipboard
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne, Holger Schwenk, Hongyu Gong, Hirofumi Inaguma, Sravya Popuri, Changhan Wang, Juan Pino, Wei-Ning Hsu, Ann Lee
| Challenge: | a new study examines speech-to-speech translation (S2ST) that translates speech from one language into another . the research area for unwritten languages remains a research area with little exploration due to the lack of training data. |
| Approach: | They propose a system that translates speech from one language into another . they use Taiwanese Hokkien as an example of an unwritten language . |
| Outcome: | The proposed system can be used to train models in languages without standard writing systems. |
UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units (2023.acl-long)
Copied to clipboard
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, Juan Pino
| Challenge: | Experimental evaluations show that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up. |
| Approach: | They propose a two-pass direct S2ST architecture which generates textual representations and predicts discrete acoustic units . they show that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up. |
| Outcome: | The proposed architecture outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up on large datasets. |
Improved Speech Representations with Multi-Target Autoregressive Predictive Coding (2020.acl-main)
Copied to clipboard
| Challenge: | Autoregressive coding targets are used to learn meaningful representations from unlabeled speech. |
| Approach: | They propose a method that trains an autoregressive RNN to generate an unseen future frame given a context such as recent past frames. |
| Outcome: | The proposed method can learn representations from unlabeled speech. |
SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding (2021.naacl-main)
Copied to clipboard
| Challenge: | Experimental results show that SPLAT improves the previous state-of-the-art performance on the Spoken SQuAD dataset by more than 10%. |
| Approach: | They propose a semi-supervised learning framework to jointly pre-train the speech and language modules using unpaired speech and text. |
| Outcome: | The proposed framework improves the previous state-of-the-art performance on the Spoken SQuAD dataset by more than 10%. |