Papers by Yong Ro
Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies show that multilingual models outperform monolingual ones. |
| Approach: | They propose a single model that can capture which language is given as input speech . they use a pre-trained model to fine-tune the model so it can recognize the language class as well as the speech with the corresponding language. |
| Outcome: | The proposed model can recognize which language is given as input speech . it can accurately recognize speech in noisy environments, such as crowded restaurants . |