Papers by Yong Ro

1 papers
Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that multilingual models outperform monolingual ones.
Approach: They propose a single model that can capture which language is given as input speech . they use a pre-trained model to fine-tune the model so it can recognize the language class as well as the speech with the corresponding language.
Outcome: The proposed model can recognize which language is given as input speech . it can accurately recognize speech in noisy environments, such as crowded restaurants .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations