Papers by Matúš Žilinec
Khan Academy Corpus: A Multilingual Corpus of Khan Academy Lectures (2024.lrec-main)
Copied to clipboard
| Challenge: | a dataset of 10122 hours in 87394 recordings is presented in a new journal . 43% of recordings have human-written subtitles, covering a total of 137 languages. |
| Approach: | They present a Khan Academy corpus with 10122 hours in 87394 recordings . 43% of recordings have human-written subtitles, and 137 languages are included . |
| Outcome: | The dataset can be used to train multilingual speech recognition and translation models. |
Backtranslation Feedback Improves User Confidence in MT, Not Quality (2021.naacl-main)
Copied to clipboard
Vilém Zouhar, Michal Novák, Matúš Žilinec, Ondřej Bojar, Mateo Obregón, Robin L. Hill, Frédéric Blain, Marina Fomicheva, Lucia Specia, Lisa Yankovskaya
| Challenge: | Inbound translation is a modern need for which the user experience has significant room for improvement, beyond the basic machine translation facility. |
| Approach: | They propose to provide cues that indicate the quality of MT output as well as suggest possible rephrasing of the source language. |
| Outcome: | The proposed feedback module increases user confidence in the produced translation, but not the objective quality. |