Papers by Matúš Žilinec

2 papers
Khan Academy Corpus: A Multilingual Corpus of Khan Academy Lectures (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of 10122 hours in 87394 recordings is presented in a new journal . 43% of recordings have human-written subtitles, covering a total of 137 languages.
Approach: They present a Khan Academy corpus with 10122 hours in 87394 recordings . 43% of recordings have human-written subtitles, and 137 languages are included .
Outcome: The dataset can be used to train multilingual speech recognition and translation models.
Backtranslation Feedback Improves User Confidence in MT, Not Quality (2021.naacl-main)

Copied to clipboard

Challenge: Inbound translation is a modern need for which the user experience has significant room for improvement, beyond the basic machine translation facility.
Approach: They propose to provide cues that indicate the quality of MT output as well as suggest possible rephrasing of the source language.
Outcome: The proposed feedback module increases user confidence in the produced translation, but not the objective quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations