Papers by Ondrej Klejch
European Language Grid: A Joint Platform for the European Language Technology Community (2021.eacl-demos)
Copied to clipboard
Georg Rehm, Stelios Piperidis, Kalina Bontcheva, Jan Hajic, Victoria Arranz, Andrejs Vasiļjevs, Gerhard Backfried, Jose Manuel Gomez-Perez, Ulrich Germann, Rémi Calizzano, Nils Feldhus, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Julian Moreno-Schneider, Dimitris Galanis, Penny Labropoulou, Miltos Deligiannis, Katerina Gkirtzou, Athanasia Kolovou, Dimitris Gkoumas, Leon Voukoutis, Ian Roberts, Jana Hamrlova, Dusan Varis, Lukas Kacena, Khalid Choukri, Valérie Mapelli, Mickaël Rigault, Julija Melnika, Miro Janosik, Katja Prinz, Andres Garcia-Silva, Cristian Berrio, Ondrej Klejch, Steve Renals
| Challenge: | Europe is a multilingual society, in which dozens of languages are spoken. |
| Approach: | They describe the European Language Grid, which is targeted to evolve into the primary platform and marketplace for LT in Europe by providing one umbrella platform for the European LT landscape. |
| Outcome: | The European Language Grid (ELG) will provide access to 1300 services for all European languages as well as thousands of data sets. |
F-Actor: Controllable Conversational Behavior in Full-Duplex Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Current spoken conversational systems lack customization capabilities, limiting their naturalness and usability. |
| Approach: | They propose an instruction-following full-duplex conversational speech model that can be trained efficiently under typical academic resource constraints. |
| Outcome: | The proposed model requires just 2,000 hours of data to be trained under typical academic resource constraints. |
Spoken Document Retrieval for an Unwritten Language: A Case Study on Gormati (2025.findings-emnlp)
Copied to clipboard
Sanjay Booshanam, Kelly Chen, Ondrej Klejch, Thomas Reitmaier, Dani Kalarikalayil Raju, Electra Wallington, Nina Markl, Jennifer Pearson, Matt Jones, Simon Robinson, Peter Bell
| Challenge: | Speakers of unwritten languages have the potential to benefit from speech-based automatic information retrieval systems. |
| Approach: | They propose a speech embedding technique that facilitates a zero-shot speech-based automatic information retrieval system for unwritten languages. |
| Outcome: | The proposed method achieves a Top 5 retrieval rate of 87.9% on a corpus of Gormati, an unwritten language, that was collected in partnership with an agrarian Banjara community in Maharashtra State, India. |
European Language Grid: An Overview (2020.lrec-1)
Copied to clipboard
Georg Rehm, Maria Berger, Ela Elsholz, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Stelios Piperidis, Miltos Deligiannis, Dimitris Galanis, Katerina Gkirtzou, Penny Labropoulou, Kalina Bontcheva, David Jones, Ian Roberts, Jan Hajič, Jana Hamrlová, Lukáš Kačena, Khalid Choukri, Victoria Arranz, Andrejs Vasiļjevs, Orians Anvari, Andis Lagzdiņš, Jūlija Meļņika, Gerhard Backfried, Erinç Dikici, Miroslav Janosik, Katja Prinz, Christoph Prinz, Severin Stampler, Dorothea Thomas-Aniola, José Manuel Gómez-Pérez, Andres Garcia Silva, Christian Berrío, Ulrich Germann, Steve Renals, Ondrej Klejch
| Challenge: | European LT business is dominated by hundreds of SMEs and a few large players, with technologies that outperform the global players. |
| Approach: | European Language Grid (ELG) project addresses this by establishing the ELG as the primary platform for LT in Europe. |
| Outcome: | European Language Grid (ELG) will be primary platform for LT in Europe . it will provide access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. |