Arkady Arkhangorodsky, Christopher Chu, Scot Fang, Yiqi Huang, Denglin Jiang, Ajay Nagesh, Boliang Zhang, Kevin Knight
| Challenge: | MeetDot is a videoconferencing system with live translation captions overlaid on screen . currently, the system supports speech and captions in 4 languages . |
| Approach: | They propose a videoconferencing system with live translation captions overlaid on screen . the system supports speech and captions in 4 languages and combines automatic speech recognition and machine translation in a cascade . |
| Outcome: | The proposed system supports speech and captions in 4 languages and has very tight latency requirements to have acceptable call quality. |
Similar Papers
ELITR Multilingual Live Subtitling: Demo and Strategy (2021.eacl-demos)
Copied to clipboard
Ondřej Bojar, Dominik Macháček, Sangeet Sagar, Otakar Smrž, Jonáš Kratochvíl, Peter Polák, Ebrahim Ansari, Mohammad Mahmoudi, Rishu Kumar, Dario Franceschini, Chiara Canton, Ivan Simonini, Thai-Son Nguyen, Felix Schneider, Sebastian Stüker, Alex Waibel, Barry Haddow, Rico Sennrich, Philip Williams
| Challenge: | Using a prototype, we present an automatic speech translation system for live subtitling of conference speech . the system is routinely tested in recognizing English, Czech, and German speech - and presenting it simultaneously into 42 target languages. |
| Approach: | They propose an automatic speech translation system aimed at live subtitling of conference presentations. |
| Outcome: | The proposed system is a working prototype that is routinely tested in recognizing English, Czech, and German speech and presenting it translated simultaneously into 42 target languages. |
A Semi-Automated Live Interlingual Communication Workflow Featuring Intralingual Respeaking: Evaluation and Benchmarking (2022.lrec-1)
Copied to clipboard
| Challenge: | Traditionally, live interlingual communication has been achieved only with the help of human interpreters. |
| Approach: | They propose a semi-automated workflow which uses a human respeaker and speaker-dependent speech recognition software to deliver punctuated same-language output of superior quality than the out-of-the-box ASR system. |
| Outcome: | The proposed workflow produces a similar quality output to the best-in-class simultaneous interpreters working with the same source speeches from the European Parliament. |
sign.mt: Real-Time Multilingual Sign Language Translation Application (2024.emnlp-demo)
Copied to clipboard
| Challenge: | open-source application for real-time multilingual bi-directional translation between spoken and signed languages. |
| Approach: | They present an open-source application for real-time multilingual bi-directional translation between spoken and signed languages. |
| Outcome: | The open-source sign.mt application aims to address the communication divide between the hearing and the deaf. |
Direct Simultaneous Speech-to-Text Translation Assisted by Synchronized Streaming ASR (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to simultaneous speech-to-text translation suffer from error propagation and extra latency. |
| Approach: | They propose a new paradigm for simultaneous speech-to-text translation using two separate decoders . they use multitask learning to jointly learn these two tasks with a shared encoder . |
| Outcome: | The proposed method achieves substantially better translation quality at similar levels of latency. |
SIMULEVAL: An Evaluation Toolkit for Simultaneous Translation (2020.emnlp-demos)
Copied to clipboard
| Challenge: | SimulEval is an evaluation toolkit for simultaneous text and speech translation. |
| Approach: | They propose a server-client scheme for simultaneous translation that uses server input and client policies to evaluate models. |
| Outcome: | The proposed evaluation toolkit is available for both text and speech translation. |
It’s Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems (2025.acl-long)
Copied to clipboard
| Challenge: | idioms are defined as words with a figurative meaning not deducible from their individual components. |
| Approach: | They compare idiom translation as compared to conventional news translation in two languages . they compare MT and SLT systems with MT, Large Language Models and cascaded alternatives . |
| Outcome: | The proposed systems show better handling of idioms than standard news translation systems. |
Neural Machine Translation Methods for Translating Text to Sign Language Glosses (2023.acl-long)
Copied to clipboard
| Challenge: | State-of-the-art techniques common to low resource Machine Translation (MT) are applied to improve MT of spoken language text to Sign Language glosses. |
| Approach: | They propose to use data augmentation, semi-supervised Neural Machine Translation, transfer learning and multilingual NMT to improve MT of spoken language to Sign Language glosses. |
| Outcome: | The proposed models outperform previous work on two German SL corpora and are confirmed by human evaluation. |
Machine Translation between Spoken Languages and Signed Languages Represented in SignWriting (2023.findings-eacl)
Copied to clipboard
| Challenge: | Yin et al. ( 2021) calls for including sign language processing (SLP) in natural language processing research. |
| Approach: | They propose to use a sign language writing system to parse, factorize, decode and evaluate signed languages. |
| Outcome: | The proposed method achieves over 30 BLEU in a bilingual setup and over 20 BLUE in two multilingual setups. |
SignCLIP: Connecting Text and Sign Language by Contrastive Learning (2024.emnlp-main)
Copied to clipboard
| Challenge: | SignCLIP is an efficient method of learning useful visual representations for sign language processing from large-scale, multilingual video-text pairs without optimizing for a specific task or sign language of limited size. |
| Approach: | They propose a method for learning visual representations for sign language processing from large-scale video-text pairs without directly optimizing for a specific task or sign language. |
| Outcome: | The proposed model can learn from multilingual video-text pairs without optimizing for a specific task or sign language of limited size. |
Japanese-to-English Simultaneous Dubbing Prototype (2023.acl-demo)
Copied to clipboard
| Challenge: | Our system translates and replaces the original speech of a live video stream in a simultaneous manner. |
| Approach: | They propose a simultaneous dubbing prototype that translates and replaces the original speech of a live video stream in a simultaneous manner. |
| Outcome: | The proposed system achieves a low average latency of 11.90 seconds and meets a smoothness criterion. |