MeetDot: Videoconferencing with Live Translation Captions (2021.emnlp-demo)

Copied to clipboard

Challenge: MeetDot is a videoconferencing system with live translation captions overlaid on screen . currently, the system supports speech and captions in 4 languages .
Approach: They propose a videoconferencing system with live translation captions overlaid on screen . the system supports speech and captions in 4 languages and combines automatic speech recognition and machine translation in a cascade .
Outcome: The proposed system supports speech and captions in 4 languages and has very tight latency requirements to have acceptable call quality.

Similar Papers

ELITR Multilingual Live Subtitling: Demo and Strategy (2021.eacl-demos)

Copied to clipboard

Challenge: Using a prototype, we present an automatic speech translation system for live subtitling of conference speech . the system is routinely tested in recognizing English, Czech, and German speech - and presenting it simultaneously into 42 target languages.
Approach: They propose an automatic speech translation system aimed at live subtitling of conference presentations.
Outcome: The proposed system is a working prototype that is routinely tested in recognizing English, Czech, and German speech and presenting it translated simultaneously into 42 target languages.
A Semi-Automated Live Interlingual Communication Workflow Featuring Intralingual Respeaking: Evaluation and Benchmarking (2022.lrec-1)

Copied to clipboard

Challenge: Traditionally, live interlingual communication has been achieved only with the help of human interpreters.
Approach: They propose a semi-automated workflow which uses a human respeaker and speaker-dependent speech recognition software to deliver punctuated same-language output of superior quality than the out-of-the-box ASR system.
Outcome: The proposed workflow produces a similar quality output to the best-in-class simultaneous interpreters working with the same source speeches from the European Parliament.
sign.mt: Real-Time Multilingual Sign Language Translation Application (2024.emnlp-demo)

Copied to clipboard

Challenge: open-source application for real-time multilingual bi-directional translation between spoken and signed languages.
Approach: They present an open-source application for real-time multilingual bi-directional translation between spoken and signed languages.
Outcome: The open-source sign.mt application aims to address the communication divide between the hearing and the deaf.
Direct Simultaneous Speech-to-Text Translation Assisted by Synchronized Streaming ASR (2021.findings-acl)

Copied to clipboard

Challenge: Existing approaches to simultaneous speech-to-text translation suffer from error propagation and extra latency.
Approach: They propose a new paradigm for simultaneous speech-to-text translation using two separate decoders . they use multitask learning to jointly learn these two tasks with a shared encoder .
Outcome: The proposed method achieves substantially better translation quality at similar levels of latency.
SIMULEVAL: An Evaluation Toolkit for Simultaneous Translation (2020.emnlp-demos)

Copied to clipboard

Challenge: SimulEval is an evaluation toolkit for simultaneous text and speech translation.
Approach: They propose a server-client scheme for simultaneous translation that uses server input and client policies to evaluate models.
Outcome: The proposed evaluation toolkit is available for both text and speech translation.
It’s Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems (2025.acl-long)

Copied to clipboard

Challenge: idioms are defined as words with a figurative meaning not deducible from their individual components.
Approach: They compare idiom translation as compared to conventional news translation in two languages . they compare MT and SLT systems with MT, Large Language Models and cascaded alternatives .
Outcome: The proposed systems show better handling of idioms than standard news translation systems.
Neural Machine Translation Methods for Translating Text to Sign Language Glosses (2023.acl-long)

Copied to clipboard

Challenge: State-of-the-art techniques common to low resource Machine Translation (MT) are applied to improve MT of spoken language text to Sign Language glosses.
Approach: They propose to use data augmentation, semi-supervised Neural Machine Translation, transfer learning and multilingual NMT to improve MT of spoken language to Sign Language glosses.
Outcome: The proposed models outperform previous work on two German SL corpora and are confirmed by human evaluation.
Machine Translation between Spoken Languages and Signed Languages Represented in SignWriting (2023.findings-eacl)

Copied to clipboard

Challenge: Yin et al. ( 2021) calls for including sign language processing (SLP) in natural language processing research.
Approach: They propose to use a sign language writing system to parse, factorize, decode and evaluate signed languages.
Outcome: The proposed method achieves over 30 BLEU in a bilingual setup and over 20 BLUE in two multilingual setups.
SignCLIP: Connecting Text and Sign Language by Contrastive Learning (2024.emnlp-main)

Copied to clipboard

Challenge: SignCLIP is an efficient method of learning useful visual representations for sign language processing from large-scale, multilingual video-text pairs without optimizing for a specific task or sign language of limited size.
Approach: They propose a method for learning visual representations for sign language processing from large-scale video-text pairs without directly optimizing for a specific task or sign language.
Outcome: The proposed model can learn from multilingual video-text pairs without optimizing for a specific task or sign language of limited size.
Japanese-to-English Simultaneous Dubbing Prototype (2023.acl-demo)

Copied to clipboard

Challenge: Our system translates and replaces the original speech of a live video stream in a simultaneous manner.
Approach: They propose a simultaneous dubbing prototype that translates and replaces the original speech of a live video stream in a simultaneous manner.
Outcome: The proposed system achieves a low average latency of 11.90 seconds and meets a smoothness criterion.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations