Papers by Ebrahim Ansari
ELITR Multilingual Live Subtitling: Demo and Strategy (2021.eacl-demos)
Copied to clipboard
Ondřej Bojar, Dominik Macháček, Sangeet Sagar, Otakar Smrž, Jonáš Kratochvíl, Peter Polák, Ebrahim Ansari, Mohammad Mahmoudi, Rishu Kumar, Dario Franceschini, Chiara Canton, Ivan Simonini, Thai-Son Nguyen, Felix Schneider, Sebastian Stüker, Alex Waibel, Barry Haddow, Rico Sennrich, Philip Williams
| Challenge: | Using a prototype, we present an automatic speech translation system for live subtitling of conference speech . the system is routinely tested in recognizing English, Czech, and German speech - and presenting it simultaneously into 42 target languages. |
| Approach: | They propose an automatic speech translation system aimed at live subtitling of conference presentations. |
| Outcome: | The proposed system is a working prototype that is routinely tested in recognizing English, Czech, and German speech and presenting it translated simultaneously into 42 target languages. |
Extracting an English-Persian Parallel Corpus from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | Existing methods to extract parallel sentences from Wikipedia are limited for some language pairs such as Persian-English. |
| Approach: | They propose a bidirectional method to extract parallel sentences from Wikipedia . they add extracted sentences to existing training data and use IR system to measure similarity . |
| Outcome: | The proposed method outperforms the one-directional approach in analyzing translation data from two translation systems and IR systems. |
LSCP: Enhanced Large Scale Colloquial Persian Language Understanding (2020.lrec-1)
Copied to clipboard
| Challenge: | a gap exists in describing low-resource formal languages such as Persian . a large scale corpus of 120M sentences is proposed to fill this gap . |
| Approach: | They propose to target a gap in describing the colloquial language for low-resource ones such as Persian . a large scale Persian corpus is hierarchically organized in a semantic taxonomy . |
| Outcome: | The proposed corpus consists of 120M sentences from 27M tweets annotated with parsing tree, part-of-speech tags, sentiment polarity and translation in five different languages. |
SLTEV: Comprehensive Evaluation of Spoken Language Translation (2021.eacl-demos)
Copied to clipboard
| Challenge: | Spoken Language Translation (SLT) evaluation of machine translation (MT) quality has been investigated for decades. |
| Approach: | They propose an open-source tool for assessing machine translation (MT) quality based on time-stamped transcripts and reference translations. |
| Outcome: | The proposed evaluation tool is based on time-stamped transcripts and reference translations into a target language. |