The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)
Copied to clipboard
| Challenge: | Using the corpus, we study the characteristics of interpreters' work and train machine translation systems. |
| Approach: | They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work. |
| Outcome: | The proposed corpus can be used for teaching interpreters and to train machine translation systems. |
Similar Papers
EPIC UdS - Creation and Applications of a Simultaneous Interpreting Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | EPIC UdS is a multilingual corpus of simultaneous interpreting for English, German and Spanish. |
| Approach: | They describe the creation and annotation of EPIC UdS, a multilingual corpus of simultaneous interpreting for English, German and Spanish. |
| Outcome: | The proposed corpus includes transcripts suitable for research on more than one language pair and on interpreting with regard to German. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)
Copied to clipboard
| Challenge: | specialized corpora on a variety of topics are available online, but online corporates are scarce. |
| Approach: | They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists . |
| Outcome: | The proposed corpus contains more than six million speeches in English and Chinese and is available for free online. |
Toward Machine Interpreting: Lessons from Human Interpreting Studies (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current speech translation systems are static and do not adapt to real-world situations in ways human interpreters do. |
| Approach: | They propose to model human interpreting using a new language model to improve usability . they argue that there is great potential to adopt many human interpreted principles . |
| Outcome: | The proposed models can be used to improve human interpreting and improve translation performance. |
LibriS2S: A German-English Speech-to-Speech Translation Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in speech-to-text translation have led to significant improvements, but the availability of appropriate training data is limiting. |
| Approach: | They propose a new text-to-speech and speech-tospech translation model that directly learns to generate the speech signal based on the pronunciation of the source language. |
| Outcome: | The proposed model learns to generate speech signal based on pronunciation of source language. |
WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse (D18-1)
Copied to clipboard
| Challenge: | a corpus of 43 million atomic edits is available for Wikipedia edit history . edits are instances in which a human editor has inserted a single contiguous phrase into, or deleted a contigous phrase from, an existing sentence. |
| Approach: | They use Wikipedia edit history to mine atomic edits across 8 languages . they find edits contain instances in which a human editor has inserted a single phrase into, or deleted a contiguous phrase from, an existing sentence. |
| Outcome: | The data show that edits differ from the language observed in standard corpora and that models trained on edits encode different aspects of semantics and discourse than models trained in raw text. |
The Abkhaz National Corpus (L18-1)
Copied to clipboard
| Challenge: | Abkhaz National Corpus is a comprehensive and open, grammatically annotated text corpus . it is currently growing and is being extended to include all important texts written in the language . |
| Approach: | They propose to use the Abkhaz National Corpus to annotate Abkhhaz texts . the corpus is a comprehensive and open, grammatically annotated text corpus . |
| Outcome: | The proposed corpus is a grammatically annotated text corpus which makes the language accessible to scientific investigations from various perspectives. |
Arabic Speech Rhythm Corpus: Read and Spontaneous Speaking Styles (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Arabic speech recordings has been built to allow comparisons between Arabic and other languages. |
| Approach: | They propose to build a corpus of Arabic speech recordings that can be compared with other languages. |
| Outcome: | The proposed corpus can be used for forensic phonetic research and casework applications. |
The GermaParl Corpus of Parliamentary Protocols (L18-1)
Copied to clipboard
| Challenge: | Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making. |
| Approach: | They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations. |
| Outcome: | The proposed corpus provides a valuable resource for research and teaching purposes. |
The EDGeS Diachronic Bible Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EDGeS is a diachronic and parallel corpus of Bible translations in Dutch, English, German and Swedish . it is intended to be used for longitudinal studies of complex verb constructions in Germanic . |
| Approach: | They present the EDGeS Diachronic Bible Corpus, a diachronic corpus of Bible translations in Dutch, English, German and Swedish . they use a synchronically and synchronly parallel corpus to study complex verb constructions in Germanic . |
| Outcome: | The EDGeS is a diachronic and parallel corpus of Bible translations in Dutch, English, German and Swedish spanning six and a half centuries. |