| Challenge: | a corpus of AZee discourse expressions which formally describe Sign Language utterances of any length is presented. |
| Approach: | They propose to build a corpus of AZee discourse expressions which formally describe Sign Language utterances of any length using the AZEe approach and language. |
| Outcome: | The proposed corpus is based on a video production of a French Sign Language speech expression containing 40 breves and is evaluated on real-life utterances. |
Similar Papers
Extending AZee with Non-manual Gesture Rules for French Sign Language (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, Sign Languages (SLs) are under-resourced and are difficult to develop. |
| Approach: | They propose to extend AZee to formally represent Sign Language discourses, but also to animate them with a virtual signer. |
| Outcome: | The proposed model allows to formally represent Sign Language discourses, but also to animate them with a virtual signer. |
Elicitation protocol and material for a corpus of long prepared monologues in Sign Language (L18-1)
Copied to clipboard
| Challenge: | elicitation of long discourses is difficult in Sign Language, and is often a problem . e.g., elicitation of long texts is a technique that can be used to collect long discourse . |
| Approach: | They propose a protocol and two tasks to collect long discourse in Sign Language . they propose to ensure both are collected and prepared in the language . |
| Outcome: | The proposed protocol improves the produced data and the results of a test with LSF informants. |
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)
Copied to clipboard
| Challenge: | Using the corpus, we study the characteristics of interpreters' work and train machine translation systems. |
| Approach: | They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work. |
| Outcome: | The proposed corpus can be used for teaching interpreters and to train machine translation systems. |
Rosetta-LSF: an Aligned Corpus of French Sign Language and French for Text-to-Sign Translation (2022.lrec-1)
Copied to clipboard
Elise Bertin-Lemée, Annelies Braffort, Camille Challant, Claire Danet, Boris Dauriac, Michael Filhol, Emmanuella Martinod, Jérémie Segouat
| Challenge: | a new corpus of french Sign Language (LSF) data is created to support future studies on the automatic translation of written French into LSF, rendered through the animation of a virtual signer. |
| Approach: | They propose to use a French Sign Language corpus called "Rosetta-LSF" it is intended to support studies on automatic translation of written French into LSF . |
| Outcome: | The proposed corpus supports future studies on automatic translation of written French into LSF, rendered through animation of a virtual signer. |
Alignment Data base for a Sign Language Concordancer (2020.lrec-1)
Copied to clipboard
| Challenge: | a new study examines the need for sign language translators to have tools similar to text-to-text translation. |
| Approach: | They propose to use a concordancer to search for parallel Franch-LSF segments . they use dozens of short news clips and 120 SL videos to align them manually . |
| Outcome: | The proposed data base will be searched using a concordancer and expand in the future. |
Dicta-Sign-LSF-v2: Remake of a Continuous French Sign Language Dialogue Corpus and a First Baseline for Automatic Sign Language Processing (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing research on automatic Sign Language Processing (SLP) has focused on recognizing lexical signs, but other gestural units like iconic structures need to be recognized. |
| Approach: | They propose a public remake of the French Sign Language part of the Dicta-Sign corpus with clean annotations and a Convolutional-Recurrent Neural Network to train and test it. |
| Outcome: | The proposed version of the publicly available SL corpus Dicta-Sign is limited to its French Sign Language part and includes lexical and non-lexical annotations over 11 hours of video recording with 35000 manual units. |
MAGPIE: A Large Corpus of Potentially Idiomatic Expressions (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora cover less than 5,000 instances of less than 100 different idiom types . large corpus allows for better evaluation of assumptions about idiomatic expressions . |
| Approach: | They propose to build the largest-to-date corpus of idioms for English using crowdsourcing methods. |
| Outcome: | The proposed corpus is larger than existing resources and contains rich metadata and is made publicly available. |
RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text (2020.lrec-1)
Copied to clipboard
| Challenge: | In literature, spoken interactions between characters are of central importance to the narrative. |
| Approach: | They propose to annotate quotations, including their interpersonal structure, for English literary text. |
| Outcome: | The proposed dataset provides a rich view of dialogue structures not available from other available corpora. |
Input Representations for Parsing Discourse Representation Structures: Comparing English with Chinese (2021.acl-short)
Copied to clipboard
| Challenge: | Neural semantic parsers have obtained acceptable results in parsing DRSs . previous studies have focused on parse of DRS in English, but have focused only on a few languages . |
| Approach: | They propose to use character sequences as input to map meaning representations to string format. |
| Outcome: | The proposed models learn the meaning of a series of semantic phenomena by taking sentences as input and outputting the corresponding DRSs, without the aid of any extra linguistic information. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |