Challenge: elicitation of long discourses is difficult in Sign Language, and is often a problem . e.g., elicitation of long texts is a technique that can be used to collect long discourse .
Approach: They propose a protocol and two tasks to collect long discourse in Sign Language . they propose to ensure both are collected and prepared in the language .
Outcome: The proposed protocol improves the produced data and the results of a test with LSF informants.

Similar Papers

A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)

Copied to clipboard

Challenge: Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications.
Approach: They analyze available dialogue corpora and propose possible approaches to create new ones.
Outcome: The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed .
A First Corpus of AZee Discourse Expressions (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of AZee discourse expressions which formally describe Sign Language utterances of any length is presented.
Approach: They propose to build a corpus of AZee discourse expressions which formally describe Sign Language utterances of any length using the AZEe approach and language.
Outcome: The proposed corpus is based on a video production of a French Sign Language speech expression containing 40 breves and is evaluated on real-life utterances.
Alignment Data base for a Sign Language Concordancer (2020.lrec-1)

Copied to clipboard

Challenge: a new study examines the need for sign language translators to have tools similar to text-to-text translation.
Approach: They propose to use a concordancer to search for parallel Franch-LSF segments . they use dozens of short news clips and 120 SL videos to align them manually .
Outcome: The proposed data base will be searched using a concordancer and expand in the future.
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
Dicta-Sign-LSF-v2: Remake of a Continuous French Sign Language Dialogue Corpus and a First Baseline for Automatic Sign Language Processing (2020.lrec-1)

Copied to clipboard

Challenge: Existing research on automatic Sign Language Processing (SLP) has focused on recognizing lexical signs, but other gestural units like iconic structures need to be recognized.
Approach: They propose a public remake of the French Sign Language part of the Dicta-Sign corpus with clean annotations and a Convolutional-Recurrent Neural Network to train and test it.
Outcome: The proposed version of the publicly available SL corpus Dicta-Sign is limited to its French Sign Language part and includes lexical and non-lexical annotations over 11 hours of video recording with 35000 manual units.
Towards Continuous Dialogue Corpus Creation: writing to corpus and generating from it (L18-1)

Copied to clipboard

Challenge: Existing methods to create dialogue corpora annotated with interoperable semantic information are based on ISO standard data models and tools.
Approach: They propose to use a corpus as a shared repository for analysis and modelling of interactive dialogue behaviour and for implementation, integration and evaluation of dialogue system components.
Outcome: The proposed method is applied to the design of two multimodal interactive applications - the Virtual Negotiation Coach and the Virtual Debate Coach.
An Exploratory Study on Long Dialogue Summarization: What Works and What’s Next (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models for dialogue summarization focus on extracting the main events of short conversations, but real-world dialogues are difficult to train.
Approach: They propose three strategies to deal with the lengthy input problem and locate relevant information using long dialogue datasets.
Outcome: The retrieve-then-summarize pipeline models yield the best performance on three long dialogue datasets.
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage.
Approach: They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages.
Outcome: The proposed index covers 120 resources across 35 sign languages.
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2025.acl-demo)

Copied to clipboard

Challenge: ACL 2025 System Demonstration Track accepted 64 papers based on reviews . short-listed 7 papers for Best System Demo award .
Approach: the ACL 2025 System Demonstration Track is a conference for papers describing system demonstrations . the track received a record 187 submissions, of which 178 papers were valid with required materials .
Outcome: the ACL 2025 System Demonstration Track received 187 submissions . 178 papers were valid with required materials .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations