| Challenge: | a new study examines the need for sign language translators to have tools similar to text-to-text translation. |
| Approach: | They propose to use a concordancer to search for parallel Franch-LSF segments . they use dozens of short news clips and 120 SL videos to align them manually . |
| Outcome: | The proposed data base will be searched using a concordancer and expand in the future. |
Similar Papers
How to Align Multiple Signed Language Corpora for Better Sign-to-Sign Translations? (2025.naacl-long)
Copied to clipboard
| Challenge: | despite the growing need for advanced signing technologies, signed language resources remain scarce. |
| Approach: | They propose a linguistically informed alignment algorithm that matches instances between signed languages . they compare similarities and differences across three signed languages to develop a model . |
| Outcome: | The proposed algorithm performs well on automatic metrics for sign-to-sign translation and generation. |
Rosetta-LSF: an Aligned Corpus of French Sign Language and French for Text-to-Sign Translation (2022.lrec-1)
Copied to clipboard
Elise Bertin-Lemée, Annelies Braffort, Camille Challant, Claire Danet, Boris Dauriac, Michael Filhol, Emmanuella Martinod, Jérémie Segouat
| Challenge: | a new corpus of french Sign Language (LSF) data is created to support future studies on the automatic translation of written French into LSF, rendered through the animation of a virtual signer. |
| Approach: | They propose to use a French Sign Language corpus called "Rosetta-LSF" it is intended to support studies on automatic translation of written French into LSF . |
| Outcome: | The proposed corpus supports future studies on automatic translation of written French into LSF, rendered through animation of a virtual signer. |
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches for aligning spoken language text to sign language videos rely on end-to-end training tied to a specific language or dataset. |
| Approach: | They propose a universal approach for aligning spoken language text with corresponding timestamps to sign language videos using a lightweight dynamic programming procedure. |
| Outcome: | The proposed method can be used on four sign language datasets and is highly efficient on CPU. |
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage. |
| Approach: | They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages. |
| Outcome: | The proposed index covers 120 resources across 35 sign languages. |
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |
BinaryAlign: Word Alignment as Binary Sequence Labeling (2024.acl-long)
Copied to clipboard
| Challenge: | State-of-the-art word alignment training methods require a different class depending on the availability of gold data for a particular language pair. |
| Approach: | They propose a novel word alignment technique based on binary sequence labeling that outperforms existing approaches in both scenarios. |
| Outcome: | The proposed method outperforms existing models on non-English language pairs and performs stratified error analysis over alignment error type. |
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)
Copied to clipboard
Sina Ahmadi, John Philip McCrae, Sanni Nimb, Fahad Khan, Monica Monachini, Bolette Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José Luis Sancho, Rafael-J. Ureña-Ruiz, Jordi Porta Zamorano, Kiril Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stanković, Andrej Perdih, Dejan Gabrovsek
| Challenge: | a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources . |
| Approach: | They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment . |
| Outcome: | The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language. |
SwissSLi: The Multi-parallel Sign Language Corpus for Switzerland (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a CC BY-NC-SA 4.0 license, this corpus contains parallel sign language videos and spoken language subtitles. |
| Approach: | They introduce SwissSLi, the first sign language corpus that contains parallel data of all three Swiss sign languages. |
| Outcome: | The proposed corpus contains parallel sign language videos and spoken language subtitles. |
Challenges with Sign Language Datasets for Sign Language Recognition and Translation (2022.lrec-1)
Copied to clipboard
Mirella De Sisto, Vincent Vandeghinste, Santiago Egea Gómez, Mathieu De Coster, Dimitar Shterionov, Horacio Saggion
| Challenge: | Sign Languages are the primary means of communication for at least half a million people in Europe . however, the development of SL recognition and translation tools is slowed down by resource scarcity and data formats are not suitable for machine learning. |
| Approach: | They propose a framework to unify available resources and facilitate SL research for different languages. |
| Outcome: | The proposed framework is based on a set of ELAN files and returns textual and visual data ready to train SL recognition and translation models. |
Word Alignment by Fine-tuning Embeddings on Parallel Corpora (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing work on word alignment has focused on unsupervised learning on parallel text. |
| Approach: | They propose to combine pre-trained contextualized word embeddings with multilingually trained language models to achieve competitive results on word alignment tasks. |
| Outcome: | The proposed model outperforms state-of-the-art models on five language pairs and can train multilingual word aligners that can obtain robust performance on different language pairs. |