Alignment Data base for a Sign Language Concordancer (2020.lrec-1)

Copied to clipboard

Challenge: a new study examines the need for sign language translators to have tools similar to text-to-text translation.
Approach: They propose to use a concordancer to search for parallel Franch-LSF segments . they use dozens of short news clips and 120 SL videos to align them manually .
Outcome: The proposed data base will be searched using a concordancer and expand in the future.

Similar Papers

How to Align Multiple Signed Language Corpora for Better Sign-to-Sign Translations? (2025.naacl-long)

Copied to clipboard

Challenge: despite the growing need for advanced signing technologies, signed language resources remain scarce.
Approach: They propose a linguistically informed alignment algorithm that matches instances between signed languages . they compare similarities and differences across three signed languages to develop a model .
Outcome: The proposed algorithm performs well on automatic metrics for sign-to-sign translation and generation.
Rosetta-LSF: an Aligned Corpus of French Sign Language and French for Text-to-Sign Translation (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of french Sign Language (LSF) data is created to support future studies on the automatic translation of written French into LSF, rendered through the animation of a virtual signer.
Approach: They propose to use a French Sign Language corpus called "Rosetta-LSF" it is intended to support studies on automatic translation of written French into LSF .
Outcome: The proposed corpus supports future studies on automatic translation of written French into LSF, rendered through animation of a virtual signer.
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches for aligning spoken language text to sign language videos rely on end-to-end training tied to a specific language or dataset.
Approach: They propose a universal approach for aligning spoken language text with corresponding timestamps to sign language videos using a lightweight dynamic programming procedure.
Outcome: The proposed method can be used on four sign language datasets and is highly efficient on CPU.
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage.
Approach: They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages.
Outcome: The proposed index covers 120 resources across 35 sign languages.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
BinaryAlign: Word Alignment as Binary Sequence Labeling (2024.acl-long)

Copied to clipboard

Challenge: State-of-the-art word alignment training methods require a different class depending on the availability of gold data for a particular language pair.
Approach: They propose a novel word alignment technique based on binary sequence labeling that outperforms existing approaches in both scenarios.
Outcome: The proposed method outperforms existing models on non-English language pairs and performs stratified error analysis over alignment error type.
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources .
Approach: They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment .
Outcome: The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language.
SwissSLi: The Multi-parallel Sign Language Corpus for Switzerland (2024.lrec-main)

Copied to clipboard

Challenge: Using a CC BY-NC-SA 4.0 license, this corpus contains parallel sign language videos and spoken language subtitles.
Approach: They introduce SwissSLi, the first sign language corpus that contains parallel data of all three Swiss sign languages.
Outcome: The proposed corpus contains parallel sign language videos and spoken language subtitles.
Challenges with Sign Language Datasets for Sign Language Recognition and Translation (2022.lrec-1)

Copied to clipboard

Challenge: Sign Languages are the primary means of communication for at least half a million people in Europe . however, the development of SL recognition and translation tools is slowed down by resource scarcity and data formats are not suitable for machine learning.
Approach: They propose a framework to unify available resources and facilitate SL research for different languages.
Outcome: The proposed framework is based on a set of ELAN files and returns textual and visual data ready to train SL recognition and translation models.
Word Alignment by Fine-tuning Embeddings on Parallel Corpora (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on word alignment has focused on unsupervised learning on parallel text.
Approach: They propose to combine pre-trained contextualized word embeddings with multilingually trained language models to achieve competitive results on word alignment tasks.
Outcome: The proposed model outperforms state-of-the-art models on five language pairs and can train multilingual word aligners that can obtain robust performance on different language pairs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations