Challenge: a new general framework for sign recognition from monocular video is presented . the framework exploits state-of-the-art learning methods while incorporating features based on what we know about the linguistic composition of lexical signs.
Approach: They propose a general framework for sign recognition from monocular video . they exploit state-of-the-art learning methods while incorporating features from linguistic information .
Outcome: The proposed framework exploits state-of-the-art learning methods while incorporating features based on what we know about linguistic composition of lexical signs.

Similar Papers

How to Align Multiple Signed Language Corpora for Better Sign-to-Sign Translations? (2025.naacl-long)

Copied to clipboard

Challenge: despite the growing need for advanced signing technologies, signed language resources remain scarce.
Approach: They propose a linguistically informed alignment algorithm that matches instances between signed languages . they compare similarities and differences across three signed languages to develop a model .
Outcome: The proposed algorithm performs well on automatic metrics for sign-to-sign translation and generation.
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage.
Approach: They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages.
Outcome: The proposed index covers 120 resources across 35 sign languages.
WLASL-LEX: a Dataset for Recognising Phonological Properties in American Sign Language (2022.acl-short)

Copied to clipboard

Challenge: Signed Language Processing (SLP) is a major form of NLP, but has been overlooked by the NLP community.
Approach: They leverage existing resources to construct a large-scale dataset of American Sign Language signs annotated with six different phonological properties.
Outcome: The proposed model outperforms existing approaches on signs unobserved during training.
Challenges with Sign Language Datasets for Sign Language Recognition and Translation (2022.lrec-1)

Copied to clipboard

Challenge: Sign Languages are the primary means of communication for at least half a million people in Europe . however, the development of SL recognition and translation tools is slowed down by resource scarcity and data formats are not suitable for machine learning.
Approach: They propose a framework to unify available resources and facilitate SL research for different languages.
Outcome: The proposed framework is based on a set of ELAN files and returns textual and visual data ready to train SL recognition and translation models.
SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale (2025.findings-acl)

Copied to clipboard

Challenge: Existing work on sign language video processing focuses on the face, hands and body posture of the signer.
Approach: They propose to learn the handshapes and rich facial expressions of sign languages in a self-supervised fashion by learning from individual frames rather than video sequences.
Outcome: The proposed model is more efficient than previous work on sign language pre-training.
MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production (2024.findings-acl)

Copied to clipboard

Challenge: Existing solutions for sign language production are limited due to phonological differences and data scarcity.
Approach: They propose a unified framework for continuous sign language production that generates sign predictions step by step from text or speech embeddings.
Outcome: The proposed model achieves competitive performance on how2sign and PHOENIX14T datasets.
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to sign language translation use gloss annotations as an intermediary . a new approach to use large language models and word embeddings to improve Gloss2Text translation is needed.
Approach: They propose to leverage large language models pre-trained on expansive and diverse corpora to improve Gloss2Text translation stage by using data augmentation and label-smoothing loss function.
Outcome: The proposed approach surpasses state-of-the-art methods on the PHOENIX Weather 2014T dataset . it shows that gloss annotations can be used to guide the translation process .
Modeling Intensification for Sign Language Generation: A Computational Approach (2022.findings-acl)

Copied to clipboard

Challenge: End-to-end sign language generation models do not accurately represent prosody in sign language.
Approach: They propose to model intensification in a data-driven manner to improve prosody in generated sign languages by modeling temporal and spatial variations.
Outcome: The proposed models improve the prosody of generated sign languages by using data-driven models.
Dynamic Feature Fusion for Sign Language Translation Using HyperNetworks (2025.findings-naacl)

Copied to clipboard

Challenge: Using RGB and keypoint streams, sign language translation is highly dependent on the brain's ability to process color, shape, and motion simultaneously.
Approach: They propose a hypernetwork-based fusion method that extracts salient features from RGB and keypoint streams and introduces self-distillation and SST contrastive learning to maintain feature advantages while aligning the global semantic space.
Outcome: The proposed method achieves state-of-the-art performance on two public sign language datasets, reducing model parameters by about two-thirds.
OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across Languages (2022.acl-long)

Copied to clipboard

Challenge: a new study examines the performance of pretraining for sign language recognition in low-resource settings.
Approach: They propose using pose extracted through pretrained models as the standard modality of data to reduce training time and enable efficient inference.
Outcome: The proposed model reduces training time and allows efficient inference in sign languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations