Challenge: Existing studies on continuous sign language recognition (CSLR) use offline models with high latency and memory usage.
Approach: They develop a sign dictionary and train an isolated sign language recognition model on the dictionary.
Outcome: The proposed model achieves state-of-the-art on three popular benchmarks across task settings.

Similar Papers

Improving Sign Recognition with Phonology (2023.eacl-main)

Copied to clipboard

Challenge: Existing work does not consider sign language phonology, but none leverages it . a recent study has shown that sign language recognition models lack structure .
Approach: They explicitly recognize the role of phonology in sign production to train models for isolated sign language recognition . they train models that take in pose estimations of a signer producing a single sign to predict its phonological characteristics .
Outcome: The proposed model improves sign recognition accuracy by 9% on the WLASL benchmark . the study could accelerate linguistic research in the domain of signed languages .
OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across Languages (2022.acl-long)

Copied to clipboard

Challenge: a new study examines the performance of pretraining for sign language recognition in low-resource settings.
Approach: They propose using pose extracted through pretrained models as the standard modality of data to reduce training time and enable efficient inference.
Outcome: The proposed model reduces training time and allows efficient inference in sign languages.
Improvement in Sign Language Translation Using Text CTC Alignment (2025.coling-main)

Copied to clipboard

Challenge: Current sign language translation (SLT) approaches rely on gloss-based supervision with Connectionist Temporal Classification (CTC) limiting their ability to handle non-monotonic alignments between sign language video and spoken text.
Approach: They propose a method that integrates CTC/Attention with the attention mechanism during decoding and integrates it with the sign language video and spoken text.
Outcome: The proposed method outperforms the pure-attention baseline and achieves comparable results to state-of-the-art methods.
BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to connectionist temporal classification (CTC) are based on pre-trained language models (LMs)
Approach: They propose a formulation of connectionist temporal classification that relaxes the conditional independence assumptions used in conventional CTC and incorporates linguistic knowledge through explicit output dependency.
Outcome: The proposed model improves over conventional approaches across variations in speaking styles and languages while maintaining CTC’s training efficiency.
Open-Domain Sign Language Translation Learned from Online Video (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on sign language translation has focused mainly on data collected in controlled environments or domains, which limits its applicability to real-world settings.
Approach: They propose to use sign search as a pretext task and fusion of mouthing and handshape features to improve sign language translation in real-world settings.
Outcome: The proposed techniques produce consistent and large improvements over baseline models based on prior work.
Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition (L18-1)

Copied to clipboard

Challenge: a new general framework for sign recognition from monocular video is presented . the framework exploits state-of-the-art learning methods while incorporating features based on what we know about the linguistic composition of lexical signs.
Approach: They propose a general framework for sign recognition from monocular video . they exploit state-of-the-art learning methods while incorporating features from linguistic information .
Outcome: The proposed framework exploits state-of-the-art learning methods while incorporating features based on what we know about linguistic composition of lexical signs.
SignCLIP: Connecting Text and Sign Language by Contrastive Learning (2024.emnlp-main)

Copied to clipboard

Challenge: SignCLIP is an efficient method of learning useful visual representations for sign language processing from large-scale, multilingual video-text pairs without optimizing for a specific task or sign language of limited size.
Approach: They propose a method for learning visual representations for sign language processing from large-scale video-text pairs without directly optimizing for a specific task or sign language.
Outcome: The proposed model can learn from multilingual video-text pairs without optimizing for a specific task or sign language of limited size.
Continual Learning in Multilingual Sign Language Translation (2025.naacl-long)

Copied to clipboard

Challenge: Despite the low translation quality of sign language, many machine learning approaches are still in its infancy.
Approach: They propose to use continual learning for mul- tilingual SLT to improve translation quality.
Outcome: The proposed methods outperform baseline and fine-tuning approaches in sign language translation.
WLASL-LEX: a Dataset for Recognising Phonological Properties in American Sign Language (2022.acl-short)

Copied to clipboard

Challenge: Signed Language Processing (SLP) is a major form of NLP, but has been overlooked by the NLP community.
Approach: They leverage existing resources to construct a large-scale dataset of American Sign Language signs annotated with six different phonological properties.
Outcome: The proposed model outperforms existing approaches on signs unobserved during training.
CISLR: Corpus for Indian Sign Language Recognition (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on natural language processing has shown promising improvements in text classification, translation and generation in widely used spoken languages.
Approach: They propose a new Indian Sign Language corpus for word-level recognition using videos . they propose CISLR model that leverages resource rich American Sign Language to learn generalized features for improving Indian Sign language predictions.
Outcome: The proposed model improves word recognition in Indian Sign Language using video . it leverages resource rich American Sign Language to learn generalized features .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations