Towards Online Continuous Sign Language Recognition and Translation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on continuous sign language recognition (CSLR) use offline models with high latency and memory usage. |
| Approach: | They develop a sign dictionary and train an isolated sign language recognition model on the dictionary. |
| Outcome: | The proposed model achieves state-of-the-art on three popular benchmarks across task settings. |
Similar Papers
Improving Sign Recognition with Phonology (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing work does not consider sign language phonology, but none leverages it . a recent study has shown that sign language recognition models lack structure . |
| Approach: | They explicitly recognize the role of phonology in sign production to train models for isolated sign language recognition . they train models that take in pose estimations of a signer producing a single sign to predict its phonological characteristics . |
| Outcome: | The proposed model improves sign recognition accuracy by 9% on the WLASL benchmark . the study could accelerate linguistic research in the domain of signed languages . |
OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across Languages (2022.acl-long)
Copied to clipboard
| Challenge: | a new study examines the performance of pretraining for sign language recognition in low-resource settings. |
| Approach: | They propose using pose extracted through pretrained models as the standard modality of data to reduce training time and enable efficient inference. |
| Outcome: | The proposed model reduces training time and allows efficient inference in sign languages. |
Improvement in Sign Language Translation Using Text CTC Alignment (2025.coling-main)
Copied to clipboard
| Challenge: | Current sign language translation (SLT) approaches rely on gloss-based supervision with Connectionist Temporal Classification (CTC) limiting their ability to handle non-monotonic alignments between sign language video and spoken text. |
| Approach: | They propose a method that integrates CTC/Attention with the attention mechanism during decoding and integrates it with the sign language video and spoken text. |
| Outcome: | The proposed method outperforms the pure-attention baseline and achieves comparable results to state-of-the-art methods. |
BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to connectionist temporal classification (CTC) are based on pre-trained language models (LMs) |
| Approach: | They propose a formulation of connectionist temporal classification that relaxes the conditional independence assumptions used in conventional CTC and incorporates linguistic knowledge through explicit output dependency. |
| Outcome: | The proposed model improves over conventional approaches across variations in speaking styles and languages while maintaining CTC’s training efficiency. |
Open-Domain Sign Language Translation Learned from Online Video (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on sign language translation has focused mainly on data collected in controlled environments or domains, which limits its applicability to real-world settings. |
| Approach: | They propose to use sign search as a pretext task and fusion of mouthing and handshape features to improve sign language translation in real-world settings. |
| Outcome: | The proposed techniques produce consistent and large improvements over baseline models based on prior work. |
Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition (L18-1)
Copied to clipboard
| Challenge: | a new general framework for sign recognition from monocular video is presented . the framework exploits state-of-the-art learning methods while incorporating features based on what we know about the linguistic composition of lexical signs. |
| Approach: | They propose a general framework for sign recognition from monocular video . they exploit state-of-the-art learning methods while incorporating features from linguistic information . |
| Outcome: | The proposed framework exploits state-of-the-art learning methods while incorporating features based on what we know about linguistic composition of lexical signs. |
SignCLIP: Connecting Text and Sign Language by Contrastive Learning (2024.emnlp-main)
Copied to clipboard
| Challenge: | SignCLIP is an efficient method of learning useful visual representations for sign language processing from large-scale, multilingual video-text pairs without optimizing for a specific task or sign language of limited size. |
| Approach: | They propose a method for learning visual representations for sign language processing from large-scale video-text pairs without directly optimizing for a specific task or sign language. |
| Outcome: | The proposed model can learn from multilingual video-text pairs without optimizing for a specific task or sign language of limited size. |
Continual Learning in Multilingual Sign Language Translation (2025.naacl-long)
Copied to clipboard
| Challenge: | Despite the low translation quality of sign language, many machine learning approaches are still in its infancy. |
| Approach: | They propose to use continual learning for mul- tilingual SLT to improve translation quality. |
| Outcome: | The proposed methods outperform baseline and fine-tuning approaches in sign language translation. |
WLASL-LEX: a Dataset for Recognising Phonological Properties in American Sign Language (2022.acl-short)
Copied to clipboard
| Challenge: | Signed Language Processing (SLP) is a major form of NLP, but has been overlooked by the NLP community. |
| Approach: | They leverage existing resources to construct a large-scale dataset of American Sign Language signs annotated with six different phonological properties. |
| Outcome: | The proposed model outperforms existing approaches on signs unobserved during training. |
CISLR: Corpus for Indian Sign Language Recognition (2022.emnlp-main)
Copied to clipboard
Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi
| Challenge: | Existing work on natural language processing has shown promising improvements in text classification, translation and generation in widely used spoken languages. |
| Approach: | They propose a new Indian Sign Language corpus for word-level recognition using videos . they propose CISLR model that leverages resource rich American Sign Language to learn generalized features for improving Indian Sign language predictions. |
| Outcome: | The proposed model improves word recognition in Indian Sign Language using video . it leverages resource rich American Sign Language to learn generalized features . |