Action-Concentrated Embedding Framework: This Is Your Captain Sign-tokening (2024.lrec-main)
Copied to clipboard
| Challenge: | ACE is a new sign token embedding framework that tracks a signer’s actions based on human posture estimation and captures the token embeds using a short-time Fourier transform. |
| Approach: | They propose a novel sign token embedding framework that tracks a signer’s actions based on human posture estimation and a dedicated notation system tailored for sign language. |
| Outcome: | The proposed framework outperforms previous studies in translation performance against a disaster sign language dataset and improves by up to 5.79% for BLEU-4 and 5.46% for ROUGE-L metric. |
Similar Papers
MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing solutions for sign language production are limited due to phonological differences and data scarcity. |
| Approach: | They propose a unified framework for continuous sign language production that generates sign predictions step by step from text or speech embeddings. |
| Outcome: | The proposed model achieves competitive performance on how2sign and PHOENIX14T datasets. |
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches for aligning spoken language text to sign language videos rely on end-to-end training tied to a specific language or dataset. |
| Approach: | They propose a universal approach for aligning spoken language text with corresponding timestamps to sign language videos using a lightweight dynamic programming procedure. |
| Outcome: | The proposed method can be used on four sign language datasets and is highly efficient on CPU. |
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual models have been released, but many of the world's languages are not covered. |
| Approach: | They propose a method that initializes the embedding matrix for a new tokenizer based on information in the source model's embeddable matrix. |
| Outcome: | The proposed method outperforms random initialization and previous work on language modeling and on a range of downstream tasks (NLI, QA, and NER). |
Sign Language Production With Avatar Layering: A Critical Use Case over Rare Words (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing vision-based sign language production approaches suffer from out-of-vocabulary (OOV) and test-time generalization problems. |
| Approach: | They propose an avatar-based sign language production system that generates sign language videos from spoken language expressions. |
| Outcome: | The proposed system achieves higher BLEU-4 and higher ROUGE-L scores on a new Korean-Korean sign language dataset. |
Explore More Guidance: A Task-aware Instruction Network for Sign Language Translation Enhanced with Data Augmentation (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies focus on the recognition step, while paying less attention to sign language translation. |
| Approach: | They propose a task-aware instruction network, namely TIN-SLT, for sign language translation, by introducing the isntruction module and the learning-based feature fuse strategy into a Transformer network. |
| Outcome: | The proposed system outperforms existing solutions on two benchmark datasets, PHOENIX-2014-T and ASLG-PC12, and outperformed previous best solutions by 1.65 and 1.42 in terms of BLEU-4. |
SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing work on sign language video processing focuses on the face, hands and body posture of the signer. |
| Approach: | They propose to learn the handshapes and rich facial expressions of sign languages in a self-supervised fashion by learning from individual frames rather than video sequences. |
| Outcome: | The proposed model is more efficient than previous work on sign language pre-training. |
Modeling Intensification for Sign Language Generation: A Computational Approach (2022.findings-acl)
Copied to clipboard
| Challenge: | End-to-end sign language generation models do not accurately represent prosody in sign language. |
| Approach: | They propose to model intensification in a data-driven manner to improve prosody in generated sign languages by modeling temporal and spatial variations. |
| Outcome: | The proposed models improve the prosody of generated sign languages by using data-driven models. |
Multilingual Gloss-free Sign Language Translation: Towards Building a Sign Language Foundation Model (2025.acl-short)
Copied to clipboard
| Challenge: | Existing studies focus on translating a single SL into a spoken language (one-to-one SLT) however, multilingual SLT remains unexplored due to language conflicts and alignment difficulties across SLs and spoken languages. |
| Approach: | They propose a multilingual gloss-free model that can be used to translate a single SL into a spoken language and generate a token-level SL identification and spoken text. |
| Outcome: | The proposed model supports 10 SLs and handles one-to-one, many-to-1, and many- to-many SLT tasks. |
Signer Diversity-driven Data Augmentation for Signer-Independent Sign Language Translation (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for sign language translation (SLT) rely on signer identity labels, which is often impractical and costly in real-world applications. |
| Approach: | They propose a signer diversity-driven data augmentation method that can generalize to signers not encountered during training. |
| Outcome: | The proposed method achieves state-of-the-art results without relying on signer identity labels. |
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage. |
| Approach: | They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages. |
| Outcome: | The proposed index covers 120 resources across 35 sign languages. |