J-Shuwa: A Large-Scale Web-Collected Japanese Sign Language-Japanese Parallel Corpus (2026.findings-acl)
Copied to clipboard
| Challenge: | Japanese Sign Language (JSL) is a low-resource sign language that has received limited attention in the AI community due to the lack of large-scale, publicly available parallel corpora. |
| Approach: | They propose a large-scale JSL-Japanese parallel corpus constructed from YouTube videos with hard-coded subtitles and closed captions. |
| Outcome: | The proposed model is effective for training models and can be used for future research across a wide range of tasks. |
Similar Papers
Deep JSLC: A Multimodal Corpus Collection for Data-driven Generation of Japanese Sign Language Expressions (L18-1)
Copied to clipboard
| Challenge: | Existing technologies for CG-supported data display are not able to depict all relevant features of a natural signing sequence such as facial expression, spatial references or inter-sign movement. |
| Approach: | They collected a corpus of Japanese Sign Language sentences for deep neural network learning. |
| Outcome: | The proposed model could be used to train language features in Japanese Sign Language (JSL) |
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them. |
| Approach: | They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus . |
| Outcome: | The proposed corpus includes a broader range of domains and can be trained with a pre-trained model. |
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | a recent study has demonstrated that patent translation accuracy improves as the amount of training data or the number of model parameters increases. |
| Approach: | They construct a bilingual corpus of Japanese-English patent application data from 2000 to 2021 . they extracted 1.4M Japanese- English document pairs and extracted 350M sentence pairs . |
| Outcome: | The proposed method improves translation accuracy by 20 bleu points . it is the first publicly available large-scale Japanese-English patent corpus . |
SwissSLi: The Multi-parallel Sign Language Corpus for Switzerland (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a CC BY-NC-SA 4.0 license, this corpus contains parallel sign language videos and spoken language subtitles. |
| Approach: | They introduce SwissSLi, the first sign language corpus that contains parallel data of all three Swiss sign languages. |
| Outcome: | The proposed corpus contains parallel sign language videos and spoken language subtitles. |
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models. |
| Approach: | They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 . |
| Outcome: | The proposed corpus boosts the accuracy of machine translation models on various domains. |
JESC: Japanese-English Subtitle Corpus (L18-1)
Copied to clipboard
| Challenge: | Existing data on Japanese-English subtitles are limited due to the high cost of manual construction. |
| Approach: | They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web. |
| Outcome: | The JESC dataset covers the underrepresented domain of conversational dialogue. |
JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries. |
| Approach: | They propose a large and highly multilingual dataset for sign language translation: JWSign. |
| Outcome: | The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers. |
JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages (P19-1)
Copied to clipboard
| Challenge: | a shortage of parallel data in low-resource languages creates a bottleneck for cross-lingual transfer . a massive collection of parallel texts for over 300 diverse languages is our main contribution . |
| Approach: | They propose a parallel corpus of over 300 languages with 100 thousand parallel sentences per language pair on average. |
| Outcome: | The proposed dataset can be used to build cross-lingual word embeddings and multi-source part-of-speech projections. |
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage. |
| Approach: | They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages. |
| Outcome: | The proposed index covers 120 resources across 35 sign languages. |
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News (2024.lrec-main)
Copied to clipboard
| Challenge: | a new dataset is being developed to enrich resources for sign language research . the dataset is 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses and 2,850 Chinese characters or 18K Chinese words. |
| Approach: | They introduce a new Hong Kong sign language dataset called TVB-HKSL-News . the dataset is collected from a TV news program and contains sign videos . they aim to support research in sign language recognition and translation . |
| Outcome: | The proposed dataset supports sign language recognition and translation research in Hong Kong . it consists of 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses and 2,850 Chinese characters or 18K Chinese words . |