Challenge: Japanese Sign Language (JSL) is a low-resource sign language that has received limited attention in the AI community due to the lack of large-scale, publicly available parallel corpora.
Approach: They propose a large-scale JSL-Japanese parallel corpus constructed from YouTube videos with hard-coded subtitles and closed captions.
Outcome: The proposed model is effective for training models and can be used for future research across a wide range of tasks.

Similar Papers

Deep JSLC: A Multimodal Corpus Collection for Data-driven Generation of Japanese Sign Language Expressions (L18-1)

Copied to clipboard

Challenge: Existing technologies for CG-supported data display are not able to depict all relevant features of a natural signing sequence such as facial expression, spatial references or inter-sign movement.
Approach: They collected a corpus of Japanese Sign Language sentences for deep neural network learning.
Outcome: The proposed model could be used to train language features in Japanese Sign Language (JSL)
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them.
Approach: They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus .
Outcome: The proposed corpus includes a broader range of domains and can be trained with a pre-trained model.
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has demonstrated that patent translation accuracy improves as the amount of training data or the number of model parameters increases.
Approach: They construct a bilingual corpus of Japanese-English patent application data from 2000 to 2021 . they extracted 1.4M Japanese- English document pairs and extracted 350M sentence pairs .
Outcome: The proposed method improves translation accuracy by 20 bleu points . it is the first publicly available large-scale Japanese-English patent corpus .
SwissSLi: The Multi-parallel Sign Language Corpus for Switzerland (2024.lrec-main)

Copied to clipboard

Challenge: Using a CC BY-NC-SA 4.0 license, this corpus contains parallel sign language videos and spoken language subtitles.
Approach: They introduce SwissSLi, the first sign language corpus that contains parallel data of all three Swiss sign languages.
Outcome: The proposed corpus contains parallel sign language videos and spoken language subtitles.
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models.
Approach: They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 .
Outcome: The proposed corpus boosts the accuracy of machine translation models on various domains.
JESC: Japanese-English Subtitle Corpus (L18-1)

Copied to clipboard

Challenge: Existing data on Japanese-English subtitles are limited due to the high cost of manual construction.
Approach: They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web.
Outcome: The JESC dataset covers the underrepresented domain of conversational dialogue.
JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries.
Approach: They propose a large and highly multilingual dataset for sign language translation: JWSign.
Outcome: The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers.
JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages (P19-1)

Copied to clipboard

Challenge: a shortage of parallel data in low-resource languages creates a bottleneck for cross-lingual transfer . a massive collection of parallel texts for over 300 diverse languages is our main contribution .
Approach: They propose a parallel corpus of over 300 languages with 100 thousand parallel sentences per language pair on average.
Outcome: The proposed dataset can be used to build cross-lingual word embeddings and multi-source part-of-speech projections.
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage.
Approach: They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages.
Outcome: The proposed index covers 120 resources across 35 sign languages.
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News (2024.lrec-main)

Copied to clipboard

Challenge: a new dataset is being developed to enrich resources for sign language research . the dataset is 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses and 2,850 Chinese characters or 18K Chinese words.
Approach: They introduce a new Hong Kong sign language dataset called TVB-HKSL-News . the dataset is collected from a TV news program and contains sign videos . they aim to support research in sign language recognition and translation .
Outcome: The proposed dataset supports sign language recognition and translation research in Hong Kong . it consists of 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses and 2,850 Chinese characters or 18K Chinese words .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations