| Challenge: | Existing data on Japanese-English subtitles are limited due to the high cost of manual construction. |
| Approach: | They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web. |
| Outcome: | The JESC dataset covers the underrepresented domain of conversational dialogue. |
Similar Papers
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them. |
| Approach: | They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus . |
| Outcome: | The proposed corpus includes a broader range of domains and can be trained with a pre-trained model. |
A-TASC: Asian TED-Based Automatic Subtitling Corpus (2025.acl-long)
Copied to clipboard
| Challenge: | Existing AS corpora and primary metric SubER focus on European languages. |
| Approach: | They propose an Asian TED-based automatic subtitling corpus derived from English TED Talks and a modification of SubER to enable reliable evaluation of subtitle quality for languages without explicit word boundaries. |
| Outcome: | The proposed corpus is based on TED Talks audio segments, transcripts, and subtitles in Chinese, Japanese, Korean, and Vietnamese. |
SumTitles: a Summarization Dataset with Low Extractiveness (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for extractive summarization of dialogue data are limited by the grammar and structure of the utterances used. |
| Approach: | They propose a low-extractive corpus of movie dialogues for abstractive text summarization . they use an alignment algorithm to construct the corpus and a baseline evaluation . |
| Outcome: | The proposed method is low-extractive and shows high performance in dialogue datasets. |
Simplified Corpus with Core Vocabulary (L18-1)
Copied to clipboard
| Challenge: | a study has found that simple Japanese is more accessible to foreigners than English. |
| Approach: | They have constructed a simplified corpus for the Japanese language and selected the core vocabulary. |
| Outcome: | The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa. |
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models. |
| Approach: | They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 . |
| Outcome: | The proposed corpus boosts the accuracy of machine translation models on various domains. |
JCoLA: Japanese Corpus of Linguistic Acceptability (2024.lrec-main)
Copied to clipboard
| Challenge: | Neural language models have exhibited outstanding performance in downstream tasks, yet there is limited understanding regarding the extent of their internalization of syntactic knowledge. |
| Approach: | They introduce a dataset that analyzes sentences annotated with binary acceptability judgments from linguistic textbooks and handbooks and splits them into in-domain and out-of-domain data. |
| Outcome: | The proposed datasets show that models can surpass human performance for in-domain data while no models can exceed human performance on out-of-domain datasets. |
A Document-Level Text Simplification Dataset for Japanese (2024.lrec-main)
Copied to clipboard
| Challenge: | Document-level text simplification tasks combine summarization and intra-sentence simplification. |
| Approach: | They devised a Japanese document-level text simplification dataset based on newspaper articles and Wikipedia. |
| Outcome: | The proposed dataset compared Japanese document-level text simplification models with English models and newspaper articles. |
Crowdsourced Corpus of Sentence Simplification with Core Vocabulary (L18-1)
Copied to clipboard
| Challenge: | a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones. |
| Approach: | They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences. |
| Outcome: | The proposed set of simplified sentences is a good quality data set for machine learning. |
J-Shuwa: A Large-Scale Web-Collected Japanese Sign Language-Japanese Parallel Corpus (2026.findings-acl)
Copied to clipboard
| Challenge: | Japanese Sign Language (JSL) is a low-resource sign language that has received limited attention in the AI community due to the lack of large-scale, publicly available parallel corpora. |
| Approach: | They propose a large-scale JSL-Japanese parallel corpus constructed from YouTube videos with hard-coded subtitles and closed captions. |
| Outcome: | The proposed model is effective for training models and can be used for future research across a wide range of tasks. |
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | a recent study has demonstrated that patent translation accuracy improves as the amount of training data or the number of model parameters increases. |
| Approach: | They construct a bilingual corpus of Japanese-English patent application data from 2000 to 2021 . they extracted 1.4M Japanese- English document pairs and extracted 350M sentence pairs . |
| Outcome: | The proposed method improves translation accuracy by 20 bleu points . it is the first publicly available large-scale Japanese-English patent corpus . |