JESC: Japanese-English Subtitle Corpus (L18-1)

Copied to clipboard

Challenge: Existing data on Japanese-English subtitles are limited due to the high cost of manual construction.
Approach: They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web.
Outcome: The JESC dataset covers the underrepresented domain of conversational dialogue.

Similar Papers

JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them.
Approach: They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus .
Outcome: The proposed corpus includes a broader range of domains and can be trained with a pre-trained model.
A-TASC: Asian TED-Based Automatic Subtitling Corpus (2025.acl-long)

Copied to clipboard

Challenge: Existing AS corpora and primary metric SubER focus on European languages.
Approach: They propose an Asian TED-based automatic subtitling corpus derived from English TED Talks and a modification of SubER to enable reliable evaluation of subtitle quality for languages without explicit word boundaries.
Outcome: The proposed corpus is based on TED Talks audio segments, transcripts, and subtitles in Chinese, Japanese, Korean, and Vietnamese.
SumTitles: a Summarization Dataset with Low Extractiveness (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for extractive summarization of dialogue data are limited by the grammar and structure of the utterances used.
Approach: They propose a low-extractive corpus of movie dialogues for abstractive text summarization . they use an alignment algorithm to construct the corpus and a baseline evaluation .
Outcome: The proposed method is low-extractive and shows high performance in dialogue datasets.
Simplified Corpus with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a study has found that simple Japanese is more accessible to foreigners than English.
Approach: They have constructed a simplified corpus for the Japanese language and selected the core vocabulary.
Outcome: The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa.
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models.
Approach: They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 .
Outcome: The proposed corpus boosts the accuracy of machine translation models on various domains.
JCoLA: Japanese Corpus of Linguistic Acceptability (2024.lrec-main)

Copied to clipboard

Challenge: Neural language models have exhibited outstanding performance in downstream tasks, yet there is limited understanding regarding the extent of their internalization of syntactic knowledge.
Approach: They introduce a dataset that analyzes sentences annotated with binary acceptability judgments from linguistic textbooks and handbooks and splits them into in-domain and out-of-domain data.
Outcome: The proposed datasets show that models can surpass human performance for in-domain data while no models can exceed human performance on out-of-domain datasets.
A Document-Level Text Simplification Dataset for Japanese (2024.lrec-main)

Copied to clipboard

Challenge: Document-level text simplification tasks combine summarization and intra-sentence simplification.
Approach: They devised a Japanese document-level text simplification dataset based on newspaper articles and Wikipedia.
Outcome: The proposed dataset compared Japanese document-level text simplification models with English models and newspaper articles.
Crowdsourced Corpus of Sentence Simplification with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones.
Approach: They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences.
Outcome: The proposed set of simplified sentences is a good quality data set for machine learning.
J-Shuwa: A Large-Scale Web-Collected Japanese Sign Language-Japanese Parallel Corpus (2026.findings-acl)

Copied to clipboard

Challenge: Japanese Sign Language (JSL) is a low-resource sign language that has received limited attention in the AI community due to the lack of large-scale, publicly available parallel corpora.
Approach: They propose a large-scale JSL-Japanese parallel corpus constructed from YouTube videos with hard-coded subtitles and closed captions.
Outcome: The proposed model is effective for training models and can be used for future research across a wide range of tasks.
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has demonstrated that patent translation accuracy improves as the amount of training data or the number of model parameters increases.
Approach: They construct a bilingual corpus of Japanese-English patent application data from 2000 to 2021 . they extracted 1.4M Japanese- English document pairs and extracted 350M sentence pairs .
Outcome: The proposed method improves translation accuracy by 20 bleu points . it is the first publicly available large-scale Japanese-English patent corpus .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations