Papers by Hitoshi Ito

4 papers
Content-Equivalent Translated Parallel News Corpus and Extension of Domain Adaptation for NMT (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to train NMT systems with noisy data are not sufficient . a recent increase in foreigners visiting Japan has created a significant information gap .
Approach: They propose a Japanese-English parallel news corpus that is content-equivalent . they extend a domain-adaptation method to train NMT models with clean corpus .
Outcome: The proposed corpus improves translation quality and is more effective than existing methods.
Neural Machine Translation System using a Content-equivalently Translated Parallel Corpus for the Newswire Translation Tasks at WAT 2019 (D19-52)

Copied to clipboard

Challenge: In addition to the JIJI Corpus, we developed a corpus of 0.22M sentence pairs by manually, translating Japanese news sentences into English content- equivalently.
Approach: They propose to use JIJI Corpus and Equivalent-style sentences to translate Japanese news sentences into English content- equivalently.
Outcome: The proposed translation models achieved the best human evaluation scores in the newswire translation tasks at WAT 2019 . they used the JIJI Corpus, which was provided by the task organizer, and the Equivalent-style translation model to translate Japanese news sentences into English content- equivalently.
Context-Driven and Reference-Guided Data Augmentation for Subtitle Translation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated strong performance in translation tasks.
Approach: They propose a method that expands source-side data by rewriting original subtitles using information that can be extracted from the context, such as character profiles and scene descriptions.
Outcome: The proposed method improves BLEU scores for film subtitle translation and achieves superior stylistic quality in human evaluation.
Effective Use of Target-side Context for Neural Machine Translation (2020.coling-main)

Copied to clipboard

Challenge: Existing methods to train NMT systems with noisy data are not sufficient . et al., 2018) found that NMT models can learn with multiple types of corpora .
Approach: They propose a Japanese-English news corpus that is content-equivalent . they extend a domain-adaptation method to train NMT models with clean corpus .
Outcome: The proposed corpus improves translation quality and is more efficient than existing methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations