Challenge: spoken-to-written style conversion is becoming an important technology to increase the readability of ASR transcriptions.
Approach: They propose to build a Japanese parallel corpus of spoken-to-written style conversions . they use crowdsourcing to convert spoken-style text into written-style texts .
Outcome: The proposed corpus can handle general and specific spoken-to-written style conversion problems in Japanese.

Similar Papers

CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)

Copied to clipboard

Challenge: Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues.
Approach: They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms.
Outcome: The proposed corpus includes parallel text and speech data of 21 Japanese dialects.
CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion Types (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets focus on a single type of spoken style, such as disfluencies.
Approach: They propose a Chinese Spoken-to-Written style conversion dataset with 7,237 spoken sentences extracted from transcribed conversational texts.
Outcome: The proposed dataset covers four major conversion problems corresponding to the majority of spoken styles.
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them.
Approach: They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus .
Outcome: The proposed corpus includes a broader range of domains and can be trained with a pre-trained model.
Designing the Business Conversation Corpus (D19-52)

Copied to clipboard

Challenge: Existing parallel corpora for machine translation of written text and monologues are limited.
Approach: They propose to introduce a Japanese-English business conversation parallel corpus into machine translation training scenarios and show how it improves machine translation quality.
Outcome: The proposed corpus is used in a Japanese-English business conversation training scenario and shows how it performs.
Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer (N18-1)

Copied to clipboard

Challenge: a lack of training and evaluation datasets, benchmarks and automatic metrics has blocked progress in this field.
Approach: They propose to use a grammarly's Yahoo Answers Formality corpus to create the largest corpus for a particular style . they also propose to apply machine translation metrics to the task .
Outcome: The proposed model can be used to train and evaluate a text in a particular style . the proposed model is based on the existing model and can be applied to other tasks .
Parallel Data Augmentation for Formality Style Transfer (2020.acl-main)

Copied to clipboard

Challenge: Formality style transfer is a task of automatically transforming text in one particular formality style into another.
Approach: They propose to augment parallel data with three specific data augmentation methods to improve the model's generalization ability and reduce the overfitting risk.
Outcome: The proposed methods significantly improve performance when used to pre-train the model and lead to the state-of-the-art results in the GYAFC benchmark dataset.
An Empirical Study on Multi-Task Learning for Text Style Transfer and Paraphrase Generation (2020.coling-industry)

Copied to clipboard

Challenge: a limited amount of style data is needed for text style transfer, but there are no convincing methods for evaluating them.
Approach: They propose an efficient method for neutral-to-style transformation using the transformer framework.
Outcome: The proposed method can train neutral-to-style transformation models using large paraphrases and a small style transfer corpus.
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models.
Approach: They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 .
Outcome: The proposed corpus boosts the accuracy of machine translation models on various domains.
A Parallel Corpus of Arabic-Japanese News Articles (L18-1)

Copied to clipboard

Challenge: a large-scale parallel corpora with manually verified subsets of sentences has been used for machine translation between major language pairs.
Approach: They describe the creation process and statistics of the Arabic-Japanese portion of the TUFS Media Corpus . they also report the first results of Arabic-japanese phrase-based machine translation trained on the corpus based on the Arabic corpus.
Outcome: The proposed corpus is a document-level parallel corpus and sentence-level parser corpus . it is the first time that Arabic-Japanese translations have been trained on it .
Formality Style Transfer for Noisy, User-generated Conversations: Extracting Labeled, Parallel Data from Unlabeled Corpora (D19-55)

Copied to clipboard

Challenge: Typical datasets used for style transfer in NLP contain aligned pairs of two opposite extremes of a style.
Approach: They propose a technique to derive a dataset of aligned pairs from an unlabeled corpus by using an auxiliary dataset, allowing for in-domain training.
Outcome: The proposed method significantly outperforms OpenNMT’s Seq2Seq model trained on the Yahoo Formality Dataset and 6 novel datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations