| Challenge: | spoken-to-written style conversion is becoming an important technology to increase the readability of ASR transcriptions. |
| Approach: | They propose to build a Japanese parallel corpus of spoken-to-written style conversions . they use crowdsourcing to convert spoken-style text into written-style texts . |
| Outcome: | The proposed corpus can handle general and specific spoken-to-written style conversion problems in Japanese. |
Similar Papers
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)
Copied to clipboard
| Challenge: | Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues. |
| Approach: | They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms. |
| Outcome: | The proposed corpus includes parallel text and speech data of 21 Japanese dialects. |
CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion Types (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets focus on a single type of spoken style, such as disfluencies. |
| Approach: | They propose a Chinese Spoken-to-Written style conversion dataset with 7,237 spoken sentences extracted from transcribed conversational texts. |
| Outcome: | The proposed dataset covers four major conversion problems corresponding to the majority of spoken styles. |
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them. |
| Approach: | They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus . |
| Outcome: | The proposed corpus includes a broader range of domains and can be trained with a pre-trained model. |
Designing the Business Conversation Corpus (D19-52)
Copied to clipboard
| Challenge: | Existing parallel corpora for machine translation of written text and monologues are limited. |
| Approach: | They propose to introduce a Japanese-English business conversation parallel corpus into machine translation training scenarios and show how it improves machine translation quality. |
| Outcome: | The proposed corpus is used in a Japanese-English business conversation training scenario and shows how it performs. |
Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer (N18-1)
Copied to clipboard
| Challenge: | a lack of training and evaluation datasets, benchmarks and automatic metrics has blocked progress in this field. |
| Approach: | They propose to use a grammarly's Yahoo Answers Formality corpus to create the largest corpus for a particular style . they also propose to apply machine translation metrics to the task . |
| Outcome: | The proposed model can be used to train and evaluate a text in a particular style . the proposed model is based on the existing model and can be applied to other tasks . |
Parallel Data Augmentation for Formality Style Transfer (2020.acl-main)
Copied to clipboard
| Challenge: | Formality style transfer is a task of automatically transforming text in one particular formality style into another. |
| Approach: | They propose to augment parallel data with three specific data augmentation methods to improve the model's generalization ability and reduce the overfitting risk. |
| Outcome: | The proposed methods significantly improve performance when used to pre-train the model and lead to the state-of-the-art results in the GYAFC benchmark dataset. |
An Empirical Study on Multi-Task Learning for Text Style Transfer and Paraphrase Generation (2020.coling-industry)
Copied to clipboard
Pawel Bujnowski, Kseniia Ryzhova, Hyungtak Choi, Katarzyna Witkowska, Jaroslaw Piersa, Tymoteusz Krumholc, Katarzyna Beksa
| Challenge: | a limited amount of style data is needed for text style transfer, but there are no convincing methods for evaluating them. |
| Approach: | They propose an efficient method for neutral-to-style transformation using the transformer framework. |
| Outcome: | The proposed method can train neutral-to-style transformation models using large paraphrases and a small style transfer corpus. |
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models. |
| Approach: | They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 . |
| Outcome: | The proposed corpus boosts the accuracy of machine translation models on various domains. |
A Parallel Corpus of Arabic-Japanese News Articles (L18-1)
Copied to clipboard
| Challenge: | a large-scale parallel corpora with manually verified subsets of sentences has been used for machine translation between major language pairs. |
| Approach: | They describe the creation process and statistics of the Arabic-Japanese portion of the TUFS Media Corpus . they also report the first results of Arabic-japanese phrase-based machine translation trained on the corpus based on the Arabic corpus. |
| Outcome: | The proposed corpus is a document-level parallel corpus and sentence-level parser corpus . it is the first time that Arabic-Japanese translations have been trained on it . |
Formality Style Transfer for Noisy, User-generated Conversations: Extracting Labeled, Parallel Data from Unlabeled Corpora (D19-55)
Copied to clipboard
| Challenge: | Typical datasets used for style transfer in NLP contain aligned pairs of two opposite extremes of a style. |
| Approach: | They propose a technique to derive a dataset of aligned pairs from an unlabeled corpus by using an auxiliary dataset, allowing for in-domain training. |
| Outcome: | The proposed method significantly outperforms OpenNMT’s Seq2Seq model trained on the Yahoo Formality Dataset and 6 novel datasets. |