| Challenge: | a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones. |
| Approach: | They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences. |
| Outcome: | The proposed set of simplified sentences is a good quality data set for machine learning. |
Similar Papers
Simplified Corpus with Core Vocabulary (L18-1)
Copied to clipboard
| Challenge: | a study has found that simple Japanese is more accessible to foreigners than English. |
| Approach: | They have constructed a simplified corpus for the Japanese language and selected the core vocabulary. |
| Outcome: | The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa. |
A Document-Level Text Simplification Dataset for Japanese (2024.lrec-main)
Copied to clipboard
| Challenge: | Document-level text simplification tasks combine summarization and intra-sentence simplification. |
| Approach: | They devised a Japanese document-level text simplification dataset based on newspaper articles and Wikipedia. |
| Outcome: | The proposed dataset compared Japanese document-level text simplification models with English models and newspaper articles. |
Evaluation Dataset for Japanese Medical Text Simplification (2024.naacl-srw)
Copied to clipboard
| Challenge: | Existing studies on medical text simplification in English have not been well explored in Japanese because of the lack of a parallel corpus of this domain. |
| Approach: | They propose a lexically constrained reranking method that allows to avoid technical terms to be output. |
| Outcome: | The proposed method improves on the weblogs of Japanese patients and reduces the need for a training corpus. |
Word Complexity Estimation for Japanese Lexical Simplification (2020.lrec-1)
Copied to clipboard
| Challenge: | Experimental results show that the proposed method achieves the highest performance of Japanese lexical simplification. |
| Approach: | They propose a large-scale word complexity lexicon, a synonym lexicone and a toolkit for developing and benchmarking Japanese lexical simplification systems. |
| Outcome: | The proposed method achieves the highest performance of Japanese lexical simplification. |
Text Simplification from Professionally Produced Corpora (L18-1)
Copied to clipboard
| Challenge: | Existing approaches to Text Simplification rely on the Wikipedia-Simple Wikipedia parallel corpus, which is used for many tasks. |
| Approach: | They propose to use the Newsela corpus to extract 550, 644 complex-simple sentence pairs from the corpus and introduce a lexical simplifier that uses the corpu to generate candidate simplifications. |
| Outcome: | The proposed model outperforms state-of-the-art approaches and generates candidate simplifications from the newsela corpus. |
A Japanese News Simplification Corpus with Faithfulness (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing simplified corpora lack faithfulness to original text, resulting in errors in translation. |
| Approach: | They propose to simplify Japanese newspaper articles to prioritize faithfulness over automated models. |
| Outcome: | The proposed corpus preserves the original text, surpassing existing corpora. |
Multi-Word Lexical Simplification (2020.coling-main)
Copied to clipboard
| Challenge: | In text simplification, individual words are replaced with their simpler equivalents, but single word substitutions do not cover the full complexity of techniques humans use to approach text simulating. |
| Approach: | They propose a task of multi-word lexical simplification in which a sentence is made easier to understand by replacing its fragment with a simpler alternative. |
| Outcome: | The proposed method is based on a purpose-trained neural language model and evaluates against human and resource-based baselines. |
SimPA: A Sentence-Level Simplification Corpus for the Public Administration Domain (L18-1)
Copied to clipboard
| Challenge: | lexical simplification is the task of reducing lexically and/or structural complexity of texts. |
| Approach: | They propose to collect manual simplifications for 1,100 original sentences using a sentence-level corpus from the Public Administration domain. |
| Outcome: | The proposed corpus contains 1,100 original sentences with manual simplifications collected through a two-stage process. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)
Copied to clipboard
Hanae Koiso, Yasuharu Den, Yuriko Iseki, Wakako Kashino, Yoshiko Kawabata, Ken’ya Nishikawa, Yayoi Tanaka, Yasuyuki Usuda
| Challenge: | a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations . |
| Approach: | They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner. |
| Outcome: | The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings. |