Crowdsourced Corpus of Sentence Simplification with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones.
Approach: They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences.
Outcome: The proposed set of simplified sentences is a good quality data set for machine learning.

Similar Papers

Simplified Corpus with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a study has found that simple Japanese is more accessible to foreigners than English.
Approach: They have constructed a simplified corpus for the Japanese language and selected the core vocabulary.
Outcome: The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa.
A Document-Level Text Simplification Dataset for Japanese (2024.lrec-main)

Copied to clipboard

Challenge: Document-level text simplification tasks combine summarization and intra-sentence simplification.
Approach: They devised a Japanese document-level text simplification dataset based on newspaper articles and Wikipedia.
Outcome: The proposed dataset compared Japanese document-level text simplification models with English models and newspaper articles.
Evaluation Dataset for Japanese Medical Text Simplification (2024.naacl-srw)

Copied to clipboard

Challenge: Existing studies on medical text simplification in English have not been well explored in Japanese because of the lack of a parallel corpus of this domain.
Approach: They propose a lexically constrained reranking method that allows to avoid technical terms to be output.
Outcome: The proposed method improves on the weblogs of Japanese patients and reduces the need for a training corpus.
Word Complexity Estimation for Japanese Lexical Simplification (2020.lrec-1)

Copied to clipboard

Challenge: Experimental results show that the proposed method achieves the highest performance of Japanese lexical simplification.
Approach: They propose a large-scale word complexity lexicon, a synonym lexicone and a toolkit for developing and benchmarking Japanese lexical simplification systems.
Outcome: The proposed method achieves the highest performance of Japanese lexical simplification.
Text Simplification from Professionally Produced Corpora (L18-1)

Copied to clipboard

Challenge: Existing approaches to Text Simplification rely on the Wikipedia-Simple Wikipedia parallel corpus, which is used for many tasks.
Approach: They propose to use the Newsela corpus to extract 550, 644 complex-simple sentence pairs from the corpus and introduce a lexical simplifier that uses the corpu to generate candidate simplifications.
Outcome: The proposed model outperforms state-of-the-art approaches and generates candidate simplifications from the newsela corpus.
A Japanese News Simplification Corpus with Faithfulness (2024.lrec-main)

Copied to clipboard

Challenge: Existing simplified corpora lack faithfulness to original text, resulting in errors in translation.
Approach: They propose to simplify Japanese newspaper articles to prioritize faithfulness over automated models.
Outcome: The proposed corpus preserves the original text, surpassing existing corpora.
Multi-Word Lexical Simplification (2020.coling-main)

Copied to clipboard

Challenge: In text simplification, individual words are replaced with their simpler equivalents, but single word substitutions do not cover the full complexity of techniques humans use to approach text simulating.
Approach: They propose a task of multi-word lexical simplification in which a sentence is made easier to understand by replacing its fragment with a simpler alternative.
Outcome: The proposed method is based on a purpose-trained neural language model and evaluates against human and resource-based baselines.
SimPA: A Sentence-Level Simplification Corpus for the Public Administration Domain (L18-1)

Copied to clipboard

Challenge: lexical simplification is the task of reducing lexically and/or structural complexity of texts.
Approach: They propose to collect manual simplifications for 1,100 original sentences using a sentence-level corpus from the Public Administration domain.
Outcome: The proposed corpus contains 1,100 original sentences with manual simplifications collected through a two-stage process.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations