Simplified Corpus with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a study has found that simple Japanese is more accessible to foreigners than English.
Approach: They have constructed a simplified corpus for the Japanese language and selected the core vocabulary.
Outcome: The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa.

Similar Papers

Crowdsourced Corpus of Sentence Simplification with Core Vocabulary (L18-1)

Copied to clipboard

Challenge: a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones.
Approach: They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences.
Outcome: The proposed set of simplified sentences is a good quality data set for machine learning.
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts.
Approach: They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent .
Outcome: The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora .
JESC: Japanese-English Subtitle Corpus (L18-1)

Copied to clipboard

Challenge: Existing data on Japanese-English subtitles are limited due to the high cost of manual construction.
Approach: They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web.
Outcome: The JESC dataset covers the underrepresented domain of conversational dialogue.
User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization (2021.naacl-main)

Copied to clipboard

Challenge: Morphological analysis (MA) and lexical normalization (LN) are important tasks for Japanese user-generated text.
Approach: They construct a publicly available Japanese UGT corpus annotated with morphological and normalization information.
Outcome: The proposed corpus shows low performance for non-general words and non-standard forms . morphological analysis is an important task in Japanese user-generated text .
Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)

Copied to clipboard

Challenge: a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks .
Approach: They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs.
Outcome: The proposed method is the most accurate and leads to lesser performance in downstream tasks.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
Building a List of Synonymous Words and Phrases of Japanese Compound Verbs (L18-1)

Copied to clipboard

Challenge: Japanese is rich in compound verbs consisting of two verbs joined together.
Approach: They built a database of Japanese "Verb + Verb" compounds semi-automatically . they extracted Japanese compound verbs from corpus and found suitable clusters .
Outcome: The proposed database extracts synonymous expressions of Japanese compound verbs from corpus . it then links the results to the "Compound Verb Lexicon"
A Japanese News Simplification Corpus with Faithfulness (2024.lrec-main)

Copied to clipboard

Challenge: Existing simplified corpora lack faithfulness to original text, resulting in errors in translation.
Approach: They propose to simplify Japanese newspaper articles to prioritize faithfulness over automated models.
Outcome: The proposed corpus preserves the original text, surpassing existing corpora.
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models.
Approach: They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 .
Outcome: The proposed corpus boosts the accuracy of machine translation models on various domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations