A Japanese Corpus for Analyzing Customer Loyalty Information (L18-1)

Copied to clipboard

Challenge: a corpus of customer loyalty information is used to analyze customer's voice . a variety of studies have focused on analyzing attitudes, opinions, sentiments of text data .
Approach: They present a corpus for analyzing customer loyalty information . they use voice of customer to capture customer's behaviors, needs and feedbacks .
Outcome: The proposed corpus analyzes customer loyalty by analyzing their voice . the study is based on a corpus of customer loyalty information .

Similar Papers

User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization (2021.naacl-main)

Copied to clipboard

Challenge: Morphological analysis (MA) and lexical normalization (LN) are important tasks for Japanese user-generated text.
Approach: They construct a publicly available Japanese UGT corpus annotated with morphological and normalization information.
Outcome: The proposed corpus shows low performance for non-general words and non-standard forms . morphological analysis is an important task in Japanese user-generated text .
A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog Domain (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently.
Approach: They extend the WRIME dataset with basic emotion intensity from both the writer's subjective and reader's perspective to include the Japanese sentiment polarity.
Outcome: The proposed dataset is the first large-scale corpus to annotate both basic emotions and sentiment polarity labels from both the writer’s and reader’s perspectives.
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts.
Approach: They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent .
Outcome: The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora .
Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)

Copied to clipboard

Challenge: a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks .
Approach: They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs.
Outcome: The proposed method is the most accurate and leads to lesser performance in downstream tasks.
A Large-Scale Japanese Dataset for Aspect-based Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Aspect-based sentiment analysis (ABSA) has not been explored in the Japanese language . there is no standard Japanese dataset available for ABSA task in the language - a paper by cnn.
Approach: They propose to use a Japanese aspect-based sentiment analysis dataset for hotel reviews domain . they propose to include 53,192 review sentences with seven aspect categories and two polarity labels .
Outcome: The proposed dataset contains 53,192 review sentences with seven aspect categories and two polarity labels.
JESC: Japanese-English Subtitle Corpus (L18-1)

Copied to clipboard

Challenge: Existing data on Japanese-English subtitles are limited due to the high cost of manual construction.
Approach: They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web.
Outcome: The JESC dataset covers the underrepresented domain of conversational dialogue.
Annotation and Quantitative Analysis of Speaker Information in Novel Conversation Sentences in Japanese (L18-1)

Copied to clipboard

Challenge: a qualitative lexicological analysis of conversation sentences in novels is performed . gender and age of conversation sentence information is not used as actual speech .
Approach: They performed a quantitative lexicological analysis using attributed speaker information . they also examined the differences between Japanese novels and translations of foreign novels .
Outcome: The results show that conversation sentences in novels are representative of spoken language . the authors conclude that conversation sentence data are not useful as speech materials .
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer (2024.lrec-main)

Copied to clipboard

Challenge: Document-level Relation Extraction (DocRE) is the task of extracting all semantic relationships from a document.
Approach: They propose to transfer an English document to Japanese to promote DocRE in other languages.
Outcome: The proposed model reduces the human edit steps by 50% compared with the previous approach.
A Corpus Study and Annotation Schema for Named Entity Recognition and Relation Extraction of Business Products (L18-1)

Copied to clipboard

Challenge: Existing annotation guidelines for non-standard entity types and relations are lacking in news and forum texts.
Approach: They propose a corpus study and an annotation schema for the annotation of product entity and company-product relation mentions.
Outcome: The proposed annotation schema and guidelines are applied to the annotation of product entities and company-product relation mentions.
Sudachi: a Japanese Tokenizer for Business (L18-1)

Copied to clipboard

Challenge: Lack of token unit compatibility is one of the critical problems of Japanese language resources.
Approach: They develop a Japanese tokenizer called Sudachi and its accompanying dictionary . they use multi-granular output and normalization of notation variations to improve tokenization .
Outcome: The proposed tokenizer and dictionary improve tokenization in Japanese for business use.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations