| Challenge: | a corpus of customer loyalty information is used to analyze customer's voice . a variety of studies have focused on analyzing attitudes, opinions, sentiments of text data . |
| Approach: | They present a corpus for analyzing customer loyalty information . they use voice of customer to capture customer's behaviors, needs and feedbacks . |
| Outcome: | The proposed corpus analyzes customer loyalty by analyzing their voice . the study is based on a corpus of customer loyalty information . |
Similar Papers
User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization (2021.naacl-main)
Copied to clipboard
| Challenge: | Morphological analysis (MA) and lexical normalization (LN) are important tasks for Japanese user-generated text. |
| Approach: | They construct a publicly available Japanese UGT corpus annotated with morphological and normalization information. |
| Outcome: | The proposed corpus shows low performance for non-general words and non-standard forms . morphological analysis is an important task in Japanese user-generated text . |
A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog Domain (2022.lrec-1)
Copied to clipboard
Haruya Suzuki, Yuto Miyauchi, Kazuki Akiyama, Tomoyuki Kajiwara, Takashi Ninomiya, Noriko Takemura, Yuta Nakashima, Hajime Nagahara
| Challenge: | Existing studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently. |
| Approach: | They extend the WRIME dataset with basic emotion intensity from both the writer's subjective and reader's perspective to include the Japanese sentiment polarity. |
| Outcome: | The proposed dataset is the first large-scale corpus to annotate both basic emotions and sentiment polarity labels from both the writer’s and reader’s perspectives. |
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts. |
| Approach: | They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent . |
| Outcome: | The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora . |
Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)
Copied to clipboard
| Challenge: | a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks . |
| Approach: | They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs. |
| Outcome: | The proposed method is the most accurate and leads to lesser performance in downstream tasks. |
A Large-Scale Japanese Dataset for Aspect-based Sentiment Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) has not been explored in the Japanese language . there is no standard Japanese dataset available for ABSA task in the language - a paper by cnn. |
| Approach: | They propose to use a Japanese aspect-based sentiment analysis dataset for hotel reviews domain . they propose to include 53,192 review sentences with seven aspect categories and two polarity labels . |
| Outcome: | The proposed dataset contains 53,192 review sentences with seven aspect categories and two polarity labels. |
JESC: Japanese-English Subtitle Corpus (L18-1)
Copied to clipboard
| Challenge: | Existing data on Japanese-English subtitles are limited due to the high cost of manual construction. |
| Approach: | They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web. |
| Outcome: | The JESC dataset covers the underrepresented domain of conversational dialogue. |
Annotation and Quantitative Analysis of Speaker Information in Novel Conversation Sentences in Japanese (L18-1)
Copied to clipboard
| Challenge: | a qualitative lexicological analysis of conversation sentences in novels is performed . gender and age of conversation sentence information is not used as actual speech . |
| Approach: | They performed a quantitative lexicological analysis using attributed speaker information . they also examined the differences between Japanese novels and translations of foreign novels . |
| Outcome: | The results show that conversation sentences in novels are representative of spoken language . the authors conclude that conversation sentence data are not useful as speech materials . |
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer (2024.lrec-main)
Copied to clipboard
| Challenge: | Document-level Relation Extraction (DocRE) is the task of extracting all semantic relationships from a document. |
| Approach: | They propose to transfer an English document to Japanese to promote DocRE in other languages. |
| Outcome: | The proposed model reduces the human edit steps by 50% compared with the previous approach. |
A Corpus Study and Annotation Schema for Named Entity Recognition and Relation Extraction of Business Products (L18-1)
Copied to clipboard
| Challenge: | Existing annotation guidelines for non-standard entity types and relations are lacking in news and forum texts. |
| Approach: | They propose a corpus study and an annotation schema for the annotation of product entity and company-product relation mentions. |
| Outcome: | The proposed annotation schema and guidelines are applied to the annotation of product entities and company-product relation mentions. |
Sudachi: a Japanese Tokenizer for Business (L18-1)
Copied to clipboard
| Challenge: | Lack of token unit compatibility is one of the critical problems of Japanese language resources. |
| Approach: | They develop a Japanese tokenizer called Sudachi and its accompanying dictionary . they use multi-granular output and normalization of notation variations to improve tokenization . |
| Outcome: | The proposed tokenizer and dictionary improve tokenization in Japanese for business use. |