JDCFC: A Japanese Dialogue Corpus with Feature Changes (L18-1)

Copied to clipboard

Challenge: Existing corpora focus on emotional expressions in conversations, but there are no large-scale corpors focusing on the relationships between emotions and utterances.
Approach: They propose a Japanese Feature Change Knowledge Base (JFCKB) that focuses on emotional expressions in conversations.
Outcome: The proposed corpus can recognize reasonableness of a given conversation.

Similar Papers

JFCKB: Japanese Feature Change Knowledge Base (L18-1)

Copied to clipboard

Challenge: constructing commonsense knowledge including connotational meanings is challenging . a recent study focused on denotation and connotations, but few studies focused on connotating meanings .
Approach: They propose to construct a Japanese knowledge base where arguments in event sentences are associated with feature changes caused by events.
Outcome: The proposed knowledge base is able to generate anaphora resolution tasks in Japanese . it is useful for computers to understand texts, but it is difficult to acquire it .
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.
Building a Dialogue Corpus Annotated with Expressed and Experienced Emotions (2022.acl-srw)

Copied to clipboard

Challenge: a human would recognize the emotion of an interlocutor and respond with an appropriate emotion, such as empathy and comfort.
Approach: They propose to build a dialogue corpus annotated with two kinds of emotions . they collect tweets and annotate them with the emotion they put into the utterance .
Outcome: The proposed method shows that it is difficult to recognize experienced emotions and multitask learning is effective.
JESC: Japanese-English Subtitle Corpus (L18-1)

Copied to clipboard

Challenge: Existing data on Japanese-English subtitles are limited due to the high cost of manual construction.
Approach: They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web.
Outcome: The JESC dataset covers the underrepresented domain of conversational dialogue.
Annotating Modality Expressions and Event Factuality for a Japanese Chess Commentary Corpus (L18-1)

Copied to clipboard

Challenge: In recent years, there has been a surge of interest in the natural language processing related to the real world . shogi commentaries are an interesting testbed for these tasks, but can be grounded in the game tree .
Approach: They propose to augment shogi commentaries with game states to generate a game commentary generator.
Outcome: The proposed system can be used to ground symbols and events with factuality . it can be compared with other systems to find out if a commentator is a human .
The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets for human-like dialogue tasks are deficient due to the complexity of human conversations.
Approach: They construct a large-scale Chinese E-commerce conversation corpus with 1 million dialogues, 20 million utterances, and 150 million words.
Outcome: The proposed dataset includes 1 million multi-turn dialogues, 20 million utterances, and 150 million words.
Self-Contained Utterance Description Corpus for Japanese Dialog (2022.lrec-1)

Copied to clipboard

Challenge: Existing task frameworks for dialog-act classification and slot filling can only interpret utterances using pre-defined types and slots.
Approach: They propose a task to describe the intent of an utterance in a dialog with multiple simple natural sentences without the context.
Outcome: The proposed task can describe the intent of an utterance in a dialog with multiple simple natural sentences without the context.
JCoLA: Japanese Corpus of Linguistic Acceptability (2024.lrec-main)

Copied to clipboard

Challenge: Neural language models have exhibited outstanding performance in downstream tasks, yet there is limited understanding regarding the extent of their internalization of syntactic knowledge.
Approach: They introduce a dataset that analyzes sentences annotated with binary acceptability judgments from linguistic textbooks and handbooks and splits them into in-domain and out-of-domain data.
Outcome: The proposed datasets show that models can surpass human performance for in-domain data while no models can exceed human performance on out-of-domain datasets.
A Method for Building a Commonsense Inference Dataset based on Basic Events (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to acquire commonsense are limited by the general-purpose language models.
Approach: They propose a method for building a commonsense inference dataset using crowdsourcing and automatic extraction from a corpus.
Outcome: The proposed method can solve 104k commonsense inference problems in a Japanese corpus with high accuracy, but low bias.
JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization (2022.emnlp-main)

Copied to clipboard

Challenge: e-commerce users express their needs using text, images, or videos . but detailed information provided by images is limited, and customer service systems cannot understand the intent of users without the input text.
Approach: They construct a large-scale multimodal multi-turn dialogue dataset from a mainstream Chinese E-commerce platform . the dataset contains about 246K dialogue sessions, 3M utterances, and 507K images .
Outcome: The proposed dataset contains 246K dialogue sessions, 3M utterances, 507K images . it also includes product knowledge bases and image category annotations .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations