Challenge: In Japan, the number of single-person households is increasing, reducing opportunities for people to narrate.
Approach: They propose to collect 148,962 responsive utterances by listeners and annotate existing narrative speech with responsive . they also propose to use robots and smart speakers to listen to narratives .
Outcome: The proposed method can be used to annotate existing narrative speech with responsive utterances.

Similar Papers

Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.
Japanese Dialogue Corpus of Information Navigation and Attentive Listening Annotated with Extended ISO-24617-2 Dialogue Act Tags (L18-1)

Copied to clipboard

Challenge: Large-scale conventional dialogue corpora are mainly built for specified tasks with specially designed dialogue states.
Approach: They propose to annotate large-scale dialogue data with an extended ISO-24617-2 dialogue act tag-set to model a natural conversation with machines.
Outcome: The proposed corpus covers a wider range of dialogue tasks than existing task-oriented systems or text-chat systems.
InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (2025.findings-emnlp)

Copied to clipboard

Challenge: Spoken Dialogue models face challenges in handling nuanced interactional phenomena, such as interruptions and backchannels.
Approach: They propose to use a 150-hour English speech interaction dialogue dataset to empower spoken dialogue models with nuanced real-time interaction capabilities.
Outcome: The proposed dataset trains and evaluates a speech understanding model that classifies key interactional events directly from audio.
Is a Knowledge-based Response Engaging?: An Analysis on Knowledge-Grounded Dialogue with Information Source Annotation (2023.acl-srw)

Copied to clipboard

Challenge: Currently, most knowledge-grounded dialogue models focus on reflecting given external knowledge.
Approach: They analyze human behavior by annotating utterances in an existing knowledge-grounded dialogue corpus and find that speaker-derived information improves dialogue engagingness.
Outcome: The proposed model cannot include speaker-derived information as often as humans do.
Self-Contained Utterance Description Corpus for Japanese Dialog (2022.lrec-1)

Copied to clipboard

Challenge: Existing task frameworks for dialog-act classification and slot filling can only interpret utterances using pre-defined types and slots.
Approach: They propose a task to describe the intent of an utterance in a dialog with multiple simple natural sentences without the context.
Outcome: The proposed task can describe the intent of an utterance in a dialog with multiple simple natural sentences without the context.
More Diverse Dialogue Datasets via Diversity-Informed Data Collection (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to generate conversational dialogue produce uninteresting, predictable responses.
Approach: They propose a method to collect and determine more diverse data from conversational participants . they use dynamically computed corpus-level statistics to determine which conversational participant to collect data from .
Outcome: The proposed method produces significantly more diverse data than baseline methods and better results on emotion classification and dialogue generation tasks.
Annotation and Analysis of Extractive Summaries for the Kyutech Corpus (L18-1)

Copied to clipboard

Challenge: Summarization of multi-party conversation requires corpora to analyze characteristics of conversations and construct a method for summary generation.
Approach: They propose to annotate a Japanese conversation corpus for a decision-making task . they compare extractive summarization methods with the annotated extractive summary .
Outcome: The proposed corpus is the first annotated for conversation summarization tasks and freely available to anyone.
Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Spoken dialogues lack explicit modeling of behavior traits that are often overlooked in language models . et al.: our work opens new possibilities for developing behaviorally-aware dialogue systems .
Approach: They propose a large-scale dataset with over 100K spoken dialogues (2,164 hours) they propose BeDLM, the first dialogue model capable of generating natural conversations .
Outcome: The proposed model outperforms baseline models in generating natural dialogues . the proposed model can generate natural conversations conditioned on behavioral and narrative contexts - a key feature of spoken language models .
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology (2022.acl-long)

Copied to clipboard

Challenge: Informal social interaction is the primordial home of human language.
Approach: They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future.
Outcome: The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure.
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations