Construction of Responsive Utterance Corpus for Attentive Listening Response Production (2022.lrec-1)
Copied to clipboard
| Challenge: | In Japan, the number of single-person households is increasing, reducing opportunities for people to narrate. |
| Approach: | They propose to collect 148,962 responsive utterances by listeners and annotate existing narrative speech with responsive . they also propose to use robots and smart speakers to listen to narratives . |
| Outcome: | The proposed method can be used to annotate existing narrative speech with responsive utterances. |
Similar Papers
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)
Copied to clipboard
Hanae Koiso, Yasuharu Den, Yuriko Iseki, Wakako Kashino, Yoshiko Kawabata, Ken’ya Nishikawa, Yayoi Tanaka, Yasuyuki Usuda
| Challenge: | a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations . |
| Approach: | They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner. |
| Outcome: | The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings. |
Japanese Dialogue Corpus of Information Navigation and Attentive Listening Annotated with Extended ISO-24617-2 Dialogue Act Tags (L18-1)
Copied to clipboard
| Challenge: | Large-scale conventional dialogue corpora are mainly built for specified tasks with specially designed dialogue states. |
| Approach: | They propose to annotate large-scale dialogue data with an extended ISO-24617-2 dialogue act tag-set to model a natural conversation with machines. |
| Outcome: | The proposed corpus covers a wider range of dialogue tasks than existing task-oriented systems or text-chat systems. |
InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Spoken Dialogue models face challenges in handling nuanced interactional phenomena, such as interruptions and backchannels. |
| Approach: | They propose to use a 150-hour English speech interaction dialogue dataset to empower spoken dialogue models with nuanced real-time interaction capabilities. |
| Outcome: | The proposed dataset trains and evaluates a speech understanding model that classifies key interactional events directly from audio. |
Is a Knowledge-based Response Engaging?: An Analysis on Knowledge-Grounded Dialogue with Information Source Annotation (2023.acl-srw)
Copied to clipboard
| Challenge: | Currently, most knowledge-grounded dialogue models focus on reflecting given external knowledge. |
| Approach: | They analyze human behavior by annotating utterances in an existing knowledge-grounded dialogue corpus and find that speaker-derived information improves dialogue engagingness. |
| Outcome: | The proposed model cannot include speaker-derived information as often as humans do. |
Self-Contained Utterance Description Corpus for Japanese Dialog (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing task frameworks for dialog-act classification and slot filling can only interpret utterances using pre-defined types and slots. |
| Approach: | They propose a task to describe the intent of an utterance in a dialog with multiple simple natural sentences without the context. |
| Outcome: | The proposed task can describe the intent of an utterance in a dialog with multiple simple natural sentences without the context. |
More Diverse Dialogue Datasets via Diversity-Informed Data Collection (2020.acl-main)
Copied to clipboard
| Challenge: | Existing approaches to generate conversational dialogue produce uninteresting, predictable responses. |
| Approach: | They propose a method to collect and determine more diverse data from conversational participants . they use dynamically computed corpus-level statistics to determine which conversational participant to collect data from . |
| Outcome: | The proposed method produces significantly more diverse data than baseline methods and better results on emotion classification and dialogue generation tasks. |
Annotation and Analysis of Extractive Summaries for the Kyutech Corpus (L18-1)
Copied to clipboard
| Challenge: | Summarization of multi-party conversation requires corpora to analyze characteristics of conversations and construct a method for summary generation. |
| Approach: | They propose to annotate a Japanese conversation corpus for a decision-making task . they compare extractive summarization methods with the annotated extractive summary . |
| Outcome: | The proposed corpus is the first annotated for conversation summarization tasks and freely available to anyone. |
Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Spoken dialogues lack explicit modeling of behavior traits that are often overlooked in language models . et al.: our work opens new possibilities for developing behaviorally-aware dialogue systems . |
| Approach: | They propose a large-scale dataset with over 100K spoken dialogues (2,164 hours) they propose BeDLM, the first dialogue model capable of generating natural conversations . |
| Outcome: | The proposed model outperforms baseline models in generating natural dialogues . the proposed model can generate natural conversations conditioned on behavioral and narrative contexts - a key feature of spoken language models . |
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology (2022.acl-long)
Copied to clipboard
| Challenge: | Informal social interaction is the primordial home of human language. |
| Approach: | They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future. |
| Outcome: | The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure. |
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)
Copied to clipboard
| Challenge: | a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information. |
| Approach: | They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants. |
| Outcome: | The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants. |