COMICORDA: Dialogue Act Recognition in Comic Books (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on dialogue act recognition from images is limited to speech balloon segmentation and optical character recognition.
Approach: They propose a novel DA recognition approach for comic books using speech balloon segmentation, optical character recognition and DA classification.
Outcome: The proposed method achieves 98% average precision for speech balloon segmentation and exceeds 70% accuracy for the DA recognition task.

Similar Papers

Speaker Turn Modeling for Dialogue Act Classification (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to DA classification model utterances without incorporating the turn changes among speakers throughout the dialogue, thus treating it no different than non-interactive written text.
Approach: They propose to integrate the turn changes in conversations among speakers when modeling DAs by learning conversation-invariant speaker turn embeddings to represent speaker turns in a conversation.
Outcome: The proposed model captures semantics from the dialogue content while accounting for different speaker turns in a conversation.
Two-level classification for dialogue act recognition in task-oriented dialogues (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for dialogue act classification are limited and feature sets are low . recognizing dialogue acts is useful for identifying type of information and knowledge to be conveyed .
Approach: They propose a 2-level classification technique, distinguishing between generic and specific dialogue acts (DA) they propose an efficient approach for specific DA, based on high-level linguistic features.
Outcome: The proposed method outperforms classical methods for DA classification by including high-level features.
ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents (C18-1)

Copied to clipboard

Challenge: Existing methods for DA annotation are incompatible with each other and do not cover all aspects necessary for open-domain human-machine interaction.
Approach: They propose to map publicly available corpora to a subset of the ISO standard and create a task-independent training corpus for DA classification.
Outcome: The proposed method can train a domain-independent DA tagger on out-of-domain conversational data and achieve robustness across different DA categories.
Hierarchical Fusion for Online Multimodal Dialog Act Classification (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing multimodal DA classification approaches are limited by ineffective audio modeling and late-stage fusion.
Approach: They propose a framework for online multimodal dialog act (DA) classification based on raw audio and ASR-generated transcriptions of current and past utterances.
Outcome: The proposed model achieves a significant increase in the F1 score relative to current state-of-the-art models on two prominent DA classification datasets, MRDA and EMOTyDA.
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to training or evaluating non-English dialogue datasets often introduce artifacts that reduce their naturalness and cultural appropriateness.
Approach: They propose a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations.
Outcome: The proposed framework outperforms translation models in Italian, German, and Chinese on cultural relevance, coherence, and situational appropriateness.
ComicScene154: A Scene Dataset for Comic Analysis (2025.emnlp-main)

Copied to clipboard

Challenge: Comics offer compelling yet under-explored domain for computational narrative analysis . authors highlight potential of comics for narrative-driven, multimodal data analysis based on novel comics .
Approach: They propose a dataset of scene-level narrative arcs derived from comic books . they highlight their potential to inform broader research on multimodal storytelling .
Outcome: The dataset provides an initial benchmark that future studies can build upon.
Dialogue Act Classification with Context-Aware Self-Attention (N19-1)

Copied to clipboard

Challenge: Recent work in Dialogue Act classification has treated the task as a sequence labeling problem using hierarchical deep neural networks.
Approach: They propose a hierarchical deep neural network to model different levels of utterance and dialogue act semantics and use contextual dependencies to improve performance.
Outcome: The proposed model improves on the Switchboard Dialogue Act Corpus while maintaining high accuracy.
What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition (2021.tacl-1)

Copied to clipboard

Challenge: Existing punctuation in the transcripts has a massive effect on the models’ performance, and specific label set specificity does not affect dialog act segmentation performance.
Approach: They apply two pre-trained transformer models to a conversation transcript as a sequence of dialog acts and achieve strong results on Switchboard Dialog Act and Meeting Recorder Dialog Act corpora.
Outcome: The proposed models achieve 8.4% and 14.2% error rates on the Switchboard Dialog Act and Meeting Recorder Dialog Act corpora.
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 TV and movie series is filled with challenging multi-party dialogues.
Approach: They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues.
Outcome: The proposed dataset is a step towards better multi-party dialogue structuring and understanding.
Augmenting Small Data to Classify Contextualized Dialogue Acts for Exploratory Visualization (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of conversations is being developed to support data visualization exploration . we use data augmentation to improve our methods for dialogue act classification .
Approach: They propose to use a corpus of conversations to annotate contextualized dialogue acts . they highlight how thinking aloud affects interpretation of dialogue acts in the context .
Outcome: The proposed AI can support visualization exploration with a small corpus of conversations . the proposed AI outperforms existing models in terms of performance and performance .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations