LIDA: Lightweight Interactive Dialogue Annotator (D19-3)

Copied to clipboard

Challenge: Dialogue systems are dependent on the quality of the data used to train them.
Approach: They propose to develop an annotation tool specifically for conversation data that handles the entire dialogue annotation pipeline from raw text to structured conversation data.
Outcome: The proposed tool handles the entire dialogue annotation pipeline from raw text to structured conversation data and has a dedicated interface to resolve inter-annotator disagreements.

Similar Papers

MATILDA - Multi-AnnoTator multi-language InteractiveLight-weight Dialogue Annotator (2021.eacl-demos)

Copied to clipboard

Challenge: MATILDA is the first multi-annotator, multi-language dialogue annotation tool . it allows the creation of corpora, the management of users, the annotation of dialogues, the quick adaptation of the user interface to any language and the resolution of interannotation disagreement.
Approach: They propose to use MATILDA to create corpora, manage users, and annotation dialogues.
Outcome: The proposed tool supports the full pipeline for dialogue annotation, and non-technical people can use it.
MIDAS: A Dialog Act Annotation Scheme for Open Domain HumanMachine Spoken Conversations (2021.eacl-main)

Copied to clipboard

Challenge: Existing dialog act schemes are designed for human-human conversations, but are not suitable for automatic speech recognition.
Approach: They propose a dialog act annotation scheme for open-domain human-machine conversations . they collected 24K utterances from a large open- domain spoken conversation dataset .
Outcome: The proposed scheme achieves an F1 score of 0.79 on a 24K spoken conversation dataset.
InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (2025.findings-emnlp)

Copied to clipboard

Challenge: Spoken Dialogue models face challenges in handling nuanced interactional phenomena, such as interruptions and backchannels.
Approach: They propose to use a 150-hour English speech interaction dialogue dataset to empower spoken dialogue models with nuanced real-time interaction capabilities.
Outcome: The proposed dataset trains and evaluates a speech understanding model that classifies key interactional events directly from audio.
Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems (2021.naacl-demos)

Copied to clipboard

Challenge: Traditional goal-oriented dialogue systems require annotations which are hard to obtain for every new domain, limiting scalability.
Approach: They propose a data-driven approach to building goal-oriented dialogue systems . they use a seed dialogue simulator to generate annotated conversations instead of collecting annotations .
Outcome: The proposed system improves turn-level action signature prediction accuracy by 50% . the system is scalable, extensible and data efficient .
EDA: Enriching Emotional Dialogue Acts using an Ensemble of Neural Annotators (2020.lrec-1)

Copied to clipboard

Challenge: Emotion recognition helps to build natural dialogue systems.
Approach: They propose to use a recurrent neural model to annotate emotion corpora with dialogue act labels and an ensemble annotator to extract the final dialogue act label.
Outcome: The proposed model annotates two accessible multi-modal emotion corpora with and without context and extracts the final dialogue act label.
Deep Learning for Dialogue Systems (C18-3)

Copied to clipboard

Challenge: Using deep learning to build robust and scalable spoken dialogue systems is still a challenging task.
Approach: tutorial focuses on an overview of dialogue system development . goal-oriented spoken dialogue systems are most prominent component in virtual personal assistants .
Outcome: This tutorial focuses on an overview of dialogue system development while summarizing the challenges.
Fora: A corpus and framework for the study of facilitated dialogue (2024.acl-long)

Copied to clipboard

Challenge: a new study of facilitated dialogues focuses on the sharing of personal experience . social media is a popular method of civic engagement but lacks the tools to analyze it .
Approach: They compile 262 facilitated conversations hosted with partner organizations . they taxonomize personal sharing behaviors and facilitation strategies in the corpus .
Outcome: The proposed framework can be used to analyze facilitated dialogues and parse spoken conversations . the data can be applied to other fields, including civic use in governance and social science .
Language Model as an Annotator: Exploring DialoGPT for Dialogue Summarization (2021.acl-long)

Copied to clipboard

Challenge: Existing dialogue summarization systems encode text with a number of general semantic features, but these are often not available in open-domain tools.
Approach: They propose to use DialoGPT to label three types of features on two datasets . they propose to employ pre-trained and non-pre-tried models as dialogue annotators .
Outcome: The proposed method improves on two dialogue summarization datasets and achieves state-of-the-art performance.
A Unifying View On Task-oriented Dialogue Annotation (2022.lrec-1)

Copied to clipboard

Challenge: Recent research attention in task-oriented dialogue systems focuses on end-to-end neural models.
Approach: They present a dataset that combines annotated corpora from four domains to provide a unified ontology and annotation schema for task-oriented dialogues.
Outcome: The proposed dataset improves language, information content and performance in dialogues with two recent models.
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI (2024.findings-eacl)

Copied to clipboard

Challenge: DialogStudio is the largest and most diverse collection of dialogue datasets . existing datasets lack diversity and comprehensiveness, authors say .
Approach: They introduce DialogStudio: the largest and most diverse collection of dialogue datasets . DialogStuio aggregates more than 80 diverse dialogue dataset .
Outcome: a new dataset is created to improve the quality and diversity of dialogue datasets . DialogStudio is the largest and most diverse collection of dialogue data .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations