Challenge: MATILDA is the first multi-annotator, multi-language dialogue annotation tool . it allows the creation of corpora, the management of users, the annotation of dialogues, the quick adaptation of the user interface to any language and the resolution of interannotation disagreement.
Approach: They propose to use MATILDA to create corpora, manage users, and annotation dialogues.
Outcome: The proposed tool supports the full pipeline for dialogue annotation, and non-technical people can use it.

Similar Papers

LIDA: Lightweight Interactive Dialogue Annotator (D19-3)

Copied to clipboard

Challenge: Dialogue systems are dependent on the quality of the data used to train them.
Approach: They propose to develop an annotation tool specifically for conversation data that handles the entire dialogue annotation pipeline from raw text to structured conversation data.
Outcome: The proposed tool handles the entire dialogue annotation pipeline from raw text to structured conversation data and has a dedicated interface to resolve inter-annotator disagreements.
TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of Texts (2020.lrec-1)

Copied to clipboard

Challenge: Existing annotation tools are not efficient for the annotation of corpora and are not error-free.
Approach: They propose to extend existing annotation tools by evaluating their flexibility and efficiency.
Outcome: The proposed system performs platform-independent multimodal annotations and annotates complex textual structures.
Dialogue Structure Annotation for Multi-Floor Interaction (L18-1)

Copied to clipboard

Challenge: Existing annotation schemes do not address dialogue structure.
Approach: They propose an annotation scheme for meso-level dialogue structure that clusters utterances from multiple participants and floors into units according to realization of an initiator's intent.
Outcome: The proposed annotation scheme is used to annotate a corpus of human-robot interaction dialogues.
xDial-Eval: A Multilingual Open-Domain Dialogue Evaluation Benchmark (2023.findings-emnlp)

Copied to clipboard

Challenge: Currently, human evaluation is the most reliable way to holistically judge the quality of the dialogue.
Approach: They propose to use English dialogue evaluation metrics to generalize them to other languages.
Outcome: The proposed metrics outperform OpenAI’s ChatGPT in terms of average Pearson correlations over all datasets and languages.
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)

Copied to clipboard

Challenge: Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora.
Approach: They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora.
Outcome: The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora.
Language Model as an Annotator: Exploring DialoGPT for Dialogue Summarization (2021.acl-long)

Copied to clipboard

Challenge: Existing dialogue summarization systems encode text with a number of general semantic features, but these are often not available in open-domain tools.
Approach: They propose to use DialoGPT to label three types of features on two datasets . they propose to employ pre-trained and non-pre-tried models as dialogue annotators .
Outcome: The proposed method improves on two dialogue summarization datasets and achieves state-of-the-art performance.
Enhanced Entity Annotations for Multilingual Corpora (2022.lrec-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a new language for natural language processing.
Approach: They propose to improve the annotation quality of the English Wikipedia tool WEXEA . they propose to use a proven NER system to annotate entities in Wikipedia .
Outcome: The proposed tool can be used to exhaustively annotate entities in Wikipedia articles.
AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a small fraction of the languages currently covered by speech technologies are mainly spoken in English.
Approach: They present an annotation toolkit that detects when a person speaks on the scene and the corresponding transcription.
Outcome: The proposed toolkit can speed up the annotation process by up to four times . it can be used in Spanish, and is available on github.
LightTag: Text Annotation Platform (2021.emnlp-demo)

Copied to clipboard

Challenge: LightTag is a text annotation tool built on the premise of global optimization by addressing annotator as well as project managers and data scientists who manage the work and enforce production quality.
Approach: They propose to use LightTag to optimize the global NLP process by addressing annotators as well as project managers and data scientists who manage the work and enforce production quality.
Outcome: The proposed tool is based on the theory of constraints and is available for free for academic use.
Commentator: A Code-mixed Multilingual Text Annotation Framework (2024.emnlp-demo)

Copied to clipboard

Challenge: Existing annotation tools fail to address multilingual datasets efficiently.
Approach: They introduce a code-mixed multilingual text annotation framework, COMMENTATOR . they perform robust qualitative human-based evaluations to showcase its effectiveness .
Outcome: The proposed framework performs faster than baseline annotations in Hinglish and Hindi.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations