WikiConv: A Corpus of the Complete Conversational History of a Large Online Collaborative Community (D18-1)
Copied to clipboard
Yiqing Hua, Cristian Danescu-Niculescu-Mizil, Dario Taraborelli, Nithum Thain, Jeffery Sorensen, Lucas Dixon
| Challenge: | Compared to large-scale collections of conversations from social media, Wikipedia talk pages only capture a subset of all discussions and only accounts for the final form of each conversation. |
| Approach: | They propose to reconstruct a corpus that encompasses the complete history of conversations between Wikipedia contributors. |
| Outcome: | The proposed corpus extracts high quality data in both Chinese and English. |
Similar Papers
WAC: A Corpus of Wikipedia Conversations for Online Abuse Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for moderation of abusive content are limited by the lack of large corpora of conversations. |
| Approach: | They propose a framework with comment-level abuse annotations based on the Wikipedia Comment corpus . they propose 'context-based' approaches to detect abusive content based upon conversational context . |
| Outcome: | The proposed framework can be used to improve the moderation process of abusive content on the Internet. |
C3KG: A Chinese Commonsense Conversation Knowledge Graph (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing commonsense knowledge bases organize tuples in an isolated manner, causing problems for chatbots . |
| Approach: | They create a Chinese commonsense conversation knowledge graph which integrates social commonsensm and dialog flow information. |
| Outcome: | The proposed graph incorporates social commonsense knowledge and dialog flow information. |
WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse (D18-1)
Copied to clipboard
| Challenge: | a corpus of 43 million atomic edits is available for Wikipedia edit history . edits are instances in which a human editor has inserted a single contiguous phrase into, or deleted a contigous phrase from, an existing sentence. |
| Approach: | They use Wikipedia edit history to mine atomic edits across 8 languages . they find edits contain instances in which a human editor has inserted a single phrase into, or deleted a contiguous phrase from, an existing sentence. |
| Outcome: | The data show that edits differ from the language observed in standard corpora and that models trained on edits encode different aspects of semantics and discourse than models trained in raw text. |
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)
Copied to clipboard
Hanae Koiso, Yasuharu Den, Yuriko Iseki, Wakako Kashino, Yoshiko Kawabata, Ken’ya Nishikawa, Yayoi Tanaka, Yasuyuki Usuda
| Challenge: | a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations . |
| Approach: | They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner. |
| Outcome: | The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings. |
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology (2022.acl-long)
Copied to clipboard
| Challenge: | Informal social interaction is the primordial home of human language. |
| Approach: | They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future. |
| Outcome: | The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure. |
The Discussion Tracker Corpus of Collaborative Argumentation (2020.lrec-1)
Copied to clipboard
| Challenge: | The Discussion Tracker corpus is an annotated dataset of transcripts of spoken, multi-party argumentation transcribed from 985 minutes of audio . |
| Approach: | They analyze 29 multi-party arguments transcribed from 985 minutes of audio . they provide descriptive statistics and code for predicting each dimension separately. |
| Outcome: | The Discussion Tracker corpus was collected in high school English classes and annotated for argument moves, specificity, specificities and collaboration dimensions. |
KGConv, a Conversational Corpus Grounded in Wikidata (2024.lrec-main)
Copied to clipboard
| Challenge: | a large corpus of 71k English conversations contains on average 8.6 questions . Unlike open domain and task-oriented dialogues, information seeking conversations are driven by the desire to acquire or evaluate knowledge. |
| Approach: | They propose a large corpus of 71k English conversations where each question-answer pair is grounded in a Wikidata fact. |
| Outcome: | The proposed dataset can be used for knowledge-based, conversational question generation . it can also be used to generate single-turn questions from Wikidata triples, question rewriting, question answering from conversation or knowledge graphs and quiz generation. |
The Niki and Julie Corpus: Collaborative Multimodal Dialogues between Humans, Robots, and Virtual Agents (L18-1)
Copied to clipboard
Ron Artstein, Jill Boberg, Alesia Gainer, Jonathan Gratch, Emmanuel Johnson, Anton Leuski, Gale Lucas, David Traum
| Challenge: | Niki and Julie corpus contains more than 600 dialogues between humans and robots . corpus includes audio and video recordings, results of ranking tasks, questionnaire responses . |
| Approach: | the corpus contains more than 600 dialogues between human participants and a robot . the dialogues are part of a collaborative item-ranking task designed to measure influence . |
| Outcome: | the corpus contains more than 600 dialogues between human participants and a robot or virtual agent . the dialogues contain conversational errors by the robot, which simulates typical of modern automated agents . |
Building and curating conversational corpora for diversity-aware language science and technology (2022.lrec-1)
Copied to clipboard
| Challenge: | Language resources that capture language use in its natural habitat of social interaction are rare despite the obvious merits of studying the very environment where we all learn and use it everyday. |
| Approach: | They propose to build an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. |
| Outcome: | The proposed pipeline can be used to collect and curate conversational corpora in 67 languages and varieties from 28 phyla. |
A Large-Scale Corpus for Conversation Disentanglement (P19-1)
Copied to clipboard
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph J. Peper, Vignesh Athreya, Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros C Polymenakos, Walter Lasecki
| Challenge: | a dataset of 77,563 messages manually annotated with reply-structure graphs disentangles conversations and defines internal conversation structure. |
| Approach: | They use a dataset of 77,563 messages manually annotated with reply-structure graphs to disentangle conversations and define internal conversation structure. |
| Outcome: | The new dataset is 16 times larger than all previous datasets combined and includes adjudication of annotation disagreements and context. |