Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)
Copied to clipboard
Hanae Koiso, Yasuharu Den, Yuriko Iseki, Wakako Kashino, Yoshiko Kawabata, Ken’ya Nishikawa, Yayoi Tanaka, Yasuyuki Usuda
| Challenge: | a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations . |
| Approach: | They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner. |
| Outcome: | The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings. |
Similar Papers
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts. |
| Approach: | They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent . |
| Outcome: | The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora . |
JAIST Annotated Corpus of Free Conversation (L18-1)
Copied to clipboard
| Challenge: | Annotated corpus of free conversations in Japanese is the first publicly available one. |
| Approach: | They propose to annotate free conversations in Japanese with dialog act and sympathy tags . they report how to construct the corpus and its statistics . |
| Outcome: | The proposed corpus is the first annotated corpus of free conversations in Japanese . it consists of 92,031 utterances in 97 dialogs. |
Simplified Corpus with Core Vocabulary (L18-1)
Copied to clipboard
| Challenge: | a study has found that simple Japanese is more accessible to foreigners than English. |
| Approach: | They have constructed a simplified corpus for the Japanese language and selected the core vocabulary. |
| Outcome: | The simplified corpus can be used for automatic text simplification and translating simple Japanese into English and vice-versa. |
Japanese Dialogue Corpus of Information Navigation and Attentive Listening Annotated with Extended ISO-24617-2 Dialogue Act Tags (L18-1)
Copied to clipboard
| Challenge: | Large-scale conventional dialogue corpora are mainly built for specified tasks with specially designed dialogue states. |
| Approach: | They propose to annotate large-scale dialogue data with an extended ISO-24617-2 dialogue act tag-set to model a natural conversation with machines. |
| Outcome: | The proposed corpus covers a wider range of dialogue tasks than existing task-oriented systems or text-chat systems. |
Crowdsourced Corpus of Sentence Simplification with Core Vocabulary (L18-1)
Copied to clipboard
| Challenge: | a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones. |
| Approach: | They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences. |
| Outcome: | The proposed set of simplified sentences is a good quality data set for machine learning. |
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)
Copied to clipboard
| Challenge: | Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues. |
| Approach: | They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms. |
| Outcome: | The proposed corpus includes parallel text and speech data of 21 Japanese dialects. |
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)
Copied to clipboard
| Challenge: | a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information. |
| Approach: | They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants. |
| Outcome: | The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants. |
User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization (2021.naacl-main)
Copied to clipboard
| Challenge: | Morphological analysis (MA) and lexical normalization (LN) are important tasks for Japanese user-generated text. |
| Approach: | They construct a publicly available Japanese UGT corpus annotated with morphological and normalization information. |
| Outcome: | The proposed corpus shows low performance for non-general words and non-standard forms . morphological analysis is an important task in Japanese user-generated text . |
WikiConv: A Corpus of the Complete Conversational History of a Large Online Collaborative Community (D18-1)
Copied to clipboard
Yiqing Hua, Cristian Danescu-Niculescu-Mizil, Dario Taraborelli, Nithum Thain, Jeffery Sorensen, Lucas Dixon
| Challenge: | Compared to large-scale collections of conversations from social media, Wikipedia talk pages only capture a subset of all discussions and only accounts for the final form of each conversation. |
| Approach: | They propose to reconstruct a corpus that encompasses the complete history of conversations between Wikipedia contributors. |
| Outcome: | The proposed corpus extracts high quality data in both Chinese and English. |
Building and curating conversational corpora for diversity-aware language science and technology (2022.lrec-1)
Copied to clipboard
| Challenge: | Language resources that capture language use in its natural habitat of social interaction are rare despite the obvious merits of studying the very environment where we all learn and use it everyday. |
| Approach: | They propose to build an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. |
| Outcome: | The proposed pipeline can be used to collect and curate conversational corpora in 67 languages and varieties from 28 phyla. |