Papers by Richard Csaky
The Gutenberg Dialogue Dataset (2021.eacl-main)
Copied to clipboard
| Challenge: | Current open-domain dialogue datasets offer a trade-off between quality and size . we build a dataset of 14.8M utterances in English and smaller datasets in german, Dutch, Spanish, Portuguese, Italian, and Hungarian . |
| Approach: | They build a high-quality dialogue corpus of 14.8M utterances in English using public-domain books from Project Gutenberg. |
| Outcome: | The proposed datasets show that the extracted dialogues are more accurate and more accurate than the larger Opensubtitles dataset. |