Papers by Richard Csaky

1 papers
The Gutenberg Dialogue Dataset (2021.eacl-main)

Copied to clipboard

Challenge: Current open-domain dialogue datasets offer a trade-off between quality and size . we build a dataset of 14.8M utterances in English and smaller datasets in german, Dutch, Spanish, Portuguese, Italian, and Hungarian .
Approach: They build a high-quality dialogue corpus of 14.8M utterances in English using public-domain books from Project Gutenberg.
Outcome: The proposed datasets show that the extracted dialogues are more accurate and more accurate than the larger Opensubtitles dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations