Papers by Han-Gyu Kim
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing multi-modal dialogue datasets that focus on image-based dialogues have low quality and limited diversity of images per dialogue. |
| Approach: | They propose to construct a multi-modal dialogue dataset that guarantees both dialogue quality and image diversity without requiring minimum human effort. |
| Outcome: | The proposed dataset outperforms existing datasets in terms of quality and diversity in human evaluation. |