Papers by Ruo-Ping Dong
Reading Between the Lines: Exploring Infilling in Visual Narratives (2020.emnlp-main)
Copied to clipboard
| Challenge: | Generating long form narratives from multiple modalities requires a model to learn surrounding contextual information by masking spans of input while decoding attempts in generating the entire text. |
| Approach: | They propose to use infilling techniques to generate textual descriptions from images that are rich in contextual dependencies. |
| Outcome: | The proposed model outperforms existing models in visual storytelling by generating text from a large scale dataset of 46,200 procedures and 340k pairwise images and textual descriptions. |