Papers with ResNet
A Visually-grounded First-person Dialogue Dataset with Verbal and Non-verbal Responses (2020.emnlp-main)
Copied to clipboard
| Challenge: | In visual-grounded dialogue systems, first-person visual information about where the other speakers are and what they are paying attention to is crucial to understand their intentions. |
| Approach: | They propose a visually-grounded first-person dialogue (VFD) dataset with verbal and non-verbal responses. |
| Outcome: | The proposed dataset provides verbal and non-verbal responses for first-person visual information and recent neural network models. |
ESPVR: Entity Spans Position Visual Regions for Multimodal Named Entity Recognition (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for acquiring local visual information are limited . existing methods for named entity recognition are redundant or insufficient . |
| Approach: | They propose an Entity Spans Position Visual Regions module to obtain visual regions corresponding to entities in the text. |
| Outcome: | The proposed method achieves the SOTA on Twitter-2017 and competitive results on Twitter 2015 . previous efforts have yielded promising results, but they still fall short in selecting visual information. |