Papers with ResNet

2 papers
A Visually-grounded First-person Dialogue Dataset with Verbal and Non-verbal Responses (2020.emnlp-main)

Copied to clipboard

Challenge: In visual-grounded dialogue systems, first-person visual information about where the other speakers are and what they are paying attention to is crucial to understand their intentions.
Approach: They propose a visually-grounded first-person dialogue (VFD) dataset with verbal and non-verbal responses.
Outcome: The proposed dataset provides verbal and non-verbal responses for first-person visual information and recent neural network models.
ESPVR: Entity Spans Position Visual Regions for Multimodal Named Entity Recognition (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for acquiring local visual information are limited . existing methods for named entity recognition are redundant or insufficient .
Approach: They propose an Entity Spans Position Visual Regions module to obtain visual regions corresponding to entities in the text.
Outcome: The proposed method achieves the SOTA on Twitter-2017 and competitive results on Twitter 2015 . previous efforts have yielded promising results, but they still fall short in selecting visual information.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations