Challenge: Name tagging is a key task for language understanding, but is often limited by the short textual components.
Approach: They propose a novel model architecture based on visual attention that outperforms other methods . they use multimodal datasets to analyze the name tagging task on social media .
Outcome: The proposed model outperforms existing methods and significantly outperformed existing methods.

Similar Papers

Multimodal Named Entity Recognition for Short Social Media Posts (N18-1)

Copied to clipboard

Challenge: Social media posts often contain inconsistent or incomplete syntax and lexical notations with limited textual contexts.
Approach: They propose a task called Multimodal Named Entity Recognition (MNER) for noisy user-generated data . they use a dataset called SnapCaptions to build upon the state-of-the-art NER models .
Outcome: The proposed model outperforms existing models on noisy user-generated data . it uses a deep image network and generic modality attention module .
Multimodal Named Entity Disambiguation for Noisy Social Media Posts (P18-1)

Copied to clipboard

Challenge: Social media posts often contain unstructured text or images, making opinion mining challenging.
Approach: They propose a new task for multimodal social media captions with named entities annotated and linked to external knowledge bases.
Outcome: The proposed model outperforms state-of-the-art text-only NED models . it predicts correct entities in knowledge graph embeddings space, showing its efficacy and potentials .
Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks (2024.findings-eacl)

Copied to clipboard

Challenge: Prior work on multimodal content classification has not addressed these challenges.
Approach: They propose to use two auxiliary tasks to fine-tune multimodal models to address hidden cross-modal semantics and weak image-text relationships when modeling text and images.
Outcome: The proposed model improves by up to 2.6 F1 score across five diverse social media datasets.
Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification (2022.emnlp-main)

Copied to clipboard

Challenge: Social media users are using images and text to voice opinions and share ideas.
Approach: They propose to use user comments to extract hinting features from user comments and explore them via self-training.
Outcome: The proposed framework improves on four social media benchmarks for image-text relation classification, sarcasm detection, sentiment classification, and hate speech detection.
Leveraging Hashtag Networks for Multimodal Popularity Prediction of Instagram Posts (2022.lrec-1)

Copied to clipboard

Challenge: Existing popularity prediction approaches reduce hashtags to simple features such as hashtag length or number of hashtags in a post.
Approach: They propose a multimodal framework to predict popular influencer posts on Instagram using post captions, image, hashtag network and topic model.
Outcome: The proposed framework outperforms baseline models and unimodal models on popular influencer posts in Taiwan . it uses post captions, image, hashtag network, and topic model to predict popular influence post .
RIVA: A Pre-trained Tweet Multimodal Model Based on Text-image Relation for Multimodal NER (2020.coling-main)

Copied to clipboard

Challenge: Named entity recognition (MNER) for tweets is a key task of many applications.
Approach: They propose a pre-trained multimodal named entity recognition model based on Relationship Inference and Visual Attention (RIVA) for tweets.
Outcome: The proposed model improves on the multimodal named entity recognition (MNER) task on tweets with the aid of visual clues.
A Visual Attention Grounding Neural Model for Multimodal Machine Translation (D18-1)

Copied to clipboard

Challenge: Existing approaches to multimodal machine translation do not integrate visual information into the translation process.
Approach: They propose a multimodal machine translation model that utilizes parallel visual and textual information.
Outcome: The proposed model outperforms existing methods on the Multi30K and Ambiguous COCO datasets.
Integrating Text and Image: Determining Multimodal Document Intent in Instagram Posts (D19-1)

Copied to clipboard

Challenge: Existing studies on text-image content have focused on image as primary content, and text as secondary content.
Approach: They propose a multimodal dataset of 1299 Instagram posts labeled for three orthogonal taxonomies . they show that employing both text and image improves intent detection by 9.6 .
Outcome: The proposed model shows that using both text and image improves intent detection by 9.6 compared to using only the image modality.
Grounded Multimodal Named Entity Recognition on Social Media (2023.acl-long)

Copied to clipboard

Challenge: Existing studies on Multimodal Named Entity Recognition only extract entity-type pairs in text, which is useless for multimodal knowledge graph construction.
Approach: They propose a task to identify named entities in text and their bounding box groundings in image . they extend four well-known MNER methods to establish a number of baseline systems .
Outcome: The proposed framework outperforms baseline systems on the GMNER task.
MRE-MI: A Multi-image Dataset for Multimodal Relation Extraction in Social Media Posts (2025.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to Multimodal Relation Extraction focus on single image scenarios . current approaches focus on text paired with a single image, ignoring valuable insights provided by remaining images.
Approach: They propose a human-annotated dataset that includes multi-image and single-image instances for relation extraction.
Outcome: The proposed model significantly improves relation extraction in multi-image scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations