Challenge: Existing methods for name-based entity recognition neglect the integrity of entity semantics and conduct cross-modal interaction at token-level.
Approach: They propose a multimodal named entity recognition model that captures visual information and fuses it into tokens to rid non-entity tokens of visual noise.
Outcome: The proposed model captures entity-related visual information and fuses it into tokens . it eliminates visual noise and makes non-entity tokens easily misidentified as entities .

Similar Papers

Flat Multi-modal Interaction Transformer for Named Entity Recognition (2022.coling-1)

Copied to clipboard

Challenge: Named entity recognition (MNER) aims at identifying entity spans and recognizing their categories in social media posts with the aid of images.
Approach: They propose to use sentences and general domain words to obtain visual cues to transform the fine-grained semantic representation of vision and text into a unified lattice structure and leverage entity boundary detection as an auxiliary task to alleviate visual bias.
Outcome: The proposed method achieves state-of-the-art on two benchmark datasets.
Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal Transformer (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for named entity recognition ignore visual context bias . NER is a key component of many information extraction tasks .
Approach: They propose to use a multimodal interaction module to generate word-aware visual representations and leverage purely text-based entity span detection as an auxiliary module to guide the final predictions.
Outcome: The proposed approach achieves state-of-the-art on two benchmark datasets.
An Effective Span-based Multimodal Named Entity Recognition with Consistent Cross-Modal Alignment (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to name entity recognition rely on word-based sequence labeling and align image and text at inconsistent semantic levels.
Approach: They propose a span-based method which achieves a more consistent multimodal alignment from the perspectives of information-theoretic and cross-modal interaction.
Outcome: Experiments on two datasets show that SMNER outperforms the state-of-the-art methods.
Multimodal Named Entity Recognition for Short Social Media Posts (N18-1)

Copied to clipboard

Challenge: Social media posts often contain inconsistent or incomplete syntax and lexical notations with limited textual contexts.
Approach: They propose a task called Multimodal Named Entity Recognition (MNER) for noisy user-generated data . they use a dataset called SnapCaptions to build upon the state-of-the-art NER models .
Outcome: The proposed model outperforms existing models on noisy user-generated data . it uses a deep image network and generic modality attention module .
Grounded Multimodal Named Entity Recognition on Social Media (2023.acl-long)

Copied to clipboard

Challenge: Existing studies on Multimodal Named Entity Recognition only extract entity-type pairs in text, which is useless for multimodal knowledge graph construction.
Approach: They propose a task to identify named entities in text and their bounding box groundings in image . they extend four well-known MNER methods to establish a number of baseline systems .
Outcome: The proposed framework outperforms baseline systems on the GMNER task.
Decompose, Prioritize, and Eliminate: Dynamically Integrating Diverse Representations for Multimodal Named Entity Recognition (2024.lrec-main)

Copied to clipboard

Challenge: Existing research on multi-modal Named Entity Recognition (MNER) does not integrate all multi-modal representations to provide rich contextual information to improve NER.
Approach: They propose an iterative reasoning framework that integrates all the diverse multi-modal representations following the strategy of "decompose, prioritize, and eliminate" . they propose to use hierarchically connected fusion layers to prioritize transitions from "easy-to-hard" and "coarse-to fine"
Outcome: The proposed framework integrates all the diverse multi-modal representations following the strategy of "decompose, prioritize, and eliminate".
Nested Named Entity Recognition with Span-level Graphs (2022.acl-long)

Copied to clipboard

Challenge: Named entity recognition is one of the major subtasks of information extraction for extracting categorized named entities from unstructured text.
Approach: They propose to use retrieval-based span-level graphs to connect spans and entities in the training data based on n-gram features to integrate information of similar neighbor entities into the span representation.
Outcome: The proposed method achieves general improvements on all three benchmarks and special superiority on low frequency entities.
A Graph Interaction Framework on Relevance for Multimodal Named Entity Recognition with Multiple Images (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to determine whether images are related to named entities are not effective in multi-image scenarios.
Approach: They propose a graph interaction framework on relevance for Multimodal Named Entity Recognition with multiple images to integrate human abilities into the model.
Outcome: The proposed framework achieves state-of-the-art on benchmark datasets and compares with CLIP and CLIP-based approaches.
MNER-MI: A Multi-image Dataset for Multimodal Named Entity Recognition in Social Media (2024.lrec-main)

Copied to clipboard

Challenge: Recent research has focused on multimodal named entity recognition (MNER) but current approaches focus on text and a single accompanying image, leaving a significant research gap in multi-image scenarios.
Approach: They propose to construct a human-annotated MNER dataset with multiple images called MNER-MI and a temporal prompt model with multiple image to address the new challenges in multi-image scenarios.
Outcome: The proposed method achieves state-of-the-art results on both MNER-MI and MNER -MI-Plus, demonstrating its effectiveness.
Granular Entity Mapper: Advancing Fine-grained Multimodal Named Entity Recognition and Grounding (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for fine-grained content extraction are limited by long-tailed distribution of textual entity categories and performance of object detectors.
Approach: They propose a multi-granularity entity recognition module and a reranking module to integrate hierarchical information of entity categories, visual cues, and external textual resources collectively.
Outcome: The proposed framework achieves state-of-the-art on the fine-grained content extraction task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations