Challenge: Using a large-scale dataset, we explore Chinese named entity recognition (NER) with both textual and acoustic contents.
Approach: They propose a Chinese multimodal named entity recognition dataset . their corpus contains 42,987 annotated sentences and 71 hours of speech data .
Outcome: The proposed model yields state-of-the-art (SoTA) results on Chinese multimodal named entity recognition (NER) based on 42,987 annotated sentences and 71 hours of speech data.

Similar Papers

Breaking the Boundaries: A Unified Framework for Chinese Named Entity Recognition Across Text and Speech (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to Named Entity Recognition (NER) tasks are limited by the complexity of the data and the potential connections between tasks.
Approach: They propose a task to break the boundaries between different modal NER tasks by using a unified data format for inputs from different modalités.
Outcome: The proposed task breaks the boundaries between different modal NER tasks and is a unified implementation of them.
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them .
Approach: They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition .
Outcome: The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs .
MultiCoNER: A Large-scale Multilingual Dataset for Complex Named Entity Recognition (2022.coling-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a core task in Natural Language Processing.
Approach: They present a large multilingual dataset for Named Entity Recognition that covers 3 domains across 11 languages and multilingual and code-mixing subsets.
Outcome: The proposed dataset is large and multilingual, covering 11 languages and subsets.
Chinese Spoken Named Entity Recognition in Real-world Scenarios: Dataset and Approaches (2024.findings-acl)

Copied to clipboard

Challenge: Current Chinese Spoken NER datasets are laboratory-controlled and are limited in topics.
Approach: They propose to use Chinese Spoken NER datasets to extract entities from speech to help voice assistants better grasp the intent behind user's questions and instructions.
Outcome: The proposed methods improve on self-training-asr and mapping then distilling, and even compared with GPT4.0, they achieve better results.
M-CNER: A Corpus for Chinese Named Entity Recognition in Multi-Domains (L18-1)

Copied to clipboard

Challenge: NER is one of the most important natural language processing tasks.
Approach: They propose to annotate sentences from human-computer interaction, social media, and e-commerce using two rounds of annotation.
Outcome: The proposed system performs the best on all the data sets.
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks.
Approach: They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data.
Outcome: The proposed model improves on lip reading sentences 2 by 30% even without an external language model.
Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach (2025.acl-long)

Copied to clipboard

Challenge: Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals.
Approach: They propose a Chinese multimodal coreference dataset based on Douyin short-video platform to help researchers understand multimodal content.
Outcome: The proposed dataset pairs short videos with corresponding textual dialogues from user comments and includes manually annotated coreference clusters for person mentions in the text and the coreferential person head regions in the corresponding video frames.
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation.
Approach: They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks.
Outcome: The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results.
Chinese Whispers: A Multimodal Dataset for Embodied Language Grounding (2020.lrec-1)

Copied to clipboard

Challenge: In this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture.
Approach: They propose a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture.
Outcome: The proposed method avoids implicit experimenter biases by allowing subjects to instruct each other on the nature of the task: the process of the furniture assembly.
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language.
Approach: They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language.
Outcome: The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations