Multilingual Coreference Resolution in Multiparty Dialogue (2023.tacl-1)

Copied to clipboard

Challenge: Existing datasets for entity coreference resolution are limited to English and other languages are rare.
Approach: They propose to use TV transcripts to create multilingual multiparty coreference datasets that leverage existing subtitles in Chinese and Farsi.
Outcome: The proposed dataset re-annotates for coreference on TV transcripts and then leverages existing subtitle translations to create a multilingual corpus.

Similar Papers

Investigating Multilingual Coreference Resolution by Universal Annotations (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing systems for multilingual coreference resolution have been challenging due to linguistic diversity and complexity of different languages.
Approach: They propose a multilingual coreference dataset with universal morphosyntactic and coreference annotations.
Outcome: The proposed dataset improves the baseline system by 0.9% . the proposed dataset is based on the framework of Universal Dependencies 2 .
Multilingual Coreference Resolution in Low-resource South Asian Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing coreference resolution models for South Asian languages are limited . a a sanity check for the prediction of translations is required to ensure accuracy of the model, authors say .
Approach: They evaluate an end-to-end coreference resolution model on a Hindi golden set . they use translation and word-alignment tools to translate a translated dataset into 31 languages .
Outcome: The proposed model scored 64 and 68 on a Hindi golden set.
CorefUD 1.0: Coreference Meets Universal Dependencies (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in standardization for annotated language resources have led to successful large scale efforts, such as the Universal Dependencies (UD) project for multilingual syntactically annotized data.
Approach: They propose a multilingual collection of corpora and a standardized format for coreference resolution compatible with morphosyntactic annotations in the UD framework.
Outcome: The proposed framework is compatible with morphosyntactic annotations and includes facilities for related tasks such as named entity recognition.
ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution (2023.findings-eacl)

Copied to clipboard

Challenge: Existing datasets vary in definition of coreferences and are curated for linguistic experts.
Approach: They propose to use ezCoref to create a crowdsourcing-friendly coreference annotation methodology that teaches annotators only cases that are treated similarly across existing datasets.
Outcome: The proposed method reannotates 240 passages from seven existing english coreference datasets while teaching annotators only cases that are treated similarly across them.
NovelCR: A Large-Scale Bilingual Dataset Tailored for Long-Span Coreference Resolution (2025.findings-acl)

Copied to clipboard

Challenge: Existing coreference resolution datasets are either small in scale or restrict coreference to a limited text span.
Approach: They present a large-scale bilingual benchmark for long-span coreference resolution . they find that NovelCR is notably rich in long-spanning coreference pairs .
Outcome: The proposed benchmark is rich in long-span coreference pairs and notably low baselines.
Joint Coreference Resolution and Character Linking for Multiparty Conversation (2021.eacl-main)

Copied to clipboard

Challenge: Character linking is the task of linking mentioned people in conversations to the real world . human use of pronouns or normal entities makes it difficult to link mentioned people to real people . a critical step towards understanding conversations is grounding mentioned people - a goal of the natural language processing community .
Approach: They propose to integrate richer context from the coreference relations among different mentions to help the linking task.
Outcome: The proposed model outperforms all previous models on both tasks.
MultiMUC: Multilingual Template Filling on MUC-4 (2024.eacl-long)

Copied to clipboard

Challenge: We present multilingual parallel template filling datasets for MUCs . systems were required to extract one template per incident, containing details about perpetrators, victims, weapons used .
Approach: They introduce MultiMUC, the first multilingual parallel corpus for template filling . they obtain automatic translations from a strong multilingual machine translation system .
Outcome: The proposed dataset includes translations of the classic MUC-4 template filling benchmark into Arabic, Chinese, Farsi, Korean, and Russian.
Neural Cross-Lingual Coreference Resolution And Its Application To Entity Linking (P18-2)

Copied to clipboard

Challenge: a cross-lingual coreference model is based on multi-lingual embeddings and language independent features.
Approach: They propose a crosslingual coreference model that builds on multi-lingual embeddings and language independent features.
Outcome: The proposed model outperforms the existing models on Chinese and Spanish test sets.
Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach (2025.acl-long)

Copied to clipboard

Challenge: Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals.
Approach: They propose a Chinese multimodal coreference dataset based on Douyin short-video platform to help researchers understand multimodal content.
Outcome: The proposed dataset pairs short videos with corresponding textual dialogues from user comments and includes manually annotated coreference clusters for person mentions in the text and the coreferential person head regions in the corresponding video frames.
Conundrums in Entity Coreference Resolution: Making Sense of the State of the Art (2020.emnlp-main)

Copied to clipboard

Challenge: despite significant progress on entity coreference resolution, there is a general lack of understanding of what has been improved.
Approach: They present an empirical analysis of entity coreference resolvers to provide an understanding of what has been improved.
Outcome: The proposed model improves the performance of entity coreference resolvers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations