Challenge: High lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging.
Approach: They present a large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit.
Outcome: The proposed dataset is aligned with an English knowledge base to support cross-lingual linking.

Similar Papers

Entity Linking in 100 Languages (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to multilingual entity linking are cross-lingual, with a focus on zero-shot evaluation.
Approach: They propose a new formulation for multilingual entity linking where language-specific mentions resolve to a language-agnostic Knowledge Base.
Outcome: The proposed model outperforms state-of-the-art models on a large multilingual dataset and shows that frequency-based analysis provided key insights for the model and training enhancements.
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking (2026.tacl-1)

Copied to clipboard

Challenge: Existing methods for multilingual entity linking are limited by textual contexts and limited resources.
Approach: They propose a testbed system for multilingual multimodal entity linking using BBC news articles paired with corresponding images in five languages.
Outcome: The proposed system improves accuracy for entities with ambiguous textual contexts and models with weak multilingual abilities.
mReFinED: An Efficient End-to-End Multilingual Entity Linking System (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work assumed that entity mentions were given and skipped the entity mention detection step due to a lack of high-quality multilingual training corpora.
Approach: They propose a bootstrapping mention detection framework that enhances the quality of training corpora.
Outcome: The proposed framework outperforms existing work in the end-to-end MEL task while being 44 times faster.
An annotated dataset of literary entities (N19-1)

Copied to clipboard

Challenge: Existing datasets built on news focus on non-named entities, but not literary texts.
Approach: They propose to annotate 210,532 tokens from 100 different English-language literary texts for ACE entity categories (person, location, geo-political entity, facility, organization, and vehicle).
Outcome: The proposed dataset includes 210,532 tokens drawn from 100 different English-language literary texts.
AELC: Adaptive Entity Linking with LLM-Driven Contextualization (2025.findings-emnlp)

Copied to clipboard

Challenge: Entity linking (EL) focuses on associating ambiguous mentions in text with corresponding entities in a knowledge graph.
Approach: Entity linking (EL) focuses on associating ambiguous mentions in text with corresponding entities in a knowledge graph.
Outcome: Experiments on four public benchmark datasets show that AELC achieves state-of-the-art performance.
EDIN: An End-to-end Benchmark and Pipeline for Unknown Entity Discovery and Indexing (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on Entity Linking assumes that the knowledge base is complete and all mentions can be linked.
Approach: They propose a temporally segmented Unknown Entity Discovery and Indexing (EDIN) benchmark where unknown entities have to be integrated into existing entity linking systems.
Outcome: The proposed system detects, clusters, and indexes mentions of unknown entities in context.
ZELDA: A Comprehensive Benchmark for Supervised Entity Disambiguation (2023.eacl-main)

Copied to clipboard

Challenge: Entity disambiguation (ED) is the task of disambiguating named entity mentions in text to unique entries in a knowledge base.
Approach: They propose a benchmark for entity disambiguation that includes a unified training data set, entity vocabulary, candidate lists and challenging evaluation splits covering 8 different domains.
Outcome: The proposed benchmark is based on a unified training data set, entity vocabulary, candidate lists and evaluation splits covering 8 different domains.
ELISA-EDL: A Cross-lingual Entity Extraction, Linking and Localization System (N18-5)

Copied to clipboard

Challenge: ELISA-EDL is a cross-lingual entity extraction, linking and localization system for Wikipedia languages.
Approach: They propose a cross-lingual entity extraction, linking and localization system for English speakers . it extracts entities from unstructured text in any of 282 Wikipedia languages and links them to English knowledge bases .
Outcome: The proposed system extracts entity mentions from Wikipedia and links them to English knowledge bases and visualizes locations related to disaster topics on a world heatmap.
NovelCR: A Large-Scale Bilingual Dataset Tailored for Long-Span Coreference Resolution (2025.findings-acl)

Copied to clipboard

Challenge: Existing coreference resolution datasets are either small in scale or restrict coreference to a limited text span.
Approach: They present a large-scale bilingual benchmark for long-span coreference resolution . they find that NovelCR is notably rich in long-spanning coreference pairs .
Outcome: The proposed benchmark is rich in long-span coreference pairs and notably low baselines.
COMETA: A Corpus for Medical Entity Linking in the Social Media (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for Entity Linking (EL) fail to address the complex nature of health terminology in layman’s language.
Approach: They propose to use a corpus of 20k English biomedical entity mentions from Reddit expert-annotated with links to a widely-used medical knowledge graph to investigate the ability of these systems to perform complex inference on entities and concepts.
Outcome: The proposed corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by existing resources in the field.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations