Challenge: Previously, work on toponym resolution has focused on identifying and resolving individual toponyms in text like Adrano, S.Maria di Licodia or Catania.
Approach: They propose a method that parses a set of coordinates and a collection of 360,187 uncurated complex geolocation descriptions to automate the process.
Outcome: The proposed approach automates most of the process by combining Wikipedia and OpenStreetMap.

Similar Papers

GeospaCy: A tool for extraction and geographical referencing of spatial expressions in textual data (2024.eacl-demo)

Copied to clipboard

Challenge: Spatial information in text enables to understand the geographical context and relationships within text for location-sensitive applications.
Approach: They propose to use spatial information extracted from textual data to perform geoparsing and geocoding tasks.
Outcome: The GeospaCy software tool is designed for the extraction and georeferencing of spatial information present in textual data.
Arukikata Travelogue Dataset with Geographic Entity Mention, Coreference, and Link Annotation (2024.findings-eacl)

Copied to clipboard

Challenge: et al., 2006) considers geographic relatedness among geo-entity mentions in document-level geoparsing.
Approach: They present a Japanese travelogue dataset that considers geographic relatedness among geo-entity mentions.
Outcome: The proposed dataset includes 200 travelogue documents with rich geo-entity information . it shows that human activities, mobility, and events are often described with natural language expressions of locations or geographic entities (geo-entities)
Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to evaluate geoparsing systems are small-scale and lack coverage of location expressions on general domains.
Approach: They propose a method to construct a large-scale corpus for geoparsing from Wikipedia articles.
Outcome: The proposed method can annotate multiple location expressions with coordinates even with ambiguous expressions.
Improving Toponym Resolution by Predicting Attributes to Constrain Geographical Ontology Entries (2024.naacl-short)

Copied to clipboard

Challenge: Existing approaches to geocoding only encode location mentions and their context .
Approach: They propose a prompt-based approach to geocoding where the machine learning algorithm encodes only the location mention and its context.
Outcome: The proposed model achieves state-of-the-art performance on multiple datasets.
Coordinates from Context: Using LLMs to Ground Complex Location References (2026.eacl-long)

Copied to clipboard

Challenge: Existing geocoding tools can only link locations already in a geographic database, which often do not include compositional locations.
Approach: They propose a geocoding strategy that leverages LLMs' geospatial knowledge versus reasoning skills to improve performance for the task.
Outcome: The proposed model improves performance and is comparable to larger models.
Basreh or Basra? Geoparsing Historical Locations in the Svoboda Diaries (2024.acl-srw)

Copied to clipboard

Challenge: In the historical domain, many geoparsing corpora are from large news collections.
Approach: They propose a pipeline employing named entity recognition for geotagging and a map-based generate-and-rank approach incorporating candidate name augmentation and clustering of location context words for geocoding.
Outcome: The proposed pipeline outperforms existing map-based geoparsers in terms of accuracy, lowest mean distance error, and number of locations correctly identified.
Geo-Encoder: A Chunk-Argument Bi-Encoder Framework for Chinese Geographic Re-Ranking (2024.eacl-long)

Copied to clipboard

Challenge: Chinese geographic re-ranking task aims to find the most relevant addresses among retrieved candidates.
Approach: They propose a framework to integrate Chinese geographic semantics into re-ranking pipelines.
Outcome: The proposed framework improves on two Chinese benchmark datasets.
SpatialWebAgent: Leveraging Large Language Models for Automated Spatial Information Extraction and Map Grounding (2025.acl-demo)

Copied to clipboard

Challenge: Understanding and extracting spatial information from text is vital for a wide range of applications, says nielsen . inherent complexity of geographic expressions in natural language presents significant hurdles for traditional extraction methods.
Approach: They propose a system that leverages large language models to extract spatial information from natural language.
Outcome: SpatialWebAgent is designed to extract, standardize, and ground spatial information from natural language text directly onto maps.
Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense Knowledge (2025.acl-long)

Copied to clipboard

Challenge: a new task of answering geographic reasoning questions based on the given image is proposed . the task requires identifying the objects in the image and understanding the background context .
Approach: They propose a task of answering geographic reasoning questions based on the given image . they analyze the image and describe its fine-grained content by text and keywords .
Outcome: The proposed method can be used to answer geographic reasoning questions based on an image . it can be applied to a large-scale dataset with 41,329 samples .
Tagging Location Phrases in Text (2020.lrec-1)

Copied to clipboard

Challenge: a number of studies have focused on detecting named entities in written language.
Approach: They describe a Location Phrase Detection task to detect non-named locations . they use sequential tagging and an annotation approach to create annotated datasets .
Outcome: The proposed task can detect non-named locations in English and Russian news . the authors develop a sequential tagging approach and annotate datasets for English and Russia .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations