Challenge: a new task uses explicit knowledge from human-written guidebooks to improve geolocation accuracy . a state-of-the-art image-only method is unable to predict the location of an image .
Approach: They propose a task that uses streetview images and a guidebook to predict a country for each image . they add clues from the guidebook and supervise attention with country-level pseudo labels .
Outcome: The proposed method outperforms state-of-the-art image-only geolocation methods with 5% improvement in Top-1 accuracy.

Similar Papers

Illustrative Language Understanding: Large-Scale Visual Grounding with Image Search (P18-1)

Copied to clipboard

Challenge: a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search is currently used to learn word representations.
Approach: They propose a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search.
Outcome: The proposed model is based on a large-scale lookup operation to ground language using image search.
Where are you from? Geolocating Speech and Applications to Language Identification (2024.naacl-long)

Copied to clipboard

Challenge: Language identification (LID) is a critical component in many modern multilingual speech technologies.
Approach: They propose to use radio broadcasts with known origin to train regression models . they also propose to explore using geolocation as a proxy task for LID .
Outcome: The proposed model outperforms pretrained models on the FLEURS benchmark and on the VoxLingua benchmark.
AI Knows Where You Are: Exposure, Bias, and Inference in Multimodal Geolocation with KoreaGEO (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks show coarse granularity, linguistic bias, and a neglect of multimodal privacy risks.
Approach: They propose a benchmark for visual-language models that analyzes social photos to assess location privacy risks.
Outcome: The proposed benchmarks show coarse granularity, linguistic bias, and neglect of privacy risks.
Improving Toponym Resolution by Predicting Attributes to Constrain Geographical Ontology Entries (2024.naacl-short)

Copied to clipboard

Challenge: Existing approaches to geocoding only encode location mentions and their context .
Approach: They propose a prompt-based approach to geocoding where the machine learning algorithm encodes only the location mention and its context.
Outcome: The proposed model achieves state-of-the-art performance on multiple datasets.
HeGeL: A Novel Dataset for Geo-Location from Hebrew Text (2023.findings-acl)

Copied to clipboard

Challenge: Existing datasets in English for textual geolocation are limited because of the location of the place is implicit.
Approach: They propose to use a Hebrew place description corpus to analyze lingual geospatial reasoning.
Outcome: The Hebrew Geo-Location corpus collects literal Hebrew place descriptions and analyzes lingual geospatial reasoning.
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction (2021.naacl-main)

Copied to clipboard

Challenge: Scholarly work in this area uses toy worlds and synthetic linguistic data, but grounded language learning offers several practical and scientific advantages.
Approach: They propose to model teacher-learner dynamics through natural interactions occurring between users and search engines.
Outcome: The proposed model is better than non-grounded models on compositionality and zero-shot inference tasks.
Location Name Extraction from Targeted Text Streams using Gazetteer-based Statistical Language Models (C18-1)

Copied to clipboard

Challenge: Location name extraction tool (LNEx) is a statistical language for extracting location names from informal and unstructured social media data.
Approach: They propose a location name extraction tool that extracts location names from social media data . they use n-gram statistics and location-related dictionaries to evaluate an observed n in targeted text .
Outcome: The proposed tool outperforms state-of-the-art taggers on 4,500 event-specific tweets . it improves the average F-Score by 33-179%, outperforming all tagger .
Domain-Specific Lexical Grounding in Noisy Visual-Textual Documents (2020.emnlp-main)

Copied to clipboard

Challenge: Existing image-text grounding approaches require detailed annotations, authors say . existing methods are difficult to adapt to unlabeled multi-image, multi-sentence documents, they say .
Approach: They propose a method that can learn contextual meanings from unlabeled documents . they demonstrate that a simple unsupervised clustering-based method can be useful .
Outcome: The proposed method is particularly effective for local contextual meanings of a word . existing image-text grounding methods are difficult to adapt to unlabeled multi-image, multi-sentence documents .
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for visual grounding rely on the assumption that the given expression must be literal . this impedes the practical deployment of agents in real-world scenarios.
Approach: They propose a visual grounding task that uses intention expressions to locate foreground entities . they build a large-scale IVG dataset with free-form intention expression to promote VG .
Outcome: The proposed method is based on a large-scale intention-driven visual-language (V-L) dataset with free-form intention expressions.
The PhotoBook Dataset: Building Common Ground through Visually-Grounded Dialogue (P19-1)

Copied to clipboard

Challenge: Using the PhotoBook dataset, we investigate shared dialogue history accumulating during conversation . human interlocutors are known to collaboratively establish a shared repository of mutual information during a conversation - this common ground is then used to optimise understanding and communication efficiency.
Approach: They propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context and previously established referring expressions.
Outcome: The proposed model takes into account shared information accumulated in a reference chain and is important to resolve later descriptions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations