| Challenge: | a recent study shows that social media posts are often given as continuous coordinates . but, the resulting discrete coordinates do not always correspond to existing linguistic areas . |
| Approach: | They propose an algorithm for clustering coordinates and associating them with towns using point-to-city (P2C) they compare accuracy of a state-of-the-art geolocation model with P2C labels to one with regular k-d tree labels. |
| Outcome: | The proposed method improves accuracy at 100 miles, but degrades for finer-grained distinctions . iterative k-d tree-based method can cluster coordinates and associate them with towns . |
Similar Papers
Geographically-Informed Language Identification (2024.lrec-main)
Copied to clipboard
| Challenge: | a paper develops a method to identify languages based on geographic origin of text . the model is based in regions where languages are widely spoken and may occur anywhere . |
| Approach: | They propose to incorporate geographic information into a language identification model to ensure coverage of linguae francae regardless of location. |
| Outcome: | The proposed model includes 31 widely-spoken international languages . the model improves on social media data and improves performance on 916 languages compared to baseline models . |
Where are you from? Geolocating Speech and Applications to Language Identification (2024.naacl-long)
Copied to clipboard
Patrick Foley, Matthew Wiesner, Bismarck Odoom, Leibny Paola Garcia Perera, Kenton Murray, Philipp Koehn
| Challenge: | Language identification (LID) is a critical component in many modern multilingual speech technologies. |
| Approach: | They propose to use radio broadcasts with known origin to train regression models . they also propose to explore using geolocation as a proxy task for LID . |
| Outcome: | The proposed model outperforms pretrained models on the FLEURS benchmark and on the VoxLingua benchmark. |
Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting (D18-1)
Copied to clipboard
| Challenge: | Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools. |
| Approach: | They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities. |
| Outcome: | The proposed method matches dialect areas at different granularities against an existing dialect map. |
AI Knows Where You Are: Exposure, Bias, and Inference in Multimodal Geolocation with KoreaGEO (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks show coarse granularity, linguistic bias, and a neglect of multimodal privacy risks. |
| Approach: | They propose a benchmark for visual-language models that analyzes social photos to assess location privacy risks. |
| Outcome: | The proposed benchmarks show coarse granularity, linguistic bias, and neglect of privacy risks. |
HeGeL: A Novel Dataset for Geo-Location from Hebrew Text (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets in English for textual geolocation are limited because of the location of the place is implicit. |
| Approach: | They propose to use a Hebrew place description corpus to analyze lingual geospatial reasoning. |
| Outcome: | The Hebrew Geo-Location corpus collects literal Hebrew place descriptions and analyzes lingual geospatial reasoning. |
Which Melbourne? Augmenting Geocoding with Maps (P18-1)
Copied to clipboard
| Challenge: | Existing methods to associate geographic information in text with coordinates are limited by lexical features and cartesian coordinates. |
| Approach: | They propose a geocoder that exploits implicit lexical clues to associate coordinates with text . they propose encoding of geographic metadata to generate two distinct views of the same text. |
| Outcome: | The proposed method improves state-of-the-art results on three datasets and an open-source dataset for disease outbreaks and epidemics. |
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)
Copied to clipboard
| Challenge: | linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages . |
| Approach: | They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions . |
| Outcome: | The proposed model is representative of the language diversity and coverage of natural language processing systems. |
Tagging Location Phrases in Text (2020.lrec-1)
Copied to clipboard
| Challenge: | a number of studies have focused on detecting named entities in written language. |
| Approach: | They describe a Location Phrase Detection task to detect non-named locations . they use sequential tagging and an annotation approach to create annotated datasets . |
| Outcome: | The proposed task can detect non-named locations in English and Russian news . the authors develop a sequential tagging approach and annotate datasets for English and Russia . |
A Dataset and Evaluation Framework for Complex Geographical Description Parsing (2020.coling-main)
Copied to clipboard
| Challenge: | Previously, work on toponym resolution has focused on identifying and resolving individual toponyms in text like Adrano, S.Maria di Licodia or Catania. |
| Approach: | They propose a method that parses a set of coordinates and a collection of 360,187 uncurated complex geolocation descriptions to automate the process. |
| Outcome: | The proposed approach automates most of the process by combining Wikipedia and OpenStreetMap. |
Recognition of Implicit Geographic Movement in Text (2020.lrec-1)
Copied to clipboard
| Challenge: | a growing field of research is analyzing the geographic movement of humans, animals, and other entities. |
| Approach: | They created a corpus of sentences labeled as describing geographic movement or not . they used hand labeling, crowd voting and machine learning to predict more labels . |
| Outcome: | a new method uses hand labeling, crowd voting and machine learning to predict more labels. |