Identifying Linguistic Areas for Geolocation (D19-55)

Copied to clipboard

Challenge: a recent study shows that social media posts are often given as continuous coordinates . but, the resulting discrete coordinates do not always correspond to existing linguistic areas .
Approach: They propose an algorithm for clustering coordinates and associating them with towns using point-to-city (P2C) they compare accuracy of a state-of-the-art geolocation model with P2C labels to one with regular k-d tree labels.
Outcome: The proposed method improves accuracy at 100 miles, but degrades for finer-grained distinctions . iterative k-d tree-based method can cluster coordinates and associate them with towns .

Similar Papers

Geographically-Informed Language Identification (2024.lrec-main)

Copied to clipboard

Challenge: a paper develops a method to identify languages based on geographic origin of text . the model is based in regions where languages are widely spoken and may occur anywhere .
Approach: They propose to incorporate geographic information into a language identification model to ensure coverage of linguae francae regardless of location.
Outcome: The proposed model includes 31 widely-spoken international languages . the model improves on social media data and improves performance on 916 languages compared to baseline models .
Where are you from? Geolocating Speech and Applications to Language Identification (2024.naacl-long)

Copied to clipboard

Challenge: Language identification (LID) is a critical component in many modern multilingual speech technologies.
Approach: They propose to use radio broadcasts with known origin to train regression models . they also propose to explore using geolocation as a proxy task for LID .
Outcome: The proposed model outperforms pretrained models on the FLEURS benchmark and on the VoxLingua benchmark.
Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting (D18-1)

Copied to clipboard

Challenge: Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools.
Approach: They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities.
Outcome: The proposed method matches dialect areas at different granularities against an existing dialect map.
AI Knows Where You Are: Exposure, Bias, and Inference in Multimodal Geolocation with KoreaGEO (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks show coarse granularity, linguistic bias, and a neglect of multimodal privacy risks.
Approach: They propose a benchmark for visual-language models that analyzes social photos to assess location privacy risks.
Outcome: The proposed benchmarks show coarse granularity, linguistic bias, and neglect of privacy risks.
HeGeL: A Novel Dataset for Geo-Location from Hebrew Text (2023.findings-acl)

Copied to clipboard

Challenge: Existing datasets in English for textual geolocation are limited because of the location of the place is implicit.
Approach: They propose to use a Hebrew place description corpus to analyze lingual geospatial reasoning.
Outcome: The Hebrew Geo-Location corpus collects literal Hebrew place descriptions and analyzes lingual geospatial reasoning.
Which Melbourne? Augmenting Geocoding with Maps (P18-1)

Copied to clipboard

Challenge: Existing methods to associate geographic information in text with coordinates are limited by lexical features and cartesian coordinates.
Approach: They propose a geocoder that exploits implicit lexical clues to associate coordinates with text . they propose encoding of geographic metadata to generate two distinct views of the same text.
Outcome: The proposed method improves state-of-the-art results on three datasets and an open-source dataset for disease outbreaks and epidemics.
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)

Copied to clipboard

Challenge: linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages .
Approach: They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions .
Outcome: The proposed model is representative of the language diversity and coverage of natural language processing systems.
Tagging Location Phrases in Text (2020.lrec-1)

Copied to clipboard

Challenge: a number of studies have focused on detecting named entities in written language.
Approach: They describe a Location Phrase Detection task to detect non-named locations . they use sequential tagging and an annotation approach to create annotated datasets .
Outcome: The proposed task can detect non-named locations in English and Russian news . the authors develop a sequential tagging approach and annotate datasets for English and Russia .
A Dataset and Evaluation Framework for Complex Geographical Description Parsing (2020.coling-main)

Copied to clipboard

Challenge: Previously, work on toponym resolution has focused on identifying and resolving individual toponyms in text like Adrano, S.Maria di Licodia or Catania.
Approach: They propose a method that parses a set of coordinates and a collection of 360,187 uncurated complex geolocation descriptions to automate the process.
Outcome: The proposed approach automates most of the process by combining Wikipedia and OpenStreetMap.
Recognition of Implicit Geographic Movement in Text (2020.lrec-1)

Copied to clipboard

Challenge: a growing field of research is analyzing the geographic movement of humans, animals, and other entities.
Approach: They created a corpus of sentences labeled as describing geographic movement or not . they used hand labeling, crowd voting and machine learning to predict more labels .
Outcome: a new method uses hand labeling, crowd voting and machine learning to predict more labels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations