Challenge: Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools.
Approach: They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities.
Outcome: The proposed method matches dialect areas at different granularities against an existing dialect map.

Similar Papers

Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method (L18-1)

Copied to clipboard

Challenge: a mixed corpus composed of different dialects is sufficiently resourced to cluster them into dialects.
Approach: They propose a pipeline to derive clusters of dialects from a mixed corpus when their standard counterpart is sufficiently resourced.
Outcome: The proposed pipeline can identify dialectal content when its standard counterpart is sufficiently resourced and can then cluster it into four dialects.
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum (2025.findings-naacl)

Copied to clipboard

Challenge: Recent work on dialect variation in NLP treats dialects as discrete categories . dialect variation is a focus of increasing interest in the field .
Approach: They examine performance differences between Italian dialects by incorporating performance data from different regions of the world.
Outcome: The results show that performance disparities are due to dialects that are more similar to the standard variety.
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Historically, studies investigating minority variants of languages have been limited to a select few languages.
Approach: They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors.
Outcome: The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap.
Identifying Linguistic Areas for Geolocation (D19-55)

Copied to clipboard

Challenge: a recent study shows that social media posts are often given as continuous coordinates . but, the resulting discrete coordinates do not always correspond to existing linguistic areas .
Approach: They propose an algorithm for clustering coordinates and associating them with towns using point-to-city (P2C) they compare accuracy of a state-of-the-art geolocation model with P2C labels to one with regular k-d tree labels.
Outcome: The proposed method improves accuracy at 100 miles, but degrades for finer-grained distinctions . iterative k-d tree-based method can cluster coordinates and associate them with towns .
Building Knowledge-Guided Lexica to Model Cultural Variation (2024.naacl-long)

Copied to clipboard

Challenge: Cultural variation exists between nations, but also within regions . Historically, it has been difficult to computationally model cultural variation due to a lack of training data and scalability constraints.
Approach: They propose a method to measure cultural variation using a knowledge-guided lexical model using geolocated tweets.
Outcome: The proposed method could help us better understand the way people communicate and build more culturally-aware NLP systems.
Crowdsourcing Regional Variation Data and Automatic Geolocalisation of Speakers of European French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French.
Approach: They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French.
Outcome: The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French.
Retrofitting Light-weight Language Models for Emotions using Supervised Contrastive Learning (2023.emnlp-main)

Copied to clipboard

Challenge: a novel retrofitting method to induce emotion aspects into pre-trained language models is proposed . the models are computationally less expensive and open, but do not capture affective aspects of human communication well.
Approach: They propose a retrofitting method to induce emotion aspects into pre-trained language models . they retrofit text fragments exhibiting similar emotions into pretrained networks .
Outcome: The proposed method produces emotion-aware text representations for sentiment analysis and sarcasm detection tasks.
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)

Copied to clipboard

Challenge: linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages .
Approach: They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions .
Outcome: The proposed model is representative of the language diversity and coverage of natural language processing systems.
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)

Copied to clipboard

Challenge: a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well.
Approach: They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data .
Outcome: The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP .
Automated Tone Transcription and Clustering with Tone2Vec (2024.findings-emnlp)

Copied to clipboard

Challenge: Lexical tones play a crucial role in Sino-Tibetan languages, but current phonetic fieldwork relies on manual effort.
Approach: They propose a pitch-based similarity representations for tone transcription called Tone2Vec . they propose an open-source package that facilitates automated fieldwork and analysis .
Outcome: Experiments on dialect clustering and variance show that Tone2Vec captures fine-grained tone variation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations