Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting (D18-1)
Copied to clipboard
| Challenge: | Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools. |
| Approach: | They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities. |
| Outcome: | The proposed method matches dialect areas at different granularities against an existing dialect map. |
Similar Papers
Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method (L18-1)
Copied to clipboard
| Challenge: | a mixed corpus composed of different dialects is sufficiently resourced to cluster them into dialects. |
| Approach: | They propose a pipeline to derive clusters of dialects from a mixed corpus when their standard counterpart is sufficiently resourced. |
| Outcome: | The proposed pipeline can identify dialectal content when its standard counterpart is sufficiently resourced and can then cluster it into four dialects. |
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent work on dialect variation in NLP treats dialects as discrete categories . dialect variation is a focus of increasing interest in the field . |
| Approach: | They examine performance differences between Italian dialects by incorporating performance data from different regions of the world. |
| Outcome: | The results show that performance disparities are due to dialects that are more similar to the standard variety. |
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Historically, studies investigating minority variants of languages have been limited to a select few languages. |
| Approach: | They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors. |
| Outcome: | The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap. |
Identifying Linguistic Areas for Geolocation (D19-55)
Copied to clipboard
| Challenge: | a recent study shows that social media posts are often given as continuous coordinates . but, the resulting discrete coordinates do not always correspond to existing linguistic areas . |
| Approach: | They propose an algorithm for clustering coordinates and associating them with towns using point-to-city (P2C) they compare accuracy of a state-of-the-art geolocation model with P2C labels to one with regular k-d tree labels. |
| Outcome: | The proposed method improves accuracy at 100 miles, but degrades for finer-grained distinctions . iterative k-d tree-based method can cluster coordinates and associate them with towns . |
Building Knowledge-Guided Lexica to Model Cultural Variation (2024.naacl-long)
Copied to clipboard
| Challenge: | Cultural variation exists between nations, but also within regions . Historically, it has been difficult to computationally model cultural variation due to a lack of training data and scalability constraints. |
| Approach: | They propose a method to measure cultural variation using a knowledge-guided lexical model using geolocated tweets. |
| Outcome: | The proposed method could help us better understand the way people communicate and build more culturally-aware NLP systems. |
Crowdsourcing Regional Variation Data and Automatic Geolocalisation of Speakers of European French (L18-1)
Copied to clipboard
Jean-Philippe Goldman, Yves Scherrer, Julie Glikman, Mathieu Avanzi, Christophe Benzitoun, Philippe Boula de Mareüil
| Challenge: | a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French. |
| Approach: | They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French. |
| Outcome: | The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French. |
Retrofitting Light-weight Language Models for Emotions using Supervised Contrastive Learning (2023.emnlp-main)
Copied to clipboard
| Challenge: | a novel retrofitting method to induce emotion aspects into pre-trained language models is proposed . the models are computationally less expensive and open, but do not capture affective aspects of human communication well. |
| Approach: | They propose a retrofitting method to induce emotion aspects into pre-trained language models . they retrofit text fragments exhibiting similar emotions into pretrained networks . |
| Outcome: | The proposed method produces emotion-aware text representations for sentiment analysis and sarcasm detection tasks. |
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)
Copied to clipboard
| Challenge: | linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages . |
| Approach: | They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions . |
| Outcome: | The proposed model is representative of the language diversity and coverage of natural language processing systems. |
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)
Copied to clipboard
Piroska Lendvai, Maarten van Gompel, Anna Jouravel, Elena Renje, Uwe Reichel, Achim Rabus, Eckhart Arnold
| Challenge: | a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well. |
| Approach: | They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data . |
| Outcome: | The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP . |
Automated Tone Transcription and Clustering with Tone2Vec (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Lexical tones play a crucial role in Sino-Tibetan languages, but current phonetic fieldwork relies on manual effort. |
| Approach: | They propose a pitch-based similarity representations for tone transcription called Tone2Vec . they propose an open-source package that facilitates automated fieldwork and analysis . |
| Outcome: | Experiments on dialect clustering and variance show that Tone2Vec captures fine-grained tone variation. |