Challenge: a paper examines the degree to which dialect classifiers remain stable over time . it finds that the models remain robust over time with a fixed decay rate .
Approach: They construct a test set for 12 dialects of English that spans three years at monthly intervals with a fixed spatial distribution across 1,120 cities.
Outcome: The proposed model can reveal linguistic variation over space and time.

Similar Papers

Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are powerful representations that form the foundation of many natural language processing architectures.
Approach: They explore word embedding stability in a wide range of languages to gain insight into their stability.
Outcome: The proposed results provide insights into word embedding stability in English and other languages.
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum (2025.findings-naacl)

Copied to clipboard

Challenge: Recent work on dialect variation in NLP treats dialects as discrete categories . dialect variation is a focus of increasing interest in the field .
Approach: They examine performance differences between Italian dialects by incorporating performance data from different regions of the world.
Outcome: The results show that performance disparities are due to dialects that are more similar to the standard variety.
Diachronic degradation of language models: Insights from social media (P18-2)

Copied to clipboard

Challenge: Existing studies have explored whether and how language models degrade over time, i.e. why they fail to work on contemporary language.
Approach: They investigate the accuracy of pre-trained language models for downstream tasks in machine learning and user profiling.
Outcome: The results show that it is possible to measure diachronic drifts within social media and within the span of a few years.
Measuring and Modeling Language Change (N19-5)

Copied to clipboard

Challenge: This tutorial will help researchers answer questions fundamental to the social sciences and humanities .
Approach: This tutorial is designed to help researchers answer questions in the social sciences and humanities . it synthesizes recent computational techniques for handling and modeling temporal data .
Outcome: The tutorial will synthesize recent techniques for handling and modeling temporal data, such as dynamic word embeddings, and identify useful tools for social scientists and digital humanities scholars.
Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation Models (P19-1)

Copied to clipboard

Challenge: Recent studies show that document classifiers can become more stable over time when trained in ways that account for temporal variations.
Approach: They propose a method for embedding diachronic word embedds into document classification models . they propose 'time-driven neural classification model' that accounts for temporal variations .
Outcome: The proposed model can be trained on six corpora and make it more robust over time.
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English (2024.lrec-main)

Copied to clipboard

Challenge: Syntactic acceptance dataset is a resource being designed for syntax and computational linguistics research.
Approach: They propose to use the Syntactic Acceptability Dataset to examine the syntactical discourse.
Outcome: The proposed dataset is the largest of its kind that is publicly accessible.
Modeling language evolution and feature dynamics in a realistic geographic environment (2020.coling-main)

Copied to clipboard

Challenge: a number of studies have examined the stability or biases of typological features within language families .
Approach: They propose a model for simulating languages and their features over time in a realistic geographic environment.
Outcome: The proposed model is flexible and realistic, and can be used to answer questions.
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Historically, studies investigating minority variants of languages have been limited to a select few languages.
Approach: They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors.
Outcome: The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap.
Examining Temporality in Document Classification (P18-2)

Copied to clipboard

Challenge: a recent study examines how document classification models trained during one time period perform on documents trained during other time periods.
Approach: They propose to use a domain adaptation approach to adjust for changes in time to improve document classification.
Outcome: The proposed model improves on documents trained on time intervals even on future time interval intervals.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations