| Challenge: | Using both linguistically-motivated features and the characteristics of the social media outlet, we obtain high accuracy on this challenging task. |
| Approach: | They propose to use linguistically-motivated features and social media characteristics to obtain high accuracy on this task. |
| Outcome: | The proposed method is highly accurate on a social media content where authors are highly-fluent nonnative speakers. |
Similar Papers
Native-like Expression Identification by Contrasting Native and Proficient Second Language Speakers (2020.coling-main)
Copied to clipboard
| Challenge: | a novel task of native-like expression identification is proposed by contrasting texts written by native speakers and those by proficient second language speakers. |
| Approach: | They propose a task of native-like expression identification by contrasting texts written by native speakers and those by proficient second language speakers. |
| Outcome: | The proposed method uncovers linguistically interesting usages distinctive of native speech. |
A Deep Generative Approach to Native Language Identification (2020.coling-main)
Copied to clipboard
| Challenge: | Native language identification (NLI) is a multi-class classification task involving multiple features that capture the systematic fingerprints of the first language in the second language writing. |
| Approach: | They propose a deep generative language modelling approach to NLI that fine-tunes a GPT-2 model separately on texts written by the authors with the same L1 and assigns n-grams to an unseen text. |
| Outcome: | The proposed method outperforms traditional machine learning approaches and currently achieves the best results on the benchmark NLI datasets. |
Native Language Identification in Texts: A Survey (2024.naacl-long)
Copied to clipboard
| Challenge: | Native language identification is the task of automatically identifying an author’s native language (L1) based on their second language production. |
| Approach: | They present a survey of native language identification applied to texts . authors describe several text representations and computational techniques used in the task . |
| Outcome: | The proposed task has been widely studied for both text and speech, particularly for L2 English due to the availability of suitable corpora. |
Using Classifier Features to Determine Language Transfer on Morphemes (N18-4)
Copied to clipboard
| Challenge: | Using native English data, we identify an English learner’s native language background based solely on the learner's English writing samples. |
| Approach: | They perform a Native Language Identification task where they identify an English learner’s native language background based only on the learner's English writing samples. |
| Outcome: | The proposed task is connected to a position in second language acquisition research that holds all learners acquire English grammatical morphemes in the same order, regardless of native language background. |
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages. |
| Approach: | They found that the correlation between eye movements and native language similarity may be more complex than the original study found. |
| Outcome: | The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study. |
BigNLI: Native Language Identification with Big Bird Embeddings (2024.lrec-main)
Copied to clipboard
| Challenge: | Native Language Identification (NLI) is a task that relies on time-consuming linguistic feature engineering and current transformer models are limited by input size. |
| Approach: | They propose to train a logistic regression classifier which only uses Big Bird embeddings to overcome this limitation. |
| Outcome: | The proposed method outperforms linguistic feature engineering models on the Reddit-L2 dataset and shows consistent out-of-sample and out-off-domain performance. |
Multilingual Native Language Identification with Large Language Models (2025.naacl-srw)
Copied to clipboard
| Challenge: | Native Language Identification (NLI) is the task of automatically identifying the native language (L1) of individuals based on their second language production. |
| Approach: | They evaluated the performance of several LLMs on non-English NLI corpora compared to traditional statistical machine learning models and language-specific BERT-based models. |
| Outcome: | The proposed models outperform statistical models and language-specific BERT-based models on English, Italian, Norwegian, and Portuguese. |
What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)
Copied to clipboard
| Challenge: | Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems. |
| Approach: | They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples. |
| Outcome: | The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache . |
Words are the Window to the Soul: Language-based User Representations for Fake News Detection (2020.coling-main)
Copied to clipboard
| Challenge: | Existing studies on fake news classification focus on textual content, but also social context in which news are consumed. |
| Approach: | They propose a model that creates representations of individuals on social media based only on the language they produce and uses them to detect fake news. |
| Outcome: | The proposed model exploits the relationship between language use and connections in the social graph to assess the presence of the Echo Chamber effect in the data. |
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |