Papers by Kai North

6 papers
Target-Based Offensive Language Identification (2023.acl-short)

Copied to clipboard

Challenge: Popular social media annotation taxonomies focus on the post level and token-level annotations are not available.
Approach: They propose a new dataset for Target-based Offensive language identification that uses post-level and token-level annotations to identify offensive language on Twitter.
Outcome: The proposed taxonomy can be used to annotate offensive language on English Twitter posts.
Native Language Identification in Texts: A Survey (2024.naacl-long)

Copied to clipboard

Challenge: Native language identification is the task of automatically identifying an author’s native language (L1) based on their second language production.
Approach: They present a survey of native language identification applied to texts . authors describe several text representations and computational techniques used in the task .
Outcome: The proposed task has been widely studied for both text and speech, particularly for L2 English due to the availability of suitable corpora.
Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations (2025.emnlp-main)

Copied to clipboard

Challenge: high-quality learner corpora are rarely available for studies of second language acquisition and language transfer.
Approach: They propose to curate a corpus of adult learners with longitudinal data that includes 15 different L1s.
Outcome: The proposed corpus contains 687 texts written by adult learners in the USA . authors show that the corpus can be used to explore language learning trajectories over time.
Language Variety Identification with True Labels (2024.lrec-main)

Copied to clipboard

Challenge: Language identification datasets are compiled with the assumption that the gold label of each instance is determined by where texts are retrieved from.
Approach: They present a human-annotated multilingual dataset for language variety identification . they use a model to train multiple models to discriminate between different languages .
Outcome: The proposed dataset provides a reliable benchmark toward robust and fairer language variety identification systems.
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification (2022.coling-1)

Copied to clipboard

Challenge: Lexical simplification (LS) is the task of replacing complex words with simpler alternatives to make texts more accessible to various target populations.
Approach: They propose to use a Brazilian Portuguese multi-candidate dataset to test LS systems.
Outcome: The proposed model outperforms existing models on Brazilian Portuguese and Brazilian newspaper articles.
Multilingual Native Language Identification with Large Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Native Language Identification (NLI) is the task of automatically identifying the native language (L1) of individuals based on their second language production.
Approach: They evaluated the performance of several LLMs on non-English NLI corpora compared to traditional statistical machine learning models and language-specific BERT-based models.
Outcome: The proposed models outperform statistical models and language-specific BERT-based models on English, Italian, Norwegian, and Portuguese.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations