Papers by Kai North
Target-Based Offensive Language Identification (2023.acl-short)
Copied to clipboard
Marcos Zampieri, Skye Morgan, Kai North, Tharindu Ranasinghe, Austin Simmmons, Paridhi Khandelwal, Sara Rosenthal, Preslav Nakov
| Challenge: | Popular social media annotation taxonomies focus on the post level and token-level annotations are not available. |
| Approach: | They propose a new dataset for Target-based Offensive language identification that uses post-level and token-level annotations to identify offensive language on Twitter. |
| Outcome: | The proposed taxonomy can be used to annotate offensive language on English Twitter posts. |
Native Language Identification in Texts: A Survey (2024.naacl-long)
Copied to clipboard
| Challenge: | Native language identification is the task of automatically identifying an author’s native language (L1) based on their second language production. |
| Approach: | They present a survey of native language identification applied to texts . authors describe several text representations and computational techniques used in the task . |
| Outcome: | The proposed task has been widely studied for both text and speech, particularly for L2 English due to the availability of suitable corpora. |
Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations (2025.emnlp-main)
Copied to clipboard
Poorvi Acharya, J. Elizabeth Liebl, Dhiman Goswami, Kai North, Marcos Zampieri, Antonios Anastasopoulos
| Challenge: | high-quality learner corpora are rarely available for studies of second language acquisition and language transfer. |
| Approach: | They propose to curate a corpus of adult learners with longitudinal data that includes 15 different L1s. |
| Outcome: | The proposed corpus contains 687 texts written by adult learners in the USA . authors show that the corpus can be used to explore language learning trajectories over time. |
Language Variety Identification with True Labels (2024.lrec-main)
Copied to clipboard
Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice, Neha Kumari, Nishant Nair, Yash Mahesh Bangera
| Challenge: | Language identification datasets are compiled with the assumption that the gold label of each instance is determined by where texts are retrieved from. |
| Approach: | They present a human-annotated multilingual dataset for language variety identification . they use a model to train multiple models to discriminate between different languages . |
| Outcome: | The proposed dataset provides a reliable benchmark toward robust and fairer language variety identification systems. |
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification (2022.coling-1)
Copied to clipboard
| Challenge: | Lexical simplification (LS) is the task of replacing complex words with simpler alternatives to make texts more accessible to various target populations. |
| Approach: | They propose to use a Brazilian Portuguese multi-candidate dataset to test LS systems. |
| Outcome: | The proposed model outperforms existing models on Brazilian Portuguese and Brazilian newspaper articles. |
Multilingual Native Language Identification with Large Language Models (2025.naacl-srw)
Copied to clipboard
| Challenge: | Native Language Identification (NLI) is the task of automatically identifying the native language (L1) of individuals based on their second language production. |
| Approach: | They evaluated the performance of several LLMs on non-English NLI corpora compared to traditional statistical machine learning models and language-specific BERT-based models. |
| Outcome: | The proposed models outperform statistical models and language-specific BERT-based models on English, Italian, Norwegian, and Portuguese. |