Challenge: Complex Word Identification (CWI) is the task of identifying which words or phrases in a sentence are difficult to understand by a specific type of reader.
Approach: They propose to use monolingual and cross-lingual CWI models to make predictions for languages not seen during training.
Outcome: The proposed models perform as well as (or better than) most models submitted to the latest CWI Shared Task.

Similar Papers

Domain Adaptation in Multilingual and Multi-Domain Monolingual Settings for Complex Word Identification (2022.acl-long)

Copied to clipboard

Challenge: Existing datasets for complex word identification (CWI) are limited and the difficulty of the task is augmented by the scarcity of input examples.
Approach: They propose a novel training technique for the complex word identification task based on domain adaptation to improve character and context representations.
Outcome: The proposed training technique improves the target character and context representations and also smooths differences between datasets.
Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are popular in the Natural Language Processing community because of their versatility and capability to solve unseen tasks in zero/few-shot settings.
Approach: They investigate the use of large language models in CWI, LCP, and MWE settings by evaluating their use in zero-shot, few-shot and fine-tuning settings.
Outcome: The proposed models struggle in certain conditions or achieve comparable results against existing methods.
One Size Does Not Fit All: The Case for Personalised Word Complexity Models (2022.findings-naacl)

Copied to clipboard

Challenge: Complex word identification (CWI) aims to identify words in a text that are difficult for a reader to understand and therefore benefit from simplification.
Approach: They propose to use a novel active learning framework to tailor models to individual readers and release a dataset of complexity annotations and models as a benchmark for further research.
Outcome: The proposed model can be tailored to individual readers and released as a benchmark for future research.
Complex Word Identification as a Sequence Labelling Task (P19-1)

Copied to clipboard

Challenge: Complex Word Identification (CWI) is a crucial first step in a simplification pipeline.
Approach: They propose a system that performs CWI in context without extensive feature engineering and outperforms state-of-the-art systems on this task.
Outcome: The proposed system outperforms state-of-the-art systems on complex word identification.
Easy as PIE? Identifying Multi-Word Expressions with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Multiword expressions (MWEs) are a semantically non-compositional subclass of multiword expression . authors show that prompt-based LLMs can perform competitively with supervised models .
Approach: They propose a prompt-based approach to identify idiomatic expressions in running text . they find prompt-driven LLMs can perform competitively with supervised models .
Outcome: The proposed approach can perform well with supervised models on annotated data.
Multilingual Native Language Identification with Large Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Native Language Identification (NLI) is the task of automatically identifying the native language (L1) of individuals based on their second language production.
Approach: They evaluated the performance of several LLMs on non-English NLI corpora compared to traditional statistical machine learning models and language-specific BERT-based models.
Outcome: The proposed models outperform statistical models and language-specific BERT-based models on English, Italian, Norwegian, and Portuguese.
Unsupervised Cross-Lingual Representation Learning (P19-4)

Copied to clipboard

Challenge: a comprehensive survey of cutting-edge weakly-supervised and unsupervised cross-lingual word representations is presented .
Approach: This tutorial provides a comprehensive survey of recent work on weakly-supervised and unsupervised cross-lingual word representations.
Outcome: This tutorial provides a comprehensive survey of cutting-edge weakly-supervised and unsupervised word representations.
Identifying Open Challenges in Language Identification (2025.acl-long)

Copied to clipboard

Challenge: Existing work on language identification has focused on cross-domain setups, but no systematic comparison is available.
Approach: They propose to train an accurate multi-domain languageidentification model on 2,034 languages and analyze the remaining errors.
Outcome: The proposed model performs well on 2,034 languages with training with 1,000 instances per language and a maximum input length of 100 characters.
Complex Word Identification: A Comparative Study between ChatGPT and a Dedicated Model for This Task (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to assess lexical complexity are used to evaluate the difficulty of vocabulary for language learners.
Approach: They propose to use pre-trained language models to assess the complexity of a word based on its context.
Outcome: The proposed method outperforms the best systems in SemEval-2021.
Cross-type French Multiword Expression Identification with Pre-trained Masked Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Multiword expressions (MWEs) have linguistic features that distinguish them from regular word groupings.
Approach: They propose a combination of two systems that learn verbal multiword expressions and non-verbal MWEs to improve performance on a cross-type dataset .
Outcome: The proposed system improves the F1 score on a french treebank with VMWEs and nVMWES training data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations