Strong Baselines for Complex Word Identification across Multiple Languages (N19-1)
Copied to clipboard
Pierre Finnimore, Elisabeth Fritzsch, Daniel King, Alison Sneyd, Aneeq Ur Rehman, Fernando Alva-Manchego, Andreas Vlachos
| Challenge: | Complex Word Identification (CWI) is the task of identifying which words or phrases in a sentence are difficult to understand by a specific type of reader. |
| Approach: | They propose to use monolingual and cross-lingual CWI models to make predictions for languages not seen during training. |
| Outcome: | The proposed models perform as well as (or better than) most models submitted to the latest CWI Shared Task. |
Similar Papers
Domain Adaptation in Multilingual and Multi-Domain Monolingual Settings for Complex Word Identification (2022.acl-long)
Copied to clipboard
| Challenge: | Existing datasets for complex word identification (CWI) are limited and the difficulty of the task is augmented by the scarcity of input examples. |
| Approach: | They propose a novel training technique for the complex word identification task based on domain adaptation to improve character and context representations. |
| Outcome: | The proposed training technique improves the target character and context representations and also smooths differences between datasets. |
Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups (2024.emnlp-main)
Copied to clipboard
Răzvan-Alexandru Smădu, David-Gabriel Ion, Dumitru-Clementin Cercel, Florin Pop, Mihaela-Claudia Cercel
| Challenge: | Large language models (LLMs) are popular in the Natural Language Processing community because of their versatility and capability to solve unseen tasks in zero/few-shot settings. |
| Approach: | They investigate the use of large language models in CWI, LCP, and MWE settings by evaluating their use in zero-shot, few-shot and fine-tuning settings. |
| Outcome: | The proposed models struggle in certain conditions or achieve comparable results against existing methods. |
One Size Does Not Fit All: The Case for Personalised Word Complexity Models (2022.findings-naacl)
Copied to clipboard
| Challenge: | Complex word identification (CWI) aims to identify words in a text that are difficult for a reader to understand and therefore benefit from simplification. |
| Approach: | They propose to use a novel active learning framework to tailor models to individual readers and release a dataset of complexity annotations and models as a benchmark for further research. |
| Outcome: | The proposed model can be tailored to individual readers and released as a benchmark for future research. |
Complex Word Identification as a Sequence Labelling Task (P19-1)
Copied to clipboard
| Challenge: | Complex Word Identification (CWI) is a crucial first step in a simplification pipeline. |
| Approach: | They propose a system that performs CWI in context without extensive feature engineering and outperforms state-of-the-art systems on this task. |
| Outcome: | The proposed system outperforms state-of-the-art systems on complex word identification. |
Easy as PIE? Identifying Multi-Word Expressions with LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multiword expressions (MWEs) are a semantically non-compositional subclass of multiword expression . authors show that prompt-based LLMs can perform competitively with supervised models . |
| Approach: | They propose a prompt-based approach to identify idiomatic expressions in running text . they find prompt-driven LLMs can perform competitively with supervised models . |
| Outcome: | The proposed approach can perform well with supervised models on annotated data. |
Multilingual Native Language Identification with Large Language Models (2025.naacl-srw)
Copied to clipboard
| Challenge: | Native Language Identification (NLI) is the task of automatically identifying the native language (L1) of individuals based on their second language production. |
| Approach: | They evaluated the performance of several LLMs on non-English NLI corpora compared to traditional statistical machine learning models and language-specific BERT-based models. |
| Outcome: | The proposed models outperform statistical models and language-specific BERT-based models on English, Italian, Norwegian, and Portuguese. |
Unsupervised Cross-Lingual Representation Learning (P19-4)
Copied to clipboard
| Challenge: | a comprehensive survey of cutting-edge weakly-supervised and unsupervised cross-lingual word representations is presented . |
| Approach: | This tutorial provides a comprehensive survey of recent work on weakly-supervised and unsupervised cross-lingual word representations. |
| Outcome: | This tutorial provides a comprehensive survey of cutting-edge weakly-supervised and unsupervised word representations. |
Identifying Open Challenges in Language Identification (2025.acl-long)
Copied to clipboard
| Challenge: | Existing work on language identification has focused on cross-domain setups, but no systematic comparison is available. |
| Approach: | They propose to train an accurate multi-domain languageidentification model on 2,034 languages and analyze the remaining errors. |
| Outcome: | The proposed model performs well on 2,034 languages with training with 1,000 instances per language and a maximum input length of 100 characters. |
Complex Word Identification: A Comparative Study between ChatGPT and a Dedicated Model for This Task (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to assess lexical complexity are used to evaluate the difficulty of vocabulary for language learners. |
| Approach: | They propose to use pre-trained language models to assess the complexity of a word based on its context. |
| Outcome: | The proposed method outperforms the best systems in SemEval-2021. |
Cross-type French Multiword Expression Identification with Pre-trained Masked Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Multiword expressions (MWEs) have linguistic features that distinguish them from regular word groupings. |
| Approach: | They propose a combination of two systems that learn verbal multiword expressions and non-verbal MWEs to improve performance on a cross-type dataset . |
| Outcome: | The proposed system improves the F1 score on a french treebank with VMWEs and nVMWES training data. |