Challenge: The Common European Framework of Reference for Languages defines six levels of learner proficiency and links them to particular communicative abilities.
Approach: They propose to compile lexical resources that link single words and multi-word expressions to specific CEFR levels.
Outcome: The results show that the English CEFRLex resource is in accordance with external resources that are gold standard.

Similar Papers

EFLLex: A Graded Lexical Resource for Learners of English as a Foreign Language (L18-1)

Copied to clipboard

Challenge: EFLLex describes the use of 15,280 English words in pedagogical materials across proficiency levels.
Approach: They propose to use a part-of-speech tagger and a robust estimator to compute frequency and do manual post-editing work to improve the resource.
Outcome: The proposed resource describes the use of 15,280 English words across proficiency levels of the European Framework of Reference for Languages.
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI.
Approach: They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages.
Outcome: The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment.
An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)

Copied to clipboard

Challenge: a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures .
Approach: They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners .
Outcome: The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied .
SW4ALL: a CEFR Classified and Aligned Corpus for Language Learning (L18-1)

Copied to clipboard

Challenge: Learning a second language requires exposition to texts, especially for the acquisition of vocabulary.
Approach: They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia .
Outcome: The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency.
CEFR-based Lexical Simplification Dataset (L18-1)

Copied to clipboard

Challenge: Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective.
Approach: They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile .
Outcome: The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus.
Reproducing Monolingual, Multilingual and Cross-Lingual CEFR Predictions (2020.lrec-1)

Copied to clipboard

Challenge: POStag and dependency n-grams are more effective than text length and global linguistic indices for this kind of task.
Approach: They propose to use POStag and dependency n-grams to predict the quality of a text written by learners of another language to categorize texts according to their CEFR level.
Outcome: The proposed model is more effective than POStag and dependency n-grams in cross-lingual experiments than the previous models.
Multilingual Dependency Parsing for Low-Resource Languages: Case Studies on North Saami and Komi-Zyrian (L18-1)

Copied to clipboard

Challenge: Developing systems for low-resource languages is a crucial issue for Natural Language Processing (NLP).
Approach: They propose a method for parsing low-resource languages with very small training corpora using multilingual word embeddings and annotated corporata of larger languages.
Outcome: The proposed method improves dependency parsing for low-resource languages with very small training corpora compared to previous work . it also explores whether contemporary contact languages or genetically related languages would be the most fruitful starting point for multilingual parsers.
Language Proficiency Scoring (2020.lrec-1)

Copied to clipboard

Challenge: a new paper evaluates and extends the results of an automated proficiency classification system for different languages.
Approach: They propose to extend an automated essay scoring system proposed by CEFR . they compare results with those from previous paper and add a new corpus for english .
Outcome: The proposed approach does not scale well with the added English corpus.
Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure (2022.lrec-1)

Copied to clipboard

Challenge: a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT were selected and annotated by two Arabic native speakers independently.
Approach: They propose to use Arabic as an annotation framework to extend PARSEME to modern standard Arabic by measuring inter-annotator agreement.
Outcome: The proposed framework is based on a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT and is already exceeding the smallest corpus of the PARSEME suite.
A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)

Copied to clipboard

Challenge: Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics .
Approach: They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative.
Outcome: The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations