Papers with CEFR
ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | ZAEBUC is an annotated Arabic-English bilingual writer corpus . it is a corpus of short essays written by first-year university students . |
| Approach: | They propose to use a standard Arabic-English bilingual writer corpus to match comparable texts written by the same writer on different occasions. |
| Outcome: | The ZAEBUC corpus is an annotated Arabic-English bilingual writer corpus by first-year university students at Zayed University in the United Arab Emirates. |
Using Crowdsourced Exercises for Vocabulary Training to Expand ConceptNet (2020.lrec-1)
Copied to clipboard
Christos Rodosthenous, Verena Lyding, Federico Sangati, Alexander König, Umair ul Hassan, Lionel Nicolas, Jolita Horbacauskiene, Anisia Katinskaia, Lavinia Aparaschivei
| Challenge: | Language resources (LRs) are expensive to create and maintain, and this makes it difficult to create or extend LRs. |
| Approach: | They propose to use a Telegram chatbot interface to gather knowledge on word relations suitable for expanding ConceptNet with new words. |
| Outcome: | The proposed model allows to gather 12,000 answers from learners on different question types over 16 days and shows that it is a potential tool for crowdsourcing and fostering vocabulary skills. |
Using Multilingual Resources to Evaluate CEFRLex for Learner Applications (2020.lrec-1)
Copied to clipboard
| Challenge: | The Common European Framework of Reference for Languages defines six levels of learner proficiency and links them to particular communicative abilities. |
| Approach: | They propose to compile lexical resources that link single words and multi-word expressions to specific CEFR levels. |
| Outcome: | The results show that the English CEFRLex resource is in accordance with external resources that are gold standard. |
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation (2026.findings-eacl)
Copied to clipboard
| Challenge: | Practicing conversations with large language models is a promising alternative to traditional in-person language learning. |
| Approach: | They propose a new token-level evaluation metric, Token Miss Rate, that measures the proportion of incomprehensible tokens per utterance and correlates strongly with human judgments. |
| Outcome: | The proposed methods improve comprehensibility for beginner speakers from 39.4% to 83.3%, compared with prompting alone and a token-level evaluation metric, Token Miss Rate (TMR). |
Jump-Starting Item Parameters for Adaptive Language Tests (2021.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has addressed ‘cold start’ estimation of item difficulties without piloting, but a multi-task generalized linear model with BERT features is needed to jump-start new items without pilot. |
| Approach: | They propose a multi-task generalized linear model with BERT features to jump-start new item difficulties without piloting them first. |
| Outcome: | The proposed model compares test-taker proficiency, item difficulty, and language proficiency frameworks like the Common European Framework of Reference (CEFR). |
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution (2024.findings-naacl)
Copied to clipboard
| Challenge: | Automated speaking assessment (ASA) typically involves automatic speech recognition (ASR) and hand-crafted feature extraction from the transcript of a learner’s speech. |
| Approach: | They propose to use metric-based classification and loss re-weighting to model the impact of different SSL-based embedding features on the CEFR score. |
| Outcome: | The proposed model outperforms baselines on the ICNALE benchmark dataset, achieving a significant improvement of more than 10% in CEFR prediction accuracy. |
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Domain experts in engineering, healthcare, and education follow strict standards for producing quality content. |
| Approach: | They propose a retrieval-style in-context learning-based framework to guide large language models to align with expert-defined standards. |
| Outcome: | The proposed framework shows that models can gain 45% to 100% increase in precise accuracy across open and commercial LLMs evaluated. |
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments (2025.acl-industry)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) systems aim to evaluate the quality of candidate writing using computational methods. |
| Approach: | They propose a model that assigns a confidence score to each automated score to ensure it meets high reliability standards. |
| Outcome: | The proposed model achieves an F1 score of 0.97 and releases 47% of predicted scores with 100% CEFR agreement and 99% with at least 95% CEFR agreeance compared to the standalone model where all predicted scores are released. |
Automatic Extraction of Nominal Phrases from German Learner Texts of Different Proficiency Levels (2024.lrec-main)
Copied to clipboard
| Challenge: | a pilot study has found that inflecting determiners and adjectives correctly is a challenge for learners of German. |
| Approach: | They propose to use dependency parsing to extract nouns, grammatical heads and dependents that have to agree with the noun in German. |
| Outcome: | The proposed method performs well on CEFR levels A1-B1 but not level B2 texts. |
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) aims to automatically assess the quality of essays. |
| Approach: | They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French. |
| Outcome: | The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam. |
CEPOC: The Cambridge Exams Publishing Open Cloze dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | This paper presents the first dataset of open cloze tests for language learners at different proficiency levels. |
| Approach: | They present the Cambridge Exams Publishing Open Cloze (CEPOC) dataset . they perform a set of experiments on three tasks: gap filling, gap prediction, and CEFR text classification. |
| Outcome: | The results of the study are promising for a number of NLP tasks. |
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)
Copied to clipboard
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
| Challenge: | Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI. |
| Approach: | They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages. |
| Outcome: | The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment. |
CEFR-based Lexical Simplification Dataset (L18-1)
Copied to clipboard
| Challenge: | Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective. |
| Approach: | They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile . |
| Outcome: | The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus. |
An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)
Copied to clipboard
| Challenge: | a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures . |
| Approach: | They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners . |
| Outcome: | The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied . |
Reproducing Monolingual, Multilingual and Cross-Lingual CEFR Predictions (2020.lrec-1)
Copied to clipboard
| Challenge: | POStag and dependency n-grams are more effective than text length and global linguistic indices for this kind of task. |
| Approach: | They propose to use POStag and dependency n-grams to predict the quality of a text written by learners of another language to categorize texts according to their CEFR level. |
| Outcome: | The proposed model is more effective than POStag and dependency n-grams in cross-lingual experiments than the previous models. |
Language Proficiency Scoring (2020.lrec-1)
Copied to clipboard
| Challenge: | a new paper evaluates and extends the results of an automated proficiency classification system for different languages. |
| Approach: | They propose to extend an automated essay scoring system proposed by CEFR . they compare results with those from previous paper and add a new corpus for english . |
| Outcome: | The proposed approach does not scale well with the added English corpus. |
Exploring Paraphrasing Strategies for CEFR A1-Level Constraints in LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a new study compares prompt engineering approaches to rephrase general-domain texts . it compares 4 approaches to meet CEFR A1-level constraints in english and italian . |
| Approach: | They compare prompt engineering approaches to rephrase general-domain texts to meet CEFR A1-level constraints in English and Italian. |
| Outcome: | The proposed approaches meet CEFR A1-level constraints in English and Italian. |
MALT-IT2: A New Resource to Measure Text Difficulty in Light of CEFR Levels for Italian L2 Learning (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to assess text difficulty in second or foreign language classrooms are subjective . formal and quantitative characteristics of a text have a major role in determining comprehensibility . |
| Approach: | They propose a system that automatically classifies inputted texts according to CEFR levels . they describe the rationale of the project and the corpus and computational system it is based on . |
| Outcome: | The proposed system is able to predict text difficulty in Italian, and it is reliable, the authors say . they also identify the features which most influenced the predictions . |
Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification (2026.acl-long)
Copied to clipboard
| Challenge: | Existing large language model-based readability control methods rely on pre-labeled sentence corpora and primarily target English. |
| Approach: | They propose a framework for adaptive multilingual text simplification without parallel corpora supervision that integrates three reward modules: vocabulary coverage, semantic preservation, and coherence. |
| Outcome: | The proposed framework achieves higher lexical coverage at target proficiency levels while maintaining original meaning and fluency compared to stronger LLMs. |