Challenge: Learning a second language requires exposition to texts, especially for the acquisition of vocabulary.
Approach: They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia .
Outcome: The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency.

Similar Papers

An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)

Copied to clipboard

Challenge: a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures .
Approach: They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners .
Outcome: The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied .
Using Multilingual Resources to Evaluate CEFRLex for Learner Applications (2020.lrec-1)

Copied to clipboard

Challenge: The Common European Framework of Reference for Languages defines six levels of learner proficiency and links them to particular communicative abilities.
Approach: They propose to compile lexical resources that link single words and multi-word expressions to specific CEFR levels.
Outcome: The results show that the English CEFRLex resource is in accordance with external resources that are gold standard.
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI.
Approach: They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages.
Outcome: The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment.
CEFR-Based Sentence Difficulty Annotation and Assessment (2022.emnlp-main)

Copied to clipboard

Challenge: Controllable text simplification is a crucial assistive technique for language learning and teaching.
Approach: They propose a sentence-level assessment model to handle unbalanced level distribution . previous studies have suggested that controllable text simplification is difficult to apply .
Outcome: The proposed method outperforms baselines in readability assessment by scoring macro-F1 on the level assessment.
Reproducing Monolingual, Multilingual and Cross-Lingual CEFR Predictions (2020.lrec-1)

Copied to clipboard

Challenge: POStag and dependency n-grams are more effective than text length and global linguistic indices for this kind of task.
Approach: They propose to use POStag and dependency n-grams to predict the quality of a text written by learners of another language to categorize texts according to their CEFR level.
Outcome: The proposed model is more effective than POStag and dependency n-grams in cross-lingual experiments than the previous models.
CEFR-based Lexical Simplification Dataset (L18-1)

Copied to clipboard

Challenge: Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective.
Approach: They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile .
Outcome: The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus.
Language Proficiency Scoring (2020.lrec-1)

Copied to clipboard

Challenge: a new paper evaluates and extends the results of an automated proficiency classification system for different languages.
Approach: They propose to extend an automated essay scoring system proposed by CEFR . they compare results with those from previous paper and add a new corpus for english .
Outcome: The proposed approach does not scale well with the added English corpus.
A Corpus for Multilingual Document Classification in Eight Languages (L18-1)

Copied to clipboard

Challenge: a subset of the Reuters corpus volume 2 is used to evaluate cross-lingual document classification . current best practice is to evaluate document classification on resources in one language and transfer it to another without additional resources.
Approach: They propose to use a subset of the Reuters corpus to evaluate cross-lingual document classification . they propose to add Italian, Russian, Japanese and Chinese to the subset .
Outcome: The proposed subset of the Reuters corpus has balanced class priors for eight languages.
WikiBank: Using Wikidata to Improve Multilingual Frame-Semantic Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Frame-semantic annotations exist for a tiny fraction of the world’s languages, however, Wikidata provides a common, distant supervision signal for semantic parsers.
Approach: They propose a multilingual resource with partial semantic dependency structures that can be used to extend pre-existing resources rather than creating new man-made resources from scratch.
Outcome: The proposed resource can be used to augment pre-existing resources or reduce the annotation effort for low-resource languages.
Cross-Lingual Learning-to-Rank with Shared Representations (N18-2)

Copied to clipboard

Challenge: Cross-lingual information retrieval (CLIR) is a document retrieval task where the documents are written in a language different from that of the user's query.
Approach: They propose a large-scale dataset derived from Wikipedia to support CLIR research in 25 languages.
Outcome: The proposed model can improve the results of Swahili-English CLIR in Japanese and Japanese.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations