| Challenge: | Learning a second language requires exposition to texts, especially for the acquisition of vocabulary. |
| Approach: | They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia . |
| Outcome: | The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency. |
Similar Papers
An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)
Copied to clipboard
| Challenge: | a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures . |
| Approach: | They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners . |
| Outcome: | The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied . |
Using Multilingual Resources to Evaluate CEFRLex for Learner Applications (2020.lrec-1)
Copied to clipboard
| Challenge: | The Common European Framework of Reference for Languages defines six levels of learner proficiency and links them to particular communicative abilities. |
| Approach: | They propose to compile lexical resources that link single words and multi-word expressions to specific CEFR levels. |
| Outcome: | The results show that the English CEFRLex resource is in accordance with external resources that are gold standard. |
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)
Copied to clipboard
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
| Challenge: | Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI. |
| Approach: | They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages. |
| Outcome: | The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment. |
CEFR-Based Sentence Difficulty Annotation and Assessment (2022.emnlp-main)
Copied to clipboard
| Challenge: | Controllable text simplification is a crucial assistive technique for language learning and teaching. |
| Approach: | They propose a sentence-level assessment model to handle unbalanced level distribution . previous studies have suggested that controllable text simplification is difficult to apply . |
| Outcome: | The proposed method outperforms baselines in readability assessment by scoring macro-F1 on the level assessment. |
Reproducing Monolingual, Multilingual and Cross-Lingual CEFR Predictions (2020.lrec-1)
Copied to clipboard
| Challenge: | POStag and dependency n-grams are more effective than text length and global linguistic indices for this kind of task. |
| Approach: | They propose to use POStag and dependency n-grams to predict the quality of a text written by learners of another language to categorize texts according to their CEFR level. |
| Outcome: | The proposed model is more effective than POStag and dependency n-grams in cross-lingual experiments than the previous models. |
CEFR-based Lexical Simplification Dataset (L18-1)
Copied to clipboard
| Challenge: | Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective. |
| Approach: | They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile . |
| Outcome: | The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus. |
Language Proficiency Scoring (2020.lrec-1)
Copied to clipboard
| Challenge: | a new paper evaluates and extends the results of an automated proficiency classification system for different languages. |
| Approach: | They propose to extend an automated essay scoring system proposed by CEFR . they compare results with those from previous paper and add a new corpus for english . |
| Outcome: | The proposed approach does not scale well with the added English corpus. |
A Corpus for Multilingual Document Classification in Eight Languages (L18-1)
Copied to clipboard
| Challenge: | a subset of the Reuters corpus volume 2 is used to evaluate cross-lingual document classification . current best practice is to evaluate document classification on resources in one language and transfer it to another without additional resources. |
| Approach: | They propose to use a subset of the Reuters corpus to evaluate cross-lingual document classification . they propose to add Italian, Russian, Japanese and Chinese to the subset . |
| Outcome: | The proposed subset of the Reuters corpus has balanced class priors for eight languages. |
WikiBank: Using Wikidata to Improve Multilingual Frame-Semantic Parsing (2020.lrec-1)
Copied to clipboard
| Challenge: | Frame-semantic annotations exist for a tiny fraction of the world’s languages, however, Wikidata provides a common, distant supervision signal for semantic parsers. |
| Approach: | They propose a multilingual resource with partial semantic dependency structures that can be used to extend pre-existing resources rather than creating new man-made resources from scratch. |
| Outcome: | The proposed resource can be used to augment pre-existing resources or reduce the annotation effort for low-resource languages. |
Cross-Lingual Learning-to-Rank with Shared Representations (N18-2)
Copied to clipboard
| Challenge: | Cross-lingual information retrieval (CLIR) is a document retrieval task where the documents are written in a language different from that of the user's query. |
| Approach: | They propose a large-scale dataset derived from Wikipedia to support CLIR research in 25 languages. |
| Outcome: | The proposed model can improve the results of Swahili-English CLIR in Japanese and Japanese. |