An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)
Copied to clipboard
| Challenge: | a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures . |
| Approach: | They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners . |
| Outcome: | The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied . |
Similar Papers
SW4ALL: a CEFR Classified and Aligned Corpus for Language Learning (L18-1)
Copied to clipboard
| Challenge: | Learning a second language requires exposition to texts, especially for the acquisition of vocabulary. |
| Approach: | They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia . |
| Outcome: | The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency. |
Using Multilingual Resources to Evaluate CEFRLex for Learner Applications (2020.lrec-1)
Copied to clipboard
| Challenge: | The Common European Framework of Reference for Languages defines six levels of learner proficiency and links them to particular communicative abilities. |
| Approach: | They propose to compile lexical resources that link single words and multi-word expressions to specific CEFR levels. |
| Outcome: | The results show that the English CEFRLex resource is in accordance with external resources that are gold standard. |
Building a TOCFL Learner Corpus for Chinese Grammatical Error Diagnosis (L18-1)
Copied to clipboard
| Challenge: | Annotated learner corpus is valuable for research in second language acquisition, foreign language teaching, and contrastive interlanguage analysis. |
| Approach: | They construct a TOCFL learner corpus and use it for Chinese grammatical error diagnosis. |
| Outcome: | The constructed corpus is available to the public and will be used for shared tasks on Chinese grammatical error diagnosis. |
Construction of an Evaluation Corpus for Grammatical Error Correction for Learners of Japanese as a Second Language (2020.lrec-1)
Copied to clipboard
| Challenge: | The Lang-8 corpus is suitable as a training dataset for machine translation-based grammatical error correction systems but it is not suitable as an evaluation dataset because corrected sentences sometimes include inappropriate sentences. |
| Approach: | They created an evaluation corpus for correcting grammatical errors made by Japanese as a second language learners using neural machine translation and statistical machine translation techniques. |
| Outcome: | The proposed corpus has less noise and its annotation scheme reflects the characteristics of the dataset, making it ideal for correcting grammatical errors in sentences written by learners of Japanese as a Second Language (JSL). |
CEFR-Based Sentence Difficulty Annotation and Assessment (2022.emnlp-main)
Copied to clipboard
| Challenge: | Controllable text simplification is a crucial assistive technique for language learning and teaching. |
| Approach: | They propose a sentence-level assessment model to handle unbalanced level distribution . previous studies have suggested that controllable text simplification is difficult to apply . |
| Outcome: | The proposed method outperforms baselines in readability assessment by scoring macro-F1 on the level assessment. |
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)
Copied to clipboard
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
| Challenge: | Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI. |
| Approach: | They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages. |
| Outcome: | The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment. |
Reproducing Monolingual, Multilingual and Cross-Lingual CEFR Predictions (2020.lrec-1)
Copied to clipboard
| Challenge: | POStag and dependency n-grams are more effective than text length and global linguistic indices for this kind of task. |
| Approach: | They propose to use POStag and dependency n-grams to predict the quality of a text written by learners of another language to categorize texts according to their CEFR level. |
| Outcome: | The proposed model is more effective than POStag and dependency n-grams in cross-lingual experiments than the previous models. |
Semi-automatically Annotated Learner Corpus for Russian (2022.lrec-1)
Copied to clipboard
| Challenge: | Revita Learner Corpus is a semi-automatically annotated learner corpus for Russian . it is used for research in second language acquisition and foreign language teaching . |
| Approach: | They propose a semi-automatically annotated learner corpus for Russian that detects errors automatically and annotates errors by type. |
| Outcome: | The proposed corpus detects errors automatically and is annotated by type . the data is made public and the process is much cheaper and faster . |
Schema Learning Corpus: Data and Annotation Focused on Complex Events (2024.lrec-main)
Copied to clipboard
| Challenge: | The Schema Learning Corpus is a linguistic resource designed to support research into the structure of complex events in multilingual data. |
| Approach: | The Schema Learning Corpus is a linguistic resource that includes large volumes of background data in English, Spanish and Russian. |
| Outcome: | The SLC defines 100 complex events (CEs) across 12 domains and multiple documents labeled for each . multiple documents contain evidence for each step, plus labeles events and relations along with their arguments across a large tag set. |
Language Proficiency Scoring (2020.lrec-1)
Copied to clipboard
| Challenge: | a new paper evaluates and extends the results of an automated proficiency classification system for different languages. |
| Approach: | They propose to extend an automated essay scoring system proposed by CEFR . they compare results with those from previous paper and add a new corpus for english . |
| Outcome: | The proposed approach does not scale well with the added English corpus. |