Challenge: EFLLex describes the use of 15,280 English words in pedagogical materials across proficiency levels.
Approach: They propose to use a part-of-speech tagger and a robust estimator to compute frequency and do manual post-editing work to improve the resource.
Outcome: The proposed resource describes the use of 15,280 English words across proficiency levels of the European Framework of Reference for Languages.

Similar Papers

Using Multilingual Resources to Evaluate CEFRLex for Learner Applications (2020.lrec-1)

Copied to clipboard

Challenge: The Common European Framework of Reference for Languages defines six levels of learner proficiency and links them to particular communicative abilities.
Approach: They propose to compile lexical resources that link single words and multi-word expressions to specific CEFR levels.
Outcome: The results show that the English CEFRLex resource is in accordance with external resources that are gold standard.
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution (2024.findings-acl)

Copied to clipboard

Challenge: Lexical Substitution fails to consider substitutes of equal or higher proficiency than the target word.
Approach: They propose a task to find appropriate substitutes for a given word in a context sentence but not those that are of equal or higher proficiency than the target.
Outcome: The proposed model outperforms ChatGPT by an average of 3.2% in F-score and achieves comparable results with GPT-4 on ProLex.
Cross-lingual Semantic Representation for NLP with UCCA (2020.coling-tutorials)

Copied to clipboard

Challenge: introductory tutorial to UCCA, a symbolic meaning representation for semantic representations.
Approach: This tutorial introduces UCCA, a cross-linguistically applicable framework for semantic representation . it will provide a detailed introduction to the UCca annotation guidelines, design philosophy and available resources .
Outcome: The tutorial will provide a detailed introduction to the UCCA framework and compare it to other meaning representations.
CEFR-based Lexical Simplification Dataset (L18-1)

Copied to clipboard

Challenge: Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective.
Approach: They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile .
Outcome: The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus.
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
On Modelling Corpus Citations in Computational Lexical Resources (2024.lrec-main)

Copied to clipboard

Challenge: TEI and OntoLex deal with corpus citations in lexicons.
Approach: They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding .
Outcome: The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary .
A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)

Copied to clipboard

Challenge: Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics .
Approach: They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative.
Outcome: The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy .
Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0 (2020.lrec-1)

Copied to clipboard

Challenge: Diachronic lexical information is increasingly used in historical linguistics and in NLP . etymological resources need to be fine-grained, large-coverage and accurate .
Approach: They propose guidelines to generate etymological lexical resources for each step of the life-cycle of an ethymology . they introduce EtymDB 2.0, an 'etiological database' generated from the Wiktionary .
Outcome: The proposed resources are generated for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation.
CILex: An Investigation of Context Information for Lexical Substitution Methods (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for lexical substitution rely on manually curated lexicals and contextual word embedding models.
Approach: They propose a method that uses contextual sentence embeddings to generate substitutes for a target word given a context and a model that captures additional context information complimenting contextual word embedders.
Outcome: The proposed method is state-of-the-art on the widely used LS07 and CoInCo datasets with P@1 scores of 55.96% and 57.25% for lexical substitution.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations