Challenge: Language resources (LRs) are expensive to create and maintain, and this makes it difficult to create or extend LRs.
Approach: They propose to use a Telegram chatbot interface to gather knowledge on word relations suitable for expanding ConceptNet with new words.
Outcome: The proposed model allows to gather 12,000 answers from learners on different question types over 16 days and shows that it is a potential tool for crowdsourcing and fostering vocabulary skills.

Similar Papers

Building an English Vocabulary Knowledge Dataset of Japanese English-as-a-Second-Language Learners Using Crowdsourcing (L18-1)

Copied to clipboard

Challenge: a dataset for analyzing the English vocabulary of English-as-a-second language learners is available . a vocabulary size test was performed by 100 test takers hired via crowdsourcing .
Approach: They propose a dataset for analyzing the English vocabulary of English-as-a-second language learners.
Outcome: a dataset for analyzing the English vocabulary of English-as-a-second language learners is available online . the results show that the test is reliable and can be predicted with high accuracy .
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)

Copied to clipboard

Challenge: Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones.
Approach: They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises .
Outcome: The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs.
Crowdsourcing in the Development of a Multilingual FrameNet: A Case Study of Korean FrameNet (2020.lrec-1)

Copied to clipboard

Challenge: Using current methods, the construction of multilingual FrameNets is expensive and complex.
Approach: They evaluated whether crowdsourcing approaches captured cross-cultural and cross-linguistic meanings . they found that crowd workers made intuitive choices comparable to trained FrameNet experts .
Outcome: The results are now available in Korean FrameNet 1.1.
DVAGen: Dynamic Vocabulary Augmented Generation (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing dynamic vocabulary approaches struggle to generalize to novel or out-of-vocabulary words, limiting their flexibility in handling diverse token combinations.
Approach: They propose an open-source framework for training, evaluation, and visualization of dynamic vocabulary-augmented language models.
Outcome: The proposed framework validates the effectiveness of dynamic vocabulary-augmented language models on modern LLMs and shows support for batch inference significantly improving inference throughput.
Training on Lexical Resources (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we fine-tune pretrained deep nets such as BERT and ERNIE . at inference time, these nets can be used to distinguish synonyms from antonyms .
Approach: They propose to use lexical resources to fine-tune pretrained deep nets such as BERT and ERNIE to distinguish synonyms from antonyms.
Outcome: The proposed method can be applied to multiword expressions, out of vocabulary words, morphological variants and more.
Semantic Frame Induction from a Real-World Corpus (2025.acl-srw)

Copied to clipboard

Challenge: Existing studies on semantic frame induction have demonstrated that pre-trained language models (PLMs) have led to more accurate results.
Approach: They conduct semantic frame induction using the Colossal Clean Crawled Corpus and assess the applicability of existing frame inducing methods to real-world data.
Outcome: The proposed methods outperform existing methods on real-world data and can induce frames corresponding to novel concepts.
Crowdsourcing Beyond Annotation: Case Studies in Benchmark Data Collection (2021.emnlp-tutorials)

Copied to clipboard

Challenge: Developing a theory of crowdsourcing for practical language problems remains an open challenge .
Approach: This tutorial exposes NLP researchers to data collection crowdsourcing methods and principles through case studies.
Outcome: This tutorial exposes NLP researchers to various data collection crowdsourcing methods and practices through case studies.
Evaluating Pretrained Causal Language Models for Synonymy (2025.findings-acl)

Copied to clipboard

Challenge: Despite the scaling of causal language models, the underlying basis of complex skills remains unclear.
Approach: They propose that subjacent skills such as synonymy might be explained using linguistic concepts.
Outcome: The proposed model recognizes synonymy but struggles to generate synonyms when prompted with relevant context.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
WordNet under Scrutiny: Dictionary Examples in the Era of Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Lexical resources are a repository of knowledge and are used for many tasks, including word sense disambiguation and etymology.
Approach: They compare WordNet, the most commonly used lexical resource in NLP, with a variety of dictionaries and examples that were generated by ChatGPT.
Outcome: The most commonly used lexical resource in NLP, with a variety of dictionaries and examples that were generated by ChatGPT.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations