English WordNet Random Walk Pseudo-Corpora (2020.lrec-1)

Copied to clipboard

Challenge: a random walk over the WordNet taxonomy generates a set of pseudo-corpora . a resource description paper describes the creation and properties of such pseudo-corporates .
Approach: They propose to use random walk to generate a set of pseudo-corpora over the English WordNet taxonomy.
Outcome: The proposed pseudo-corpora can be used to train taxonomic word embeddings . the proposed pseudo corpora are generated from a random walk over the English wordnet taxonomy .

Similar Papers

A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
Browsing and Supporting Pluricentric Global Wordnet, or just your Wordnet of Interest (L18-1)

Copied to clipboard

Challenge: a wordnet browser that allows to consult wordnet content is presented in this paper . the paper presents a browser that meets design requirements and complies with the most ample range of design features.
Approach: They propose a wordnet browser that meets design requirements for wordnets . they use existing browsers to analyze their functionalities and build a new browser .
Outcome: The proposed browser meets design requirements and complies with the most ample range of design features.
A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers (2020.acl-main)

Copied to clipboard

Challenge: Existing MWP corpora are limited in language patterns and problem types . a new corpus of 2,305 MWps is proposed that is more diverse in terms of lexicon usage .
Approach: They propose to use ASDiv to measure lexicon usage diversity of a given MWP corpus.
Outcome: The proposed corpus covers more problem types and text patterns than existing corpora and reflects the true capability of solvers more faithfully.
ENGLAWI: From Human- to Machine-Readable Wiktionary (2020.lrec-1)

Copied to clipboard

Challenge: ENGLAWI is a structured and normalized version of the English Wiktionary encoded into a workable XML format.
Approach: They introduce ENGLAWI, a large, versatile, XML-encoded machine-readable dictionary extracted from Wiktionary.
Outcome: The proposed lexicographic word embeddings are based on the ENGLAWI definitions and are available for download and are supplied with G-PeTo scripts.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Semantic Frame Induction from a Real-World Corpus (2025.acl-srw)

Copied to clipboard

Challenge: Existing studies on semantic frame induction have demonstrated that pre-trained language models (PLMs) have led to more accurate results.
Approach: They conduct semantic frame induction using the Colossal Clean Crawled Corpus and assess the applicability of existing frame inducing methods to real-world data.
Outcome: The proposed methods outperform existing methods on real-world data and can induce frames corresponding to novel concepts.
Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus (L18-1)

Copied to clipboard

Challenge: et al. (2017): WiFiNE annotated with fine-grained entity types . lack of a well-established training corpus makes it difficult to manually annotate the amount of data needed for training.
Approach: They propose an English corpus annotated with fine-grained entity types based on Wikipedia . they use heuristics to build a large, high quality, annotating corpus using 2 manually annotized benchmarks .
Outcome: The proposed system outperforms the existing systems with two datasets and gains a 2.8 macro F1 score.
Inferences for Lexical Semantic Resource Building with Less Supervision (2020.lrec-1)

Copied to clipboard

Challenge: lexical semantic resources may be built using various approaches such as extraction from corpora, integration of relevant pieces of knowledge from pre-existing knowledge resources and endogenous inference.
Approach: They propose a method where the resource building process appears as a self learning process . they propose lexical and semantic resource building based on inference .
Outcome: The proposed method reduces the human effort needed for lexical semantic resource building.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations