Corpus-based Check-up for Thesaurus (P19-1)

Copied to clipboard

Challenge: Existing thesaurus descriptions are expensive and time-consuming, but there are ways to maintain and improve them.
Approach: They propose to apply a checking procedure to existing thesaurus to find errors . they found errors in word sense descriptions, including inaccurate relationships .
Outcome: The proposed method can reveal errors in thesaurus descriptions, but it is much harder to detect them than with manual methods.

Similar Papers

Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora (2020.acl-main)

Copied to clipboard

Challenge: comparing two corpus texts and searching for words that differ in their usage between them is a common problem in digital humanities and computational social science.
Approach: They propose an alternative approach that does not use vector space alignment, and instead considers the neighbors of each word.
Outcome: The proposed method is interpretable and stable in 9 different setups and is highly reliable.
Aligning Wikipedia with WordNet:a Review and Evaluation of Different Techniques (2020.lrec-1)

Copied to clipboard

Challenge: a reliable alignment between WordNet and Wikipedia is a valuable resource for the creation of new wordnets in other languages and for the development of existing wordnet.
Approach: They evaluate methods for aligning Wikipedia articles with WordNet synsets . they use a new gold and silver standard and a method that creates wordnets in other languages .
Outcome: The proposed methods can be used to evaluate the quality of alignments between Wikipedia and WordNet synsets.
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)

Copied to clipboard

Challenge: Terms are notoriously difficult to identify, both automatically and manually.
Approach: They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information .
Outcome: The proposed method provides a tool for evaluation and rich source of information about terms.
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
Some Issues with Building a Multilingual Wordnet (2020.lrec-1)

Copied to clipboard

Challenge: Notable extensions include: confidence, corpus frequency, orthographic variants, lexicalized and non-lexicalised synsets and lemmas, new parts of speech, and more.
Approach: They propose to integrate a new open multilingual wordnet format that tests the extensions introduced by the new format and integrates a set of tools to ensure the integrity of the Collaborative Interlingual Index.
Outcome: The proposed format integrates a set of tools that test the extensions while ensuring the integrity of the Collaborative Interlingual Index (CILI).
LanguageNet: Learning to Find Sense Relevant Example Sentences (C18-2)

Copied to clipboard

Challenge: LanguageNet is a system that can help second language learners to search for different meanings and usages of a word . the polysemy of words, namely words with more than one sense, is one of the major challenges for ESOL learners .
Approach: They propose a system which can help second language learners to search for different meanings of a word.
Outcome: The proposed system can help second language learners to search for different meanings and usages of a word.
Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment (2023.emnlp-main)

Copied to clipboard

Challenge: a new method for finding semantic differences in words appears in two corpora, but it requires a variance of word vectors . a word covers more meanings in a corpus, and its mean word vector becomes shorter .
Approach: They propose a method to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector.
Outcome: The proposed methods rival the best-performing system in the SemEval-2020 Task 1 . they are robust for the skew in corpus sizes and capable of detecting infrequent words .
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
Extrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri (2020.lrec-1)

Copied to clipboard

Challenge: Using a frequency-based method, we can identify subsets of the same word contexts without any reference data.
Approach: They compare 11 different French dependency parsers on a specialized corpus to generate distributional thesauri using a frequency-based method.
Outcome: The proposed method can identify relevant subsets without reference data and the similarity is confirmed on a restricted distributional benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations