| Challenge: | Existing thesaurus descriptions are expensive and time-consuming, but there are ways to maintain and improve them. |
| Approach: | They propose to apply a checking procedure to existing thesaurus to find errors . they found errors in word sense descriptions, including inaccurate relationships . |
| Outcome: | The proposed method can reveal errors in thesaurus descriptions, but it is much harder to detect them than with manual methods. |
Similar Papers
Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora (2020.acl-main)
Copied to clipboard
| Challenge: | comparing two corpus texts and searching for words that differ in their usage between them is a common problem in digital humanities and computational social science. |
| Approach: | They propose an alternative approach that does not use vector space alignment, and instead considers the neighbors of each word. |
| Outcome: | The proposed method is interpretable and stable in 9 different setups and is highly reliable. |
Aligning Wikipedia with WordNet:a Review and Evaluation of Different Techniques (2020.lrec-1)
Copied to clipboard
| Challenge: | a reliable alignment between WordNet and Wikipedia is a valuable resource for the creation of new wordnets in other languages and for the development of existing wordnet. |
| Approach: | They evaluate methods for aligning Wikipedia articles with WordNet synsets . they use a new gold and silver standard and a method that creates wordnets in other languages . |
| Outcome: | The proposed methods can be used to evaluate the quality of alignments between Wikipedia and WordNet synsets. |
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)
Copied to clipboard
| Challenge: | Terms are notoriously difficult to identify, both automatically and manually. |
| Approach: | They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information . |
| Outcome: | The proposed method provides a tool for evaluation and rich source of information about terms. |
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)
Copied to clipboard
| Challenge: | WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them. |
| Approach: | This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings. |
| Outcome: | The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets. |
Some Issues with Building a Multilingual Wordnet (2020.lrec-1)
Copied to clipboard
| Challenge: | Notable extensions include: confidence, corpus frequency, orthographic variants, lexicalized and non-lexicalised synsets and lemmas, new parts of speech, and more. |
| Approach: | They propose to integrate a new open multilingual wordnet format that tests the extensions introduced by the new format and integrates a set of tools to ensure the integrity of the Collaborative Interlingual Index. |
| Outcome: | The proposed format integrates a set of tools that test the extensions while ensuring the integrity of the Collaborative Interlingual Index (CILI). |
LanguageNet: Learning to Find Sense Relevant Example Sentences (C18-2)
Copied to clipboard
| Challenge: | LanguageNet is a system that can help second language learners to search for different meanings and usages of a word . the polysemy of words, namely words with more than one sense, is one of the major challenges for ESOL learners . |
| Approach: | They propose a system which can help second language learners to search for different meanings of a word. |
| Outcome: | The proposed system can help second language learners to search for different meanings and usages of a word. |
Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment (2023.emnlp-main)
Copied to clipboard
| Challenge: | a new method for finding semantic differences in words appears in two corpora, but it requires a variance of word vectors . a word covers more meanings in a corpus, and its mean word vector becomes shorter . |
| Approach: | They propose a method to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector. |
| Outcome: | The proposed methods rival the best-performing system in the SemEval-2020 Task 1 . they are robust for the skew in corpus sizes and capable of detecting infrequent words . |
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
Extrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a frequency-based method, we can identify subsets of the same word contexts without any reference data. |
| Approach: | They compare 11 different French dependency parsers on a specialized corpus to generate distributional thesauri using a frequency-based method. |
| Outcome: | The proposed method can identify relevant subsets without reference data and the similarity is confirmed on a restricted distributional benchmark. |