Annotating Chinese Word Senses with English WordNet: A Practice on OntoNotes Chinese Sense Inventories (2024.lrec-main)
Copied to clipboard
| Challenge: | a recent study has shown that large language models can be useful for cross-lingual applications. |
| Approach: | They propose to annotate Chinese word senses using English WordNet synsets . they examine the relationship between two annotators and find patterns among synset . |
| Outcome: | The proposed method shows that the annotators agree on 38% of the synsets compared with the original synset . the results highlight similarities between the synnotated synset and the WordNet structure . |
Similar Papers
Automatic Wordnet Mapping: from CoreNet to Princeton WordNet (L18-1)
Copied to clipboard
| Challenge: | Existing mappings focus on identifying the semantic categories of CoreNet, but not the word senses. |
| Approach: | They propose to map the word senses of CoreNet into Princeton WordNet synsets by lexical relations by a taxonomy. |
| Outcome: | The proposed mapping bridging the gap between CoreNet and WordNet shows that the word senses of CoreNet are mapped with precision of 91.2%. |
LanguageNet: Learning to Find Sense Relevant Example Sentences (C18-2)
Copied to clipboard
| Challenge: | LanguageNet is a system that can help second language learners to search for different meanings and usages of a word . the polysemy of words, namely words with more than one sense, is one of the major challenges for ESOL learners . |
| Approach: | They propose a system which can help second language learners to search for different meanings of a word. |
| Outcome: | The proposed system can help second language learners to search for different meanings and usages of a word. |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
Incorporating Chinese Characters of Words for Lexical Sememe Prediction (P18-1)
Copied to clipboard
| Challenge: | Existing methods of lexical sememe prediction rely on external context information of words to represent meaning. |
| Approach: | They propose a character-enhanced sememe prediction framework for Chinese language that takes advantage of internal character information and external context information. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on a Chinese sememe knowledge base and maintains robust performance even for low-frequency words. |
WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations (N19-1)
Copied to clipboard
| Challenge: | Existing word embeddings cannot model the dynamic nature of words’ semantics, i.e., the property of words to correspond to potentially different meanings. |
| Approach: | They propose a large-scale Word in Context dataset, called WiC, which is curated by experts and can be used to evaluate context-sensitive representations. |
| Outcome: | The proposed models outperform the standard evaluation dataset for the purpose and highlight their shortcomings. |
A Dataset of Translational Equivalents Built on the Basis of plWordNet-Princeton WordNet Synset Mapping (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 11,000 Polish-English translational equivalents is presented . the dataset is a novum in the wordnet domain and can facilitate the precision of bilingual NLP tasks. |
| Approach: | They present a dataset of Polish-English translational equivalents linked by three types of equivalence links. |
| Outcome: | The proposed dataset contains 11,000 Polish-English translational equivalents . the resulting subsets are based on a manual annotation process and a set of formal features . |
Latent semantic network induction in the context of linked example senses (D19-55)
Copied to clipboard
| Challenge: | Using the Princeton WordNet, we construct a network using the entirety of Wiktionary. |
| Approach: | They propose to use Wiktionary to construct a wordnet using the entirety of the open-source dictionary. |
| Outcome: | The proposed network induction process is similar to the Princeton WordNet, but with a more data-driven approach. |
Sense and Sentiment (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing sentiment lexicons and concept-based sentiment-tagged corpora are not accurate, and it is difficult to map sentiment scores accurately to different languages. |
| Approach: | They examine existing sentiment lexicons and sense-based sentiment-tagged corpora to find out how sense and concept-based semantic relations effect sentiment scores. |
| Outcome: | The proposed lexicon can be used to generate sentiment lexicos for English using the Open Multilingual Wordnet. |
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)
Copied to clipboard
| Challenge: | WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them. |
| Approach: | This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings. |
| Outcome: | The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets. |
WordNet under Scrutiny: Dictionary Examples in the Era of Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Lexical resources are a repository of knowledge and are used for many tasks, including word sense disambiguation and etymology. |
| Approach: | They compare WordNet, the most commonly used lexical resource in NLP, with a variety of dictionaries and examples that were generated by ChatGPT. |
| Outcome: | The most commonly used lexical resource in NLP, with a variety of dictionaries and examples that were generated by ChatGPT. |