| Challenge: | a new lexical resource called CzEngClass is being built to help define synonyms in a bilingual context. |
| Approach: | They propose to group verb senses into bilingual verbal synonym groups and use a parallel dependency corpus to explore semantic 'equivalence' they argue that existence of core argument mappings and adjunct mappings to a common set of semantic roles is a suitable criterion for a reasonable verb synonymy definition . |
| Outcome: | The proposed resource will be available by mid-2018 . |
Similar Papers
Synonymy in Bilingual Context: The CzEngClass Lexicon (C18-1)
Copied to clipboard
| Challenge: | Existing lexical resources for semantic annotation of synonyms are lacking in computational language processing. |
| Approach: | They describe a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and relate semantic roles common to one synonym class to verb arguments. |
| Outcome: | The proposed resource is based on English and Czech WordNet, FrameNet, PropBank, VerbNet (SemLink), and valency lexicons for Czech and English (PDT-Vallex, Vallex, and EngValleX). |
Tools for Building an Interlinked Synonym Lexicon Network (L18-1)
Copied to clipboard
| Challenge: | a new lexicon is being developed for cross-lingual (Czech and English) synonyms based on their syntactic and semantic behavior in (bilingual) context. |
| Approach: | They propose to build a new interlinked verbal synonym lexicon called CzEngClass using a tool that helps to keep cross-lingual synonym classes consistent. |
| Outcome: | The proposed lexicon captures cross-lingual (Czech and English) synonyms . the tool, called Synonym Class Editor -SynEd, is customized to build and edit entries . |
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)
Copied to clipboard
| Challenge: | WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them. |
| Approach: | This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings. |
| Outcome: | The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets. |
A Broad-Coverage Deep Semantic Lexicon for Verbs (2020.lrec-1)
Copied to clipboard
| Challenge: | a lack of a broad-coverage deep semantic lexicon hinders deep language understanding . we have developed a resource for verbs with the coverage of WordNet and syntactic and semantic details . |
| Approach: | They propose a deep lexical resource for verbs with the coverage of WordNet and syntactic and semantic details that meet or exceed existing resources. |
| Outcome: | The proposed resource has the coverage of WordNet and syntactic and semantic details that exceed existing resources. |
Using Wiktionary to Create Specialized Lexical Resources and Datasets (2022.lrec-1)
Copied to clipboard
| Challenge: | Using Wiktionary data to build specialized lexical datasets can be used for evaluating or improving NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Machine Translation (MT). |
| Approach: | They propose to use Wiktionary data to create specialized lexical datasets that can be used for evaluating or improving NLP tasks. |
| Outcome: | The proposed datasets can be used to improve and/or evaluate NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Sense Linking (SL), or machine translation (MT). |
SynET: Synonym Expansion using Transitivity (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to find synonyms from text corpora are distributed and pattern based, but they suffer from low precision and low recall. |
| Approach: | They propose a task of synonym expansion using transitivity and propose auxiliary task to reduce the impact of noisy sentences. |
| Outcome: | The proposed approach reduces the impact of noisy sentences and reduces noise in a real-world dataset. |
A Dataset of Translational Equivalents Built on the Basis of plWordNet-Princeton WordNet Synset Mapping (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 11,000 Polish-English translational equivalents is presented . the dataset is a novum in the wordnet domain and can facilitate the precision of bilingual NLP tasks. |
| Approach: | They present a dataset of Polish-English translational equivalents linked by three types of equivalence links. |
| Outcome: | The proposed dataset contains 11,000 Polish-English translational equivalents . the resulting subsets are based on a manual annotation process and a set of formal features . |
Translation-based Lexicalization Generation and Lexical Gap Detection: Application to Kinship Terms (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for identifying lexical gaps have been limited . kinship terms are well-suited for investigations into lexicons and lexicals . |
| Approach: | They propose an algorithm to automatically generate concept lexicalizations based on machine translation and hypernymy relations between concepts. |
| Outcome: | Empirical evaluations show that the proposed method is more accurate than BabelNet and ChatGPT. |
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |
Leveraging Meta-Embeddings for Bilingual Lexicon Extraction from Specialized Comparable Corpora (C18-1)
Copied to clipboard
| Challenge: | Recent studies on bilingual lexicon extraction from specialized comparable corpora show differences in performance . lack of large specialized corporan to build efficient representations can be partially explained . |
| Approach: | They propose to use character-based embedding models to combine different embeddable models . they emphasize how character-driven embeddance models outperform other models on quality . |
| Outcome: | The proposed model outperforms other models on quality of extracted bilingual lexicons . comparable corpora are an interesting and practical alternative to parallel corporation . |