Challenge: WordNet-like resources are lexical databases with highly relevance information and data that could be exploited in more complex computational linguistics research and applications.
Approach: They propose to build a WordNet database for a low-resourced and indigenous language in Peru . they propose to use word2vec similarity to compare definition glosses in a dictionary with the content of a Spanish WordNet .
Outcome: The proposed database is based on a bilingual dictionary written in Spanish and an automatic evaluation process using a manually annotated Gold Standard in Shipibo-Koniba.

Similar Papers

WordNet-QU: Development of a Lexical Database for Quechua Varieties (2022.coling-1)

Copied to clipboard

Challenge: Quechua is a low-resource language from south America but lacks resources to build high-performance computational systems.
Approach: They propose to include Quechua in a lexical database called wordnet . they propose a synset alignment algorithm to compare Quechuan to its nearest high-resource language .
Outcome: The proposed system compares Quechua to its nearest high-resource language, Spanish . it uses a synset alignment algorithm to find Quechuan resources in a lexical database .
No Data to Crawl? Monolingual Corpus Creation from PDF Files of Truly low-Resource Languages in Peru (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for extracting text from PDF files are expensive and limited by the absence of web content of endangered languages.
Approach: They propose a method for creating monolingual corpora for four endangered languages . they use a PDF file format with multilingual sentences and noisy pages .
Outcome: The proposed method allows the creation of clean corpora for the four languages, a key resource for natural language processing tasks nowadays.
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
A Dataset of Translational Equivalents Built on the Basis of plWordNet-Princeton WordNet Synset Mapping (2020.lrec-1)

Copied to clipboard

Challenge: a dataset of 11,000 Polish-English translational equivalents is presented . the dataset is a novum in the wordnet domain and can facilitate the precision of bilingual NLP tasks.
Approach: They present a dataset of Polish-English translational equivalents linked by three types of equivalence links.
Outcome: The proposed dataset contains 11,000 Polish-English translational equivalents . the resulting subsets are based on a manual annotation process and a set of formal features .
ChAnot: An Intelligent Annotation Tool for Indigenous and Highly Agglutinative Languages in Peru (L18-1)

Copied to clipboard

Challenge: Linguistic corpus annotation is one of the most important phases for addressing natural language processing (NLP) tasks.
Approach: They propose a web-based annotation tool for Peruvian indigenous and highly agglutinative languages that supports a variety of linguistic annotation tasks.
Outcome: The proposed tool supports a diverse set of linguistic annotation tasks, such as morphological segmentation markup, POS-tag markup and other.
Using Wiktionary to Create Specialized Lexical Resources and Datasets (2022.lrec-1)

Copied to clipboard

Challenge: Using Wiktionary data to build specialized lexical datasets can be used for evaluating or improving NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Machine Translation (MT).
Approach: They propose to use Wiktionary data to create specialized lexical datasets that can be used for evaluating or improving NLP tasks.
Outcome: The proposed datasets can be used to improve and/or evaluate NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Sense Linking (SL), or machine translation (MT).
Evaluating Word Embeddings in Extremely Under-Resourced Languages: A Case Study in Bribri (2022.coling-1)

Copied to clipboard

Challenge: a case study of word embeddings in an under-resourced context is presented . Embeddings are critical for many NLP tasks, but their evaluation in actual under-sourced settings needs further investigation.
Approach: They adapt word embeddings to an under-resourced indigenous language in Bribri . they find that the best models find the appropriate semantic target 60% of the time .
Outcome: The proposed models were adapted to an under-resourced English language . the best models found the appropriate semantic target 60% of the time .
Aligning Wikipedia with WordNet:a Review and Evaluation of Different Techniques (2020.lrec-1)

Copied to clipboard

Challenge: a reliable alignment between WordNet and Wikipedia is a valuable resource for the creation of new wordnets in other languages and for the development of existing wordnet.
Approach: They evaluate methods for aligning Wikipedia articles with WordNet synsets . they use a new gold and silver standard and a method that creates wordnets in other languages .
Outcome: The proposed methods can be used to evaluate the quality of alignments between Wikipedia and WordNet synsets.
Empowering Low-Resource Regional Languages with Lexicons : A Comparative Study of NLP Tools for Morphosyntactic Analysis (2024.lrec-main)

Copied to clipboard

Challenge: a lack of human and financial resources makes integrating lexicon information to low-resource languages challenging.
Approach: They propose to use a bilingual lexicon to integrate lexical information to low-resource language . they compare a lexiconal approach to a neural approach that uses a larger lexicone .
Outcome: The proposed approach improves POS tagging while using different lexicon sizes.
Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0 (2020.lrec-1)

Copied to clipboard

Challenge: Diachronic lexical information is increasingly used in historical linguistics and in NLP . etymological resources need to be fine-grained, large-coverage and accurate .
Approach: They propose guidelines to generate etymological lexical resources for each step of the life-cycle of an ethymology . they introduce EtymDB 2.0, an 'etiological database' generated from the Wiktionary .
Outcome: The proposed resources are generated for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations