Challenge: TEI and OntoLex deal with corpus citations in lexicons.
Approach: They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding .
Outcome: The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary .

Similar Papers

Modelling Frequency, Attestation, and Corpus-Based Information with OntoLex-FrAC (2022.coling-1)

Copied to clipboard

Challenge: OntoLex-Lemon has become a de facto standard for lexical resources in the web of data.
Approach: This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information.
Outcome: This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information (OntoLx-FrAC) it is intended to complement OntoLemon with the vocabulary to represent major types of information found in or automatically derived from corpora, for applications in both language technology and the language sciences.
Bridging Computational Lexicography and Corpus Linguistics: A Query Extension for OntoLex-FrAC (2024.lrec-main)

Copied to clipboard

Challenge: OntoLex is the dominant community standard for machine-readable lexical resources . it is currently extended with a designated module for Frequency, Attestations and Corpus-based Information .
Approach: They propose a module for Frequency, Attestations and Corpus-based Information for OntoLex . the module enables RDF-based web services to exchange corpus queries dynamically .
Outcome: The proposed module addresses the incorporation of corpus queries for linking dictionaries with corpus engines and enabling RDF-based web services to exchange corpus query data dynamically.
Representing Compounding with OntoLex. An Evaluation of Vocabularies for Word Formation Resources (2024.lrec-main)

Copied to clipboard

Challenge: OntoLex is a de facto standard for the modelling of lexical resources in the framework of Linguistic Linked Open Data.
Approach: They propose to use OntoLex to convert Linked Open Data into compounds by using the RDF model.
Outcome: The proposed model can be applied to all resources harmonized in that format, potentially allowing for the conversion into Linked Open Data of a large amount of structured data.
Modelling Etymology in LMF/TEI: The Grande Dicionário Houaiss da Língua Portuguesa Dictionary as a Use Case (2020.lrec-1)

Copied to clipboard

Challenge: In this article, we will introduce two of the new parts of the Lexical Markup Framework (LMF) ISO standard . part 3 deals with etymological and diachronic data and part 4 consists of a TEI serialisation of all of the prior parts of TEIS model.
Approach: They introduce two parts of the Lexical Markup Framework (LMF) ISO standard, part 3 dealing with etymological and diachronic data and part 4 containing TEI serialisation of all prior parts of a model.
Outcome: The proposed models are based on examples taken from a Portuguese dictionary conversion and are then compared with TEI-XML models.
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
Inferences for Lexical Semantic Resource Building with Less Supervision (2020.lrec-1)

Copied to clipboard

Challenge: lexical semantic resources may be built using various approaches such as extraction from corpora, integration of relevant pieces of knowledge from pre-existing knowledge resources and endogenous inference.
Approach: They propose a method where the resource building process appears as a self learning process . they propose lexical and semantic resource building based on inference .
Outcome: The proposed method reduces the human effort needed for lexical semantic resource building.
The Return of Lexical Dependencies: Neural Lexicalized PCFGs (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches to grammar induction focus on discovering constituents or dependencies.
Approach: They propose to model lexical dependencies using context free grammars instead of lexicals . they show that this unified framework induces both constituents and dependencies .
Outcome: The proposed model overcomes sparsity problems and induces constituents and dependencies better than the current methods.
Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0 (2020.lrec-1)

Copied to clipboard

Challenge: Diachronic lexical information is increasingly used in historical linguistics and in NLP . etymological resources need to be fine-grained, large-coverage and accurate .
Approach: They propose guidelines to generate etymological lexical resources for each step of the life-cycle of an ethymology . they introduce EtymDB 2.0, an 'etiological database' generated from the Wiktionary .
Outcome: The proposed resources are generated for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation.
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)

Copied to clipboard

Challenge: OntoLex is a widely used community standard for machine-readable lexical resources on the web.
Approach: They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis.
Outcome: The proposed module can be used to represent morphological resources on a unified basis.
Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus (L18-1)

Copied to clipboard

Challenge: et al. (2017): WiFiNE annotated with fine-grained entity types . lack of a well-established training corpus makes it difficult to manually annotate the amount of data needed for training.
Approach: They propose an English corpus annotated with fine-grained entity types based on Wikipedia . they use heuristics to build a large, high quality, annotating corpus using 2 manually annotized benchmarks .
Outcome: The proposed system outperforms the existing systems with two datasets and gains a 2.8 macro F1 score.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations