On Modelling Corpus Citations in Computational Lexical Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | TEI and OntoLex deal with corpus citations in lexicons. |
| Approach: | They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding . |
| Outcome: | The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary . |
Similar Papers
Modelling Frequency, Attestation, and Corpus-Based Information with OntoLex-FrAC (2022.coling-1)
Copied to clipboard
| Challenge: | OntoLex-Lemon has become a de facto standard for lexical resources in the web of data. |
| Approach: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information. |
| Outcome: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information (OntoLx-FrAC) it is intended to complement OntoLemon with the vocabulary to represent major types of information found in or automatically derived from corpora, for applications in both language technology and the language sciences. |
Bridging Computational Lexicography and Corpus Linguistics: A Query Extension for OntoLex-FrAC (2024.lrec-main)
Copied to clipboard
| Challenge: | OntoLex is the dominant community standard for machine-readable lexical resources . it is currently extended with a designated module for Frequency, Attestations and Corpus-based Information . |
| Approach: | They propose a module for Frequency, Attestations and Corpus-based Information for OntoLex . the module enables RDF-based web services to exchange corpus queries dynamically . |
| Outcome: | The proposed module addresses the incorporation of corpus queries for linking dictionaries with corpus engines and enabling RDF-based web services to exchange corpus query data dynamically. |
Representing Compounding with OntoLex. An Evaluation of Vocabularies for Word Formation Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | OntoLex is a de facto standard for the modelling of lexical resources in the framework of Linguistic Linked Open Data. |
| Approach: | They propose to use OntoLex to convert Linked Open Data into compounds by using the RDF model. |
| Outcome: | The proposed model can be applied to all resources harmonized in that format, potentially allowing for the conversion into Linked Open Data of a large amount of structured data. |
Modelling Etymology in LMF/TEI: The Grande Dicionário Houaiss da Língua Portuguesa Dictionary as a Use Case (2020.lrec-1)
Copied to clipboard
| Challenge: | In this article, we will introduce two of the new parts of the Lexical Markup Framework (LMF) ISO standard . part 3 deals with etymological and diachronic data and part 4 consists of a TEI serialisation of all of the prior parts of TEIS model. |
| Approach: | They introduce two parts of the Lexical Markup Framework (LMF) ISO standard, part 3 dealing with etymological and diachronic data and part 4 containing TEI serialisation of all prior parts of a model. |
| Outcome: | The proposed models are based on examples taken from a Portuguese dictionary conversion and are then compared with TEI-XML models. |
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)
Copied to clipboard
| Challenge: | WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them. |
| Approach: | This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings. |
| Outcome: | The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets. |
Inferences for Lexical Semantic Resource Building with Less Supervision (2020.lrec-1)
Copied to clipboard
| Challenge: | lexical semantic resources may be built using various approaches such as extraction from corpora, integration of relevant pieces of knowledge from pre-existing knowledge resources and endogenous inference. |
| Approach: | They propose a method where the resource building process appears as a self learning process . they propose lexical and semantic resource building based on inference . |
| Outcome: | The proposed method reduces the human effort needed for lexical semantic resource building. |
The Return of Lexical Dependencies: Neural Lexicalized PCFGs (2020.tacl-1)
Copied to clipboard
| Challenge: | Existing approaches to grammar induction focus on discovering constituents or dependencies. |
| Approach: | They propose to model lexical dependencies using context free grammars instead of lexicals . they show that this unified framework induces both constituents and dependencies . |
| Outcome: | The proposed model overcomes sparsity problems and induces constituents and dependencies better than the current methods. |
Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0 (2020.lrec-1)
Copied to clipboard
| Challenge: | Diachronic lexical information is increasingly used in historical linguistics and in NLP . etymological resources need to be fine-grained, large-coverage and accurate . |
| Approach: | They propose guidelines to generate etymological lexical resources for each step of the life-cycle of an ethymology . they introduce EtymDB 2.0, an 'etiological database' generated from the Wiktionary . |
| Outcome: | The proposed resources are generated for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation. |
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)
Copied to clipboard
| Challenge: | OntoLex is a widely used community standard for machine-readable lexical resources on the web. |
| Approach: | They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis. |
| Outcome: | The proposed module can be used to represent morphological resources on a unified basis. |
Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus (L18-1)
Copied to clipboard
| Challenge: | et al. (2017): WiFiNE annotated with fine-grained entity types . lack of a well-established training corpus makes it difficult to manually annotate the amount of data needed for training. |
| Approach: | They propose an English corpus annotated with fine-grained entity types based on Wikipedia . they use heuristics to build a large, high quality, annotating corpus using 2 manually annotized benchmarks . |
| Outcome: | The proposed system outperforms the existing systems with two datasets and gains a 2.8 macro F1 score. |