Towards the Construction of a WordNet for Old English (2022.lrec-1)

Copied to clipboard

Challenge: In this paper we discuss our preliminary work towards the construction of a WordNet for Old English, taking our inspiration from other similar WN construction projects for ancient languages such as Ancient Greek, Latin and Sanskrit.
Approach: They propose to use a legacy Old English dictionary to build a WordNet for Old English using a lexicographic resource and the naisc system to automatically compile a provisional version of the WordNet.
Outcome: The proposed OldEWN will be based on lemmas and definitions extracted from a legacy Old English dictionary and will be automatically compile and enriched by experts using the naisc system.

Similar Papers

A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
Building the Old Javanese Wordnet (2020.lrec-1)

Copied to clipboard

Challenge: Existing wordnets for Old Javanese are limited and lack of an open-source version of the language is a barrier to its development.
Approach: They propose to build a machine readable resource for Old Javanese using the Princeton Wordnet's synsets and semantic hierarchy.
Outcome: The wordnet contains 2,054 concepts or synsets and 5,911 senses.
Browsing and Supporting Pluricentric Global Wordnet, or just your Wordnet of Interest (L18-1)

Copied to clipboard

Challenge: a wordnet browser that allows to consult wordnet content is presented in this paper . the paper presents a browser that meets design requirements and complies with the most ample range of design features.
Approach: They propose a wordnet browser that meets design requirements for wordnets . they use existing browsers to analyze their functionalities and build a new browser .
Outcome: The proposed browser meets design requirements and complies with the most ample range of design features.
Enriching Linguistic Representation in the Cantonese Wordnet and Building the New Cantonese Wordnet Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Currently, our wordnet includes a little over 5,200 concepts and 16,300 senses .
Approach: They propose to improve the Cantonese Wordnet by increasing the general coverage, adding functional categories, enriching verbal representations and creating the Cannese WordNet Corpus .
Outcome: The new version includes a little over 5,200 concepts and 16,300 senses .
Preserving Semantic Information from Old Dictionaries: Linking Senses of the ‘Altfranzösisches Wörterbuch’ to WordNet (2020.lrec-1)

Copied to clipboard

Challenge: Historical dictionaries of the pre-digital period are important resources for the study of older languages.
Approach: They propose to use printed dictionaries to create a more easily accessible and more sustainable lexical database by automating the conversion process.
Outcome: The ‘Altfranzösisches Wörterbuch’, an Old French dictionary published from 1925 onwards, shows how the printed dictionaries can be turned into a more easily accessible and more sustainable lexical database.
Some Issues with Building a Multilingual Wordnet (2020.lrec-1)

Copied to clipboard

Challenge: Notable extensions include: confidence, corpus frequency, orthographic variants, lexicalized and non-lexicalised synsets and lemmas, new parts of speech, and more.
Approach: They propose to integrate a new open multilingual wordnet format that tests the extensions introduced by the new format and integrates a set of tools to ensure the integrity of the Collaborative Interlingual Index.
Outcome: The proposed format integrates a set of tools that test the extensions while ensuring the integrity of the Collaborative Interlingual Index (CILI).
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)

Copied to clipboard

Challenge: Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available.
Approach: They propose to use a contextualised language model to analyse historical states of language in French.
Outcome: The proposed model is based on a corpus of historical texts and is evaluated with an NLP task.
Latent semantic network induction in the context of linked example senses (D19-55)

Copied to clipboard

Challenge: Using the Princeton WordNet, we construct a network using the entirety of Wiktionary.
Approach: They propose to use Wiktionary to construct a wordnet using the entirety of the open-source dictionary.
Outcome: The proposed network induction process is similar to the Princeton WordNet, but with a more data-driven approach.
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
IndoUKC: A Concept-Centered Indian Multilingual Lexical Resource (2022.lrec-1)

Copied to clipboard

Challenge: a new multilingual lexical database for Indian languages is proposed . the database provides words and crosslingually mapped word meanings specific to Indian languages and cultures.
Approach: They propose to create a multilingual lexical database for Indian languages called IndoUKC . the database is based on existing IndoWordNet resources and is available for browsing .
Outcome: The proposed database is based on the existing IndoWordNet resource and is available for download through the LiveLanguage data catalogue.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations