Fahad Khan, Francisco J. Minaya Gómez, Rafael Cruz González, Harry Diakoff, Javier E. Diaz Vera, John P. McCrae, Ciara O’Loughlin, William Michael Short, Sander Stolk
| Challenge: | In this paper we discuss our preliminary work towards the construction of a WordNet for Old English, taking our inspiration from other similar WN construction projects for ancient languages such as Ancient Greek, Latin and Sanskrit. |
| Approach: | They propose to use a legacy Old English dictionary to build a WordNet for Old English using a lexicographic resource and the naisc system to automatically compile a provisional version of the WordNet. |
| Outcome: | The proposed OldEWN will be based on lemmas and definitions extracted from a legacy Old English dictionary and will be automatically compile and enriched by experts using the naisc system. |
Similar Papers
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)
Copied to clipboard
| Challenge: | WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them. |
| Approach: | This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings. |
| Outcome: | The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets. |
Building the Old Javanese Wordnet (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing wordnets for Old Javanese are limited and lack of an open-source version of the language is a barrier to its development. |
| Approach: | They propose to build a machine readable resource for Old Javanese using the Princeton Wordnet's synsets and semantic hierarchy. |
| Outcome: | The wordnet contains 2,054 concepts or synsets and 5,911 senses. |
Browsing and Supporting Pluricentric Global Wordnet, or just your Wordnet of Interest (L18-1)
Copied to clipboard
| Challenge: | a wordnet browser that allows to consult wordnet content is presented in this paper . the paper presents a browser that meets design requirements and complies with the most ample range of design features. |
| Approach: | They propose a wordnet browser that meets design requirements for wordnets . they use existing browsers to analyze their functionalities and build a new browser . |
| Outcome: | The proposed browser meets design requirements and complies with the most ample range of design features. |
Enriching Linguistic Representation in the Cantonese Wordnet and Building the New Cantonese Wordnet Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, our wordnet includes a little over 5,200 concepts and 16,300 senses . |
| Approach: | They propose to improve the Cantonese Wordnet by increasing the general coverage, adding functional categories, enriching verbal representations and creating the Cannese WordNet Corpus . |
| Outcome: | The new version includes a little over 5,200 concepts and 16,300 senses . |
Preserving Semantic Information from Old Dictionaries: Linking Senses of the ‘Altfranzösisches Wörterbuch’ to WordNet (2020.lrec-1)
Copied to clipboard
| Challenge: | Historical dictionaries of the pre-digital period are important resources for the study of older languages. |
| Approach: | They propose to use printed dictionaries to create a more easily accessible and more sustainable lexical database by automating the conversion process. |
| Outcome: | The ‘Altfranzösisches Wörterbuch’, an Old French dictionary published from 1925 onwards, shows how the printed dictionaries can be turned into a more easily accessible and more sustainable lexical database. |
Some Issues with Building a Multilingual Wordnet (2020.lrec-1)
Copied to clipboard
| Challenge: | Notable extensions include: confidence, corpus frequency, orthographic variants, lexicalized and non-lexicalised synsets and lemmas, new parts of speech, and more. |
| Approach: | They propose to integrate a new open multilingual wordnet format that tests the extensions introduced by the new format and integrates a set of tools to ensure the integrity of the Collaborative Interlingual Index. |
| Outcome: | The proposed format integrates a set of tools that test the extensions while ensuring the integrity of the Collaborative Interlingual Index (CILI). |
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
Latent semantic network induction in the context of linked example senses (D19-55)
Copied to clipboard
| Challenge: | Using the Princeton WordNet, we construct a network using the entirety of Wiktionary. |
| Approach: | They propose to use Wiktionary to construct a wordnet using the entirety of the open-source dictionary. |
| Outcome: | The proposed network induction process is similar to the Princeton WordNet, but with a more data-driven approach. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
IndoUKC: A Concept-Centered Indian Multilingual Lexical Resource (2022.lrec-1)
Copied to clipboard
| Challenge: | a new multilingual lexical database for Indian languages is proposed . the database provides words and crosslingually mapped word meanings specific to Indian languages and cultures. |
| Approach: | They propose to create a multilingual lexical database for Indian languages called IndoUKC . the database is based on existing IndoWordNet resources and is available for browsing . |
| Outcome: | The proposed database is based on the existing IndoWordNet resource and is available for download through the LiveLanguage data catalogue. |