Challenge: The central word register for Danish is an open source lexicon project for general AI purposes funded and initiated by the Danish Agency for Digitisation in 2020.
Approach: They propose to use existing fine-grained sense inventory to compile a more AI-appropriate sense granularity level of the vocabulary.
Outcome: The proposed lexical resource is based on the fine-grained sense inventory from Den Danske Ordbog (DDO) it is designed to be more practical and suitable for AI, omitting outdated language and slang, merging subtle and rare sub-senses with their main sense, disregarding sub-domains, etc.

Similar Papers

CEFR-based Lexical Simplification Dataset (L18-1)

Copied to clipboard

Challenge: Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective.
Approach: They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile .
Outcome: The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
DocCGen: Document-based Controlled Code Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) produce state-of-the-art performance on natural language to code generation for resource-rich general-purpose languages like C++, Java, and Python.
Approach: They propose a framework that breaks the NL-to-Code generation task into two steps . they use library documentation to detect the correct libraries and schema rules extracted from the documentation to constrain the decoding .
Outcome: The proposed framework improves different sized language models across all six evaluation metrics, reducing syntactic and semantic errors in structured code.
Towards a Danish Semantic Reasoning Benchmark - Compiled from Lexical-Semantic Resources for Assessing Selected Language Understanding Capabilities of Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: a semantic reasoning benchmark for Danish is compiled from human-curated lexical-semantic resources.
Approach: They present a semantic reasoning benchmark for Danish compiled semi-automatically from a number of human-curated lexical-semantic resources.
Outcome: The proposed datasets are compiled semi-automatically from human-curated lexical-semantic resources.
Annotating the French Wiktionary with supersenses for large scale lexical analysis: a use case to assess form-meaning relationships within the nominal lexicon (2025.coling-main)

Copied to clipboard

Challenge: Conducting large-scale empirical studies in lexical semantics remains an elusive goal for many languages lacking comprehensive semantic resources.
Approach: They propose to use the Princeton WordNet to enrich the French Wiktionary with general semantic classes, known as supersenses, using a limited amount of manually annotated data.
Outcome: The proposed method can be extended to other languages provided an electronic lexicon and manually annotated senses are available.
MoNoise: A Multi-lingual and Easy-to-use Lexical Normalization Tool (P19-3)

Copied to clipboard

Challenge: In this paper, we demonstrate the online demo and command line interface of a lexical normalization system (MoNoise) for a variety of languages.
Approach: They propose to bundle seven datasets in six languages to form a new benchmark and a novel evaluation metric which is particularly suitable for cross-dataset comparisons.
Outcome: The proposed model is based on the original word and features from the original language for each normalization candidate.
Let’s Play Mono-Poly: BERT Can Reveal Words’ Polysemy Level and Partitionability into Senses (2021.tacl-1)

Copied to clipboard

Challenge: Pre-trained language models encode rich information about linguistic structure but their knowledge about lexical polysemy remains unclear.
Approach: They propose a setup for analyzing lexical polysemy knowledge in pre-trained language models and multilingual BERT models by analyzing different sense distributions and controlling for parameters that are highly correlated with polysyntax.
Outcome: The proposed model can be used to analyze lexical polysemy in English, French, Spanish, and Greek and in multilingual BERT.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text (2024.emnlp-main)

Copied to clipboard

Challenge: Existing resources for standardized, easily accessible IGT data limit their applicability to linguistic research.
Approach: They compile the largest existing corpus of interlinear glossed text data from a variety of sources and use it to generate annotated text.
Outcome: The proposed model outperforms SOTA models on monolingual corpora by 6.6%.
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources .
Approach: They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment .
Outcome: The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations