Papers with Danish
Compiling a Suitable Level of Sense Granularity in a Lexicon for AI Purposes: The Open Source COR Lexicon (2022.lrec-1)
Copied to clipboard
Bolette Pedersen, Nathalie Carmen Hau Sørensen, Sanni Nimb, Ida Flørke, Sussi Olsen, Thomas Troelsgård
| Challenge: | The central word register for Danish is an open source lexicon project for general AI purposes funded and initiated by the Danish Agency for Digitisation in 2020. |
| Approach: | They propose to use existing fine-grained sense inventory to compile a more AI-appropriate sense granularity level of the vocabulary. |
| Outcome: | The proposed lexical resource is based on the fine-grained sense inventory from Den Danske Ordbog (DDO) it is designed to be more practical and suitable for AI, omitting outdated language and slang, merging subtle and rare sub-senses with their main sense, disregarding sub-domains, etc. |
What about “em”? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns (2023.acl-long)
Copied to clipboard
| Challenge: | Wrong pronoun translations can discriminate against marginalized groups, e.g., non-binary individuals. |
| Approach: | They compare 3rd-person pronoun translations to five other languages . they propose to address gender exclusivity in future research . |
| Outcome: | The proposed method compares translations of gendered vs. gender-neutral pronouns from english to five other languages and vice versa. |
Fact from Fiction: Finding Serialized Novels in Newspapers (2025.acl-srw)
Copied to clipboard
| Challenge: | Among underrepresented but widely read forms are serialized fiction and feuilleton novels embedded in newspapers rather than published as standalone volumes. |
| Approach: | They propose to annotate 1,394 articles and evaluate classification pipelines using both selected linguistic features and embeddings to identify serialized fiction and feuilleton fiction. |
| Outcome: | The proposed methods achieve F1-scores of 0.91 in an annotated dataset of 1,394 articles and support the construction of alternative literary corpora and contribute to work on modeling the fiction–nonfiction boundary at scale. |
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)
Copied to clipboard
| Challenge: | linguistic typology is the classification of languages according to their linguistic properties. |
| Approach: | They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale. |
| Outcome: | The proposed model can predict typological properties on a massively multilingual scale. |
Can Demographic Factors Improve Text Classification? Revisiting Demographic Adaptation in the Age of Transformers (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies show that incorporating demographic factors in language representations improves performance on downstream NLP tasks. |
| Approach: | They use continuous language modeling and dynamic multi-task learning to adapt pre-trained Transformers to incorporate demographic information into their representations. |
| Outcome: | The proposed model shows that the results are consistent with previous studies. |
Cross-Lingual Cross-Domain Nested Named Entity Evaluation on English Web Texts (2021.findings-acl)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a key task in Natural Language Processing, but most existing work on NER ignores the recognition of nested entities. |
| Approach: | They propose to annotate five web domains for nested named entities on top of the English Web Treebank (EWT) . they propose to use the English web treebank to perform cross-domain evaluations. |
| Outcome: | The proposed dataset covers five domains and includes transfer results from German and Danish. |
Noise, Novels, Numbers. A Framework for Detecting and Categorizing Noise in Danish and Norwegian Literature (2024.emnlp-main)
Copied to clipboard
Ali Al-Laith, Daniel Hershcovich, Jens Bjerring-Hansen, Jakob Parby, Alexander Conroy, Timothy Tangherlini
| Challenge: | This study examines the literary perceptions of noise during the Scandinavian "Modern Breakthrough" period (1870-1899). |
| Approach: | They propose a framework for detecting and categorizing noise in literary texts from the late 19th century. |
| Outcome: | The proposed framework can be applied to Danish and Norwegian literature from the late 19th century. |
A Dataset of Offensive Language in Kosovo Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media are a central part of people’s lives but are rife with bullying and offensive language, creating an unsafe environment for their users. |
| Approach: | They propose to use user-generated comments on Facebook and YouTube from selected Kosovo news platforms to annotate offensive language in Albanian. |
| Outcome: | The proposed system improves on Danish but not Albanian, on offensive language recognition and distinguishing targeted and untargeted offence. |
MECI: A Multilingual Dataset for Event Causality Identification (2022.coling-1)
Copied to clipboard
| Challenge: | Event Causality Identification (ECI) is a task of detecting causal relations between events mentioned in text. |
| Approach: | They propose a multilingual dataset that provides consistent annotations for event causality relations in five languages. |
| Outcome: | The proposed dataset provides consistent annotation guidelines for five languages . the dataset can provide ample research challenges and directions for future research . |
How Conservative are Language Models? Adapting to the Introduction of Gender-Neutral Pronouns (2022.naacl-main)
Copied to clipboard
| Challenge: | a recent study shows that gender-neutral pronouns are not associated with processing difficulties . linguistic scholars have observed how technology has altered the course of language evolution . |
| Approach: | They show that gender-neutral pronouns in Danish, English and Swedish are not associated with processing difficulties. |
| Outcome: | a new study shows that gender-neutral pronouns are not associated with human processing difficulties . the findings suggest that such conservativity in language models may limit widespread adoption . |
Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing classification models only consider the temporal variations of existing data . current models focus on English corpora, leaving time as domains unexplored . |
| Approach: | They propose a framework to generalize classifiers over time on four languages, English, Danish, French, and German. |
| Outcome: | The proposed framework can generalize classifiers over time on four languages, English, Danish, French, and German. |
Offensive Language and Hate Speech Detection for Danish (2020.lrec-1)
Copied to clipboard
| Challenge: | a growing number of social media platforms are detecting and dealing with offensive language . a recent study found that the best performing system for English is best for Danish . |
| Approach: | They propose automatic methods to detect offensive language on social media platforms . they use user-generated comments from various social media sites to find offensive language . |
| Outcome: | The proposed system performs best for both English and Danish language . it achieves a macro averaged F1-score of 0.74 and a best for Danish achieves 0.73 . |
Development and Evaluation of Pre-trained Language Models for Historical Danish and Norwegian Literary Texts (2024.lrec-main)
Copied to clipboard
| Challenge: | et al., 2019) develop and evaluate the first pre-trained language models specifically tailored for historical Danish and Norwegian texts. |
| Approach: | They develop and evaluate pre-trained language models specifically tailored for historical Danish and Norwegian texts. |
| Outcome: | The proposed model outperforms models trained on historical Danish and Norwegian literature in two downstream NLP tasks. |
Evaluating Word Expansion for Multilingual Sentiment Analysis of Parliamentary Speech (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent efforts to create and format data sets of parliamentary speech material have facilitated cross-lingual comparisons and highlighted the need for methods that are computationally efficient and language-agnostic. |
| Approach: | They propose a word expansion method for sentiment lexicon generation that leverages word embeddings and vector similarity to expand synonym seed lists with domain-specific terms from the speech corpora. |
| Outcome: | The proposed method is compared with other multilingual lexica and is highly sensitive to processing and scoring techniques. |
Towards a Gold Standard for Evaluating Danish Word Embeddings (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing word embedding models resemble semantic similarity solely by distribution, but there seems to be a need for future judgments to measure similarity in full context and along more than a single spectrum. |
| Approach: | They propose a model-agnostic similarity goal standard for evaluating Danish word embeddings based on human judgments made by 42 native speakers of Danish. |
| Outcome: | The goal standard is applied to evaluate Danish word embeddings on 42 native speakers of Danish. |
DaNewsroom: A Large-scale Danish Summarisation Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets for automatic summarisation are English-oriented . however, only very limited datasets exist in languages other than English . |
| Approach: | They present the first large-scale non-English dataset specifically curated for automatic summarisation. |
| Outcome: | The proposed dataset is the first for the Danish language and is compared with existing datasets. |
Towards a Danish Semantic Reasoning Benchmark - Compiled from Lexical-Semantic Resources for Assessing Selected Language Understanding Capabilities of Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | a semantic reasoning benchmark for Danish is compiled from human-curated lexical-semantic resources. |
| Approach: | They present a semantic reasoning benchmark for Danish compiled semi-automatically from a number of human-curated lexical-semantic resources. |
| Outcome: | The proposed datasets are compiled semi-automatically from human-curated lexical-semantic resources. |