Challenge: euphemisms are a linguistic device used to soften discussions of uncomfortable topics . euphorias are used to refer to death in a less direct manner during a period of secularization .
Approach: They propose to use a corpus of Danish and Norwegian novels to detect death-related euphemisms . they use pre-trained language models to detect euphoric and literal references to death .
Outcome: The proposed method improves on state-of-the-art language models.

Similar Papers

Fact from Fiction: Finding Serialized Novels in Newspapers (2025.acl-srw)

Copied to clipboard

Challenge: Among underrepresented but widely read forms are serialized fiction and feuilleton novels embedded in newspapers rather than published as standalone volumes.
Approach: They propose to annotate 1,394 articles and evaluate classification pipelines using both selected linguistic features and embeddings to identify serialized fiction and feuilleton fiction.
Outcome: The proposed methods achieve F1-scores of 0.91 in an annotated dataset of 1,394 articles and support the construction of alternative literary corpora and contribute to work on modeling the fiction–nonfiction boundary at scale.
Development and Evaluation of Pre-trained Language Models for Historical Danish and Norwegian Literary Texts (2024.lrec-main)

Copied to clipboard

Challenge: et al., 2019) develop and evaluate the first pre-trained language models specifically tailored for historical Danish and Norwegian texts.
Approach: They develop and evaluate pre-trained language models specifically tailored for historical Danish and Norwegian texts.
Outcome: The proposed model outperforms models trained on historical Danish and Norwegian literature in two downstream NLP tasks.
Noise, Novels, Numbers. A Framework for Detecting and Categorizing Noise in Danish and Norwegian Literature (2024.emnlp-main)

Copied to clipboard

Challenge: This study examines the literary perceptions of noise during the Scandinavian "Modern Breakthrough" period (1870-1899).
Approach: They propose a framework for detecting and categorizing noise in literary texts from the late 19th century.
Outcome: The proposed framework can be applied to Danish and Norwegian literature from the late 19th century.
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry (2024.eacl-short)

Copied to clipboard

Challenge: a corpus of late antique and medieval Hebrew poetry is rich in metaphors and similes . scholars in the humanities need to distinguish between figurative and literal language .
Approach: They present a corpus of late antique and medieval Hebrew poetry with expert annotations of metaphor . they hope to facilitate further research in this area .
Outcome: The proposed dataset includes a corpus of late antique and medieval Hebrew poetry with expert annotations of metaphor.
A corpus of metaphors as register markers (2023.findings-eacl)

Copied to clipboard

Challenge: Using corpus annotation, we show huge differences in metaphor usage between different registers and specific properties of registers.
Approach: They present their work on corpus annotation for metaphor in germany . they focus on metaphors that can serve as register markers and be reliably indentified .
Outcome: The proposed corpus annotations show huge differences in metaphor usage between different registers and specific properties of registers.
CATs are Fuzzy PETs: A Corpus and Analysis of Potentially Euphemistic Terms (2022.lrec-1)

Copied to clipboard

Challenge: Euphemisms are a difficult topic because they are subject to language change and humans may not agree on what is a euphemist.
Approach: They analyze a corpus of potentially euphemistic terms (PETs) and examples from the GloWbE corpus to examine their meanings.
Outcome: The proposed corpus of potentially euphemistic terms and examples from the GloWbE corpus show that PETs generally decrease negative and offensive sentiment.
Metaphor annotation for German (2022.lrec-1)

Copied to clipboard

Challenge: a corpus annotated for metaphors denotes entities or situations that are in some sense similar to the literal referent, but we believe it is of interest to research on metaphor in general.
Approach: They present a German corpus annotated for metaphor in a project on register and propose to broaden the annotation to include metonymy.
Outcome: The proposed corpus is compiled and annotated in a project on the interdependence of metaphors and register.
A Diachronic Corpus for Literary Style Analysis (L18-1)

Copied to clipboard

Challenge: Temporal style analysis is not widely taken into account, says aaron daelemans . he says it is important to consider the possibility of an author's style frequently changing over time . daelemens: synchronic style analysis requires accurate time-stamped data .
Approach: They propose a resource for diachronic style analysis in particular the analysis of literary authors over time.
Outcome: The proposed resource can be used to analyze literary authors over time.
Metaphors in Online Religious Communication: A Detailed Dataset and Cross-Genre Metaphor Detection (2024.lrec-main)

Copied to clipboard

Challenge: figurative language plays a particularly important role in religious communication . linguistic metaphors relate entities from different semantic domains by drawing on an implicit similarity between them.
Approach: They present a dataset of fine-grained metaphor annotations for online religious communication . they show that cross-genre transfer metaphor detection leads to a drop in performance .
Outcome: The proposed dataset shows that adding in-genre data improves performance . the authors show that the proposed system can detect metaphors in religious forums .
Euphemistic Phrase Detection by Masked Language Model (2021.findings-emnlp)

Copied to clipboard

Challenge: euphemisms are ordinary-sounding words with a secret meaning that are used to conceal information . a primary motive of their use on social media is to evade content moderation efforts .
Approach: They propose to use social media to detect euphemisms without human effort . they first perform phrase mining on a raw text corpus to extract quality phrases . then they use word embedding similarities to select a set of euphoristic phrase candidates .
Outcome: The proposed algorithm shows 20-50% higher detection accuracies than baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations