Challenge: a dataset with difficulty ratings for 1,030 closed noun compounds is presented . authors use a simple compound splitter to identify compound types in domain-specific texts .
Approach: They present a German closed noun compound dataset with difficulty ratings . they used a simple compound splitter to identify compounds in texts .
Outcome: The proposed dataset has difficulty ratings for 1,030 closed noun compounds extracted from domain-specific texts for do-it-ourself, cooking and automotive.

Similar Papers

Towards a Standardized Dataset for Noun Compound Interpretation (L18-1)

Copied to clipboard

Challenge: Noun compounds are interesting constructs in Natural Language Processing . lack of standardized set of relation inventories and annotated datasets hinders interpretation .
Approach: They propose a dataset that uses FrameNet as its semantic relation inventory to examine noun compounds.
Outcome: The proposed dataset is linguistically grounded and uses FrameNet as its semantic relation inventory.
A Laypeople Study on Terminology Identification across Domains and Task Definitions (N18-2)

Copied to clipboard

Challenge: Existing studies on term annotation show that even experts differ in their understanding of termhood .
Approach: They propose a new dataset of term annotation that examines the common understanding of what constitutes a term.
Outcome: The proposed datasets show that even experts differ in their understanding of termhood . the findings suggest that there is a common understanding of what constitutes a term .
A Joint Approach to Compound Splitting and Idiomatic Compound Detection (2020.lrec-1)

Copied to clipboard

Challenge: Compounding is a common word-formation process in Germanic languages . high productivity and low corpus frequency of compounds increase vocabulary size .
Approach: They develop a deep learning-based approach to noun compound splitting and idiomatic compound detection for the German language.
Outcome: The proposed approach outperforms the current state of the art in noun compound splitting and idiomatic compound detection for the German language.
What Kind of Language Is Hard to Language-Model? (P19-1)

Copied to clipboard

Challenge: a recent study suggests that language models perform poorly across languages.
Approach: They propose a model that fits a paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at least-pairwise parallel corpora.
Outcome: The proposed model is able to handle missing data and is aware of inter-sentence variation.
Revisiting Generalization Across Difficulty Levels: It’s Not So Easy (2026.eacl-long)

Copied to clipboard

Challenge: Existing research is mixed regarding whether training on easier or harder data leads to better results.
Approach: They examine how well large language models generalize across different task difficulties by using a large dataset and a well-established difficulty metric.
Outcome: The results show that training on hard data can't achieve consistent improvements across the full range of difficulties.
On the Impact of Cross-Domain Data on German Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Traditionally, large language models have been trained on general web crawls or domain-specific data.
Approach: They present a German dataset and a dataset aimed at containing high-quality data to examine the importance of data diversity over quality.
Outcome: The proposed model outperforms models trained on quality data on multiple downstream tasks.
To Split or Not to Split: Composing Compounds in Contextual Vector Spaces (2023.emnlp-main)

Copied to clipboard

Challenge: Contextual word embedding models rely on sub-word tokenization to represent single orthographic words but are often suboptimal in under-resourced contexts.
Approach: They propose to use a masked language modelling task to evaluate the model's performance . they use re-trained tokenizers to pre-split compounds into constituents .
Outcome: The proposed models improve on the masked language modelling task and compositionality prediction by pre-splitting compounds into constituents.
Klexikon: A German Dataset for Joint Summarization and Simplification (2022.lrec-1)

Copied to clipboard

Challenge: Traditionally, Text Simplification is a monolingual translation task where individual sentences are "translated" into a simplified version.
Approach: They propose to use a dataset to jointly simplify long source documents by combining sentences from a source and their simplified counterparts.
Outcome: The proposed system can summarize and simplify long source documents using almost 2,900 documents.
Modeling the Evolution of English Noun Compounds with Feature-Rich Diachronic Compositionality Prediction (2025.acl-long)

Copied to clipboard

Challenge: Empirical research directly addressing these issues is limited to a small number of studies suggesting that compounding is a highly productive process.
Approach: They represent English noun compounds as vectors of time-specific values and implement a set of features to classify them for present-day compositionality and assess the informativeness of the corresponding linguistic patterns.
Outcome: The proposed method captures relevant and complementary information across approaches and shows that low-compositional meanings are reflected by a parallel drop in compositionality and sustained semantic change.
DETECT: Determining Ease and Textual Clarity of German Text Simplifications (2026.eacl-long)

Copied to clipboard

Challenge: Current evaluation of German automatic text simplification relies on general-purpose metrics such as SARI, BLEU, and BERTScore.
Approach: They propose a German-specific metric that holistically evaluates ATS quality across all three dimensions of simplicity, meaning preservation, and fluency.
Outcome: The proposed metric achieves higher correlations with human judgments than widely used ATS metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations