Papers by Michael Dorna
A Domain-Specific Dataset of Difficulty Ratings for German Noun Compounds in the Domains DIY, Cooking and Automotive (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset with difficulty ratings for 1,030 closed noun compounds is presented . authors use a simple compound splitter to identify compound types in domain-specific texts . |
| Approach: | They present a German closed noun compound dataset with difficulty ratings . they used a simple compound splitter to identify compounds in texts . |
| Outcome: | The proposed dataset has difficulty ratings for 1,030 closed noun compounds extracted from domain-specific texts for do-it-ourself, cooking and automotive. |
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study has focused on term technicality, but there are still few studies on it. |
| Approach: | They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces . |
| Outcome: | The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons. |
Compound or Term Features? Analyzing Salience in Predicting the Difficulty of German Noun Compounds across Domains (2021.starsem-1)
Copied to clipboard
| Challenge: | Using domain-specific vocabulary, it is important to analyse domain-related characteristics to improve the communication between lay people and experts. |
| Approach: | They focus on the interaction of compound-based lexical features (such as frequency and productivity) and terminology-based features (contrasting domain-specific and general language) across word representations and classifiers. |
| Outcome: | The proposed model shows that the interaction of compound-based lexical features and terminology-based features across word representations and classifiers is important for a broad binary distinction into ‘easy’ vs. ‘difficult’ general-language compound frequency is sufficient, but for . a more fine-grained four-class distinction it is crucial to include contrastive termhood features and compound and constituent features. |