Papers by Anna Hätty
PAP2PAT: Benchmarking Outline-Guided Long-Text Patent Generation with Patent-Paper Pairs (2025.findings-acl)
Copied to clipboard
| Challenge: | In patents, the description constitutes more than 90% of the document on average, yet its automatic generation remains understudied. |
| Approach: | They propose a method to generate patent documents using a research paper as an invention specification. |
| Outcome: | The proposed model can generate 1.8k patent-paper pairs describing the same inventions, but it's difficult to provide the level of detail required. |
Multi-Step Generation of Test Specifications using Large Language Models for System-Level Requirements (2025.acl-industry)
Copied to clipboard
| Challenge: | System-level testing is a critical phase in the development of large, safety-dependent systems, such as those in the automotive industry. |
| Approach: | They propose an AI-powered assistant to aid users in creating test specifications for system-level requirements. |
| Outcome: | The proposed system reduces the effort required to derive test specifications by 30% in ROUGE-L. |
A Laypeople Study on Terminology Identification across Domains and Task Definitions (N18-2)
Copied to clipboard
| Challenge: | Existing studies on term annotation show that even experts differ in their understanding of termhood . |
| Approach: | They propose a new dataset of term annotation that examines the common understanding of what constitutes a term. |
| Outcome: | The proposed datasets show that even experts differ in their understanding of termhood . the findings suggest that there is a common understanding of what constitutes a term . |
A Domain-Specific Dataset of Difficulty Ratings for German Noun Compounds in the Domains DIY, Cooking and Automotive (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset with difficulty ratings for 1,030 closed noun compounds is presented . authors use a simple compound splitter to identify compound types in domain-specific texts . |
| Approach: | They present a German closed noun compound dataset with difficulty ratings . they used a simple compound splitter to identify compounds in texts . |
| Outcome: | The proposed dataset has difficulty ratings for 1,030 closed noun compounds extracted from domain-specific texts for do-it-ourself, cooking and automotive. |
A Cost-Efficient Modular Sieve for Extracting Product Information from Company Websites (2024.emnlp-industry)
Copied to clipboard
Anna Hätty, Dragan Milchevski, Kersten Döring, Marko Putnikovic, Mohsen Mesgar, Filip Novović, Maximilian Braun, Karina Borimann, Igor Stranjanac
| Challenge: | Existing methods for extracting product information are resource-intensive and computationally prohibitive due to website structure differences and numerous non-product pages. |
| Approach: | They propose a modular method that leverages low-cost classification models to filter out company web pages. |
| Outcome: | The proposed method improves on a new dataset of 7000 product and non-product web pages and reduces computational time and costs. |
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study has focused on term technicality, but there are still few studies on it. |
| Approach: | They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces . |
| Outcome: | The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons. |
Compound or Term Features? Analyzing Salience in Predicting the Difficulty of German Noun Compounds across Domains (2021.starsem-1)
Copied to clipboard
| Challenge: | Using domain-specific vocabulary, it is important to analyse domain-related characteristics to improve the communication between lay people and experts. |
| Approach: | They focus on the interaction of compound-based lexical features (such as frequency and productivity) and terminology-based features (contrasting domain-specific and general language) across word representations and classifiers. |
| Outcome: | The proposed model shows that the interaction of compound-based lexical features and terminology-based features across word representations and classifiers is important for a broad binary distinction into ‘easy’ vs. ‘difficult’ general-language compound frequency is sufficient, but for . a more fine-grained four-class distinction it is crucial to include contrastive termhood features and compound and constituent features. |
A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains (P19-1)
Copied to clipboard
| Challenge: | Existing models for diachronic and synchronic detection of lexical semantic divergences are superficial and lack of comparison. |
| Approach: | They propose to extend benchmark models on a common state-of-the-art evaluation task . they also demonstrate that the same evaluation task and modelling approaches can be utilised for synchronic detection of domain-specific sense divergences in the field of term extraction. |
| Outcome: | The proposed model can be utilised for the detection of domain-specific sense divergences in the field of term extraction. |
Varying Vector Representations and Integrating Meaning Shifts into a PageRank Model for Automatic Term Extraction (2020.lrec-1)
Copied to clipboard
| Challenge: | a comparative study for automatic term extraction from domain-specific language using a PageRank graph algorithm with different edge-weighting methods. |
| Approach: | They propose to use a PageRank algorithm to extract automatic terms from domain-specific language using different edge-weighting methods. |
| Outcome: | The proposed model is compared with a PageRank model with different edge-weighting methods. |