Papers by Anna Hätty

9 papers
PAP2PAT: Benchmarking Outline-Guided Long-Text Patent Generation with Patent-Paper Pairs (2025.findings-acl)

Copied to clipboard

Challenge: In patents, the description constitutes more than 90% of the document on average, yet its automatic generation remains understudied.
Approach: They propose a method to generate patent documents using a research paper as an invention specification.
Outcome: The proposed model can generate 1.8k patent-paper pairs describing the same inventions, but it's difficult to provide the level of detail required.
Multi-Step Generation of Test Specifications using Large Language Models for System-Level Requirements (2025.acl-industry)

Copied to clipboard

Challenge: System-level testing is a critical phase in the development of large, safety-dependent systems, such as those in the automotive industry.
Approach: They propose an AI-powered assistant to aid users in creating test specifications for system-level requirements.
Outcome: The proposed system reduces the effort required to derive test specifications by 30% in ROUGE-L.
A Laypeople Study on Terminology Identification across Domains and Task Definitions (N18-2)

Copied to clipboard

Challenge: Existing studies on term annotation show that even experts differ in their understanding of termhood .
Approach: They propose a new dataset of term annotation that examines the common understanding of what constitutes a term.
Outcome: The proposed datasets show that even experts differ in their understanding of termhood . the findings suggest that there is a common understanding of what constitutes a term .
A Domain-Specific Dataset of Difficulty Ratings for German Noun Compounds in the Domains DIY, Cooking and Automotive (2020.lrec-1)

Copied to clipboard

Challenge: a dataset with difficulty ratings for 1,030 closed noun compounds is presented . authors use a simple compound splitter to identify compound types in domain-specific texts .
Approach: They present a German closed noun compound dataset with difficulty ratings . they used a simple compound splitter to identify compounds in texts .
Outcome: The proposed dataset has difficulty ratings for 1,030 closed noun compounds extracted from domain-specific texts for do-it-ourself, cooking and automotive.
A Cost-Efficient Modular Sieve for Extracting Product Information from Company Websites (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for extracting product information are resource-intensive and computationally prohibitive due to website structure differences and numerous non-product pages.
Approach: They propose a modular method that leverages low-cost classification models to filter out company web pages.
Outcome: The proposed method improves on a new dataset of 7000 product and non-product web pages and reduces computational time and costs.
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)

Copied to clipboard

Challenge: a recent study has focused on term technicality, but there are still few studies on it.
Approach: They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces .
Outcome: The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons.
Compound or Term Features? Analyzing Salience in Predicting the Difficulty of German Noun Compounds across Domains (2021.starsem-1)

Copied to clipboard

Challenge: Using domain-specific vocabulary, it is important to analyse domain-related characteristics to improve the communication between lay people and experts.
Approach: They focus on the interaction of compound-based lexical features (such as frequency and productivity) and terminology-based features (contrasting domain-specific and general language) across word representations and classifiers.
Outcome: The proposed model shows that the interaction of compound-based lexical features and terminology-based features across word representations and classifiers is important for a broad binary distinction into ‘easy’ vs. ‘difficult’ general-language compound frequency is sufficient, but for . a more fine-grained four-class distinction it is crucial to include contrastive termhood features and compound and constituent features.
A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains (P19-1)

Copied to clipboard

Challenge: Existing models for diachronic and synchronic detection of lexical semantic divergences are superficial and lack of comparison.
Approach: They propose to extend benchmark models on a common state-of-the-art evaluation task . they also demonstrate that the same evaluation task and modelling approaches can be utilised for synchronic detection of domain-specific sense divergences in the field of term extraction.
Outcome: The proposed model can be utilised for the detection of domain-specific sense divergences in the field of term extraction.
Varying Vector Representations and Integrating Meaning Shifts into a PageRank Model for Automatic Term Extraction (2020.lrec-1)

Copied to clipboard

Challenge: a comparative study for automatic term extraction from domain-specific language using a PageRank graph algorithm with different edge-weighting methods.
Approach: They propose to use a PageRank algorithm to extract automatic terms from domain-specific language using different edge-weighting methods.
Outcome: The proposed model is compared with a PageRank model with different edge-weighting methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations