Papers by Hanna Suominen

9 papers
CILex: An Investigation of Context Information for Lexical Substitution Methods (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for lexical substitution rely on manually curated lexicals and contextual word embedding models.
Approach: They propose a method that uses contextual sentence embeddings to generate substitutes for a target word given a context and a model that captures additional context information complimenting contextual word embedders.
Outcome: The proposed method is state-of-the-art on the widely used LS07 and CoInCo datasets with P@1 scores of 55.96% and 57.25% for lexical substitution.
Contextual Diversity Measure (CDM) for Controllable Story Generation in Large Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies on controllable text generation focus on controlling attributes such as sentiment, writing style, and writing style.
Approach: They introduce a metric that quantifies semantic diversity for scenario generation under fixed abstract semantic constraints and validate it through controlled experiments.
Outcome: The proposed metric achieves excellent discrimination accuracy (100% and 91.9%, respectively), with discriminative power up to 5.5 greater than the best baseline.
To compress or not to compress? A Finite-State approach to Nen verbal morphology (2020.acl-srw)

Copied to clipboard

Challenge: a transitive verb takes up to 1,740 unique features and is highly complex, with a morphological complexity of 80.3% . a finite-state approach has been used to build morphology and phonology resources for Nen, an underresourced language in Papua New Guinea.
Approach: They propose to use Finite-State methods to build a verbal morphological parser for an under-resourced Papuan language, Nen.
Outcome: The proposed model is half the size of the full decomposed model, while the 'Chunking' model is under half the scale of the decomposer, with an overall accuracy of 80.3%.
Text-to-Text Automatic Story Generation: A Survey (2026.eacl-srw)

Copied to clipboard

Challenge: Automated story generation aims to produce coherent, engaging, and contextually consistent narratives with minimal or no human involvement . despite advances in large language models, maintaining narrative coherence, character consistency, storyline diversity, and plot controllability in generating stories is still challenging.
Approach: They propose to develop new evaluation metrics and better data sets to support automatic story generation.
Outcome: The proposed evaluation metrics and better datasets will improve narrative coherence and consistency and explore practical applications of story generation.
Tulun: Transparent and Adaptable Low-resource Machine Translation (2025.acl-demo)

Copied to clipboard

Challenge: a low-resource language that is the lingua franca in Timor-Leste lacks available corpora in the health domain.
Approach: They propose a solution that combines neural MT with large language model-based post-editing guided by existing glossaries and translation memories.
Outcome: The proposed system outperforms both standalone MT and LLM approaches across six low-resource languages on the FLORES dataset.
Automatic Gloss Dictionary for Sign Language Learners (2022.acl-demo)

Copied to clipboard

Challenge: 430 million people worldwide have developed hearing loss and 700 million more are learning a sign language as a second language . sign language learners have limited means of seeking assistance and are restricted to class offerings or relying on a webcam to look up the sign.
Approach: They propose an online tool supporting 2, 000 signs to assist language learners in determining the meaning of given signs.
Outcome: The proposed system can lower the barrier in sign language learning by addressing the common problem of sign finding and make it accessible to the wider community.
PostAc : A Visual Interactive Search, Exploration, and Analysis Platform for PhD Intensive Job Postings (P19-3)

Copied to clipboard

Challenge: Employers’ low awareness and interest in attracting PhD graduates means that the term “PhD” is rarely used as a keyword in job advertisements.
Approach: They propose an online platform that makes the job market visible to job seekers by analyzing the key factors that identify what an employer is looking for when they hire a highly skilled researcher.
Outcome: The proposed platform makes visible the geographic location, industry sector, job title, working hours, continuity, and wage of the research intensive jobs.
Pretrained Knowledge Base Embeddings for improved Sentential Relation Extraction (2022.acl-srw)

Copied to clipboard

Challenge: Existing models that perform explicit on-task training of graph embeddings are inadequate.
Approach: They propose to combine pretrained knowledge base graph embeddings with transformer based language models to improve performance on sentential Relation Extraction task.
Outcome: The proposed model outperforms state-of-the-art models on the sentential Relation Extraction task.
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)

Copied to clipboard

Challenge: Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate.
Approach: They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs.
Outcome: The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations