Papers by Hanna Suominen
CILex: An Investigation of Context Information for Lexical Substitution Methods (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for lexical substitution rely on manually curated lexicals and contextual word embedding models. |
| Approach: | They propose a method that uses contextual sentence embeddings to generate substitutes for a target word given a context and a model that captures additional context information complimenting contextual word embedders. |
| Outcome: | The proposed method is state-of-the-art on the widely used LS07 and CoInCo datasets with P@1 scores of 55.96% and 57.25% for lexical substitution. |
Contextual Diversity Measure (CDM) for Controllable Story Generation in Large Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing studies on controllable text generation focus on controlling attributes such as sentiment, writing style, and writing style. |
| Approach: | They introduce a metric that quantifies semantic diversity for scenario generation under fixed abstract semantic constraints and validate it through controlled experiments. |
| Outcome: | The proposed metric achieves excellent discrimination accuracy (100% and 91.9%, respectively), with discriminative power up to 5.5 greater than the best baseline. |
To compress or not to compress? A Finite-State approach to Nen verbal morphology (2020.acl-srw)
Copied to clipboard
| Challenge: | a transitive verb takes up to 1,740 unique features and is highly complex, with a morphological complexity of 80.3% . a finite-state approach has been used to build morphology and phonology resources for Nen, an underresourced language in Papua New Guinea. |
| Approach: | They propose to use Finite-State methods to build a verbal morphological parser for an under-resourced Papuan language, Nen. |
| Outcome: | The proposed model is half the size of the full decomposed model, while the 'Chunking' model is under half the scale of the decomposer, with an overall accuracy of 80.3%. |
Text-to-Text Automatic Story Generation: A Survey (2026.eacl-srw)
Copied to clipboard
| Challenge: | Automated story generation aims to produce coherent, engaging, and contextually consistent narratives with minimal or no human involvement . despite advances in large language models, maintaining narrative coherence, character consistency, storyline diversity, and plot controllability in generating stories is still challenging. |
| Approach: | They propose to develop new evaluation metrics and better data sets to support automatic story generation. |
| Outcome: | The proposed evaluation metrics and better datasets will improve narrative coherence and consistency and explore practical applications of story generation. |
Tulun: Transparent and Adaptable Low-resource Machine Translation (2025.acl-demo)
Copied to clipboard
| Challenge: | a low-resource language that is the lingua franca in Timor-Leste lacks available corpora in the health domain. |
| Approach: | They propose a solution that combines neural MT with large language model-based post-editing guided by existing glossaries and translation memories. |
| Outcome: | The proposed system outperforms both standalone MT and LLM approaches across six low-resource languages on the FLORES dataset. |
Automatic Gloss Dictionary for Sign Language Learners (2022.acl-demo)
Copied to clipboard
| Challenge: | 430 million people worldwide have developed hearing loss and 700 million more are learning a sign language as a second language . sign language learners have limited means of seeking assistance and are restricted to class offerings or relying on a webcam to look up the sign. |
| Approach: | They propose an online tool supporting 2, 000 signs to assist language learners in determining the meaning of given signs. |
| Outcome: | The proposed system can lower the barrier in sign language learning by addressing the common problem of sign finding and make it accessible to the wider community. |
PostAc : A Visual Interactive Search, Exploration, and Analysis Platform for PhD Intensive Job Postings (P19-3)
Copied to clipboard
| Challenge: | Employers’ low awareness and interest in attracting PhD graduates means that the term “PhD” is rarely used as a keyword in job advertisements. |
| Approach: | They propose an online platform that makes the job market visible to job seekers by analyzing the key factors that identify what an employer is looking for when they hire a highly skilled researcher. |
| Outcome: | The proposed platform makes visible the geographic location, industry sector, job title, working hours, continuity, and wage of the research intensive jobs. |
Pretrained Knowledge Base Embeddings for improved Sentential Relation Extraction (2022.acl-srw)
Copied to clipboard
| Challenge: | Existing models that perform explicit on-task training of graph embeddings are inadequate. |
| Approach: | They propose to combine pretrained knowledge base graph embeddings with transformer based language models to improve performance on sentential Relation Extraction task. |
| Outcome: | The proposed model outperforms state-of-the-art models on the sentential Relation Extraction task. |
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)
Copied to clipboard
| Challenge: | Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate. |
| Approach: | They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs. |
| Outcome: | The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis. |