Papers by Yulia Otmakhova

9 papers
M3: Multi-level dataset for Multi-document summarisation of Medical studies (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing summarisation systems are not up to such complex tasks, yet limited tools exist to determine where and why they are failing.
Approach: They propose to use a dataset to evaluate the quality of summarisation systems in the biomedical domain.
Outcome: The proposed model can be used to evaluate the quality of summarisation systems in the biomedical domain.
The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature (2022.acl-long)

Copied to clipboard

Challenge: Existing evaluation approaches to multi-document summarization of biomedical literature lack consistency and transparency.
Approach: They propose a systematic approach to human evaluation of biomedical summaries and apply it to analyze the summary generated by two current evaluation models.
Outcome: The proposed evaluation framework is based on two state-of-the-art models and examines the summaries generated by the two models to understand the deficiencies of existing evaluation approaches.
Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation (2022.aacl-main)

Copied to clipboard

Challenge: Negation is an important linguistic phenomenon which denotes non-existence, denial, or contradiction.
Approach: They propose a natural language inference test suite to test models for negation . they use a linguistic framework to analyze negation types and constructions .
Outcome: The proposed test suite is more challenging than existing benchmarks on negation . it includes annotation of negation types and constructions grounded in linguistic theory .
Revisiting subword tokenization: A case study on affixal negation in large language models (2024.naacl-long)

Copied to clipboard

Challenge: Negation is central to language understanding but is not properly captured by modern NLP methods.
Approach: They propose to use subword tokenization methods to detect negation in large language models . they find that models can reliably recognize negation, despite mismatches in tokenization accuracy .
Outcome: The proposed models can detect negation in English using subword tokenization methods despite some mismatches in tokenization accuracy and negation detection performance.
FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation (2026.findings-eacl)

Copied to clipboard

Challenge: FLUKE introduces controlled variations across linguistic levels and leverages large language models with human validation to generate modifications.
Approach: They propose a framework for assessing model robustness through systematic minimal variations of test data.
Outcome: The proposed framework evaluates models and LLMs across six diverse NLP tasks and shows that they are more robust to natural, fluent modifications than base models.
Not all ANIMALs are equal: metaphorical framing through source domains and semantic frames (2026.findings-acl)

Copied to clipboard

Challenge: a computational framework allows to derive discourse metaphors through their source domains and semantic frames.
Approach: They propose a computational framework that allows to derive salient discourse metaphors through their source domains and semantic frames.
Outcome: The proposed framework uncovers well-known source domains and reveals nuanced frame-level associations that distinguish how the issue is portrayed.
Narrative Media Framing in Political Discourse (2025.findings-acl)

Copied to clipboard

Challenge: Narrative frames are a powerful way of conceptualizing and communicating complex ideas.
Approach: They propose a framework which formalizes and operationalizes elements of narrative framing . they annotate news articles in the climate change domain and test their framework .
Outcome: The proposed framework formalizes and operationalizes elements of narrative framing . it is applied to climate change crisis data, showing generalizability of the framework .
Article and Comment Frames Shape the Quality of Online Comments (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has focused on predicting comment toxicity or quality, but it ignores audience reactions.
Approach: They propose a frame-aware system to mitigate unhealthy discourse . they analysed 1M comments across 2.7K news articles .
Outcome: The proposed system can mitigate unhealthy discourses by analyzing 1M comments across 2.7K news articles.
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations (2023.acl-long)

Copied to clipboard

Challenge: Prior work has shown that models may exploit shortcuts that are difficult to detect using standard n-gram similarity metrics such as ROUGE.
Approach: They propose to use human-assessed summary quality facets and pairwise preferences to improve MDS evaluation methods.
Outcome: The proposed methods improve the quality of literature review summarization models . they use human-assessed summary quality facets and pairwise preferences .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations