Papers by Ana Marasovic

11 papers
On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization (2022.findings-emnlp)

Copied to clipboard

Challenge: Combining visual modality with pretrained language models has been effective for descriptive tasks such as image captioning.
Approach: They ask: do multimodal models combine visual and visual adapted language models? they find that CLIP image representations and scaling of language models do not consistently improve self-rationalization in multimodal tasks.
Outcome: The proposed model types do not consistently improve self-rationalization in multimodal tasks.
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs (2025.findings-acl)

Copied to clipboard

Challenge: a core part of legal work that has been underexplored in Legal NLP is the writing and editing of legal briefs.
Approach: They propose to use large language models to help legal professionals with writing briefs by capturing and evaluating their abilities in language models.
Outcome: The proposed tasks show that the models perform well on arguments summarization, argument completion, and case retrieval tasks.
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps (2025.emnlp-main)

Copied to clipboard

Challenge: Language models (LMs) produce a chain of thought (CoT) when prompted to think step-by-step, but it is unclear whether the reasoning encoded in the CoT is faithful to the models’ parametric beliefs.
Approach: They propose a framework for measuring parametric faithfulness of generated reasoning by unlearning reasoning steps (FUR) they propose to erase information contained in reasoning steps from model parameters and measure faithfulness as the resulting effect on the model’s prediction.
Outcome: The proposed framework erases information contained in reasoning steps from model parameters and measures faithfulness as the resulting effect on the model’s prediction.
Few-Shot Self-Rationalization with Natural Language Prompts (2022.findings-naacl)

Copied to clipboard

Challenge: Existing models that generate free-text explanations for tasks are limited by human-written explanations.
Approach: They propose to use a standardized collection of natural language prompts to create a model that generates free-text explanations for tasks.
Outcome: The proposed model can predict task labels and generate free-text explanations for predictions . plausibility of human explanations is 76%, while human explanation is 51% .
Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest (2023.acl-long)

Copied to clipboard

Challenge: Large neural networks can generate jokes, but do they really “understand” humor? a new challenge challenges AI models to match a joke to a cartoon, identify a winning caption, and explain why a winner is funny.
Approach: They propose three tasks based on the New Yorker Cartoon Caption Contest . they aim to match a joke to a cartoon, identify a winning caption and explain why it's funny .
Outcome: The proposed tasks are based on the New Yorker Cartoon Caption Contest . they include matching a joke to a cartoon, identifying a winning caption, and explaining why a funny caption is funny.
Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness (2024.naacl-long)

Copied to clipboard

Challenge: Existing approaches to measure robustness are problematic, and out-of-domain evaluations are no longer relevant.
Approach: They examine models of different sizes spanning different architectural choices and pretraining objectives.
Outcome: The results show that not all out-of-domain tests provide insight into robustness . merely scaling models does not make them adequately robust .
Does Self-Rationalization Improve Robustness to Spurious Correlations? (2022.emnlp-main)

Copied to clipboard

Challenge: Rationalization is fundamental to human reasoning and learning.
Approach: They evaluate robustness to spurious correlations in encoder-decoder and decoder-only models . authors say explanations can come at the cost of robustness .
Outcome: The proposed model outputs are more interpretable and easier to interact with for end-users than nonrationalizing models.
Explanation in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events.
Approach: They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations.
Outcome: The proposed methods are based on the models of large language models (LLMs) and their opaque nature.
CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation (2022.emnlp-main)

Copied to clipboard

Challenge: Negation is fundamental to human communication.
Approach: They propose a dataset which requires reasoning about implications of negated statements in paragraphs . they collect paragraphs with diverse negation cues and crowdworkers ask questions about implications .
Outcome: The first dataset in english requires reasoning about implications of negated statements in paragraphs . it features 14,182 question-answer pairs with over 200 unique negation cues based on crowd-workers . the best performing model achieves only 42% on consistency metric, well below human performance of 81%.
On Evaluating Explanation Utility for Human-AI Decision Making in NLP (2024.findings-emnlp)

Copied to clipboard

Challenge: a lack of evidence that explanations help people in situations they are introduced for is a problem in NLP . prior work on explainability has focused on overcoming technical challenges and used proxy evaluations.
Approach: They propose to use existing metrics to evaluate the effectiveness of explanations in NLP . they argue that providing AI predictions does not cause decision makers to speed up work .
Outcome: The proposed evaluations show that providing AI predictions does not cause decision makers to speed up their work without compromising performance.
What Has Been Lost with Synthetic Evaluation? (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study evaluated the validity and difficulty of large language models for evaluation benchmarks . large language model evaluation benchmarking is challenging and requires specific phenomena to be addressed .
Approach: They compare LLM-generated reasoning-over-text benchmarks to those generated through crowdsourcing . they find they are *less challenging for LLMs* than their human-authored counterparts .
Outcome: The results show that LLMs can generate variants that are valid according to annotation guidelines, but less challenging than human-authored counterparts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations