Papers by Oana Ignat

15 papers
WhyAct: Identifying Action Reasons in Lifestyle Vlogs (2021.emnlp-main)

Copied to clipboard

Challenge: Existing systems for action recognition rely on pattern memorization and do not understand the action.
Approach: They propose a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video.
Outcome: The proposed model leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video.
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages (2025.acl-long)

Copied to clipboard

Challenge: Emotion recognition is an umbrella term for several NLP tasks, but most work on high-resource languages has focused on low-resourced languages.
Approach: They propose to use emotion recognition to describe perceived emotions in 28 different languages and across several domains to identify and annotate the datasets.
Outcome: The proposed datasets cover low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers.
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data (2024.lrec-main)

Copied to clipboard

Challenge: Synthetic data generation has the potential to impact domains with scarce data, but we need to understand how different demographics are represented in it.
Approach: They develop a procedure to generate depression data using GPT-3 and analyze it to uncover the types of stressors it assigns to demographic groups.
Outcome: The proposed procedure produces depression data using GPT-3, and compares it to a human-generated dataset.
FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework (2022.acl-long)

Copied to clipboard

Challenge: Existing video understanding evaluation frameworks that use fill-in-the-blanks do not reflect real-world tasks.
Approach: They propose to use fill-in-the-blanks as a video understanding evaluation framework and introduce a novel dataset that collects multiple perspectives on the same video.
Outcome: The proposed framework does not share the weaknesses of the current state-of-the-art language-informed video understanding tasks, namely: (1) video question answering using multiple-choice questions, where models perform relatively well because they exploit linguistic biases in the task formulation; (2) video captioning, which relies on an open-ended evaluation framework that is often inaccurate because system answers may be perceived as incorrect if they differ in form from the ground truth.
When Cultures Meet: Multicultural Text-to-Image Generation (2026.findings-acl)

Copied to clipboard

Challenge: a new task to evaluate text-to-image generation models for multicultural scenes is unexplored.
Approach: They propose a benchmark task to evaluate text-to-image models in multicultural settings . they use a dataset of 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages to analyze behavior .
Outcome: The proposed benchmark analyzes the behavior of state-of-the-art models across multiple dimensions including alignment, image quality, aesthetics, knowledge, and fairness.
Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Among the minority groups under-represented in AI, data from low-income households are often overlooked in data collection and model evaluation.
Approach: They evaluate the performance of a vision-language model on a geo-diverse dataset . they highlight insights that can help mitigate these issues and propose actionable steps for economic-level inclusive AI development.
Outcome: The proposed model performs lower for the poorer groups than the wealthier groups across topics and countries.
Identifying Visible Actions in Lifestyle Vlogs (P19-1)

Copied to clipboard

Challenge: Existing methods for identifying human actions in videos are limited by the number of visual depictions in the videos.
Approach: They propose a multimodal algorithm that leverages visual and linguistic clues to automatically infer which actions are visible in a video.
Outcome: The proposed algorithm can identify actions visible in video while verbally describing them.
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost (2024.lrec-main)

Copied to clipboard

Challenge: Current foundation models have shown impressive performance across various tasks, but they are not effective for everyone due to the imbalanced geographical and economic representation of the data used in the training process.
Approach: They propose to identify the data to be annotated to balance model performance and annotation costs by finding countries with visual similarity for the topics.
Outcome: The proposed methods improve model performance and reduce annotation costs by using data from countries with higher visual similarity for these topics.
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning (2025.naacl-long)

Copied to clipboard

Challenge: Large Multimodal Models exhibit impressive performance across multimodal tasks . effectiveness in cross-cultural contexts limited due to predominantly Western-centric nature of data and models . multi-agent models have shown significant capability in solving complex tasks despite limitations in crosscultural context .
Approach: They propose to use a multi-agent framework to enhance cross-cultural image captioning using LMMs with distinct cultural personas to evaluate cultural information within image captions.
Outcome: The proposed model outperforms single-agent models across different metrics and offers valuable insights for future research.
Scalable Performance Analysis for Vision-Language Models (2023.starsem-1)

Copied to clipboard

Challenge: a new method to probe vision-language models is proposed that does not require data annotation and makes use of existing datasets.
Approach: They propose a method that extracts features from a vision-language benchmark and measures their correlation with the output of the target model.
Outcome: The proposed method is scalable and does not require data annotation . it can be used with other models and benchmarks, and is available at https://github.com/MichiganNLP/Scalable-VLM-Probing.
OCR Improves Machine Translation for Low-Resource Languages (2022.findings-acl)

Copied to clipboard

Challenge: Despite many recent successes, Machine Translation still lacks support or fails to achieve good performance for most low-resource languages.
Approach: They propose a benchmark to evaluate OCR systems on low-resource languages and low- resource scripts.
Outcome: The proposed benchmark evaluates state-of-the-art OCR systems on low-resource languages and low-rural scripts.
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models (2025.naacl-long)

Copied to clipboard

Challenge: Unequal representation of cultures and socioeconomic groups in training data leads to biased Large Multi-modal (LMM) models.
Approach: They propose and evaluate several prompting strategies that use non-English, geographic, and socioeconomic attributes to improve LMM model performance on underrepresented data.
Outcome: The proposed prompts favor retrieving topic appearances from low-income data on lower-income datasets.
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models have led to misleading public discourse that “it’s all been solved.”
Approach: They identify 14 research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs.
Outcome: The research areas identified are 45 research directions that require new research and are not directly solvable by LLMs.
MAiDE-up: Multilingual Deception Detection of AI-generated Hotel Reviews (2025.findings-naacl)

Copied to clipboard

Challenge: Deceptive reviews are becoming more common, especially given the increase in performance and the prevalence of LLMs.
Approach: They compile and make publicly available a dataset of 10,000 real and 10,000 AI-generated fake hotel reviews in ten languages.
Outcome: The proposed model can detect real reviews and fake reviews in 10 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations