Papers by Michael Saxon
Can Vision Language Models Understand Mimed Actions? (2025.findings-acl)
Copied to clipboard
Hyundong Justin Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May
| Challenge: | Nonverbal communication (NVC) is an integral part of human language, but it has been overlooked in natural language processing research. |
| Approach: | They propose a multimodal multimodal recognition task that uses a corpus of mimed gestures to evaluate their understanding of NVC. |
| Outcome: | The proposed task is based on 86 unique gestures with perturbations applied to avatar, background, and viewpoint for evaluating recognition robustness. |
Not All Errors are Equal: Learning Text Generation Metrics using Stratified Error Synthesis (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing learning metrics are limited to tasks where large human ratings are available. |
| Approach: | They propose a model-based natural language generation (NLG) evaluation metric that is highly correlated with human judgements without requiring human annotation. |
| Outcome: | The proposed metric outperforms all prior unsupervised metrics on multiple NLG tasks including translation, image captioning, and WebNLG text generation. |
Let’s Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought (2023.emnlp-main)
Copied to clipboard
Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He, Alex Mei, Yujie Lu, Chinmay Sonar, Michael Saxon, William Wang
| Challenge: | Existing studies show vision-language systems can reason about images using natural language, but their capacity for video reasoning remains underexplored. |
| Approach: | They propose to frame video reasoning as the sequential understanding of a small number of keyframes, thereby leveraging the power and robustness of vision-language systems' capacity to reason about images using natural language. |
| Outcome: | The proposed models can generate multiple intermediate keyframes and predict future keyframe, and they perform poorly on GPT-4, GPT-3, and VICUNA. |
Multilingual Conceptual Coverage in Text-to-Image Models (2023.acl-long)
Copied to clipboard
| Challenge: | Neural text-to-image systems generate coherent, visually-appealing images with novel combinations of objects, scenarios, and styles. |
| Approach: | They propose a technique to benchmark the degree to which a generative text-to-image system provides multilingual parity to its training language in terms of tangible nouns. |
| Outcome: | The proposed technique can be used to benchmark T2I models in terms of multilinguality and identify model-specific weaknesses, spurious correlations, and biases without a-priori assumptions. |
CausalDialogue: Modeling Utterance-level Causality in Conversations (2023.findings-acl)
Copied to clipboard
| Challenge: | Despite widespread adoption, neural conversation models have yet to exhibit natural chat capabilities with humans . despite their widespread adoption in society, chatbots have yet not shown natural chat capability . |
| Approach: | They propose a causality-enhanced method to enhance the impact of causality at the utterance level in training neural conversation models. |
| Outcome: | The proposed method improves diversity and agility of loss functions and still needs improvement . the proposed method is based on a CausalDialogue dataset . |
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts (2024.naacl-short)
Copied to clipboard
| Challenge: | With growth in the popularity of text-to-image models has come interest in assessing their multilingual capabilities, including multilingual accessibility. |
| Approach: | They propose to correct translation errors in a concept list translated to seven languages and compare the outputs of the benchmark to those conditioned on the old. |
| Outcome: | The proposed benchmark contains translation errors in Spanish, Japanese, and Chinese. |
Modeling Disclosive Transparency in NLP Application Descriptions (2021.emnlp-main)
Copied to clipboard
| Challenge: | Broader disclosive transparency is difficult to define and quantify, authors say . previous work has demonstrated trade-offs and negative consequences to disclosing transparency . |
| Approach: | They propose to use neural language model-based probabilistic metrics to model disclosive transparency . they demonstrate that they correlate with user and expert opinions of system transparency a valid objective proxy . |
| Outcome: | The proposed metrics correlate with user and expert opinions of system transparency, making them a valid objective proxy. |
Do You Know About My Nation? Investigating Multilingual Language Models’ Cultural Literacy Through Factual Knowledge (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual question-answering benchmarks do not factor in regional diversity in the information they capture and tend to be Western-centric. |
| Approach: | They propose to benchmark eight standard multilingual LLMs on XNationQA and evaluate them using two novel transference metrics. |
| Outcome: | The proposed model shows greater knowledge of cultural information in English than in the dominant language of the respective culture. |
Culture is Everywhere: A Call for Intentionally Cultural Evaluation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to evaluate cultural alignment of large language models are too trivial and focus on static facts and values. |
| Approach: | They argue for intentionally cultural evaluation: an approach that examines cultural assumptions . they characterize what, how, and circumstances by which culturally contingent considerations arise in evaluation . |
| Outcome: | The authors argue for intentionally cultural evaluation: an approach that examines cultural assumptions embedded in all aspects of evaluation, not just in explicitly cultural tasks. |
CAIRE: Cultural Attribution of Images with Retrieval (2026.eacl-long)
Copied to clipboard
| Challenge: | Current text-to-image models produce homogeneous outputs given under-specified prompts and their outputs are disproportionately biased toward Western cultures. |
| Approach: | They propose a framework that assesses the degree of cultural relevance of an image, given a user-defined set of labels. |
| Outcome: | The proposed evaluation metric surpasses baselines on a manually curated dataset of culturally salient but rare items built using language models by 22% F1 points. |
PECO: Examining Single Sentence Label Leakage in Natural Language Inference Datasets through Progressive Evaluation of Cluster Outliers (2023.eacl-main)
Copied to clipboard
| Challenge: | Efforts to debias NLI have led to datasets that exhibit different kinds of bias than those shown before. |
| Approach: | They propose a new technique to detect and reduce single sentence label leakage . leakage is a problem with many modern NLI datasets, they argue . future work must prioritize reducing this problem, they write . |
| Outcome: | a new model-driven technique can detect leakage and detect subpopulations in the datasets which exhibit it . the proposed technique is based on the progressive evaluation of cluster outliers (PECO) . it allows objective measurement of leakage, and automatic detection of subpopulations in the data which exhibit leakage. |
Investigating Memorization of Conspiracy Theories in Text Generation (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies examine conspiracy theories in social media, but they have not evaluated their presence in generative language models. |
| Approach: | They examine the ability of generative language models to generate conspiracy theory text . they highlight the difficulties of this task and discuss the drawbacks . |
| Outcome: | The proposed model can generate conspiracy theories without access to training data. |
TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing video generation models struggle to interpret compositional changes and synthesize components across different time steps. |
| Approach: | They propose a temporal compositionality benchmark that uses text prompts and ground truth videos to evaluate compositional changes in video. |
| Outcome: | The proposed benchmark can be used for text-to-video and image-to video generation. |
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | evaluators of long-context vision language models (VLMs) have not kept up with the rapid development of open-weight long-constraint language models. |
| Approach: | They propose a dynamic benchmark generator for evaluating long-context reasoning in vision language models. |
| Outcome: | The proposed model can ignore irrelevant information when answering queries, showing that current models lack this capability. |