Papers by Boyi Li
From Wrong To Right: A Recursive Approach Towards Vision-Language Explanation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for generating insightful explanations with limited annotations are limited. |
| Approach: | They propose a method that iteratively computes visual features, an answer, and an explanation to improve the explanation quality step by step until the answer converges. |
| Outcome: | The proposed method outperforms previous methods while utilizing 5% of the human-annotated explanations across 10 metrics, showing up to 4.2 and 1.3 increases in BLEU-1 score on the VCR and VQA-X datasets. |
Re-evaluating the Need for Visual Signals in Unsupervised Grammar Induction (2024.findings-naacl)
Copied to clipboard
Boyi Li, Rodolfo Corona, Karttikeya Mangalam, Catherine Chen, Daniel Flaherty, Serge Belongie, Kilian Weinberger, Jitendra Malik, Trevor Darrell, Dan Klein
| Challenge: | Recent studies show multimodal inputs can improve grammar induction, but weak textual baselines are needed for training. |
| Approach: | They use a fixed grammar family to compare multimodal grammar induction methods . they find multimodal inputs can improve grammar induction by grounding textual inputs to the visual world . |
| Outcome: | The proposed model outperforms weaker baselines on four benchmark datasets. |
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences (2025.acl-industry)
Copied to clipboard
Yusuke Hirota, Boyi Li, Ryo Hachiuma, Yueh-Hua Wu, Boris Ivanovic, Marco Pavone, Yejin Choi, Yu-Chiang Frank Wang, Yuta Nakashima, Chao-Han Huck Yang
| Challenge: | Large Vision-Language Models (LVLMs) have transformed image captioning . existing evaluations lack standardized criteria and a standardized evaluation framework . |
| Approach: | They propose a leaderboard for evaluating detailed captions that addresses three main gaps in existing evaluations: lack of standardized criteria, bias-aware assessments, and user preference considerations. |
| Outcome: | The proposed model evaluates caption quality, descriptiveness, risks, and societal biases while tailoring criteria to user preferences. |