| Challenge: | Low-quality captions are common in scientific articles and can decrease understanding . this paper aims to develop an end-to-end neural framework to generate informative, high-quality figure captions for scientific figures and charts. |
| Approach: | They propose an end-to-end neural framework to automatically generate captions for scientific figures from a large-scale dataset . they used figure-type classification, sub-figure identification, text normalization, and caption text selection to build models that caption graph plots, the dominant figure type. |
| Outcome: | The proposed model can generate high-quality captions for scientific figures and charts from a large figure-caption dataset from arXiv. |
Similar Papers
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023 (2026.tacl-1)
Copied to clipboard
Ting-Yao Hsu, Yi-Li Hsu, Shaurya Rohatgi, Chieh-Yang Huang, Ho Yin Sam Ng, Ryan Rossi, Sungchul Kim, Tong Yu, Lun-Wei Ku, Clyde Lee Giles, Ting-Hao Huang
| Challenge: | SciCap dataset launched in 2021 aims to generate high-quality captions for scientific figures. |
| Approach: | They propose to use the SciCap dataset to develop models for captioning diverse figure types across various academic fields. |
| Outcome: | The proposed models showed impressive performance on the SciCap dataset and in various vision-and-language tasks. |
GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing algorithms that generate captions for scientific figures are costly and dependent on author-written captions. |
| Approach: | They constructed a human evaluation dataset that contains human judgments for 3,600 scientific figure captions for 600 arXiv figures. |
| Outcome: | The proposed model outperforms all other models and outperformed undergraduates in achieving a Kendall correlation score of 0.401 with Ph.D. students’ rankings. |
Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing datasets for evaluating text-to-image generation focus mostly on real-life images, which poses challenges for assessing academic figure generation given real scientific captions. |
| Approach: | They propose a dataset that first provides a Holistic Evaluation for Academic caption-to-Figure Generation (HE4AFG) they collect real figure captions from 8 scientific domains and generate 3,900 evaluation samples . |
| Outcome: | The proposed model provides high-quality human ratings in terms of three aspects—scientific aesthetic (SA), topic relevance (TR), and attribute correctness (AC). |
SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Generating texts in scientific papers requires not only capturing the content contained within the given input but also frequently acquiring the external information called context. |
| Approach: | They propose a task of context-aware text generation in the scientific domain to exploit the contributions of context in generated texts. |
| Outcome: | The proposed dataset comprehensively benchmarks the efficacy of the proposed dataset in generating description and paragraph. |
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles (2025.findings-emnlp)
Copied to clipboard
Ho Yin Sam Ng, Edward Hsu, Aashish Anantha Ramakrishnan, Branislav Kveton, Nedim Lipka, Franck Dernoncourt, Dongwon Lee, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ting-Hao Kenneth Huang
| Challenge: | Figure captions are crucial for helping readers understand and remember a figure’s key message. |
| Approach: | They propose a dataset for personalized figure caption generation with multimodal figure profiles that provide inputs and profiles for each figure . |
| Outcome: | The proposed dataset provides inputs and profiles for personalized figure caption generation with multimodal figure profiles. |
TLDR: Extreme Summarization of Scientific Documents (2020.findings-emnlp)
Copied to clipboard
| Challenge: | TLDR generation requires expert background knowledge and understanding of complex domain-specific language. |
| Approach: | They propose a learning strategy that exploits titles as an auxiliary training signal. |
| Outcome: | The proposed method improves upon strong baselines under both automated metrics and human evaluations. |
AudioCaps: Generating Captions for Audios in The Wild (N19-1)
Copied to clipboard
| Challenge: | a dataset of 46K audio clips with human-written text pairs is used to generate captions for audio . the task of translating a multimedia input source into natural language has been extensively studied over the past few years . |
| Approach: | They propose a top-down multi-scale encoder and aligned semantic attention for audio captioning. |
| Outcome: | The proposed captions are faithful to audio inputs and better than existing models. |
FigEx: Aligned Extraction of Scientific Figures and Captions (2025.findings-emnlp)
Copied to clipboard
| Challenge: | FigEx is a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers. |
| Approach: | They propose a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers. |
| Outcome: | The proposed model improves subfigure detection APb over Grounding DINO by 0.023 and boosts caption separation BLEU over Llama-2-13B by 0.465. |
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning (P18-1)
Copied to clipboard
| Challenge: | Practical applications of automatic image description systems include leveraging descriptions for image indexing or retrieval, and helping those with visual impairments by transforming visual signals into information that can be communicated via text-to-speech technology. |
| Approach: | They propose to extract and filter image caption annotations from billions of webpages and use them to train models. |
| Outcome: | The proposed model architectures perform better when trained on the Conceptual Captions dataset. |
Longform Multimodal Lay Summarization of Scientific Papers: Towards Automatically Generating Science Blogs from Research Articles (2024.lrec-main)
Copied to clipboard
| Challenge: | Science blogs and lay-speak are critical to communicating scientific information to the general public and policymakers. |
| Approach: | They propose to use presentation transcripts and slides to generate a scientific blog from a research article in layperson's terms. |
| Outcome: | The proposed approach can generate a blog text and select the most relevant figures to explain a research article in layperson’s terms, essentially a science blog. |