Papers with Captioning
Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning (2020.emnlp-main)
Copied to clipboard
| Challenge: | Observable changes in the scene are reflected in captions, but actions are also linked to social aspects such as intentions, effects, and attributes that describe the agent. |
| Approach: | They propose to generate captions from videos that describe latent aspects of the human agent's actions. |
| Outcome: | The proposed model can be used to describe latent aspects of human actions in video clips and answer questions about videos. |
Afrispeech Semantics: Evaluating Audio–Semantic Reasoning in Spoken Language Models Across Domains and Accents (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent multimodal models are trained on large collections of audio-text pairs using contrastive learning or nexttoken prediction objectives. |
| Approach: | They evaluate audio language models across five semantic and paralinguistic reasoning tasks: entailment, consistency, plausibility, accent drift, and accent restraint. |
| Outcome: | The evaluations assess models across five tasks including entailment, consistency, plausibility, accent drift, and accent restraint. |
CapWAP: Image Captioning with a Purpose (2020.emnlp-main)
Copied to clipboard
| Challenge: | a traditional image captioning task uses generic reference captions to provide textual information about images. |
| Approach: | They propose a task that uses question-answer pairs to provide visual information instead of generic reference captions. |
| Outcome: | The proposed captioning with a purpose task can be tailored to meet user needs . question-answer pairs are used as a source of supervision for learning visual information needs a new task is proposed . |