| Challenge: | a dataset of 46K audio clips with human-written text pairs is used to generate captions for audio . the task of translating a multimedia input source into natural language has been extensively studied over the past few years . |
| Approach: | They propose a top-down multi-scale encoder and aligned semantic attention for audio captioning. |
| Outcome: | The proposed captions are faithful to audio inputs and better than existing models. |
Similar Papers
Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing AAC datasets suffer from short and simplistic captions, limiting expressiveness and semantic depth. |
| Approach: | They propose a multi-modal dataset that pairs audio with corresponding video and leverages large language models to generate rich, descriptive captions. |
| Outcome: | The proposed framework outperforms existing benchmarks in caption length, lexical diversity, and human-rated quality. |
Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for audio captioning lack fine-grained detail and contextual accuracy due to limited unimodal or superficial information. |
| Approach: | They propose a two-stage automated pipeline that uses pretrained models to extract contextual cues from video . a large language model synthesizes these inputs to generate detailed and context-aware captions . |
| Outcome: | The proposed method is scalable and generates detailed and context-aware captions on large-scale audio datasets. |
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies (2025.findings-naacl)
Copied to clipboard
| Challenge: | Audio descriptions (ADs) are acoustic commentaries designed to assist blind and visually impaired individuals in accessing digital media content. |
| Approach: | They examine how state-of-the-art NLP and CV technologies can be applied to generate ADs . they identify essential research directions for the future . |
| Outcome: | The proposed technologies can be applied to generate audio descriptions (ADs) the process is time-consuming and costly, and requires significant human effort . the authors identify key research directions for the future . |
Text-Free Image-to-Speech Synthesis Using Learned Segmental Units (2021.acl-long)
Copied to clipboard
| Challenge: | Existing models for synthesising fluent, natural-sounding spoken audio captions do not require natural language text as an intermediate representation or source of supervision. |
| Approach: | They propose a model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supervision. |
| Outcome: | The proposed model captures diverse visual semantics of images and can replace text with a set of discrete, sub-word speech units. |
Knowledge-Enriched Natural Language Generation (2021.emnlp-tutorials)
Copied to clipboard
| Challenge: | Knowledge-enriched text generation poses unique challenges in modeling and learning . a roadmap will outline the state-of-the-art methods to tackle these challenges . |
| Approach: | They propose a roadmap to tackle the challenges of knowledge-enriched text generation . they will dive deep into various technical components to illustrate how to represent knowledge . |
| Outcome: | This tutorial outlines the state-of-the-art methods to tackle the problem . it aims to show how to represent knowledge, feed knowledge into a generation model, evaluate results . |
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)
Copied to clipboard
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, Emmanuel Dupoux
| Challenge: | Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language. |
| Approach: | They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels. |
| Outcome: | The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation. |
ReCAP: Semantic Role Enhanced Caption Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Current vision language models lack specificity and overlook various aspects of the image. |
| Approach: | They propose to use semantic roles as control signals to guide captions to specific argument structures by focusing on specific objects and their associated semantic roles instead of general descriptions. |
| Outcome: | The proposed framework produces captions that exhibit enhanced quality, diversity, and controllability. |
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis. |
| Approach: | They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
| Outcome: | The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023 (2026.tacl-1)
Copied to clipboard
Ting-Yao Hsu, Yi-Li Hsu, Shaurya Rohatgi, Chieh-Yang Huang, Ho Yin Sam Ng, Ryan Rossi, Sungchul Kim, Tong Yu, Lun-Wei Ku, Clyde Lee Giles, Ting-Hao Huang
| Challenge: | SciCap dataset launched in 2021 aims to generate high-quality captions for scientific figures. |
| Approach: | They propose to use the SciCap dataset to develop models for captioning diverse figure types across various academic fields. |
| Outcome: | The proposed models showed impressive performance on the SciCap dataset and in various vision-and-language tasks. |
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling (2024.acl-long)
Copied to clipboard
| Challenge: | Subtitling is a crucial task for enhancing the accessibility of audiovisual content and relying on automatic transcripts for the three subtasks is uncharted territory. |
| Approach: | They propose a model capable of producing automatic subtitles, completely eliminating any dependence on intermediate transcripts also for timestamp prediction. |
| Outcome: | Experimental results show that the proposed model eliminates the need for intermediate transcripts for timestamp prediction across multiple language pairs and diverse conditions. |