AudioCaps: Generating Captions for Audios in The Wild (N19-1)

Copied to clipboard

Challenge: a dataset of 46K audio clips with human-written text pairs is used to generate captions for audio . the task of translating a multimedia input source into natural language has been extensively studied over the past few years .
Approach: They propose a top-down multi-scale encoder and aligned semantic attention for audio captioning.
Outcome: The proposed captions are faithful to audio inputs and better than existing models.

Similar Papers

Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing AAC datasets suffer from short and simplistic captions, limiting expressiveness and semantic depth.
Approach: They propose a multi-modal dataset that pairs audio with corresponding video and leverages large language models to generate rich, descriptive captions.
Outcome: The proposed framework outperforms existing benchmarks in caption length, lexical diversity, and human-rated quality.
Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for audio captioning lack fine-grained detail and contextual accuracy due to limited unimodal or superficial information.
Approach: They propose a two-stage automated pipeline that uses pretrained models to extract contextual cues from video . a large language model synthesizes these inputs to generate detailed and context-aware captions .
Outcome: The proposed method is scalable and generates detailed and context-aware captions on large-scale audio datasets.
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies (2025.findings-naacl)

Copied to clipboard

Challenge: Audio descriptions (ADs) are acoustic commentaries designed to assist blind and visually impaired individuals in accessing digital media content.
Approach: They examine how state-of-the-art NLP and CV technologies can be applied to generate ADs . they identify essential research directions for the future .
Outcome: The proposed technologies can be applied to generate audio descriptions (ADs) the process is time-consuming and costly, and requires significant human effort . the authors identify key research directions for the future .
Text-Free Image-to-Speech Synthesis Using Learned Segmental Units (2021.acl-long)

Copied to clipboard

Challenge: Existing models for synthesising fluent, natural-sounding spoken audio captions do not require natural language text as an intermediate representation or source of supervision.
Approach: They propose a model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supervision.
Outcome: The proposed model captures diverse visual semantics of images and can replace text with a set of discrete, sub-word speech units.
Knowledge-Enriched Natural Language Generation (2021.emnlp-tutorials)

Copied to clipboard

Challenge: Knowledge-enriched text generation poses unique challenges in modeling and learning . a roadmap will outline the state-of-the-art methods to tackle these challenges .
Approach: They propose a roadmap to tackle the challenges of knowledge-enriched text generation . they will dive deep into various technical components to illustrate how to represent knowledge .
Outcome: This tutorial outlines the state-of-the-art methods to tackle the problem . it aims to show how to represent knowledge, feed knowledge into a generation model, evaluate results .
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)

Copied to clipboard

Challenge: Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language.
Approach: They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels.
Outcome: The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.
ReCAP: Semantic Role Enhanced Caption Generation (2024.lrec-main)

Copied to clipboard

Challenge: Current vision language models lack specificity and overlook various aspects of the image.
Approach: They propose to use semantic roles as control signals to guide captions to specific argument structures by focusing on specific objects and their associated semantic roles instead of general descriptions.
Outcome: The proposed framework produces captions that exhibit enhanced quality, diversity, and controllability.
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis.
Approach: They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Outcome: The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023 (2026.tacl-1)

Copied to clipboard

Challenge: SciCap dataset launched in 2021 aims to generate high-quality captions for scientific figures.
Approach: They propose to use the SciCap dataset to develop models for captioning diverse figure types across various academic fields.
Outcome: The proposed models showed impressive performance on the SciCap dataset and in various vision-and-language tasks.
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling (2024.acl-long)

Copied to clipboard

Challenge: Subtitling is a crucial task for enhancing the accessibility of audiovisual content and relying on automatic transcripts for the three subtasks is uncharted territory.
Approach: They propose a model capable of producing automatic subtitles, completely eliminating any dependence on intermediate transcripts also for timestamp prediction.
Outcome: Experimental results show that the proposed model eliminates the need for intermediate transcripts for timestamp prediction across multiple language pairs and diverse conditions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations