Bridge Video and Text with Cascade Syntactic Structure (C18-1)

Copied to clipboard

Challenge: Using LSTM-CSS, we construct basic syntactic structure by completing syntastic structure.
Approach: They propose a video captioning approach that progressively completes syntactic structure by a conditional random field to construct basic syntaktic structure.
Outcome: The proposed method produces natural sentences with 42.3% and 28.5% accuracy compared to state-of-the-art methods.

Similar Papers

The Role of Syntactic Planning in Compositional Image Captioning (2021.eacl-main)

Copied to clipboard

Challenge: Image captioning is a core task in multimodal NLP, where the aim is to automatically describe the content of an image in natural language.
Approach: They propose to use syntactic tags and tokens to improve caption generalization . they also propose to model the syntakic structure of a caption to improve generalization.
Outcome: The proposed models improve generalization and performance on standard metrics while requiring syntactic and semantic knowledge of the language.
ReCAP: Semantic Role Enhanced Caption Generation (2024.lrec-main)

Copied to clipboard

Challenge: Current vision language models lack specificity and overlook various aspects of the image.
Approach: They propose to use semantic roles as control signals to guide captions to specific argument structures by focusing on specific objects and their associated semantic roles instead of general descriptions.
Outcome: The proposed framework produces captions that exhibit enhanced quality, diversity, and controllability.
End-to-end Dense Video Captioning as Sequence Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for dense video captioning use a two-stage generative process . but, more complex tasks are not able to fully utilize this powerful paradigm .
Approach: They propose to model two subtasks of dense video captioning as one sequence generation task and predict the events and the corresponding descriptions.
Outcome: Experiments on YouCook2 and ViTT show that the proposed model can be used on any video platform.
A hierarchical approach to vision-based language generation: from simple sentences to complex natural language (2020.coling-main)

Copied to clipboard

Challenge: Automating video to language translation is a challenging problem, but it is unclear what the neural network learns .
Approach: They propose a hierarchical approach to automatically describing videos in natural language . they propose generating video descriptions as sequences of simple sentences followed by a more complex and fluent description in natural languages.
Outcome: The proposed approach generates video descriptions as sequences of simple sentences, followed by a more complex and fluent description in natural language.
Language-Driven Region Pointer Advancement for Controllable Image Captioning (2020.coling-main)

Copied to clipboard

Challenge: Controllable Image Captioning is a recent sub-task of Image Captions wherein constraints are placed on which regions in an image should be described in the generated natural language caption.
Approach: They propose a method for predicting the timing of region pointer advancement by treating the advancement step as a natural part of the language structure via a NEXT-token.
Outcome: The proposed method agrees with ground-truth timing in the Flickr30k Entities test data with a precision of 86.55% and a recall of 97.92%.
ViLL-E: Video LLM Embeddings for Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Video Large Language Models excel at video understanding tasks where outputs are textual . however, they underperform specialized embedding-based models in Retrieval tasks .
Approach: They propose a video-LLM-based model with an embedding generation mechanism that allows the model to "think longer" for complex videos and stop early for easy ones.
Outcome: The proposed model outperforms specialized embedding-based models in video understanding tasks while remaining competitive on VideoQA tasks.
Pretrained Image-Text Models are Secretly Video Captioners (2025.naacl-short)

Copied to clipboard

Challenge: Current video captioning methods often incorporate intricate designs tailored to video inputs.
Approach: They adapt an image-based captioning model to address dynamic video sequences without modifications.
Outcome: The proposed model outperforms specialised captioning systems on major benchmarks.
O2NA: An Object-Oriented Non-Autoregressive Approach for Controllable Video Captioning (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for video captioning consider a sequence of frames and biases towards focused objects.
Approach: They propose an Object-Oriented Non-Autoregressive approach to video captioning . it performs three steps: 1) identify the focused objects and predict their locations . 2) generate related attribute words and relation words of these focused objects to form a draft caption .
Outcome: The proposed method achieves competitive results with the state-of-the-art methods but with higher diversity and faster inference speed.
Simpler but More Accurate Semantic Dependency Parsing (P18-2)

Copied to clipboard

Challenge: Syntactic dependency parsing is the most popular method for automatically extracting low-level relationships between words in a sentence.
Approach: They extend a syntactic dependency parser to train on and generate graph-structured representations that capture between-word relationships that are more closely related to the meaning of a sentence.
Outcome: The proposed system beats the current state-of-the-art system by 0.6% and linguistically richer representations push the margin even higher.
Grafting Pre-trained Models for Multimodal Headline Generation (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to generate video headlines with pre-trained language models are labor intensive and impractical.
Approach: They propose to graft the encoder from the pre-trained video-language model on the generative pre-trainer model and propose a consensus fusion mechanism for the integration of different components.
Outcome: The proposed model achieves strong results on a brand-new dataset collected from real-world applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations