| Challenge: | Existing methods for video captioning focus on generating generic descriptions that lack contextual knowledge. |
| Approach: | They propose a method that uses video meta-data to retrieve topically related news documents for a video and extracts the events and named entities from these documents. |
| Outcome: | The proposed model is based on a news video dataset and is evaluated on it. |
Similar Papers
VIEWS: Entity-Aware News Video Captioning (2024.emnlp-main)
Copied to clipboard
Hammad Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, Shih-Fu Chang
| Challenge: | Existing video captioning benchmarks and models produce generic captions for videos that lack specific identification of individuals, locations, or organizations. |
| Approach: | They propose a task of directly summarizing news videos into captions that are entity-aware . they validate the effectiveness of their approach across three video captioning models . |
| Outcome: | The proposed approach is effective across three video captioning models. |
Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge (2026.acl-long)
Copied to clipboard
| Challenge: | Existing text-to-video models struggle to accurately simulate real-world physics and dynamic entity interactions. |
| Approach: | They propose a framework that integrates graph-structured temporal knowledge into video latent diffusion models to enhance compositional generation and interaction fidelity. |
| Outcome: | The proposed framework enhances compositional generation and interaction fidelity by integrating graph-structured temporal knowledge into video latent diffusion models. |
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on temporal sentence grounding rely on expensive video-query paired annotations . despite this, there are no ground-truth annotations in the current work . |
| Approach: | They propose to use paired video-query and segment boundary annotations to generate temporal sentence grounding without training. |
| Outcome: | The proposed model outperforms existing unsupervised methods and beats supervised ones on two challenging datasets. |
Guiding the Flowing of Semantics: Interpretable Video Captioning via POS Tag (D19-1)
Copied to clipboard
| Challenge: | Existing models of video captioning use a network and semantics are mixed into one feature. |
| Approach: | They propose an Adaptive Semantic Guidance Network which instantiates whole video semantics to different POS-aware semantics with supervision of part of speech (POS) tag. |
| Outcome: | Extensive experiments show that the proposed model is more efficient than state-of-the-art models. |
Knowledge-Enriched Natural Language Generation (2021.emnlp-tutorials)
Copied to clipboard
| Challenge: | Knowledge-enriched text generation poses unique challenges in modeling and learning . a roadmap will outline the state-of-the-art methods to tackle these challenges . |
| Approach: | They propose a roadmap to tackle the challenges of knowledge-enriched text generation . they will dive deep into various technical components to illustrate how to represent knowledge . |
| Outcome: | This tutorial outlines the state-of-the-art methods to tackle the problem . it aims to show how to represent knowledge, feed knowledge into a generation model, evaluate results . |
Towards Exploiting Background Knowledge for Building Conversation Systems (D18-1)
Copied to clipboard
| Challenge: | Existing dialog datasets contain a sequence of utterances without any explicit background knowledge associated with them. |
| Approach: | They propose to use movie chats to generate responses by copying unstructured background knowledge . they use a dataset of 9K conversations to test whether responses are generated by copy-and-modify models . |
| Outcome: | The proposed model mimics human process of conversing by copying and/or modifying sentences from unstructured background knowledge. |
MovieUN: A Dataset for Movie Understanding and Narrating (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Automatic movie narration generation and narration grounding are important to provide a true movie experience for the blind and visually impaired. |
| Approach: | They propose to use movie clips as a benchmark to support automatic movie narration generation and narration grounding tasks. |
| Outcome: | The proposed methods are effective in supporting two movie-based tasks for the blind and visually impaired. |
Towards Content Transfer through Grounded Text Generation (N19-1)
Copied to clipboard
| Challenge: | Recent work in neural natural language generation has attracted significant interest in controlling the form of text, such as style, persona, and wordiness. |
| Approach: | They propose a task where the task is to generate a next sentence in a document that fits its context and is grounded in . external textual source such as a news story. |
| Outcome: | The proposed task is based on 640k Wikipedia referenced sentences paired with the source articles to show significant improvements against baselines. |
Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning (2020.emnlp-main)
Copied to clipboard
| Challenge: | Observable changes in the scene are reflected in captions, but actions are also linked to social aspects such as intentions, effects, and attributes that describe the agent. |
| Approach: | They propose to generate captions from videos that describe latent aspects of the human agent's actions. |
| Outcome: | The proposed model can be used to describe latent aspects of human actions in video clips and answer questions about videos. |
DESCGEN: A Distantly Supervised Datasetfor Generating Entity Descriptions (2021.acl-long)
Copied to clipboard
| Challenge: | Short textual descriptions of entities provide summaries of their key attributes but generating entity descriptions can be challenging since information is scattered across multiple sources with varied content and style. |
| Approach: | They propose to generate an entity summary description from 37K entities from Wikipedia and Fandom, paired with nine evidence documents on average. |
| Outcome: | The proposed task is entity-centric, more abstractive, and covers a wide range of domains. |