Incorporating Background Knowledge into Video Description Generation (D18-1)

Copied to clipboard

Challenge: Existing methods for video captioning focus on generating generic descriptions that lack contextual knowledge.
Approach: They propose a method that uses video meta-data to retrieve topically related news documents for a video and extracts the events and named entities from these documents.
Outcome: The proposed model is based on a news video dataset and is evaluated on it.

Similar Papers

VIEWS: Entity-Aware News Video Captioning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing video captioning benchmarks and models produce generic captions for videos that lack specific identification of individuals, locations, or organizations.
Approach: They propose a task of directly summarizing news videos into captions that are entity-aware . they validate the effectiveness of their approach across three video captioning models .
Outcome: The proposed approach is effective across three video captioning models.
Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge (2026.acl-long)

Copied to clipboard

Challenge: Existing text-to-video models struggle to accurately simulate real-world physics and dynamic entity interactions.
Approach: They propose a framework that integrates graph-structured temporal knowledge into video latent diffusion models to enhance compositional generation and interaction fidelity.
Outcome: The proposed framework enhances compositional generation and interaction fidelity by integrating graph-structured temporal knowledge into video latent diffusion models.
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work on temporal sentence grounding rely on expensive video-query paired annotations . despite this, there are no ground-truth annotations in the current work .
Approach: They propose to use paired video-query and segment boundary annotations to generate temporal sentence grounding without training.
Outcome: The proposed model outperforms existing unsupervised methods and beats supervised ones on two challenging datasets.
Guiding the Flowing of Semantics: Interpretable Video Captioning via POS Tag (D19-1)

Copied to clipboard

Challenge: Existing models of video captioning use a network and semantics are mixed into one feature.
Approach: They propose an Adaptive Semantic Guidance Network which instantiates whole video semantics to different POS-aware semantics with supervision of part of speech (POS) tag.
Outcome: Extensive experiments show that the proposed model is more efficient than state-of-the-art models.
Knowledge-Enriched Natural Language Generation (2021.emnlp-tutorials)

Copied to clipboard

Challenge: Knowledge-enriched text generation poses unique challenges in modeling and learning . a roadmap will outline the state-of-the-art methods to tackle these challenges .
Approach: They propose a roadmap to tackle the challenges of knowledge-enriched text generation . they will dive deep into various technical components to illustrate how to represent knowledge .
Outcome: This tutorial outlines the state-of-the-art methods to tackle the problem . it aims to show how to represent knowledge, feed knowledge into a generation model, evaluate results .
Towards Exploiting Background Knowledge for Building Conversation Systems (D18-1)

Copied to clipboard

Challenge: Existing dialog datasets contain a sequence of utterances without any explicit background knowledge associated with them.
Approach: They propose to use movie chats to generate responses by copying unstructured background knowledge . they use a dataset of 9K conversations to test whether responses are generated by copy-and-modify models .
Outcome: The proposed model mimics human process of conversing by copying and/or modifying sentences from unstructured background knowledge.
MovieUN: A Dataset for Movie Understanding and Narrating (2022.findings-emnlp)

Copied to clipboard

Challenge: Automatic movie narration generation and narration grounding are important to provide a true movie experience for the blind and visually impaired.
Approach: They propose to use movie clips as a benchmark to support automatic movie narration generation and narration grounding tasks.
Outcome: The proposed methods are effective in supporting two movie-based tasks for the blind and visually impaired.
Towards Content Transfer through Grounded Text Generation (N19-1)

Copied to clipboard

Challenge: Recent work in neural natural language generation has attracted significant interest in controlling the form of text, such as style, persona, and wordiness.
Approach: They propose a task where the task is to generate a next sentence in a document that fits its context and is grounded in . external textual source such as a news story.
Outcome: The proposed task is based on 640k Wikipedia referenced sentences paired with the source articles to show significant improvements against baselines.
Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning (2020.emnlp-main)

Copied to clipboard

Challenge: Observable changes in the scene are reflected in captions, but actions are also linked to social aspects such as intentions, effects, and attributes that describe the agent.
Approach: They propose to generate captions from videos that describe latent aspects of the human agent's actions.
Outcome: The proposed model can be used to describe latent aspects of human actions in video clips and answer questions about videos.
DESCGEN: A Distantly Supervised Datasetfor Generating Entity Descriptions (2021.acl-long)

Copied to clipboard

Challenge: Short textual descriptions of entities provide summaries of their key attributes but generating entity descriptions can be challenging since information is scattered across multiple sources with varied content and style.
Approach: They propose to generate an entity summary description from 37K entities from Wikipedia and Fandom, paired with nine evidence documents on average.
Outcome: The proposed task is entity-centric, more abstractive, and covers a wide range of domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations