Papers with TVQA

5 papers
TVQA: Localized, Compositional Video Question Answering (D18-1)

Copied to clipboard

Challenge: Recent studies have focused on image-based question-answering (QA) tasks, but little has been done on video-based QA.
Approach: They present a large-scale video QA dataset based on 6 popular TV shows . they provide analysis of the new dataset and trainable neural network framework .
Outcome: The proposed dataset includes 152,545 QA pairs from 21,793 clips spanning over 460 hours of video.
Continuous Language Generative Flow (2021.acl-long)

Copied to clipboard

Challenge: Recent years have witnessed various types of generative models for natural language generation (NLG), especially RNNs or transformers.
Approach: They propose a flow-based language generation model that adapts flow-derived generative models to language generation via continuous input embeddings, adapted affine coupling structures, and a novel architecture for autoregressive text generation.
Outcome: The proposed model improves on QG and NMT and improves performance over baselines on SQuAD and TVQA and NML16.
MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering (2020.findings-emnlp)

Copied to clipboard

Challenge: MMFT-BERT is a multimodal fusion transformer that decomposes input modalities into different BERT instances with similar architectures, but variable weights.
Approach: They propose a multimodal fusion transformer with BERT encodings to solve Visual Question Answering (VQA) .
Outcome: The proposed method achieves SOTA results on the TVQA dataset and TVQA-Visual, an isolated diagnostic subset of TVQA, which strictly requires the knowledge of visual (V) modality based on a human annotator’s judgment.
Episodic Memory Reader: Learning What to Remember for Question Answering from Streaming Data (P19-1)

Copied to clipboard

Challenge: Existing QA methods lack scalability and performance is difficult to solve with document-level contexts.
Approach: They propose an end-to-end deep network model that sequentially reads the input contexts into an external memory while replacing memories that are less important for answering unseen questions.
Outcome: The proposed model improves on a synthetic dataset and real-world large-scale textual and video QA datasets.
Question-Instructed Visual Descriptions for Zero-Shot Video Answering (2024.findings-acl)

Copied to clipboard

Challenge: Existing models for video QA rely on complex architectures, expensive pipelines or closed models like GPTs.
Approach: They propose a single instruction-aware open vision-language model to tackle videoQA using frame descriptions.
Outcome: The proposed framework achieves higher performance than current state-of-the-art models on videoQA benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations