Partners in Crime: Multi-view Sequential Inference for Movie Understanding (D19-1)

Copied to clipboard

Challenge: Existing multi-view learning approaches are tested in unsupervised setups, allowing for learning of representation for monolithic data points, not sequences.
Approach: They propose a neural architecture paired with a novel objective for incremental inference that integrates multi-view information for sequence prediction problems.
Outcome: The proposed model outperforms previous work and strong baselines on two crime cases and speaker type tagging tasks that contribute to movie understanding.

Similar Papers

Weakly-Supervised Learning of Visual Relations in Multimodal Pretraining (2023.emnlp-main)

Copied to clipboard

Challenge: Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations.
Approach: They propose two approaches to contextualise visual entities in a multimodal setup by using verbalised scene graphs and masked relation prediction.
Outcome: The proposed models can learn better representations from weakly-supervised relations data.
Multi-view Story Characterization from Movie Plot Synopses and Reviews (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for characterizing stories by generating tags from synopses suffer from coverage issues.
Approach: They propose to use synopses and reviews to characterize stories by inferring attributes such as theme and style from written synopsis and reviews.
Outcome: The proposed model improves over methods that only use synopses and reviews . it can extract a complementary set of story attributes from reviews without supervision .
Multi-view and Cross-view Brain Decoding (2022.coling-1)

Copied to clipboard

Challenge: a recent study has shown that brain decoding models can decode concepts from single view . a multi-view decoder can take brain recordings for any view as input and predict the concept .
Approach: They propose to build a multi-view decoder that can take brain recordings for any view as input and predict the concept.
Outcome: The proposed systems can decode concepts from brain recordings from any view . the proposed systems have 0.68 pairwise accuracy across view pairs and 0.8 average pairwise precision across tasks.
M3: A Multi-View Fusion and Multi-Decoding Network for Multi-Document Reading Comprehension (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for multi-document reading comprehension cannot make full of the advantages of both approaches.
Approach: They propose a multi-view fusion and multi-decoding method that integrates multiple documents for answering questions.
Outcome: The proposed method improves on two mainstream multi-document reading comprehension datasets.
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks (2026.findings-acl)

Copied to clipboard

Challenge: Survey aims to identify challenges of multimodal unlearning for vision, language, audio and video . retraining after deletion requests or policy updates is often impractical, survey finds .
Approach: They propose to enable selective removal across modalities while retaining overall utility.
Outcome: This study compares models with existing models to identify weaknesses and improves performance.
Learning a Simple and Effective Model for Multi-turn Response Generation with Auxiliary Tasks (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to multi-turn response generation for open-domain dialogues have a complexity problem . auxiliary tasks that relate to context understanding can guide the learning of the generation model .
Approach: They propose a multi-turn response generation model that has a simple structure yet can effectively leverage conversation contexts for response generation.
Outcome: The proposed model outperforms state-of-the-art models in response quality and human judgment . it also enjoys a faster decoding process .
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 TV and movie series is filled with challenging multi-party dialogues.
Approach: They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues.
Outcome: The proposed dataset is a step towards better multi-party dialogue structuring and understanding.
Tutorial on Multimodal Machine Learning (2022.naacl-tutorials)

Copied to clipboard

Challenge: Multimodal machine learning is a challenging but crucial area with numerous applications in multimedia, affective computing, robotics, finance, HCI, and healthcare.
Approach: This tutorial will describe an updated taxonomy on multimodal machine learning synthesizing its core technical challenges and major directions for future research.
Outcome: The proposed taxonomy synthesizes the core technical challenges and major directions for future research.
Multi-component Causal Tracing in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are prone to various forms of safety risks, such as learning and propagating societal biases and even creating harmful or deceptive content through jailbreak attacks.
Approach: They propose a framework for causally tracing multiple components simultaneously that systematically identifies the subsets of components most critical to a desired performance metric.
Outcome: The proposed method outperforms existing methods in identifying components critical to a desired performance metric.
A Multi-source Graph Representation of the Movie Domain for Recommendation Dialogues Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Graph databases are well-suited for crossreferencing information from multiple sources to support machine learning tasks.
Approach: They propose a graph-based structure of multiple resources enriched with graph analytics approaches to provide an encompassing view of the movie recommendation domain and of the way people talk about it during the recommendation task.
Outcome: The proposed graph-based structure provides an encompassing view of the domain and of the way people talk about it during the recommendation task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations