Proceedings of the Beyond Vision and LANguage: inTEgrating Real-world kNowledge (LANTERN)

9 papers
Structure Learning for Neural Module Networks (D19-64)

Copied to clipboard

Challenge: Neural Module Networks are a class of neural networks that involve human-specified neural modules . current models only learn the parameters of the modules and/or the order of their execution .
Approach: They propose to learn internal structure and sequence without extra supervisory signals . they use dynamically composable modules which are then assembled into a layout .
Outcome: The proposed model performs comparable to models using hand-designed modules.
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-64)

Copied to clipboard

Challenge: Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English.
Approach: They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages.
Outcome: The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments.
Big Generalizations with Small Data: Exploring the Role of Training Samples in Learning Adjectives of Size (D19-64)

Copied to clipboard

Challenge: In previous work, models have been shown to fail in generalizing to unseen adjective-noun combinations.
Approach: They propose a visual reasoning task dealing with quantities that challenges models to learn the meaning of size adjectives from visually-grounded contexts.
Outcome: The proposed task is based on a visual reasoning task dealing with quantities and shows that seeing some of the cases during training helps a model understand the rule subtending the task.
Eigencharacter: An Embedding of Chinese Character Orthography (D19-64)

Copied to clipboard

Challenge: Chinese characters encode world knowledge through thousands of years evolution .
Approach: They propose an embedding approach to encode Chinese orthography knowledge using eigencharacter space.
Outcome: The proposed representations encode lexical knowledge embedded in Chinese characters and integrate with other computational models.
On the Role of Scene Graphs in Image Captioning (D19-64)

Copied to clipboard

Challenge: Recent captioning approaches rely on ad-hoc approaches to obtain graphs for images, but they introduce noise and it is unclear the effect of parser errors on captioning accuracy.
Approach: They investigate whether scene graphs can help image captioning . they show that a scene graph parser can boost performance almost as much as ground truth graphs .
Outcome: The proposed parser can boost performance almost as much as ground truth graphs .
Understanding the Effect of Textual Adversaries in Multimodal Machine Translation (D19-64)

Copied to clipboard

Challenge: Existing studies show that multimodal machine translation systems are better than text-only systems at translating phrases that have a direct correspondence in the image.
Approach: They conduct experiments with both visual and textual adversaries to understand the role of textual inputs in multimodal machine translation.
Outcome: The proposed model can recover masked tokens in the source sentences . the proposed model is based on a model with a visual modality .
Learning to request guidance in emergent language (D19-64)

Copied to clipboard

Challenge: Previous research into agent communication has shown that a pre-trained guide can speed up the learning process of an imitation learning agent.
Approach: They extend one-directional communication by a one-bit communication channel from the learner back to the guide and limit the guidance by penalizing the learners for these requests.
Outcome: The proposed guide can speed up the learning process of an imitation learning agent by providing the learner with discrete messages in an emerged language about how to solve the task.
At a Glance: The Impact of Gaze Aggregation Views on Syntactic Tagging (D19-64)

Copied to clipboard

Challenge: Recent work uses gaze data at the type level or at the token level and mostly from a single eye-tracking corpus.
Approach: They propose to use gaze data to capture central tendency or variability of gaze data and to integrate binary phrase chunking and part-of-speech tagging.
Outcome: The proposed approaches capture the central tendency or variability of gaze data better than proposed local views which retain individual participant information.
Seeded self-play for language learning (D19-64)

Copied to clipboard

Challenge: Current methods for learning human language are too data inefficient to learn it in this way.
Approach: They propose to train a meta-learning agent in simulation to interact with populations of pre-trained agents, each with their own distinct communication protocol.
Outcome: The proposed algorithm minimizes the number of on-policy interactions while learning human language while minimizing the number on-political interactions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations