Challenge: Visual features are promising for learning bootstrap textual models, but blackbox learning models make it difficult to isolate the specific contribution of visual components.
Approach: They propose to use alignments between phrases and images as a learning signal for syntax acquisition.
Outcome: The proposed model performs better than the previous model, but it is significantly less expressive.

Similar Papers

Visually Grounded Neural Syntax Acquisition (P19-1)

Copied to clipboard

Challenge: a visually grounded neural syntax learner is an approach for learning syntactic representations without any supervision.
Approach: They propose a visually grounded neural syntax learner that acquires syntax by looking at images and reading captions.
Outcome: The proposed model outperforms unsupervised approaches on the MSCOCO data set . it is more stable with choice of initialization and amount of training data, the authors show .
Visually Grounded Compound PCFGs (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work on visual groundings for language understanding has been drawing much attention.
Approach: They propose to use an extension of probabilistic context-free grammar model to do fully-differentiable end-to-end visually grounded learning.
Outcome: The proposed model outperforms the previous grounded model and significantly outperformed the previous model on the MSCOCO test captions.
Learning Visually Grounded Sentence Representations (N18-1)

Copied to clipboard

Challenge: Unsupervised sentence representation models suffer from the grounding problem because of lack of association between symbols and external information.
Approach: They train a sentence encoder to predict image features of a caption and use them as sentence representations.
Outcome: The proposed model improves on word embeddings and word representations on standard benchmarks.
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners? (2025.coling-main)

Copied to clipboard

Challenge: Neural language models (LMs) are arguably less data-efficient than humans from a language acquisition perspective.
Approach: They investigate the advantage of grounded language acquisition over visual input to improve syntactic generalization.
Outcome: The proposed model is less efficient than humans in language acquisition . it shows that visual input helps syntactic generalization, but not vision .
Grounded PCFG Induction with Images (2020.aacl-main)

Copied to clipboard

Challenge: Recent work in unsupervised parsing has tried to incorporate visual information into learning, but results suggest that these models need linguistic bias to compete against models that only rely on text.
Approach: They propose to use visual information from images for labeled parsing and compare them to existing models which only use text.
Outcome: The proposed models achieve state-of-the-art results on multilingual induction datasets even without help from linguistic knowledge or pretrained image encoders.
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes (2024.naacl-long)

Copied to clipboard

Challenge: Modern neural language models (LMs) require distinctly un-human-like ways to achieve these results.
Approach: They train a diverse set of LM architectures with and without auxiliary visual supervision on datasets of varying scales.
Outcome: The proposed models exhibit better learning of syntactic categories, lexical relations, semantic features, word similarity and alignment with human neural representations.
How Well Do Text Embedding Models Understand Syntax? (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing text embedding models have not addressed syntactic understanding challenges, highlighting ineffectiveness and enhancing generalization ability.
Approach: They propose to examine the ability of text embedding models to generalize across syntactic contexts.
Outcome: The proposed models exhibit high similarity socres at this simple task.
Beyond Instructional Videos: Probing for More Diverse Visual-Textual Grounding on YouTube (2020.emnlp-main)

Copied to clipboard

Challenge: a representative pretraining model is fit to a diverse YouTube8M dataset . a priori, this domain is relatively easy for instructional videos .
Approach: They fit a representative pretraining model to a YouTube8M dataset and examine its success and failure cases.
Outcome: The proposed model can be trained on more diverse video corpora and achieve high performance on many video understanding tasks.
Textual Supervision for Visually Grounded Spoken Language Understanding (2020.findings-emnlp)

Copied to clipboard

Challenge: a new approach to spoken language understanding extracts semantic information directly from speech without relying on transcriptions.
Approach: They propose to use textual supervision to train visually-grounded models of spoken language understanding without relying on transcriptions.
Outcome: The proposed model improves when enough text is available, the study shows . compared with pipeline-based models, the pipeline approach performs better when enough data is available .
The Limitations of Limited Context for Constituency Parsing (2021.acl-long)

Copied to clipboard

Challenge: a language model that is syntax-aware can produce better samples, authors say . a recent study shows that neural approaches to syntax can perform unsupervised syntactic parsing .
Approach: They propose to incorporate syntax into neural approaches in NLP to produce better samples . they find that the first time neural approaches were able to perform unsupervised syntactic parsing .
Outcome: The proposed models can perform unsupervised syntactic parsing, but they are lagging behind . the proposed models are based on a sandbox of probabilistic context-free-grammars .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations