Challenge: Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition.
Approach: They propose a multimodal language acquisition model trained from image-caption pairs on naturalistic data using cross-modal self-supervision.
Outcome: The proposed model learns word categories and object recognition abilities, the authors show . their model is trained from image-caption pairs on naturalistic data using cross-modal self-supervision .

Similar Papers

Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in generative language models have demonstrated their ability to memorize knowledge from documents and recall knowledge to respond to user queries effectively.
Approach: They propose to enable multimodal large language models to memorize and recall images within their parameters.
Outcome: The proposed model performs well even with large-scale image candidate sets.
Word Acquisition in Neural Language Models (2022.tacl-1)

Copied to clipboard

Challenge: Language models acquire individual words during training, based on unigram token frequencies, before transitioning loosely to bigram probabilities, eventually converging on more nuanced predictions.
Approach: They examine how neural language models acquire individual words during training, extracting learning curves and ages of acquisition for over 600 words on the MacArthur-Bates Communicative Development Inventory.
Outcome: The models follow consistent patterns during training for both unidirectional and bidirectional models, and for both LSTM and Transformer architectures.
Is Word Segmentation Child’s Play in All Languages? (P19-1)

Copied to clipboard

Challenge: Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development .
Approach: They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants.
Outcome: The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid .
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.
Cross-Modal Taxonomic Generalization in (Vision-) Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that language models learn from surface form to learn from more grounded evidence.
Approach: They propose to use a vision-language model to learn hypernyms from images . they find that the model can recover this knowledge and generalize even when there is no hypernomia in the image.
Outcome: The proposed model can recover this knowledge and generalize even when the model receives no evidence of hypernyms during training.
Human Inspired Progressive Alignment and Comparative Learning for Grounded Word Acquisition (2023.acl-long)

Copied to clipboard

Challenge: a recent study shows that word acquisition is an efficient, supervised, and continual process.
Approach: They develop a computational process for word acquisition through comparative learning . they frame the acquisition of words as representation-symbol mapping .
Outcome: The proposed method can be used to learn the meaning of a word efficiently and efficiently.
On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization (2022.findings-emnlp)

Copied to clipboard

Challenge: Combining visual modality with pretrained language models has been effective for descriptive tasks such as image captioning.
Approach: They ask: do multimodal models combine visual and visual adapted language models? they find that CLIP image representations and scaling of language models do not consistently improve self-rationalization in multimodal tasks.
Outcome: The proposed model types do not consistently improve self-rationalization in multimodal tasks.
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-66)

Copied to clipboard

Challenge: Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English.
Approach: They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages.
Outcome: The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments.
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-64)

Copied to clipboard

Challenge: Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English.
Approach: They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages.
Outcome: The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations