Challenge: Multimodal in-context learning (ICL) is a key mechanism for harnessing the capabilities of large vision–language models.
Approach: They propose a transformer-based model with task-aware attention that dynamically configures ICL sequences.
Outcome: Experiments on five LVLMs and nine datasets show that TACO surpasses baselines across diverse ICL tasks.

Similar Papers

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks (2026.acl-long)

Copied to clipboard

Challenge: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations.
Approach: They propose a simple inference-stage enhancement method that reinforces task mapping transfer.
Outcome: The proposed method strengthens task mapping transfer in multimodal models . it performs comparable to text-only ICL in zero-shot settings but degrades significantly under few-shot demonstrations.
In-Context Learning Creates Task Vectors (2023.findings-emnlp)

Copied to clipboard

Challenge: In-context learning (ICL) is a powerful new learning paradigm for Large Language Models (LLMs).
Approach: They propose to use a model with a prompt and a query to learn a mapping based on two examples to produce the output.
Outcome: The proposed model can learn functions from a simple structure based on a training set and a single task vector calculated from the training set.
How does Multi-Task Training Affect Transformer In-Context Capabilities? Investigations with Function Classes (2024.naacl-short)

Copied to clipboard

Challenge: Multi-task learning (MTL) for generalist models is a promising direction that offers transfer learning potential.
Approach: They propose to combine multi-task learning (MTL) with in-context learning (ICL) to build models that can generalize to multiple tasks while being robust to out-of-distribution examples.
Outcome: The proposed training strategies enable models to learn difficult tasks while mixing in prior tasks, denoted as mixed curriculum.
One Task Vector is not Enough: A Large-Scale Study for In-Context Learning (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies limit comprehensive analysis of large language models based on task vectors . recent work points to "task vectors" as mechanism for encoding task rules .
Approach: They propose a novel task vector with 30 input-output pairs for in-context learning . they use a few prompt-based examples to adapt to new tasks without weight updates .
Outcome: Experiments with Llama-3-8B on QAF show task vector performance peaks at intermediate layer . complex tasks rely on multiple, subtask-specific vectors rather than a single vector .
Label Words as Local Task Vectors in In-Context Learning (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable abilities, one of the most important being in-context learning (ICL).
Approach: They hypothesized that the network creates a task vector in specific positions during ICL, which can be computed by averaging across the dataset.
Outcome: The proposed model can achieve zero-shot performance with dummy inputs comparable to few-shot learning by patching the global task vector.
What In-Context Learning “Learns” In-Context: Disentangling Task Recognition and Task Learning (2023.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) can perform in-context learning (ICL) with only a few demonstrations, but its mechanisms are not well-understood.
Approach: They characterize two ways in which LLMs leverage demonstrations to solve tasks with a few demonstrations.
Outcome: The proposed model achieves non-trivial performance with only TR, and TR does not improve with larger models or more demonstrations.
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning (2025.naacl-long)

Copied to clipboard

Challenge: Motivated by in-context learning capabilities of Large Language Models (LLMs), multimodal LLMs with additional visual modality are also exhibited with similar ICL abilities when multiple image-text pairs are provided as demonstrations.
Approach: They conduct systematic and principled evaluation of multimodal ICL for models of different scales on a broad spectrum of new yet critical tasks.
Outcome: The proposed model performance improves on a broad spectrum of new yet critical tasks.
Understanding In-Context Learning Beyond Transformers: An Investigation of State Space and Hybrid Architectures (2026.findings-acl)

Copied to clipboard

Challenge: In-context learning is an emergent ability from pretrained Large Language Models (LLMs).
Approach: They perform in-depth evaluations of in-context learning on transformers and hybrid large language models using behavioral probing and intervention-based methods.
Outcome: The proposed model performs well on state-of-the-art transformer, state-space, and hybrid large language models.
Contextual Representation Learning beyond Masked Language Modeling (2022.acl-long)

Copied to clipboard

Challenge: masked language models adopt sampled embeddings as anchors to estimate and inject contextual semantics to representations.
Approach: They propose a representation learning approach that uses embeddings as anchors to model contextual representations.
Outcome: The proposed model achieves 5x speedup and 1.2 points average improvement over MLM.
Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results Alignment (2024.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on the alignment of multimodal sequential learning using transformers.
Approach: They propose a constrained scheme to align the multiple attentional results from both local and global perspectives.
Outcome: The proposed scheme could align the multiple attentional results from both local and global perspectives, making the information capture more efficient.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations