Challenge: Large language models (LLMs) are capable of complex reasoning when given a few input-output demos.
Approach: They use fewer input-output demos for each test query to study ICL . they do not observe significant degradation when using only one randomly chosen demo .
Outcome: The proposed model outperforms multi-demo models on the tasks in 2022.

Similar Papers

More Samples or More Prompts? Exploring Effective Few-Shot In-Context Learning for LLMs with In-Context Sampling (2024.findings-naacl)

Copied to clipboard

Challenge: Existing studies on LLM prompting focus on selecting a better set of data samples inside one single prompt input, but why not design and leverage multiple ICL prompts together to further improve the LLM’s performance?
Approach: They propose a low-resource LLM prompting technique to optimize the construction of multiple ICL prompt inputs to produce confident predictions.
Outcome: The proposed technique can produce confident predictions by optimizing the construction of multiple ICL prompt inputs on four NLI datasets and one QA dataset.
When natural language is not enough: The limits of in-context learning demonstrations in multilingual reasoning (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies have demonstrated the effectiveness of reasoning methods in eliciting multi-step reasoned answers from Large Language Models (LLMs) by leveraging in-context demonstrations.
Approach: They investigate how well CoT and PAL perform across languages for arithmetic and symbolic reasoning tasks.
Outcome: The proposed methods perform well in monolingual contexts, primarily in English, but have been limited in other languages.
In-Context Learning with Iterative Demonstration Selection (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing literature has highlighted the importance of selecting examples that are diverse or semantically similar to the test sample . Existing studies have shown that the optimal selection dimension, i.e., diversity or similarity, is task-specific.
Approach: They propose to use zero-shot chain-of-thought reasoning to iteratively select examples that are diverse but still strongly correlated with the test sample as ICL demonstrations.
Outcome: The proposed method outperforms existing demonstration selection methods on reasoning, question answering, and topic classification tasks.
Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning (2025.emnlp-main)

Copied to clipboard

Challenge: In-context learning (ICL) is a critical emerging capability of large language models (LLMs), enabling few-shot learning during inference by including a few demonstrations in the prompt.
Approach: They propose to use positional bias to study ICL's performance for the first time by examining the positional variation in demos, system prompt, and user message in LLM input.
Outcome: The proposed model can predict accuracy and accuracy when demos are placed at different positions in the input prompt and in the user message.
The Whole is Better than the Sum: Using Aggregated Demonstrations in In-Context Learning for Sequential Recommendation (2024.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) have shown excellent performance on various NLP tasks.
Approach: They propose a method that integrates multiple demonstration users into one aggregated demonstration to improve sequential recommendation.
Outcome: The proposed method outperforms state-of-the-art LLM-based sequential recommendation methods on three recommendation datasets.
In-Context Learning with Long-Context Models: An In-Depth Exploration (2025.naacl-long)

Copied to clipboard

Challenge: In-context learning is limited by context length, but it can be used for many tasks.
Approach: They study the behavior of in-context learning at an extreme context length . example retrieval shows excellent performance at low context lengths but has diminished gains .
Outcome: The proposed model can perform many tasks with reasonable accuracy when a few examples are provided in-context.
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning (2025.naacl-long)

Copied to clipboard

Challenge: Motivated by in-context learning capabilities of Large Language Models (LLMs), multimodal LLMs with additional visual modality are also exhibited with similar ICL abilities when multiple image-text pairs are provided as demonstrations.
Approach: They conduct systematic and principled evaluation of multimodal ICL for models of different scales on a broad spectrum of new yet critical tasks.
Outcome: The proposed model performance improves on a broad spectrum of new yet critical tasks.
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study (2025.coling-main)

Copied to clipboard

Challenge: Existing evaluation approaches to evaluate Large Language Models are affected by potential biases within LLMs.
Approach: They propose two many-shot In-Context Learning (ICL) prompt templates to help LLM evaluators mitigate potential biases.
Outcome: The proposed templates reduce biases by using in-context examples with model-generated rationales as references.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? (2022.emnlp-main)

Copied to clipboard

Challenge: Large language models can in-context learn by conditioning on a few input-label pairs and making predictions for new inputs.
Approach: They propose to use ground truth demonstrations to replace labels in demonstrations . they also show that other aspects of the demonstrations are key drivers of endtask performance .
Outcome: The proposed model outperforms zeroshot inference on a wide range of tasks using ground truth demonstrations.
Rethinking the Evaluation of In-Context Learning for LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies evaluate In-context learning methods based on task performance . however, this evaluation protocol overlooks the significant cost associated with the demonstration configuration process .
Approach: They propose a two-dimensional evaluation paradigm that considers both configuration costs and task performance.
Outcome: The proposed evaluation paradigm can be applied to any ICL method as a plugin.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations