Task-aware Retrieval with Instructions (2023.findings-acl)

Copied to clipboard

Challenge: Existing models that learn intents from labeled data are complicated and require a vast number of annotated examples to train a model.
Approach: They propose a general-purpose task-aware retrieval system with instructions that can adapt to a new task without any parameter updates.
Outcome: The proposed system outperforms two benchmarks on a set of domains and tasks on X2-Retrieval.

Similar Papers

MAIR: A Massive Benchmark for Evaluating Instructed Retrieval (2024.emnlp-main)

Copied to clipboard

Challenge: Existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models.
Approach: They propose a multi-task instruction-tuned IR benchmark that includes 126 distinct IR tasks across 6 domains.
Outcome: The proposed model performs better on instruction-tuned models than non-instruction-tunned models on MAIR.
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks (2024.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that instruction tuning improves zero-shot generalization across various tasks and improves performance of specific tasks.
Approach: They propose a task selection method that leverages instruction information alone to identify relevant tasks and optimize instruction tuning for specific tasks.
Outcome: The proposed method is significantly more efficient than traditional approaches, which require complex measurements of pairwise transferability between tasks or the creation of data samples for the target task.
Fine-Tuning Large Language Models with Sequential Instructions (2025.naacl-long)

Copied to clipboard

Challenge: Existing instruction-tuned models struggle to adhere to a query with multiple intentions, which impairs their performance when the completion of several tasks is demanded by a single command.
Approach: They develop an automatic process that turns existing data into diverse and complex task chains and a new benchmark to evaluate a model’s ability to follow all the instructions in a sequence.
Outcome: The proposed model can follow instructions better and deliver higher results in coding, maths, and open-ended generation.
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions (2025.naacl-long)

Copied to clipboard

Challenge: Modern language models (LMs) are capable of following long and complex instructions that enable a large and diverse set of user requests.
Approach: They propose a dataset that contains an instruction evaluation benchmark and a training set to help IR models learn to follow instructions.
Outcome: The proposed model improves after fine-tuning on a training set and rigorous instruction evaluation benchmark.
Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning (2023.acl-short)

Copied to clipboard

Challenge: Recent studies on instruction tuning (IT) have achieved great performance with zero-shot generalizability to unseen tasks.
Approach: They analyze how models utilize instructions during IT by comparing model training with altered vs. original instructions.
Outcome: The proposed model outperforms naive models in low resource setting.
UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers (2025.naacl-long)

Copied to clipboard

Challenge: Existing information retrieval models assume a homogeneous structure for knowledge sources and user queries, limiting their applicability in real-world settings.
Approach: They propose a unified instruction-aware heterogeneous knowledge retriever that builds a heterogenous retrieval space for heterogenized knowledge and follows diverse user instructions to retrieve knowledge in specified types.
Outcome: The proposed framework outperforms state-of-the-art methods on CompMix-IR . it achieves 6.36% relative improvements and 54.23% relative improvements .
Multi-Task Retrieval for Knowledge-Intensive Tasks (2021.acl-long)

Copied to clipboard

Challenge: Knowledge-intensive tasks require large amounts of knowledge about the world . recent neural retrieval models achieve better results by learning directly from task-specific training data.
Approach: They propose a multi-task trained neural retrieval model that can be universally trained on a wide variety of problems.
Outcome: The proposed model outperforms specialised retrievers on a few-shot setting and matches or improves state-of-the-art on multiple benchmarks.
Robustness of Learning from Task Instructions (2023.findings-acl)

Copied to clipboard

Challenge: traditional supervised learning mostly works on individual tasks and requires training on a large set of task-specific examples.
Approach: a new study investigates the system robustness when instructions are manipulated and paraphrased . task instructions give the model the definition of the task and allow it to output the appropriate answer .
Outcome: a new study shows that supervised learning is robust when instructions are manipulated, paraphrased or iii from different levels of conciseness.
Efficiently Enhancing Zero-Shot Performance of Instruction Following Model via Retrieval of Soft Prompt (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that adding a instruction tuning stage to training large language models can improve zero-shot task generalization.
Approach: They propose a method that retrieves promptspecific source prompt embeddings from training instances . they train soft prompt embeds for each prompt through prompt tuning and store the samples .
Outcome: The proposed method outperforms hard prompts on unseen tasks by 2.39% points and outperformed 10 out of 11 datasets.
Disentangling Questions from Query Generation for Task-Adaptive Retrieval (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing work generates synthetic queries from domain-specific documents to jointly train the retriever.
Approach: They propose a query generator that better adapts to wide search intents expressed in the BeIR benchmark.
Outcome: The proposed query generator outperforms baselines and existing models on tasks with underexplored intents while using a query generator 47 times smaller than the previous state-of-the-art.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations