Challenge: Existing methods and benchmarks for information retrieval are inadequately representing the diversity of code in various domains and tasks.
Approach: They propose a benchmark specifically designed to assess code retrieval capabilities.
Outcome: The proposed benchmark aims to invigorate research in the code retrieval domain . it shares the same data schema as other popular benchmarks like MTEB and BEIR .

Similar Papers

CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on functional relevance while neglecting code quality.
Approach: They propose a multilingual benchmark to evaluate quality-aware code retrieval . they include fine-grained quality annotations over 42,725 queries and 134,907 code snippets .
Outcome: The proposed benchmarks show that state-of-the-art models fail to separate buggy or insecure code from robust counterparts.
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval (2025.naacl-long)

Copied to clipboard

Challenge: Current information retrieval systems struggle to handle complex instructions, despite its critical importance . current models struggle to follow complex instructions in real-world applications, resulting in user-specific tasks.
Approach: They propose a benchmark to evaluate instruction-following information retrieval in expert domains.
Outcome: The proposed method improves on existing models and provides valuable insights to guide future advancements in retrieval.
CodeInsight: A Curated Dataset of Practical Coding Solutions from Stack Overflow (2024.findings-acl)

Copied to clipboard

Challenge: Comprising 3,402 crafted examples, our dataset is designed for both model finetuning and standalone evaluation.
Approach: They propose a dataset that provides examples that include a clarified intent, code snippets associated, and an average of three related unit tests.
Outcome: The proposed dataset includes 3,402 hand-written examples and 3,121 unrefined examples.
Retrieval Models Aren’t Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) suffer from inherent inabilities to interact with the physical world and access vast, up-to-date knowledge.
Approach: They propose a tool retrieval benchmark for large language models (LLMs) that includes 7.6k diverse retrieval tasks and a corpus of 43k tools.
Outcome: The proposed model performs poorly on the heterogeneous tool retrieval benchmark, resulting in low pass rate and low retrieval quality.
CodeRAG-Bench: Can Retrieval Augment Code Generation? (2025.findings-naacl)

Copied to clipboard

Challenge: Language models excel at generating code, but many programs are difficult to generate using only parametric knowledge.
Approach: They propose a retrieval-augmented code generation benchmark that provides reproducible evaluations on retrieval and end-to-end code generation performance.
Outcome: The proposed benchmark covers programming, open-domain, and repository-level tasks and provides reproducible evaluations on retrieval and end-to-end code generation performance.
MultiConIR: Towards Multi-Condition Information Retrieval (2025.findings-emnlp)

Copied to clipboard

Challenge: MultiConIR is a benchmark designed to evaluate retrieval and reranking models under nuanced multi-condition query scenarios.
Approach: They propose a benchmark to evaluate retrieval and reranking models under nuanced multi-condition query scenarios.
Outcome: The proposed benchmark evaluates retrieval and reranking models under nuanced multi-condition query scenarios across five domains.
Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for code retrieval struggle to balance scalability and annotation quality.
Approach: They propose a method that integrates functions called within the repository and information on third-party APIs to enhance the annotation context.
Outcome: The proposed method improves the annotation context by incorporating functions called within the repository and information on third-party API functionalities.
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions (2025.naacl-long)

Copied to clipboard

Challenge: Modern language models (LMs) are capable of following long and complex instructions that enable a large and diverse set of user requests.
Approach: They propose a dataset that contains an instruction evaluation benchmark and a training set to help IR models learn to follow instructions.
Outcome: The proposed model improves after fine-tuning on a training set and rigorous instruction evaluation benchmark.
DAPR: A Benchmark on Document-Aware Passage Retrieval (2024.acl-long)

Copied to clipboard

Challenge: Recent approaches to retrieve short passages are limited due to expensive operations such as selfattention.
Approach: They propose to use a task to find passages from a large corpus of documents . they extend the SoTA passage retrievers with document context to achieve this task .
Outcome: The proposed task is called Document-Aware Passage Retrieval (DAPR) it performs best on easy and hard queries, but fails on hard queries requiring document-context understanding .
Cross-Lingual Learning-to-Rank with Shared Representations (N18-2)

Copied to clipboard

Challenge: Cross-lingual information retrieval (CLIR) is a document retrieval task where the documents are written in a language different from that of the user's query.
Approach: They propose a large-scale dataset derived from Wikipedia to support CLIR research in 25 languages.
Outcome: The proposed model can improve the results of Swahili-English CLIR in Japanese and Japanese.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations