Papers with search engines

59 papers
Knowledge Distillation based Contextual Relevance Matching for E-commerce Product Search (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to e-commerce relevance matching ignore bipartite graphs in logs . experimental results show that proposed method improves human relevance judgment .
Approach: They propose an efficient knowledge distillation framework for e-commerce relevance matching to exploit the advantages of Transformer-style and classical relevance matching models.
Outcome: The proposed method significantly improves human relevance judgment on large-scale real-world data.
EVIDENCEMINER: Textual Evidence Discovery for Life Sciences (2020.acl-demos)

Copied to clipboard

Challenge: EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences.
Approach: They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Outcome: EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Multi-lingual Entity Discovery and Linking (P18-5)

Copied to clipboard

Challenge: This tutorial reviews the framework of cross-lingual EL and motivates it as a broad paradigm for the Information Extraction task.
Approach: This tutorial will review the framework of cross-lingual EL and motivate it as a broad paradigm for the Information Extraction task.
Outcome: The aim of this tutorial is to review the framework of cross-lingual EL and motivate it as a broad paradigm for the Information Extraction task.
Benchmarking LLM’s Capability in Reasoning over Conflicting Web References (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) integrated with retrieval-augmented generation (RAG) are a dominant framework for building intelligent assistants.
Approach: They propose a benchmark to evaluate LLMs' reasoning capability over real-world conflicting documents retrieved from the web.
Outcome: The proposed benchmark evaluates LLMs' reasoning capability over real-world conflicting documents retrieved from the web.
Deep Metric Learning to Hierarchically Rank - An Application in Product Retrieval (2023.emnlp-industry)

Copied to clipboard

Challenge: e-commerce search engines use customer behavior signals to augment lexical matching and improve search relevance.
Approach: They propose a method to identify duplicate and near-duplicate products across stores . they use Hierarchical Ranked Multi Similarity Loss to learn hierarchical metric space .
Outcome: The proposed model outperforms baselines in terms of catalog coverage and precision of the mappings.
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face (2023.emnlp-demo)

Copied to clipboard

Challenge: a toolkit for reproducible information retrieval research is available for free.
Approach: They present a tool that integrates Pyserini and Hugging Face to enable the seamless construction and deployment of interactive search engines.
Outcome: The proposed tool makes state-of-the-art retrieval models more accessible to non-IR practitioners while minimizing deployment effort.
Alignment Analysis of Sequential Segmentation of Lexicons to Improve Automatic Cognate Detection (P18-3)

Copied to clipboard

Challenge: Existing studies on cognate detection only distinguish between a pair of words whether they are cognates or non-cognates.
Approach: They propose to incorporate ranking functions into search engine ranking functions . they also propose to use graphical error modelling to calculate morphological shifts .
Outcome: The proposed methods give better results than competing baselines, the authors show . they show that language modelling based retrieval functions with positional tokenization and error modelling give better outcomes .
Design Challenges for a Multi-Perspective Search Engine (2022.findings-naacl)

Copied to clipboard

Challenge: a document retrieval system fails to deliver diverse and direct responses to controversial questions . classical document retrievals provide a ranked list of references to relevant but not necessarily trustworthy web documents .
Approach: They propose a perspective-oriented document retrieval paradigm to address these challenges . they propose sponses with different perspectives within topically-related web documents .
Outcome: The proposed system is based on a user survey and a prototype . it will be used to assess the utility and understanding of the system .
A New Surprise Measure for Extracting Interesting Relationships between Persons (2021.eacl-demos)

Copied to clipboard

Challenge: Interesting facts are useful information for a variety of important tasks.
Approach: They propose a method that extracts all personal relationships from dependency trees and calculates surprise scores for distributed representations of the extracted relationships in an unsupervised manner.
Outcome: The proposed method extracts all personal relationships from dependency trees for the texts and calculates surprise scores for distributed representations of the extracted relationships in an unsupervised manner.
Learning to Attend On Essential Terms: An Enhanced Retriever-Reader Model for Open-domain Question Answering (N19-1)

Copied to clipboard

Challenge: Existing approaches to open-domain question answering struggle to retrieve indirectly related evidence when no direct evidence is provided.
Approach: They propose a retriever-reader model that learns to attend on essential terms during the question answering process.
Outcome: The proposed model achieves the state-of-the-art on multiple open-domain QA datasets and achieves a 'reader-reader' level.
Embedding-based Scientific Literature Discovery in a Text Editor Application (2020.acl-demos)

Copied to clipboard

Challenge: Despite the availability of powerful search engines and text editing software, discovering relevant papers and integrating the knowledge into a manuscript remain complex tasks associated with high cognitive load.
Approach: They propose to combine text editing and literature discovery in an interactive user interface with a search engine that couples Boolean keyword filtering with nearest neighbor search over text embeddings.
Outcome: The proposed application combines text editing and literature discovery in an interactive user interface.
Credible without Credit: Domain Experts Assess Generative Language Models (2023.acl-short)

Copied to clipboard

Challenge: ChatGPT has been criticized for its lack of accuracy and coherence . authors argue that language models could replace search engines and make college essays obsolete .
Approach: a team of 10 domain experts conducts an initial assessment of language models using 100 expert-written questions.
Outcome: The results show that language models are mixed in their accuracy.
Descriptive Knowledge Graph in Biomedical Domain (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing systems that retrieve unconnected passages do not provide efficient search for relational knowledge.
Approach: They propose a system that automatically extracts and generates informative and descriptive sentences from the biomedical corpus and facilitates efficient search for relational knowledge.
Outcome: The proposed system extracts and generates informative and descriptive sentences from the biomedical corpus and facilitates the efficient search for relational knowledge.
Evaluating Credibility and Political Bias in LLMs for News Outlets in Bangladesh (2025.acl-srw)

Copied to clipboard

Challenge: Large language models (LLMs) are widely used in search engines to provide direct an-swers, while AI chatbots retrieve updated infor-mation from the web.
Approach: They audit nine Large Language Models from OpenAI, Google, and Meta to assess their ability to eval-uate the credibility and political bias of the top20 most popular news outlets in Bangladesh.
Outcome: The proposed models show internal consistency in credibil-ity ratings, but misalignment with human experts.
Event-Centric Query Expansion in Web Search (2023.acl-industry)

Copied to clipboard

Challenge: Existing studies rely on long-term search log mining to improve search experience . EQE system is a novel event retrieval framework that can select the best expansion from a significant amount of potential events quickly and accurately.
Approach: They propose a QE system that uses a four-stage event retrieval framework . they collect news headlines and then refine a dual-tower semantic model to serve as an encoder .
Outcome: The proposed system can select the best expansion from a significant amount of potential events quickly and accurately.
Learning to Rewrite Negation Queries in Product Search (2025.coling-industry)

Copied to clipboard

Challenge: Negations in product search are often used to articulate unwanted product features or components.
Approach: They propose a query rewriting approach to enhance product search performance . they use large language models to extract query reawrites from product text . their results pave the way for further research on enhancing search performance of queries with negations .
Outcome: The proposed approach improves search performance by 3.17% for queries with negations.
Seeing Things from a Different Angle:Discovering Diverse Perspectives about Claims (N19-1)

Copied to clipboard

Challenge: a number of fact checking techniques are used to identify and eliminate biases in text data.
Approach: They propose to use search engines to expand and diversify a dataset of claims, perspectives and evidence to address a selection bias.
Outcome: The proposed approach outperforms existing methods in a language understanding task.
Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study (2026.acl-short)

Copied to clipboard

Challenge: Search-engine date filters are widely used to enforce pre-cutoff retrieval in retrospective evaluations of search-augmented forecasters.
Approach: They propose stronger retrieval safeguards or evaluation on frozen, time-stamped web snapshots to prevent post-cutoff leakage.
Outcome: The proposed approach is unreliable across two major search engines, and the results are inflated.
Evaluation of Argument Search Approaches in the Context of Argumentative Dialogue Systems (2020.lrec-1)

Copied to clipboard

Challenge: Argumentative dialogue systems and chat bots require a database of arguments that matches their requirements.
Approach: They propose a dialogue system that presents arguments by virtual avatar and synthetic speech to users and allows them to rate the presented content in four different categories.
Outcome: The proposed system evaluates arguments retrieved by two state-of-the-art argument search engines and a system based on traditional web search.
AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing systems that provide a graphical representation of QAC are limited in their ability to provide real-time data.
Approach: They introduce a new QAC dataset sourced from Amazon Search logs . they assess Prefix Trees, semantic retrieval, and Large Language Models with and without finetuning .
Outcome: The proposed system can predict search terms based on user-typed prefixes . the proposed system achieves only half of what is theoretically possible on the test data .
Large Language Models Help Humans Verify Truthfulness – Except When They Are Convincingly Wrong (2024.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used for accessing information on the web.
Approach: They conduct experiments with 80 crowdworkers to compare LLMs with search engines . they ask LLM to provide contrastive information to reduce over-reliance on LLM .
Outcome: The results show that LLMs can outperform search engines but not LLM explanations . the study shows that LMS explanations are not reliable replacements for reading retrieved passages compared to search engines alone.
ITERATE: Image-Text Enhancement, Retrieval, and Alignment for Transmodal Evolution with LLMs (2025.coling-main)

Copied to clipboard

Challenge: a new framework for visual annotation of text-based questions is needed to improve performance . obtaining corresponding images through manual annotation often entails high costs .
Approach: They propose a framework that uses visual modality to enhance the performance of text-based questions.
Outcome: The proposed framework improves the alignment between text and images by using search engines or web scraping techniques.
Digital Gatekeepers: Google’s Role in Curating Hashtags and Subreddits (2025.acl-long)

Copied to clipboard

Challenge: This study examines how search engines like Google selectively promote or suppress certain hashtags and subreddits, impacting the flow of information and impacting public conversations.
Approach: They compare search engine results with nonsampled data from Reddit and Twitter/X to examine how search engines curate content through algorithmic curation.
Outcome: The proposed algorithm suppresses subreddits related to sexually explicit material, conspiracy theories, advertisements, and cryptocurrencies while promoting content associated with higher engagement.
AtTGen: Attribute Tree Generation for Real-World Attribute Joint Extraction (2023.acl-long)

Copied to clipboard

Challenge: Attribute extraction aims to identify attribute names and the corresponding attribute values from descriptive texts.
Approach: They propose a unified formulation for real-world attribute extraction application, where closed-world, open-world and semi-open attribute extraction tasks are modeled uniformly.
Outcome: The proposed model outperforms existing methods on three datasets and outperformed existing methods by a large margin.
Bootstrapping Relation Extractors using Syntactic Search by Examples (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for supervised relation extraction still require a large quantity of training data.
Approach: They propose a process for bootstrapping training datasets which can be performed quickly by non-NLP-experts.
Outcome: The proposed method outperforms models trained on manual and distant data augmentation techniques and the search-based approach with the NLG method.
MILL: Mutual Verification with Large Language Models for Zero-Shot Query Expansion (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for query expansion lack corpus-specific knowledge and cost.
Approach: They propose a query-query-document generation method that leverages large language models for mutual verification to produce diverse sub-queries and corresponding documents.
Outcome: The proposed method is fully zero-shot and extensive experiments on three public benchmark datasets demonstrate its effectiveness over existing methods.
Translating Web Search Queries into Natural Language Questions (L18-1)

Copied to clipboard

Challenge: a new method to generate natural language questions from keyword-based queries is proposed . a synergy between query-to-question problem and standard machine translation (MT) model is found .
Approach: They propose a method to generate well-formed natural language questions from keyword-based queries.
Outcome: The proposed method is well-formed natural language question generated from keyword-based query.
Effective and Efficient Query-aware Snippet Extraction for Web Search (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to extract webpage snippets ignore contextual information of webpages, which may be sub-optimal.
Approach: They propose a query-aware webpage snippet extraction method called DeepQSE that captures contextual information of webpages.
Outcome: The proposed method can significantly improve the performance of DeepQSE without affecting its performance.
Chat-Ghosting: Methods for Auto-Completion in Dialog Systems (2026.eacl-long)

Copied to clipboard

Challenge: Ghosting is a type-ahead completion task that predicts a user's intended input for inline query auto-completion (QAC).
Approach: They propose to use ghosting to predict a user's intended input for inline query auto-completion by suggesting completions to incomplete queries.
Outcome: The proposed method outperforms deep learning and deep learning methods with and without dialog context for ghosting.
Towards Improved Multi-Source Attribution for Long-Form Answer Generation (2024.naacl-long)

Copied to clipboard

Challenge: Current LLMs struggle with attribution for long-form answers which require reasoning over multiple evidence sources.
Approach: They propose to improve attribution capability of large language models for long-form answer generation to multiple sources with multiple citations per sentence.
Outcome: The proposed model improves on a wide range of attribution benchmark datasets on PolitiICite, a multi-source attribution dataset based on PolitIcite articles .
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that LLMs are vulnerable to prompt injection attacks because of their instruction-following abilities and inability to distinguish the instructions in the data content.
Approach: They propose backdoor-powered prompt injection attacks that trick LLMs into deviating from the original input instruction and executing the attackers’ target instruction.
Outcome: The proposed attacks trick the LLMs into deviating from the input instruction and executing the attackers’ target instruction.
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to determine the knowledge an LLM already possesses and the knowledge that requires the help of a search engine are expensive and require excessive computational costs.
Approach: They propose a slim proxy model that detects missing knowledge in LLMs with a proxy model and use it to perform retrieval for the missing knowledge.
Outcome: The proposed approach detects missing knowledge in LLMs with a slim proxy model and takes its answers as heuristic answers.
CausalQA: A Benchmark for Causal Question Answering (2022.coling-1)

Copied to clipboard

Challenge: Existing causal question answering datasets are relatively small and only include one type of causal question.
Approach: They construct a benchmark corpus of 1.1 million causal questions with answers . they use a typology derived from a data-driven, manual analysis of QA datasets .
Outcome: The proposed model achieves a ROUGE-L F1 score of 0.48 on the new QA benchmark.
Aligning Black-box Language Models with Human Judgments (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used as automated judges to evaluate recommendation systems, search engines, and other subjective tasks.
Approach: They propose a framework to align LLM judgments with individual human evaluators or their aggregated judgments without retraining or fine-tuning the LLM.
Outcome: The proposed framework achieves 142% improvement in agreement across 29 tasks and exceeds inter-human agreement on four out of six tasks.
Hierarchical Trivia Fact Extraction from Wikipedia Articles (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for extracting trivia facts for Wikipedia categories are not efficient . a trivia fact is an interesting fact that is unusual, unexpected, or unique .
Approach: They propose an unsupervised algorithm that automatically mines trivia facts for a given entity . they propose to target at a single Wikipedia article and leverage its hierarchical structure .
Outcome: The proposed algorithm outperforms existing methods and is 100 times faster than existing methods.
Evaluating the Knowledge Base Completion Potential of GPT (2023.findings-emnlp)

Copied to clipboard

Challenge: Language models (LMs) have been proposed for unsupervised knowledge base completion (KBC) however, their ability to do this at scale and with high accuracy remains an open question.
Approach: They propose to use language models to complete a large public KB, Wikidata, with 90% precision.
Outcome: The proposed models can extend Wikidata by 27M facts at 90% precision.
A Semantic Search Engine for Mathlib4 (2024.findings-emnlp)

Copied to clipboard

Challenge: Lean is an interactive theorem prover that enables verification of formal proofs . however, searching for theoretical proofs in mathlib4 can be challenging for beginners . we present a semantic search engine that accepts informal queries and finds theorels .
Approach: They propose a semantic search engine for theorems in mathlib4 that accepts informal queries and finds relevant theorels.
Outcome: The proposed search engine accepts informal queries and finds theorems . it compares with other search engines that struggle to find theoretical proofs based on informal queries .
DeepREF: A Framework for Optimized Deep Learning-based Relation Classification (2022.lrec-1)

Copied to clipboard

Challenge: Existing frameworks for relation extraction (RE) are limited due to lack of implementation details.
Approach: They propose to use deep learning to develop relation extraction systems using deep learning models.
Outcome: The proposed framework is inspired by the OpenNRE and REflex existing frameworks.
WebWalker: Benchmarking LLMs in Web Traversal (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of natural language processing tasks.
Approach: They propose a benchmark to assess the ability of LLMs to perform web traversal by using an explore-critic paradigm.
Outcome: The proposed framework mimics human-like web navigation through an explore-critic paradigm and demonstrates the effectiveness of RAG combined with WebWalker in real-world scenarios.
Enhanced Facet Generation with LLM Editing (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have shown that search engines can recognize facets of a user's query.
Approach: They propose to use large language models to enhance the facets of a query to generate facets from a search engine.
Outcome: The proposed model can predict facets by taking only queries as input without a search engine.
CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic (2026.findings-acl)

Copied to clipboard

Challenge: Existing search agent pipelines rely on sparse outcome rewards, leading to inefficient exploration and unstable training.
Approach: They propose a tool-integrated reasoning framework that provides turn-level feedback via a retrospective critic mechanism.
Outcome: The proposed framework outperforms baselines in multi-hop reasoning benchmarks and achieves faster convergence and training stability.
DORA: A Dual-Objective Reinforcement Learning Framework for Effective and Efficient Multimodal Agentic Search (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to train large language models overlook quality of intermediate search results . existing methods often invoke search calls during reasoning, making inference inefficient .
Approach: They propose a dual-objective reinforcement learning framework to improve search strategies of MLLMs . DORA outperforms state-of-the-art methods, achieving up to 8.4% higher accuracy .
Outcome: The proposed model outperforms state-of-the-art methods while reducing search calls by 9.7%.
Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger (2025.acl-long)

Copied to clipboard

Challenge: Existing research expands the tool arrays of large language models (LLMs), but the necessity of using these tools is often overlooked, leading to indiscriminate tool invocation.
Approach: They propose a meta-cognition proxy proxy for LLMs self-assessment of their capabilities, reflecting the model’s awareness of its own limitations.
Outcome: The proposed strategy is fine-tuned-free and costs minimal.
EvolveSearch: An Iterative Self-Evolving Search Agent (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to enabling LLM web search proficiency struggle with data production in open-search domains, while supervised fine-tuning struggles with data utilization efficiency.
Approach: They propose an iterative self-evolution framework that combines SFT and RL to enhance agentic web search capabilities without external human-annotated reasoning data.
Outcome: EvolveSearch achieves 4.7% improvement over current state-of-the-art in seven benchmarks . supervised fine-tuning struggles with data production in open-search domains compared with RL .
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) drive scientific question-answering on search engines, yet their evaluation robustness remains underexplored.
Approach: They propose an open-source framework that combines rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators.
Outcome: The proposed framework combines fine-grained rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators.
CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for conversational question answering require specific retrievers to understand user questions.
Approach: They develop a query rewriting model CONQRR that rewrites a conversational question into a standalone question.
Outcome: The proposed model achieves state-of-the-art on an open-domain conversational question answering dataset and is effective for two different off-the shelf retrievers.
Numbers Matter! Bringing Quantity-awareness to Retrieval Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Quantitative information is important for understanding documents and interpreting them.
Approach: They propose two quantity-aware ranking techniques that rank both quantity and textual content . they use available retrieval systems to incorporate quantity information into queries .
Outcome: The proposed methods can rank both quantity and textual content, either jointly or independently.
Effective Contrastive Weighting for Dense Query Expansion (2023.acl-long)

Copied to clipboard

Challenge: Verbatim queries that do not adequately express the user's search intent are often lexical inadequacies.
Approach: They propose a contrastive weighting model that learns to select the most useful expansion embeddings for semantic search.
Outcome: The proposed model outperforms existing methods while maintaining its efficiency.
Efficient Document Embeddings via Self-Contrastive Bregman Divergence Learning (2023.findings-acl)

Copied to clipboard

Challenge: Despite recent advances in transformer-based sentence encoders, the encoding of long documents (Ks of words) is still challenging with respect to both efficiency and quality considerations.
Approach: They propose to combine a self-contrastive siamese network and a convex neural Bregman divergence network to train longfomer-based document encoders using an unsupervised contrastive learning method.
Outcome: The proposed model outperforms baseline models on three long document topic classification tasks from the legal and biomedical domains.
Triplet-Free Knowledge-Guided Response Generation (2023.findings-acl)

Copied to clipboard

Challenge: Prior work focused on constructing ”latent” knowledge and learning how to ground it based on pseudo triplets.
Approach: They propose to pretrain a response language model to measure relevance and consistency between any context and response and use search engines to collect top-ranked passages to serve as guiding knowledge without explicitly optimizing the ‘‘best’ latent knowledge.
Outcome: The proposed model pretrains a response language model to measure relevance and consistency between any context and response, then uses search engines to collect the top-ranked passages to serve as the guiding knowledge without explicitly optimizing the ‘‘best’ latent knowledge.
LitSearch: A Retrieval Benchmark for Scientific Literature Search (2024.emnlp-main)

Copied to clipboard

Challenge: Literature search questions pose significant challenges for modern retrieval systems . a lack of domain expertise and reasoning through lengthy papers is a challenge .
Approach: They propose a retrieval benchmark for literature search queries using inline citations from papers and questions about recently published papers.
Outcome: The proposed retrieval benchmarks outperform state-of-the-art retrieval models and reranking pipelines.
KidSpell: A Child-Oriented, Rule-Based, Phonetic Spellchecker (2020.lrec-1)

Copied to clipboard

Challenge: Existing spellcheckers are tuned to the needs of adults and are unsatisfactory for children due to their varying cognitive capabilities.
Approach: They propose a model that maps misspelled words and spelling suggestions to their phonetic keys and a selection process that prioritizes candidate spelling suggestions that closely align with the misspelled word.
Outcome: The proposed model outperforms existing spellcheckers in a number of offline experiments using existing and novel datasets.
Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citations (2025.emnlp-main)

Copied to clipboard

Challenge: generative search engines rely on in-line citations as the key gateway to original webpages . a recent study shows that LLMs tend to cite left-leaning sources at higher rates compared to traditional retrieval systems .
Approach: They construct a dataset of news articles labeled with left- or right-leaning stances . they find that LLMs tend to cite left-leansing sources at higher rates than traditional retrieval systems .
Outcome: The proposed dataset shows that LLMs tend to cite left-leaning sources at higher rates than traditional retrieval systems.
Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents (2023.emnlp-main)

Copied to clipboard

Challenge: Existing work utilizes generative LLMs for Information Retrieval (IR) rather than direct passage ranking.
Approach: They investigate generative LLMs such as ChatGPT and GPT-4 for relevance ranking in IR and use a test set to verify the model’s ability to rank unknown knowledge.
Outcome: The proposed model outperforms a 3B supervised model on the BEIR benchmark.
A Comprehensive Evaluation of Tool-Assisted Generation Strategies (2023.findings-emnlp)

Copied to clipboard

Challenge: Various few-shot tool-usage strategies have been proposed to overcome LMs' shortcomings.
Approach: They propose to augment language models with tools to overcome their shortcomings . they find strong no-tool baselines are competitive to tool-assisted strategies .
Outcome: The proposed strategies outperform those that refine incorrect outputs with tools in knowledge-retrieval tasks, the study finds . the findings suggest few-shot tool integration is still an open challenge .
QuBE: Question-based Belief Enhancement for Agentic LLM Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have led to an explosion of interest in their deployment as agents.
Approach: They propose a method that enhances agents’ focus on task-relevant contexts by constructing a belief state via question answering.
Outcome: The proposed method outperforms established baselines and achieves marked improvements on the BeIR zero-shot retrieval benchmark.
Speculating LLMs’ Chinese Training Data Pollution from Their Tokens (2025.emnlp-main)

Copied to clipboard

Challenge: Experiments on GPT and other 23 LLMs indicate that tokens widely exist while GPT’s vocabulary behaves the worst: more than 23% long Chinese tokens (i.e., a token with more than two Chinese characters) are either porn or online gambling.
Approach: They propose to locate Polluted Chinese (PoC) tokens in LLMs and build a PoC token detector to label them in vocabularies by considering each token’s semantics and related contents from the search engines.
Outcome: The proposed method predicts that the ratio of “*” related webpages in GPT-4o's training data is around 0.5%.
Breaking the Evaluation Paradox: Evaluating High-Entropy Search with Computationally Irreducible Constraints (2026.findings-acl)

Copied to clipboard

Challenge: a new framework for evaluation of exhaustive search capabilities is needed . high-entropy enumeration tasks make such ground truth impossible for humans to create . VERITAS is a framework built on the principle of computationally irreducible constraints .
Approach: They propose a framework that uses non-optimizable constraints to create verifiable searches . VERITAS can generate infinite number of test cases with perfect ground truth and precise difficulty control .
Outcome: a new evaluation framework for large language models is based on non-optimizable constraints . the framework can generate infinite number of test cases with perfect ground truth and precise difficulty control .
Is Your Language Model Ready for Monetization Decisions? (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on shopping-centric scenarios and user-facing data, overlooking intermediate decision stages and robustness considerations.
Approach: They propose a multi-task benchmark to evaluate large language models in real-world monetization contexts.
Outcome: The proposed benchmark covers intent understanding, commercial matching, and user behavior modeling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations