Papers with ranking

90 papers
Pretrained Transformers for Text Ranking: BERT and Beyond (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial provides an overview of text ranking using neural network architectures known as transformers.
Approach: This tutorial provides an overview of text ranking with neural network architectures known as transformers.
Outcome: This tutorial provides an overview of text ranking with neural network architectures known as transformers.
Extracting Text Representations for Terms and Phrases in Technical Domains (2023.acl-industry)

Copied to clipboard

Challenge: Large pre-trained language models are extensively used in modern NLP systems.
Approach: They propose an unsupervised approach to encoding using character-based models and pre-trained sentence encoders to reconstruct large pre-trained embedding matrices.
Outcome: The proposed approach matches the quality of sentence encoders in technical domains and is 5 times smaller and up to 10 times faster on high-end GPUs.
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Recent advances in large language models have enabled impressive zero-shot capabilities across various natural language tasks.
Approach: They propose two ways to exploit the emergent abilities of large language models for NLG assessment.
Outcome: The proposed methods improve performance and positional biases in comparisons between candidates.
Generative Product Recommendations for Implicit Superlative Queries (2025.naacl-srw)

Copied to clipboard

Challenge: Existing retrieval and ranking systems struggle with implicit superlative queries . lack of explicit attribute mentions and complexity of the query complicates ranking .
Approach: They propose a four-point schema for annotating the best product candidates for superlative queries . they propose pointwise, deliberated pointwise and pairwise methods to analyze the results .
Outcome: The proposed schema can be used to rank products with implicit attributes and reason over them.
Batch-Softmax Contrastive Loss for Pairwise Sentence Scoring Tasks (2022.naacl-main)

Copied to clipboard

Challenge: Recent advances in machine learning have led to the use of contrastive loss for representation learning.
Approach: They propose to use batch-softmax contrastive loss to train pairwise sentence embeddings . they propose to take a batch-softermax contrastitive loss and train it with different loss functions .
Outcome: The proposed model improves on a number of datasets and pairwise sentence scoring tasks.
RankME: Reliable Human Ratings for Natural Language Generation (N18-2)

Copied to clipboard

Challenge: Existing studies have shown that human evaluation for natural language generation often suffers from inconsistent user ratings.
Approach: They propose a rank-based magnitude estimation method which combines continuous scales and relative assessments to improve the reliability of human ratings.
Outcome: The proposed method significantly improves the reliability and consistency of human ratings compared to traditional evaluation methods.
InsBank: Evolving Instruction Subset for Ongoing Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies emphasize that quality and diversity of instruction data are more crucial than quantity, highlighting the need to select diverse, high-quality subsets to reduce training costs.
Approach: They propose to use a continuously updated repository to integrate the latest valuable instruction data with a progressive evolution framework to evolve InsBank over time.
Outcome: The proposed framework outperforms baselines in InsBank evolution and extracts budget-specific subsets.
Unsupervised Keyphrase Extraction by Jointly Modeling Local and Global Context (2021.emnlp-main)

Copied to clipboard

Challenge: Embedding based methods are widely used for unsupervised keyphrase extraction tasks.
Approach: They propose a method where local and global contexts are jointly modeled.
Outcome: The proposed method outperforms most models while generalizing better on input documents with different domains and length.
Evaluating Research Novelty Detection: Counterfactual Approaches (D19-53)

Copied to clipboard

Challenge: Despite its importance, this direction of research has not been explored as much.
Approach: They propose to use counterfactual simulations to evaluate paper novelty detection models . they ask models to differentiate papers at time t and counterf actual paper from future time .
Outcome: The proposed models can be compared against a set of papers with a given date and with different annotations.
Linguistically Informed Relation Extraction and Neural Architectures for Nested Named Entity Recognition in BioNLP-OST 2019 (D19-57)

Copied to clipboard

Challenge: Named Entity Recognition (NER) and Relation Extraction (RE) are essential tools in distilling knowledge from biomedical literature.
Approach: They propose to use Named Entities to perform nested entities extraction, Entity Normalization and Relation Extraction to generalize the approach to different languages.
Outcome: The proposed approach can be generalized to different languages and showed it’s effectiveness for English and Spanish text.
Tracking Brand-Associated Polarity-Bearing Topics in User Reviews (2023.tacl-1)

Copied to clipboard

Challenge: Existing models that infer brand polarity scores from reviews are not able to infer polarities directly.
Approach: They propose a dynamic Brand-Topic Model which detects and tracks brand-associated sentiment scores and polarity-bearing topics from product reviews organized in temporally ordered time intervals.
Outcome: The proposed model outperforms competitive models on a MakeupAlley and hotel review datasets.
Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to generate relevance labels for large language models have not been successful in generating relevance labels.
Approach: They propose a method to combine LLM relevance labels with ranking abilities . they take both LLM generated relevance labels and pairwise preferences .
Outcome: The proposed method balances the ranking and labeling abilities of large language models . it takes both LLM generated relevance labels and pairwise preferences .
Pre-Deployment Advertisement Ranking under Data Scarcity via Context-Aware Criteria Generation with VLMs (2026.acl-industry)

Copied to clipboard

Challenge: Existing VLMs perform well on general multimodal tasks, but limited labeled data makes them difficult to apply to real-world business decisions.
Approach: They propose a new task that aims to rank ads for a target brand prior to deployment . they propose 'brand-specific ad ranking' which uses brand-specific effectiveness .
Outcome: The proposed task outperforms baselines on 10 brands on real-world advertising data.
Leveraging Contextual Information for Effective Entity Salience Detection (2024.findings-naacl)

Copied to clipboard

Challenge: Prior work on salient entity detection focused on machine learning models that require heavy feature engineering.
Approach: They propose to fine-tune medium-sized language models with a cross-encoder style architecture to achieve significant performance gains over feature engineering approaches.
Outcome: The proposed model fine-tunes medium-sized pre-trained language models with a cross-encoder style architecture yields substantial performance gains over feature engineering approaches.
Beyond Yes and No: Improving Zero-Shot LLM Rankers via Scoring Fine-Grained Relevance Labels (2024.naacl-short)

Copied to clipboard

Challenge: Existing pointwise LLMs provide noisy or biased answers for documents that are partially relevant to the query.
Approach: They propose to incorporate fine-grained relevance labels into the LLM prompt . they propose to better differentiate between documents with different levels of relevance .
Outcome: The proposed model can differentiate between documents with different levels of relevance to the query and derive a more accurate ranking.
Identifying High Consideration E-Commerce Search Queries (2024.emnlp-industry)

Copied to clipboard

Challenge: Identifying high consideration queries is essential for e-commerce sites to better serve user needs . ecommerce sites can create or serve customized content for specific queries .
Approach: They propose an engagement-based Query Ranking approach to identify potential engagement levels with query-related shopping knowledge content during product search.
Outcome: The proposed method outperforms human-selected queries in terms of customer impact . human evaluation shows a precision of 96% for HC queries identified by the model .
Balancing Lexical and Semantic Quality in Abstractive Summarization (2023.acl-short)

Copied to clipboard

Challenge: Existing methods to reduce exposure bias in sequence-to-sequence models are underexplored.
Approach: They propose a method to re-rank sequence-to-sequence neural models to reduce exposure bias.
Outcome: The proposed method achieves an 89.67 BERTScore on the CNN/DailyMail and XSum datasets.
Document Ranking with a Pretrained Sequence-to-Sequence Model (2020.findings-emnlp)

Copied to clipboard

Challenge: Experimental results on the MS MARCO passage ranking task show that our ranking approach is superior to strong encoder-only models.
Approach: They propose to use a pretrained sequence-to-sequence model to generate relevance labels as "target tokens" they also show how the underlying logits of these target tokens can be interpreted as relevance probabilities for ranking.
Outcome: The proposed model outperforms existing models in a data-poor setting and significantly outperformed an encoder-only model on the MS MARCO passage ranking task.
Aligning Large Language Models with Recommendation Knowledge (2024.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) excel at natural language reasoning, but cannot model complex user-item interactions inherent in recommendation tasks.
Approach: They propose to equip large language models with recommendation-specific knowledge to address this gap by combining Masked Item Modeling and Bayesian Personalized Ranking (BPR) auxiliary task data samples are generated that encode item correlations and user preferences.
Outcome: Experiments on Amazon Toys & Games, Beauty, and Sports & Outdoors show that the proposed method outperforms conventional and LLM-based baselines by significant margins in retrieval.
A Dynamic Self-Evolving Extraction System (2026.acl-demo)

Copied to clipboard

Challenge: High-quality information extractions often require domain-specific accuracy, up-to-date understanding of specialized taxonomies, and the ability to incorporate emerging jargon and rare outliers.
Approach: They propose a Dynamic Self-Evolving Extraction and Curation Toolkit which continuously improves as it is used to extract structured information from raw text.
Outcome: The proposed toolkit continuously improves as it is used in medical, legal, and HR domains.
STAR: SQL Guided Pre-Training for Context-dependent Text-to-SQL Parsing (2022.findings-emnlp)

Copied to clipboard

Challenge: Extensive experiments show that STAR outperforms previous pre-training methods and ranks first on the leaderboard . text-to-SQL parsing aims to translate natural language (NL) questions into executable SQL queries .
Approach: They propose a SQL guided pre-training framework STAR for context-dependent text-to-SQL parsing . they propose two objectives that explore context-dependence of NL utterances and SQL queries .
Outcome: The proposed framework outperforms existing methods on two downstream benchmarks and ranks first on the leaderboard.
FAA: Fine-grained Attention Alignment for Cascade Document Ranking (2023.acl-long)

Copied to clipboard

Challenge: Contemporary document ranking methods focus on transforming documents into passages to handle long inputs, but intensive query-irrelevant content may lead to harmful distraction and high query latency.
Approach: They propose a fine-grained attention alignment approach to jointly optimize a cascade document ranking model.
Outcome: Experiments on MS MARCO and TREC DL show that the proposed method is effective in document ranking tasks.
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods to rank documents using large language models do not understand these challenging ranking formulations.
Approach: They propose to use Pairwise Ranking Prompting to improve ranking performance . they propose to outperform fine-tuned baseline rankers on benchmark datasets .
Outcome: The proposed technique outperforms supervised baselines on benchmark datasets and outperformed other LLM-based solutions by over 10% on average.
Key2Vec: Automatic Ranked Keyphrase Extraction from Scientific Articles using Phrase Embeddings (N18-2)

Copied to clipboard

Challenge: Keyphrase extraction is a fundamental task in natural language processing that facilitates mapping of documents to a set of representative phrases.
Approach: They propose an unsupervised technique that leverages phrase embeddings for ranking keyphrases extracted from scientific articles using theme-weighted PageRank.
Outcome: The proposed method performs better on benchmark datasets than other methods and is of high quality.
Group, Embed and Reason: A Hybrid LLM and Embedding Framework for Semantic Attribute Alignment (2025.emnlp-industry)

Copied to clipboard

Challenge: a framework to align attributes that refer to the same concept but differ across schemas is challenging in schema only settings where no instance data is available due to ambiguous names, inconsistent descriptions, and domain-specific terminologies.
Approach: They propose a framework that combines contextual reasoning and embedding-based similarity to address token limitations and hallucinations.
Outcome: The proposed framework scales to large schemas and shows strong performance on healthcare schemas.
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm . traditional methods of assessment and evaluation fail in dynamic and open-ended scenarios .
Approach: They propose a paradigm where LLMs are leveraged to perform scoring, ranking, or selection for machine learning evaluation scenarios.
Outcome: The proposed model-based judgment and evaluation paradigms are based on large language models and are compared to the current model-driven evaluation paradigm.
An Element is Worth a Thousand Words: Enhancing Legal Case Retrieval by Incorporating Legal Elements (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for legal case retrieval lack the definition of relevance for legal cases . however, the definition goes beyond the common semantic relevance of ad-hoc retrieval.
Approach: They propose a legal element dataset that incorporates legal elements into a semi-automatic method . they propose two models to enhance legal search using legal elements .
Outcome: The proposed models outperform existing methods in enhancing legal search using legal elements.
Synthesizing Human Gaze Feedback for Improved NLP Performance (2023.eacl-main)

Copied to clipboard

Challenge: Prior work on eye tracking and NLP reveals that human scanpaths can aid in understanding and performance of NLP models.
Approach: They propose a model for generating human scanpaths over text that approximates meaningful cognitive signals in human gaze patterns.
Outcome: The proposed model can approximate meaningful cognitive signals in human gaze patterns.
Learning to Retrieve Engaging Follow-Up Queries (2023.findings-eacl)

Copied to clipboard

Challenge: Open domain conversational agents can answer a wide range of targeted queries, but knowledge exploration is a lengthy task.
Approach: They propose a retrieval based system for predicting the next questions that the user might have . they train ranking models on a dataset called the Follow-up Query Bank .
Outcome: The proposed system can proactively assist users in knowledge exploration leading to a more engaging dialog.
Learning Invariant Representations of Social Media Users (D19-1)

Copied to clipboard

Challenge: Existing methods for learning to compare social media users fail to generalize to new users or even to previously known users.
Approach: They propose a procedure to learn a mapping from short episodes of user activity to a vector space in which the distance between points captures the similarity of the corresponding users’ invariant features.
Outcome: The proposed procedure can be applied to users not seen at training time and enables efficient comparisons of users in the resulting vector space.
FinMTEB: Finance Massive Text Embedding Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text embedding benchmarks for financial domains are inadequately addressing the nuanced requirements of specialized domains like finance.
Approach: They propose a finance-adapted embedding model that outperforms general-purpose models . they also introduce a new model, Fin-E5, which is also open-sourced .
Outcome: The proposed framework outperforms general-purpose models on financial embedding tasks.
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: coding tasks require generated code to be fully executable and functionally correct . current agentic approaches struggle with multi-stage planning, generating, and debugging .
Approach: They propose a framework for LLM agents to efficiently explore the search space in different stages of the code generation process.
Outcome: The proposed framework achieves top results on 7 code generation benchmarks and a 31.9% solving rate on the SWEBench benchmark.
Generate & Rank: A Multi-task Framework for Math Word Problems (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies formalize MWP as a generation task but mathematical expressions are prone to minor mistakes.
Approach: They propose a ranking task for math word problem (MWP) that learns from its own mistakes and distinguishes between correct and incorrect expressions.
Outcome: The proposed model outperforms baselines on the classical Math23k dataset and is 7% higher than the state-of-the-art.
Adversarial Learning of Poisson Factorisation Model for Gauging Brand Sentiment in User Reviews (2021.eacl-main)

Copied to clipboard

Challenge: Existing models for sentiment-topic extraction assume topics are grouped under discrete sentiment categories such as ‘positive’, ‘negative’ and ‘neural’.
Approach: They propose a Brand-Topic Model which aims to detect brand-associated polarity-bearing topics from product reviews.
Outcome: The proposed model outperforms existing models on Amazon reviews and shows that it is more coherent and unique than existing models.
Entity-Duet Neural Ranking: Understanding the Role of Knowledge Graph Semantics in Neural Information Retrieval (P18-1)

Copied to clipboard

Challenge: Entity-oriented search and neural-IR push the boundary of search engines from two different aspects.
Approach: They propose an Entity-Duet Neural Ranking Model which integrates knowledge graphs into neural search systems.
Outcome: The proposed model improves generalization ability of neural ranking models on a commercial search log.
Dealing with Typos for BERT-based Passage Retrieval and Ranking (2021.emnlp-main)

Copied to clipboard

Challenge: Current approaches to passage retrieval and ranking rely on pre-trained deep language models that model the semantic matching between queries and passages.
Approach: They propose a typos-aware training framework for DR and BERT to address this issue.
Outcome: The proposed models respond and adapt to keyword typos occurring in queries, and significantly improve their retrieval and ranking effectiveness.
Two Birds with One Stone: Unified Model Learning for Both Recall and Ranking in News Recommendation (2022.findings-acl)

Copied to clipboard

Challenge: Existing news recommender systems conduct news recall and ranking separately with different models, but maintaining multiple models leads to high computational cost and high latency.
Approach: They propose a unified method for recall and ranking in news recommendation that uses historical news click behaviors to extract user embeddings for ranking from the user's attention query.
Outcome: The proposed method improves recall and ranking efficiency and effectiveness on a benchmark dataset.
Finding Replicable Human Evaluations via Stable Ranking Probability (2024.naacl-long)

Copied to clipboard

Challenge: a recent study shows that human evaluation is the best way to rank natural language generation systems . human raters can exhibit different behaviors when rating outputs, causing ranking to be unstable . stability is the degree to which a specific evaluation methodology produces the same system ranking when repeated.
Approach: They propose to evaluate results through the lens of stability: stability is the degree to which a specific evaluation methodology produces the same system ranking when repeated.
Outcome: The proposed model is based on a dataset of multi-segment translations rated by multiple professionals . human raters can exhibit different behaviors when rating NLG outputs, the study shows .
Effective Inter-Clause Modeling for End-to-End Emotion-Cause Pair Extraction (2020.acl-main)

Copied to clipboard

Challenge: Emotion-cause pair extraction aims to extract all emotion clauses coupled with their cause clauses from a given document.
Approach: They propose a one-step neural approach which emphasizes inter-clause modeling to perform end-to-end extraction.
Outcome: The proposed method outperforms existing methods in the extraction of emotion-cause pairs . it emphasizes inter-clause modeling to perform end-to-end extraction .
Unified Language Representation for Question Answering over Text, Tables, and Images (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to answer complex questions are limited to text or structured data.
Approach: They propose a paradigm that transforms images and tables into unified language representations to simplify QA problems.
Outcome: The proposed framework outperforms existing methods on two datasets and the WebQA leaderboard.
EDIS: Entity-Driven Image Search over Multimodal Web Content (2023.emnlp-main)

Copied to clipboard

Challenge: Existing image retrieval methods require large datasets and a large candidate set.
Approach: They propose a news-domain dataset for cross-modal image search with 1 million web images . they propose combining multimodal image-text pairs with a million candidates .
Outcome: The proposed dataset challenges state-of-the-art methods with dense entities and the large-scale candidate set.
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Large reasoning models exhibit human-like behaviors such as exploration, verification, reflection, and correction.
Approach: They propose a supervised fine-tuning framework for long chain-of-thoughts reasoning . they leverage a difficulty-aware reward model to estimate the learning value of questions .
Outcome: The proposed framework performs fine-tuning on large reasoning models on 10% of the data selected.
SciRepEval: A Multi-Format Benchmark for Scientific Document Representations (2023.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for evaluating scientific document representations fail to capture the diversity of relevant tasks.
Approach: They propose a benchmark for training and evaluating scientific document representations that includes 24 challenging and realistic tasks across four formats: classification, regression, ranking and search.
Outcome: The proposed model outperforms existing models by over 2 points absolute.
Synthesizing Adversarial Negative Responses for Robust Response Ranking and Evaluation (2021.findings-acl)

Copied to clipboard

Challenge: Open-domain neural dialogue models have achieved high performance in response ranking and evaluation tasks.
Approach: They propose methods for automatically creating adversarial negative training data . they use mask-and-fill and keyword-guided approaches to generate negative examples .
Outcome: The proposed approaches outperform baseline models in providing informative negative examples for training dialogue systems.
Modularized Transfomer-based Ranking Framework (2020.emnlp-main)

Copied to clipboard

Challenge: Recent innovations in Transformer-based ranking models have advanced the state-of-the-art in information retrieval.
Approach: They propose to modularize a Transformer ranker into separate modules for text representation and interaction.
Outcome: The proposed model is faster than previous models and is easier to interpret and understand.
LOPS: Learning Order Inspired Pseudo-Label Selection for Weakly Supervised Text Classification (2022.findings-emnlp)

Copied to clipboard

Challenge: Weakly-supervised text classification methods are noisy due to their heuristic nature . selection of correct pseudo-labels has a huge potential for performance boost .
Approach: They propose a pseudo-label selection method that takes learning order into account . they propose to select samples that are learnt earlier based on their pseudo-labels .
Outcome: The proposed method is ineffective and unstable due to erroneous predictions from poorly calibrated models.
Dimension Reduction for Efficient Dense Retrieval via Conditional Autoencoder (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work reserves the principle dimensions of query and document embeddings for building more efficient retrieval systems.
Approach: They propose to use Conditional Autoencoder to compress high-dimensional embeddings to maintain the same embeddable distribution and better recover ranking features.
Outcome: The proposed algorithm achieves comparable ranking performance with its teacher model and makes the retrieval system more efficient.
Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval (2021.acl-long)

Copied to clipboard

Challenge: Existing retrieval models based on dense representations show better performance than sparse representations.
Approach: They propose a method to mimic the queries to each of the documents by an iterative clustering process and represent the documents using multiple pseudo queries.
Outcome: The proposed model achieves state-of-the-art results on a large dataset while remaining high efficiency.
Long Document Ranking with Query-Directed Sparse Transformer (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to document ranking require long documents to be broken to fit in pretrained models.
Approach: They propose a Query-Directed Sparse attention model that induces IR-axiomatic structures in transformer self-attention.
Outcome: The proposed model enforces the principle properties desired in ranking while also enjoying efficiency from sparsity.
DUQGen: Effective Unsupervised Domain Adaptation of Neural Rankers by Diversifying Synthetic Query Generation (2024.naacl-long)

Copied to clipboard

Challenge: State-of-the-art rankers pre-trained on large task-specific training data such as MS-MARCO exhibit strong performance on various ranking tasks without domain adaptation, also called zero-shot.
Approach: They propose a method to generate unsupervised domain adaptation for ranking using large-scale task-specific training data such as MS-MARCO and Wikipedia retrieval.
Outcome: The proposed method outperforms all zero-shot baselines and significantly outperfies the SOTA baselines on 16 out of 18 datasets, for an average of 4% relative improvement across all datasets.
A Human Evaluation of AMR-to-English Generation Systems (2020.coling-main)

Copied to clipboard

Challenge: a recent human evaluation of AMR generation systems is compared to automated metrics.
Approach: They propose a human evaluation which collects fluency and adequacy scores and categorization of error types for AMR generation systems.
Outcome: The results show that human evaluations are more nuanced than automated metrics.
MMQA: A Multi-domain Multi-lingual Question-Answering Framework for English and Hindi (L18-1)

Copied to clipboard

Challenge: Existing work on multi-domain, multi-lingual question answering is limited to the same language.
Approach: They curate 500 articles in six different domains from the web and create question-answer pairs . they develop a deep learning based model for classifying an input question into coarse and finer categories .
Outcome: The proposed model accuracies 90.12% and 80.30% for coarse and finer classes . the proposed model is the first attempt to create multi-domain, multi-lingual question answering evaluation involving English and Hindi.
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects (2024.naacl-long)

Copied to clipboard

Challenge: X-Eval is a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users.
Approach: They introduce a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users.
Outcome: The proposed framework improves the model’s ability to follow evaluation instructions and enhances the learning stage to better assess text quality.
Generating Query Focused Summaries from Query-Free Resources (2021.acl-long)

Copied to clipboard

Challenge: Existing datasets are small for data-hungry neural architectures and are limited to evaluation purposes.
Approach: They propose to decompose QFS into query modeling and conditional language modeling . they propose a Masked ROUGE Regression framework for evidence estimation and ranking .
Outcome: The proposed model achieves state-of-the-art performance despite weak supervision.
Evaluating Scholarly Impact: Towards Content-Aware Bibliometrics (2021.emnlp-main)

Copied to clipboard

Challenge: Scientific, engineering, and technological (SET) innovations drive many positive advances in our modern economy, society, and life.
Approach: They propose a new metric that uses the content of the paper as a source of distant-supervision to quantify how much the cited-node informs the citing-n node.
Outcome: The proposed method achieves up to 103% improvement over the second-best method.
AGRaME: Any-Granularity Ranking with Multi-Vector Embeddings (2024.emnlp-main)

Copied to clipboard

Challenge: Existing ranking algorithms restrict granularity to full passages or require a specific dense index for each desired level of granules.
Approach: They propose a multi-vector ranking approach that leverages multi-vctor embeddings to rank at varying levels of granularity while maintaining encoding at a single (coarser) level of grail.
Outcome: The proposed method surpasses prompt-driven citation generation by incorporating proposition-level ranking to post-hoc citation addition.
Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has developed algorithms for reinforcement learning from human feedback and AI-generated feedback.
Approach: They propose a method for reinforcement learning from human feedback and AI-generated feedback that incorporates weighting, ranking, and constraining to handle disparate rewards.
Outcome: The proposed method reduces disparity and enhances stability among rewards . empirical results show that the proposed method is efficient and straightforward .
Automating Document Discovery in the Systematic Review Process: How to Use Chaff to Extract Wheat (L18-1)

Copied to clipboard

Challenge: Systematic reviews address research questions by comprehensively examining the entire published literature.
Approach: They compare the impact of different schemes for choosing positive and negative examples from the different screening stages on the training of automated systems.
Outcome: The proposed ranking system achieves an AUC of 0.803 and 0.768 when relying on gold standard decisions based on title and abstracts of articles, and an AUT of 0.625 and 0.839 when based upon gold standard decision based in full text.
On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings (2021.emnlp-main)

Copied to clipboard

Challenge: a large corpus of domain-expert relevance ratings augments a corpus for compositional explanations . a writer's study shows that the evaluations of compositional inference models underestimate performance .
Approach: They construct a corpus of 126k domain-expert relevance ratings that augment explanations to standardized science exam questions.
Outcome: The results show that evaluations underestimate performance of compositional explanations . they show that models regularly discover and produce valid explanations that are different than gold explanations.
Pyramid-BERT: Reducing Complexity via Successive Core-set based Token Selection (2022.acl-long)

Copied to clipboard

Challenge: Existing models that use heuristics to shorten sequence lengths are computationally prohibitive.
Approach: They propose a new method to shorten sequence lengths by transforming tokens through encoders and a core-set based token selection method that avoids expensive pre-training and fine tuning.
Outcome: The proposed model outperforms existing models on GLUE benchmarks and Long Range Arena datasets and demonstrates that it is cost-effective and space-efficient.
Learning Robust Models for e-Commerce Product Search (2020.acl-main)

Copied to clipboard

Challenge: Existing models that understand search intent are difficult to learn due to lack of labeled datasets.
Approach: They develop a deep, end-to-end model that learns to effectively classify mismatches . they introduce a latent variable into the cross-entropy loss that alternates between real and generated samples .
Outcome: The proposed model achieves a relative gain of over 26% in F-score and 17% in Area Under PR curve on live search traffic in multiple countries.
CaRB: A Crowdsourced Benchmark for Open IE (D19-1)

Copied to clipboard

Challenge: Open Information Extraction (Open IE) systems have been evaluated traditionally via manual annotation.
Approach: They propose to use a dataset to score Open IE systems by matching system predictions with benchmark datasets.
Outcome: The proposed framework matches predictions with the benchmark dataset and is noisy and inconsistent.
Multi-hop Evidence Retrieval for Cross-document Relation Extraction (2023.findings-acl)

Copied to clipboard

Challenge: Relation Extraction (RE) is a task that seeks to identify the relation of entities described according to some context.
Approach: They propose a multi-hop evidence retrieval method based on evidence path mining and ranking to support cross-document relation extraction.
Outcome: The proposed method acquires cross-document evidence and boosts performance in both closed and open environments.
Fusion-in-T5: Unifying Variant Signals for Simple and Effective Document Ranking with Attention Fusion (2024.lrec-main)

Copied to clipboard

Challenge: Current document ranking pipelines involve multiple ranking layers to integrate different information step-by-step.
Approach: They propose a novel re-ranker Fusion-in-T5 which integrates text matching information, ranking features, and global document information into one single unified model via templated-based input and global attention.
Outcome: The proposed model significantly improves ranking performance over complex cascade pipelines.
An LLM-based Framework for Biomedical Terminology Normalization in Social Media via Multi-Agent Collaboration (2025.coling-main)

Copied to clipboard

Challenge: Experimental results indicate that our approach exhibits competitive performance.
Approach: They propose a tuning-free approach to normalize non-standard terms using large language models . they use a search engine and a domain knowledge base to expand the short texts into accurate descriptions .
Outcome: The proposed approach is based on the "Recall and Re-rank" framework . it can be used to identify the standard term in a specified termbase for non-standardized mentions .
HyQE: Ranking Contexts with Hypothetical Query Embeddings (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to rank contexts rely on similarity between contexts and queries, but these methods are limited by the number of candidate contexts.
Approach: They propose a scalable ranking framework that combines embedding similarity and large language models without fine-tuning.
Outcome: The proposed framework improves the performance across multiple benchmarks.
Improving Content Recommendation: Knowledge Graph-Based Semantic Contrastive Learning for Diversity and Cold-Start Users (2024.lrec-main)

Copied to clipboard

Challenge: Current approaches focus on improving ranking performance at the cost of escalating complexity and complicating the task.
Approach: They propose a hybrid multi-task learning approach that trains on user-item and item-i item interactions.
Outcome: The proposed approach improves accuracy, relevance, and diversity of user recommendations even for cold-start users.
R3-NL2GQL: A Model Coordination and Knowledge Graph Alignment Approach for NL2GQL (2024.findings-emnlp)

Copied to clipboard

Challenge: Adapting existing approaches for converting natural language to SQL encounters hurdles due to distinct nature of GQL compared to SQL.
Approach: They propose a method that integrates both small and large Foundation Models for ranking, rewriting, and refining tasks.
Outcome: The proposed approach integrates both small and large Foundation Models for ranking, rewriting, and refining tasks while capitalizing on the superior generalization and query generation prowess of larger models for the final transformation of natural language queries into GQL formats.
Learning to Select: Query-Aware Adaptive Dimension Selection for Dense Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for dense retrieval use pseudo-relevance feedback to model dimension importance . however, they learn global transformations shared across queries and do not model dimension-aware dimension importance.
Approach: They propose a Query-Aware Adaptive Dimension Selection framework that learns to predict per-dimension importance directly from query embedding.
Outcome: The proposed framework improves retrieval effectiveness over the full-dimensional and PRF-based models.
Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Current methods rely on ranking losses to teach reward model to assess preferences, but they are susceptible to noise and ambiguous data, often failing to deeply understand human intentions.
Approach: They propose a method that incorporates contrastive learning into the reward modeling process to enhance generalization and stabilize the reinforcement learning training process.
Outcome: The proposed method enhances generalization of the reward model, stabilizes the reinforcement learning training process, and improves the final alignment with human preferences.
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to judge code, but their reliability remains poorly understood.
Approach: They propose a benchmark to evaluate Large Language Models as code judges . they find that small reasoning models outperform larger non-reasoning models .
Outcome: The proposed benchmark evaluates LLM-as-a-Judge models across three coding tasks.
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, but reliable evaluation remains a challenge due to data contamination, opaque operation, and subjective preferences.
Approach: They propose a benchmark-free evaluation paradigm that organizes multiple LLMs into a self-governed league for multi-round mutual evaluation.
Outcome: Experiments on eight mainstream LLMs in mathematics and programming show that the proposed model can distinguish capabilities while maintaining high internal ranking stability.
Beyond Ranking: Fine-Grained Diagnostics and Self-Improvement for MLLMs (2026.acl-long)

Copied to clipboard

Challenge: Current paradigms rely on holistic scoring and static leaderboards to disentangle fine-grained competencies.
Approach: They propose a framework to shift the focus from ranking to fine-grained diagnosis.
Outcome: The proposed framework surpasses the strongest baseline by 7.92%.
PaRaDe: Passage Ranking using Demonstrations with LLMs (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that large language models can be instructed to perform zero-shot passage re-ranking . Existing work like UPR demonstrate promising results for zero- shot ranking using LLMs .
Approach: They propose a demonstration selection strategy based on difficulty rather than semantic similarity . they propose to include only one demonstration in the prompt to improve re-ranking .
Outcome: The proposed method improves LLM-based re-ranking by adding one demonstration to the prompt.
(Almost) Free Modality Stitching of Foundation Models (2025.emnlp-main)

Copied to clipboard

Challenge: Multi-modal foundation models often use modality-specific (uni-modal) models as sub-components, which are stitched together via a connector module.
Approach: They propose a framework that allows for optimal uni-modal model selection and connector training by leveraging hypernetworks.
Outcome: The proposed framework reduces the cost of searching for the best performing uni-modal model pair by 10 while matching the ranking and trained connector performance across diverse multi-modal benchmarks.
PK-ICR: Persona-Knowledge Interactive Multi-Context Retrieval for Grounded Dialogue (2023.emnlp-main)

Copied to clipboard

Challenge: Identifying relevant persona or knowledge for conversational systems is difficult, but recent work has shown that it is more realistic to optimize for concrete persona.
Approach: They propose a persona-knowledge dual context retrieval method that utilizes all dialogue contexts simultaneously.
Outcome: The proposed method performs zero-shot top-1 knowledge retrieval and precise persona scoring.
Characterizing Positional Bias in Large Language Models: A Multi-Model Evaluation of Prompt Order Effects (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models can be influenced by various forms of biases, says a new study . positional bias affects how LLMs interpret and weigh information, the authors say .
Approach: a new study examines the impact of positional bias on large language models . positional biased models prioritize items based on their position rather than content or quality .
Outcome: a new study shows that LLMs prioritize items based on their position rather than content or quality . the positional bias affects how LLM interpret and weigh information, the authors say .
Bias Beware: The Impact of Cognitive Biases on LLM-Driven Product Recommendations (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized product recommenders, but their susceptibility to adversarial manipulations is difficult to detect.
Approach: They propose to use large language models to investigate cognitive biases as adversarial strategies in product research using LLMs.
Outcome: The proposed approach is the first to tap into human psychological principles, making such manipulations hard to detect.
Explain then Rank: Scale Calibration of Neural Rankers Using Natural Language Explanations from LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Neural ranking models produce the final document scores, but they are often treated as transient information and only the relative orderings are preserved to produce a ranking.
Approach: They propose to exploit large language models (LLMs) to provide relevance and uncertainty signals for these neural text rankers to produce scale-calibrated scores through Monte Carlo sampling of natural language explanations (NLEs).
Outcome: The proposed approach outperforms previous calibration methods and LLM-based methods for ranking, calibration, and query performance prediction tasks.
PLLuM-Align: Polish Preference Dataset for Large Language Model Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models generate preferred responses while avoiding harmful or inappropriate outputs, despite their ability to generate cross-language transferability.
Approach: They introduce the first Polish preference dataset PLLuM-Align, created entirely through human annotation to reflect Polish language and cultural nuances.
Outcome: The proposed dataset lays the groundwork for more aligned Polish LLMs and contributes to the broader goal of multilingual alignment in underrepresented languages.
Beyond Contrastive Learning: Synthetic Data Enables List-wise Training with Multiple Levels of Relevance (2025.findings-emnlp)

Copied to clipboard

Challenge: a new approach to training with binary relevance labels uses synthetic data . contrastive learning with binary correlations leaves out subtle nuances useful for ranking .
Approach: They propose to use waterstein distance as a loss function for training transformer-based retrievers with graduated relevance labels instead of real documents.
Outcome: The proposed method outperforms conventional training with InfoNCE by a large margin on MARCO and BEIR benchmarks without using real documents.
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat (2025.acl-long)

Copied to clipboard

Challenge: Evaluating large language models (LLMs) is a complex task. Pairwise ranking has emerged as state-of-the-art method to evaluate human preferences.
Approach: They propose to use pairwise ranking to evaluate human preferences . they propose to evaluate the robustness of ranking algorithms in LLMs .
Outcome: The proposed methods are based on the principles of effective ranking and the robustness of several ranking algorithms in the context of LLMs.
KGE Calibrator: An Efficient Probability Calibration Method of Knowledge Graph Embedding Models for Trustworthy Link Prediction (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for probability calibration of knowledge graph embedding models are ill-suited for KGEs.
Approach: They propose a method to calibrate knowledge graph embedding models for ranking-based link prediction using a Jump Selection Strategy and Multi-Binning Scaling to enhance reliability.
Outcome: Experiments show that the KGEC outperforms existing calibration methods in terms of effectiveness and efficiency.
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models (2026.acl-long)

Copied to clipboard

Challenge: Competitive programming has become a rigorous benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs).
Approach: They propose a scalable and reproducible test-time compute framework that achieves IOI gold-level performance using open-weight models.
Outcome: The proposed framework achieves IOI gold-level performance using open-weight models . it scales consistently with available compute, narrowing the gap between open and closed systems.
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Relevance emphasizes the aboutness of a result to a query, while utility refers to the result’s usefulness or value to an information seeker.
Approach: They propose an Iterative utiliTy judgmEnt fraMework to promote each step in Retrieval-Augmented Generation (RAG) they propose to use relevance ranking, utility judgments, and answer generation to prioritize high-utility results over low-utilitity results.
Outcome: The proposed framework improves relevance, ranking, and answer generation on retrieval (TREC DL, WebAP), utility judgment task (GTI-NQ), and factoid question-answering (NQ) datasets.
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text-to-video retrieval systems use embedding models for feature extraction and compute cosine similarities for ranking.
Approach: They propose an explainable retrieval framework upon LLM CoT reasoning to replace embedding models for feature extraction and ranking.
Outcome: The proposed retrieval framework improves retrieval performance and produces detailed rationales.
Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer (2025.emnlp-main)

Copied to clipboard

Challenge: NN-Rank is an algorithm for ranking source languages for cross-lingual transfer . it leverages hidden representations from multilingual models and unlabeled target-language data .
Approach: They propose an algorithm for ranking source languages for cross-lingual transfer which leverages hidden representations from multilingual models and unlabeled target-language data.
Outcome: The proposed algorithm outperforms state-of-the-art models on in-domain data and shows that it can achieve 92.8% of the NDCG achieved using all available target data.
Adaptive Retrieval for Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing reasoning-based rerankers suffer from bounded recall.
Approach: They propose a framework that leverages adaptive retrieval to ensure sufficient "bridge" documents are retrieved for reasoning-intensive retrieval.
Outcome: The proposed method outperforms baselines on reasoning-intensive retrieval tasks by 5.6%pt.
PsychePass: Calibrating LLM Therapeutic Competence via Trajectory-Anchored Tournaments (2026.findings-acl)

Copied to clipboard

Challenge: evaluating therapeutic competence of large language models remains challenging due to unstructured and longitudinal nature of counseling.
Approach: They propose a framework that calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments.
Outcome: The proposed framework calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments.
R3-SQL: Ranking Reward and Resampling for Text-to-SQL (2026.findings-acl)

Copied to clipboard

Challenge: Existing rankers assign inconsistent scores to functionally equivalent SQL queries . ranking cannot recover when the correct SQL is absent from the pool.
Approach: They propose a Text-to-SQL framework that rewards ranking and resampling . it first groups candidates by execution result and ranks groups for consistency .
Outcome: The proposed framework achieves 75.03 execution accuracy on BIRD-dev, a new state of the art among methods using models with disclosed sizes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations