Challenge: Learning to hash via generative model is a powerful paradigm for fast similarity search in documents retrieval.
Approach: They propose a method that trains a generative model to generate hash codes by using continuous relaxation on priors.
Outcome: The proposed method outperforms other state-of-the-art methods in qualitative and quantitative experiments.

Similar Papers

Document Hashing with Mixture-Prior Generative Models (D19-1)

Copied to clipboard

Challenge: Existing generative hashing methods only consider the use of simple priors, which limits them to further improve their performance.
Approach: They propose to use Gaussian and Bernoulli priors to generate hashing codes . they propose to cast a Gausssian latent representation into binary code .
Outcome: The proposed models outperform existing methods on a benchmark dataset using Gaussian and Bernoulli priors.
NASH: Toward End-to-End Neural Architecture for Generative Semantic Hashing (P18-1)

Copied to clipboard

Challenge: Existing approaches to fast similarity search require two-stage training and the binary constraints are handled ad-hoc.
Approach: They propose an end-to-end neural architecture for semantic hashing where binary hash codes are treated as Bernoulli latent variables.
Outcome: The proposed approach outperforms state-of-the-art models on unsupervised and supervised scenarios on three public datasets.
DiffusionRet: Diffusion-Enhanced Generative Retriever using Constrained Decoding (2023.findings-emnlp)

Copied to clipboard

Challenge: Generative retrieval methods have suffered from the lack of the intermediate reasoning step . generative retrieval uses sequence-to-sequence diffusion models to map a query to relevant docids .
Approach: They propose a novel method that uses query as an intermediate step before retrieval . they propose to use sequence-to-sequence diffusion models to map a query to relevant docids .
Outcome: Experiments show that proposed method outperforms existing methods on MARCO and Natural Questions datasets.
Generative Semantic Hashing Enhanced via Boltzmann Machines (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for generative semantic hashing assume a factorized posterior distribution, enforcing independence among the bits of hash codes.
Approach: They propose to use a Boltzmann machine distribution as the variational posterior to introduce correlations among the bits of hash codes.
Outcome: The proposed method can achieve significant performance gains by combining two hash codes.
Latent-Variable Generative Models for Data-Efficient Text Classification (D19-1)

Copied to clipboard

Challenge: Generative classifiers offer potential advantages over discriminative classifications, including data efficiency and zero-shot learning.
Approach: They introduce discrete latent variables into generative story to improve classifiers' performance . they empirically characterize performance of their models on six text classification datasets .
Outcome: The proposed model outperforms discriminative and generative classifiers on six text classification datasets.
Document Hashing with Multi-Grained Prototype-Induced Hierarchical Generative Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing document hashing methods only consider flat semantics of documents, preserving hierarchical semantics.
Approach: They propose a hierarchical generative model that can model and leverage hierarchic semantics . they introduce hierarchically-based prototypes into the model to construct a Hierarchical prior distribution .
Outcome: The proposed model outperforms baseline methods on hierarchical and flat datasets.
Generative Text Modeling through Short Run Inference (2021.eacl-main)

Copied to clipboard

Challenge: Latent variable models for text capture global semantic and syntactic features when trained correctly.
Approach: They propose a short run dynamics for inference that initializes from the prior distribution of the latent variable and runs a small number of Langevin dynamics steps guided by its posterior distribution.
Outcome: The proposed model is able to generate coherent sentences with smooth transition and shows no sign of posterior collapse.
GLEN: Generative Retrieval via Lexical Index Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for document retrieval bypass auxiliary index structures and can be optimized through end-to-end learning.
Approach: They propose a method to generate a relevant document's identifier using an index learning strategy.
Outcome: The proposed method achieves state-of-the-art or competitive performance on benchmark datasets.
On Synthetic Data Strategies for Domain-Specific Generative Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Generative retrieval models can be used to generate ranked lists of potentially relevant document identifiers for a user query.
Approach: They propose a synthetic data generation strategy for a two-stage training framework that focuses on learning to decode document identifiers from queries and a strategy for mining hard negatives based on initial model's predictions.
Outcome: The proposed model can generate ranked lists of potentially relevant document identifiers for a user query and then refine ranking through preference learning.
Multi-level Relevance Document Identifier Learning for Generative Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Existing methods generate DocIDs based on textual content, which may result in weak semantic connections for similar documents due to variations in expression.
Approach: They propose a new retrieval paradigm that generates unique document identifiers . they propose to use queries as a bridge to connect documents with varying relevance levels .
Outcome: The proposed approach outperforms existing methods on multilingual e-commerce search datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations