Papers by Ronak Pradeep

6 papers
How Does Generative Retrieval Scale to Millions of Passages? (2023.emnlp-main)

Copied to clipboard

Challenge: generative retrieval is a new paradigm for information retrieval, enabling a sequence-to-sequence model with a single Transformer . generative encoders have been used on small corpora, but only on large ones .
Approach: They propose to encode an entire document corpus within a single Transformer . they find generative retrieval is competitive with state-of-the-art dual encoders on small corpora .
Outcome: The proposed approach is competitive with state-of-the-art dual encoders on small corpora, the study finds . the proposed approach only evaluates on document corporales on the order of 100K in size .
ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA Datasets with Large Language Models (2024.emnlp-industry)

Copied to clipboard

Challenge: Knowledge Graphs (KGs) are a powerful tool for capturing structured representations of the world.
Approach: They propose a scalable method for generating up-to-date and configurable conversational KGQA datasets that adheres to human interaction configurations and operates at a significantly larger scale.
Outcome: Qualitative psychometric analyses show that ConvKGYarn produces high-quality data comparable to popular conversational KGQA datasets across various metrics.
Entity Disambiguation via Fusion Entity Decoding (2024.naacl-long)

Copied to clipboard

Challenge: Existing generative approaches demonstrate improved accuracy compared to classification approaches under the standardized ZELDA benchmark.
Approach: They propose an encoder-decoder model to disambiguate entities with more detailed entity descriptions.
Outcome: The proposed model outperforms existing classification models on the ZELDA benchmark and on retrieval/reader frameworks.
Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages (2024.acl-short)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive zero-shot capabilities in various passage ranking tasks.
Approach: They analyze and compare the effectiveness of monolingual reranking using query or document translations and evaluate the effectiveness when leveraging their own generated translations.
Outcome: The proposed models perform better when using their own translations than when using query or document translations.
Document Ranking with a Pretrained Sequence-to-Sequence Model (2020.findings-emnlp)

Copied to clipboard

Challenge: Experimental results on the MS MARCO passage ranking task show that our ranking approach is superior to strong encoder-only models.
Approach: They propose to use a pretrained sequence-to-sequence model to generate relevance labels as "target tokens" they also show how the underlying logits of these target tokens can be interpreted as relevance probabilities for ranking.
Outcome: The proposed model outperforms existing models in a data-poor setting and significantly outperformed an encoder-only model on the MS MARCO passage ranking task.
Exploring Listwise Evidence Reasoning with T5 for Fact Verification (2021.acl-short)

Copied to clipboard

Challenge: Existing methods for fact verification use pretrained sequence-to-sequence transformers for sentence selection and label prediction.
Approach: They propose a framework for fact verification that leverages pretrained sequence-to-sequence transformer models for sentence selection and label prediction.
Outcome: The proposed framework scores higher than the second place approach on the blind test set . the proposed framework can be useful for a broader range of NLP tasks, the authors say .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations