Papers by Luyu Gao

11 papers
Modularized Transfomer-based Ranking Framework (2020.emnlp-main)

Copied to clipboard

Challenge: Recent innovations in Transformer-based ranking models have advanced the state-of-the-art in information retrieval.
Approach: They propose to modularize a Transformer ranker into separate modules for text representation and interaction.
Outcome: The proposed model is faster than previous models and is easier to interpret and understand.
DataFinder: Scientific Dataset Recommendation from Natural Language Descriptions (2023.acl-long)

Copied to clipboard

Challenge: Modern machine learning relies on datasets to develop and validate research ideas.
Approach: They propose a dataset recommendation system that uses a training set and an evaluation set to help people find relevant datasets.
Outcome: The proposed model finds more relevant search results than existing third-party search engines.
RARR: Researching and Revising What Language Models Say, Using Language Models (2023.acl-long)

Copied to clipboard

Challenge: Language models (LMs) excel at many tasks but often produce unsupported or misleading content.
Approach: They propose a system that finds attribution for any text generation model and post-edits it to fix unsupported content.
Outcome: The proposed system improves attribution while preserving the original output.
COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List (2021.naacl-main)

Copied to clipboard

Challenge: Recent neural IR models shift towards soft matching all query document terms, but they lose the computation efficiency of exact match systems.
Approach: They propose a contextualized exact match retrieval architecture where scoring is based on overlapping query document tokens’ contextualized representations.
Outcome: The proposed architecture outperforms classical lexical retrieval systems and state-of-the-art deep language models with smaller latency.
Condenser: a Pre-training Architecture for Dense Retrieval (2021.emnlp-main)

Copied to clipboard

Challenge: Prior work fine-tunes deep LMs to encode text sequences into single dense vector representations, but dense encoders require a lot of data and sophisticated techniques to train and suffer in low data situations.
Approach: They propose to pre-train Transformer language models (LMs) with a novel Transformer architecture, Condenser, where LM prediction CONditions on DENSE Representation.
Outcome: The proposed model improves on various text retrieval and similarity tasks by large margins over standard LMs.
Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval (2022.acl-long)

Copied to clipboard

Challenge: Recent research shows that fine-tuning dense retrievers to realize their capacity requires carefully designed fine-cuning techniques.
Approach: They propose a pre-training architecture that learns to condense information into the dense vector through LM pre-training and a coCondenser architecture which adds an unsupervised corpus-level contrastive loss to warm up the passage embedding space.
Outcome: The proposed architecture reduces the need for heavy data engineering and large batch training.
Retrieval as Attention: End-to-end Learning of Retrieval and Reading within a Single Transformer (2022.emnlp-main)

Copied to clipboard

Challenge: eschewing separate architecture and training for knowledge-intensive tasks is cumbersome . end-to-end training only based on supervision from the end task is awkward .
Approach: They propose a single Transformer that performs retrieval as attention and end-to-end training solely based on supervision from the end QA task.
Outcome: The proposed model outperforms state-of-the-art retrievers and readers on in-domain datasets.
Precise Zero-Shot Dense Retrieval without Relevance Labels (2023.acl-long)

Copied to clipboard

Challenge: Existing dense retrieval systems that use semantic embedding similarities can be effective across tasks and languages.
Approach: They propose to pivot through Hypothetical Document Embeddings (HyDE) given a query, HyDE first zero-shot prompts an instruction-following language model to generate a hypothetical document.
Outcome: The proposed method significantly outperforms the state-of-the-art unsupervised dense retriever Contriever and shows strong performance comparable to fine-tuned retrievers across tasks and languages.
Improving Target-side Lexical Transfer in Multilingual Neural Machine Translation (2020.findings-emnlp)

Copied to clipboard

Challenge: Multilingual data is more beneficial for NMT models that translate from the LRL to a target language than those that translate into the LLLs.
Approach: They propose a decoder that embeds character n-grams into NMT models that translate from an LRL to a target language.
Outcome: The proposed decoder improves the performance of NMT models that translate from an LRL to a target language.
BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for deep search agents rely on blackbox web search APIs . dynamic and opaque web APIs hinder reproducibility and fair comparisons - authors .
Approach: They propose a benchmark that employs a fixed corpus for controlled retrieval for deep search agents.
Outcome: The new benchmark shows that agents that combine large language models with retrieval tools excel at complex, reasoning-intensive queries.
Active Retrieval Augmented Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Generative language models (LMs) have a tendency to hallucinate and create inaccurate output.
Approach: They propose a method which iteratively uses a prediction of the upcoming sentence to anticipate future content.
Outcome: The proposed method achieves superior or competitive performance on all tasks . iteratively uses a prediction of the upcoming sentence to anticipate future content .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations