Papers by Daniel Cer

12 papers
Neural Retrieval for Question Answering with Cross-Attention Supervised Data Augmentation (2021.acl-short)

Copied to clipboard

Challenge: Early fusion models with cross-attention have shown better-than-human performance on some question answer benchmarks, while it is a poor fit for retrieval since it prevents pre-computation of the answer representations.
Approach: They propose a supervised data mining method to train an efficient late fusion retrieval model by using cross-attention models with cross-references.
Outcome: The proposed model outperforms retrieval models trained with gold annotations on Precision at N (P@N) and Mean Reciprocal Rank (MRR).
Multilingual Universal Sentence Encoder for Semantic Retrieval (2020.acl-demos)

Copied to clipboard

Challenge: Using a multi-task trained dual-encoder, our models embed text from 16 languages into a shared semantic space.
Approach: They propose retrieval focused multilingual sentence embedding models on TensorFlow Hub.
Outcome: The models achieve state-of-the-art on monolingual and cross-lingual retrieval (SR) and retrieval question answering (ReQA) competitive performance is obtained on related tasks of translation pair bitext retrieval and retrieving question answering.
SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer (2022.acl-long)

Copied to clipboard

Challenge: Recent studies show that pre-trained language models can be more efficient when they are larger than they are in their size.
Approach: They propose a prompt-based transfer learning approach called SPoT: Soft Prompt Transfer that learns a soft prompt on one or more source tasks and initializes it for a target task.
Outcome: The proposed approach outperforms Prompt Tuning and MODELTUNING on superGLUE benchmarks while using up to 27,000 fewer task-specific parameters.
Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models (2022.findings-acl)

Copied to clipboard

Challenge: Sentence embeddings are useful for language processing tasks, but it is unclear how to produce them from encoder-decoder models.
Approach: They investigate the effects of scaling up sentence encoders to 11B parameters on sentence embeddings from text-to-text transformers (T5) .
Outcome: The proposed models outperform the previous best models on both SentEval and SentGLUE transfer tasks.
Universal Sentence Encoder for English (D18-2)

Copied to clipboard

Challenge: TensorFlow Hub sentence embedding models have good task transfer performance . model variants allow for trade-offs between accuracy and compute resources .
Approach: They propose easy-to-use TensorFlow Hub sentence embedding models with good task transfer performance.
Outcome: The proposed models outperform models without transfer learning and those that use only word-level transfer on a number of NLP tasks.
Universal Sentence Representation Learning with Conditional Masked Language Model (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods to learn sentence representations on unlabeled corpora are difficult and expensive to obtain, making it hard to cover many domains and languages.
Approach: They propose a method to train sentence representations on large unlabeled corpora by conditioning on the encoded vectors of adjacent sentences.
Outcome: The proposed method outperforms existing models on SentEval and can be extended to a broad range of languages and domains.
Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO (2021.eacl-main)

Copied to clipboard

Challenge: Existing image captioning datasets have limited cross-modal associations, preventing researchers from examining how inter-modal learning impacts intra-modal tasks.
Approach: They propose to use image captioning data to support multi-modal retrieval training and evaluation to assess the impact of inter-modality learning.
Outcome: The proposed model is able to measure the influence of intra- and inter-modality learning.
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval (2024.naacl-long)

Copied to clipboard

Challenge: et al., 2020: performance of dense retrieval models in multilingual retrieval is limited due to uneven and scarce training data available across multiple languages.
Approach: They propose a synthetic retrieval training dataset containing 33 languages for fine-tuning multilingual retrievers without human supervision.
Outcome: The proposed model outperforms human-supervised retrieval models on three retrieval benchmarks.
ReQA: An Evaluation for End-to-End Answer Retrieval Models (D19-58)

Copied to clipboard

Challenge: Popular QA benchmarks like SQuAD have driven progress on identifying answer spans within a specific passage . retrieving relevant answers from a huge corpus of documents is still a challenging problem .
Approach: They propose a benchmark for evaluating large-scale sentence-level answer retrieval models . they establish baselines using both neural encoding models and classical retrieval techniques .
Outcome: The proposed model outperforms human models on identifying answer spans within a specific passage . the proposed model is scalable and can bypass the typical document retrieval step .
Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation (2022.emnlp-main)

Copied to clipboard

Challenge: generative multilingual models fine-tuned on English forget to generate non-English data when labeled data is only available in English . generative models fine tuned on English fail to generate multilingual summarization tasks when labeling data is available in other languages .
Approach: They propose to use prompt tuning to overcome catastrophic forgetting in a generative task in . they assume a strict setting with no parallel data or machine translation .
Outcome: The proposed method can overcome catastrophic forgetting to enable zero-shot cross-lingual generation.
A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations (2021.emnlp-main)

Copied to clipboard

Challenge: Language agnostic and semantic-language information isolation is an emerging research direction for multilingual representations models.
Approach: They propose a method that factors out language identity information from semantic related components in multilingual representations pre-trained on monolingual data.
Outcome: The proposed method improves cross-lingual transfer performance on weak alignment models.
Language-agnostic BERT Sentence Embedding (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for learning bilingual sentence embeddings are not well explored.
Approach: They propose to combine best methods for learning multilingual sentence embeddings with pre-trained models to achieve 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba.
Outcome: The proposed model achieves 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba, above the 65.5% achieved by LASER.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations