Challenge: a new study examines the retrieval effectiveness of commercial embedding models . robert mcgahey: can commercial embeds be "stolen" using distillation techniques? he says stealing models can offer benefits to different actors, including reduced costs and security .
Approach: They propose to "steal" embedding models by training thief models on text–embedding pairs . they replicate retrieval effectiveness of commercial embeddable models with a cost of under $300 .
Outcome: The proposed methods replicate retrieval effectiveness of commercial embedding models with under $300 . authors suggest measures to mitigate risk of model theft.

Similar Papers

Sentence Embedding Leaks More Information than You Expect: Generative Embedding Inversion Attack to Recover the Whole Sentence (2023.findings-acl)

Copied to clipboard

Challenge: Sentence-level representations are beneficial for various natural language processing tasks.
Approach: They propose a generative embedding inversion attack that reconstructs input sequences based only on their sentence embeddeds.
Outcome: The proposed model outperforms previous embedding inversion attacks in classification metrics and generates coherent and contextually similar sentences as original inputs.
Student Surpasses Teacher: Imitation Attack for Black-Box NLP APIs (2022.coling-1)

Copied to clipboard

Challenge: Existing MLaaS models are vulnerable to imitation attacks, but none of the stolen models can outperform the original black-box APIs.
Approach: They conduct unsupervised domain adaptation and multi-victim ensemble to show attackers could surpass victims.
Outcome: The proposed model outperforms the original black-box models on transferred domains.
Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models (2021.naacl-main)

Copied to clipboard

Challenge: Recent studies reveal a security threat to natural language processing models, called the Backdoor Attack.
Approach: They propose to hack a model by modifying one single word embedding vector without sacrificing accuracy on clean samples.
Outcome: The proposed method is more efficient and stealthier on sentiment analysis and sentence-pair classification tasks.
Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark (2023.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated exceptional abilities in both text understanding and generation.
Approach: They propose an Embedding Watermark method that implants backdoors on embeddings to protect copyright of large language models.
Outcome: The proposed method protects the copyright of large language models without compromising service quality while minimizing the adverse impact on the original embeddings’ utility.
Backdoor Attacks in Federated Learning by Rare Embeddings and Gradient Ensembling (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances in federated learning have demonstrated its promising capability to learn on decentralized datasets.
Approach: They propose a technique that allows adversaries to poison the global model . they propose 'model poisoning' for backdoor attacks using word embeddings of NLP models .
Outcome: The proposed technique improves the model poisoning performance in all experimental settings.
Privacy-preserving Neural Representations of Text (D18-1)

Copied to clipboard

Challenge: a specific type of attack is used to characterize the privacy of neural representations for NLP tasks, in the context of privacy protection.
Approach: They propose several defense methods based on modified training objectives and characterize the tradeoff between privacy and the utility of neural representations.
Outcome: The proposed defenses improve the privacy of neural representations and characterize the tradeoff between privacy and utility of representations.
Frustratingly Easy Meta-Embedding – Computing Meta-Embeddings by Averaging Source Word Embeddings (N18-2)

Copied to clipboard

Challenge: Existing methods for producing word embeddings have shown to produce accurate meta-embeddings from pre-trained source embeddables.
Approach: They propose to use arithmetic mean of two distinct word embedding sets to produce an accurate meta-embedding.
Outcome: The proposed method produces meta-embeddings comparable or better than more complex methods.
Evaluating Embedding APIs for Information Retrieval (2023.acl-industry)

Copied to clipboard

Challenge: a growing number of language models are limiting their access to the community . we evaluate existing APIs for domain generalization and multilingual retrieval .
Approach: They evaluate semantic embedding APIs in retrieval scenarios to assess their capabilities . they use BEIR and MIRACL to re-rank BM25 results using the APIs .
Outcome: The proposed model is based on semantic embedding APIs that build vector representations of a given text.
Membership Inference Attacks on Sequence-to-Sequence Models: Is My Data In Your Machine Translation System? (2020.tacl-1)

Copied to clipboard

Challenge: Data privacy is an important issue for “machine learning as a service” providers.
Approach: They propose an attack on membership inference attacks using a sequence-to-sequence model and a machine translation dataset to investigate the feasibility of a privacy attack.
Outcome: The proposed model can infer sentence-level membership from the output of the model, but it is difficult to infer it.
Better Word Embeddings by Disentangling Contextual n-Gram Information (N19-1)

Copied to clipboard

Challenge: Pre-trained word vectors are ubiquitous in Natural Language Processing applications.
Approach: They show that word embeddings with bigram and trigram embedds improve unigram embeds . they claim this removes contextual information from unigrammes, resulting in better unigraph embedders .
Outcome: The proposed model outperforms competing models on a wide variety of tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations