Papers with LSH

4 papers
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text (2024.findings-acl)

Copied to clipboard

Challenge: Recent watermarked generation algorithms inject detectable signatures during language generation to facilitate post-hoc detection.
Approach: They propose a watermark which assigns signatures to each watermarked sentence according to locality-sensitive hashing (LSH) they propose k-SemStamp, which uses kmeans clustering to partition the semantic space with awareness of inherent semantic structure.
Outcome: The proposed watermark improves its robustness and sampling efficiency while preserving the generation quality, making it more effective for machine-generated text detection.
SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation (2024.naacl-long)

Copied to clipboard

Challenge: Existing watermarked generation algorithms employ token-level designs and are vulnerable to paraphrase attacks.
Approach: They propose a sentence-level watermarking algorithm that uses locality-sensitive hashing to partition the semantic space of sentences.
Outcome: The proposed algorithm is more robust than the existing state-of-the-art method on paraphrasers and domains, while posing only minor degradations to SemStamp.
Transferable Neural Projection Representations (N19-1)

Copied to clipboard

Challenge: Neural word embeddings require lookup and a large memory footprint making it hard to deploy on-device.
Approach: They propose a skip-gram based architecture coupled with Locality-Sensitive Hashing projections to learn efficient dynamically computable representations.
Outcome: The proposed model performs better than previous models on multiple NLP tasks.
A Strong Baseline for Query Efficient Attacks in a Black Box Setting (2021.emnlp-main)

Copied to clipboard

Challenge: Existing black box search methods are inefficient as they do not consider the amount of queries required to generate adversarial attacks.
Approach: They propose a query efficient attack strategy to generate plausible adversarial examples on text classification and entailment tasks.
Outcome: The proposed attack reduces query count by 75% across all datasets and target models compared to prior attacks in a limited query setting.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations