Challenge: Existing backdoor watermarking techniques are limited to zero-bit detection . RShield enables reliable user-level attribution of large language models under model extraction attacks.
Approach: They propose a multi-bit backdoor watermarking technique that enables reliable user-level attribution of large language models under model extraction attacks.
Outcome: RShield achieves 100% multi-bit watermark recovery and high semantic fidelity under model extraction attacks compared to existing methods.

Similar Papers

Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark (2023.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated exceptional abilities in both text understanding and generation.
Approach: They propose an Embedding Watermark method that implants backdoors on embeddings to protect copyright of large language models.
Outcome: The proposed method protects the copyright of large language models without compromising service quality while minimizing the adverse impact on the original embeddings’ utility.
WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection (2024.acl-long)

Copied to clipboard

Challenge: Prior studies have shown that EaaS can be prone to model extraction attacks, however, this concern could be mitigated by adding backdoor watermarks to the text embeddings.
Approach: They propose a new method that removes backdoor watermarks while maintaining the high utility of embeddings.
Outcome: The proposed approach increases the stealthiness of watermarks and has been empirically shown to be effective against CSE attacks.
GuardEmb: Dynamic Watermark for Safeguarding Large Language Model Embedding Service Against Model Stealing Attack (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies reveal the risk of the model stealing attack, posing a financial threat to EaaS providers.
Approach: They propose a dynamic embedding watermarking method that detects watermarks in embedded text . this method is a cross-platform approach that trains a verifier to detect watermark .
Outcome: The proposed method enables an attacker to replicate the proposed method for profit without compromising embedding functionality.
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks (2025.acl-long)

Copied to clipboard

Challenge: Existing EaaS watermarks can be removed by paraphrasing when attackers clone the model.
Approach: They propose a method that integrates a target embedding into the original embeddable based on the presence of trigger words in the input text.
Outcome: The proposed technique is empirically and theoretically robust against paraphrasing.
Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing detection methods for large language models rely on fixed strategies to steal watermarks.
Approach: They propose a novel steal-based watermark algorithm that derives watermark information from watermarked texts to craft highly targeted adversarial attacks.
Outcome: The proposed system significantly increases steal efficiency against target watermarks under identical conditions.
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark (2025.findings-emnlp)

Copied to clipboard

Challenge: Embedding-as-a-Service (EaaS) is a successful business pattern but faces significant challenges related to various forms of copyright infringement.
Approach: They propose a semantic-independent watermarking scheme that exploits semantic perturbation tests to bypass verification.
Outcome: The proposed watermarking schemes possess semantic-independent characteristics and exploit semantic perturbation tests to bypass verification.
Subtle Signatures, Strong Shields: Advancing Robust and Imperceptible Watermarking in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have led to an increase in AI-generated text on the Internet, presenting a crucial challenge to differentiate AI-created content from human-written text.
Approach: They propose a novel approach to embed watermarks into LLMs that leverages token prior probabilities to improve detectability and maintain watermark imperceptibility.
Outcome: The proposed method improves detectability and imperceptibility of watermarks by partitioning tokens into two distinct groups based on prior probabilities and employing tailored strategies for each group.
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Backdoor-based LLM fingerprinting is a promising solution for intellectual property protection . however, the vulnerability of existing LLMs for the ensemble scenario is unexplored .
Approach: They propose two new fingerprinting attack methods to assess the robustness of LLM fingerprinting by token filter attack and sentence verification attack.
Outcome: The proposed methods inhibit the fingerprint response while maintaining ensemble performance.
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing watermarking methods face limitations that hinder their effectiveness in diverse and adversarial scenarios.
Approach: They propose a symbiotic watermarking framework with three strategies: serial, parallel, and hybrid.
Outcome: The proposed framework outperforms baselines and achieves state-of-the-art (SOTA) performance.
Composite Backdoor Attacks Against Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated superior performance on various tasks, but untrustworthy third-party LLMs may covertly introduce vulnerabilities for downstream tasks.
Approach: They propose a composite backdoor attack that scatters multiple trigger keys in different prompt components.
Outcome: The proposed attack achieves 100% Attack Success Rate (ASR) with a False Triggered Rate (FTR) below 2.06% and negligible model accuracy degradation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations