Profiling LLM’s Copyright Infringement Risks under Adversarial Persuasive Prompting (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models have demonstrated impressive capabilities in text generation but raise concerns regarding potential copyright infringement. |
| Approach: | They propose a structured persuasion workflow to analyze the influence of persuasive prompts on LLM outputs. |
| Outcome: | The proposed method analyzes the influence of persuasive prompts on LLM outputs. |
Similar Papers
Can You Trick the Grader? Adversarial Persuasion of LLM Judges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used as automated evaluators in practical settings . |
| Approach: | a study by the university of california reveals that persuasive language can bias large language models when scoring mathematical reasoning tasks. |
| Outcome: | The proposed model can bias judges when scoring mathematical reasoning tasks . Consistency causes the most severe distortion, with Consistencies leading to 8% distortion . |
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have transformed machine learning but have raised significant legal concerns due to their potential to produce text that infringes on copyrights. |
| Approach: | They propose a lightweight, real-time defense mechanism to prevent the generation of copyrighted text by evaluating methods and testing attack strategies. |
| Outcome: | The proposed defense significantly reduces the volume of copyrighted text generated by LLMs by effectively refusing malicious requests. |
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities (2024.naacl-long)
Copied to clipboard
| Challenge: | Rapid progress in open-source Large Language Models (LLMs) is driving AI development, but lacks sufficient trustworthiness to detect and mitigate adversarial demonstrations. |
| Approach: | They propose an extended Chain of Utterances-based (CoU) prompting strategy to attack open-source LLMs. |
| Outcome: | The proposed attack strategy is based on malicious demonstrations and toxicity tests on open-source models. |
Nine Ways to Break Copyright Law and Why Our LLM Won’t: A Fair Use Aligned Generation Framework (2025.findings-emnlp)
Copied to clipboard
Aakash Sen Sharma, Debdeep Sanyal, Priyansh Srivastava, Sundar Athreya H, Shirish Karande, Mohan Kankanhalli, Murari Mandal
| Challenge: | Large language models (LLMs) often risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications. |
| Approach: | They propose a legally-grounded framework to align LLM outputs with fair-use doctrine . LAW-LM uses a dataset containing 18,000 expert-validated examples . |
| Outcome: | The proposed framework aligns outputs with fair-use doctrine and is validated by 18,000 experts. |
LLMs and Copyright Risks: Benchmarks and Mitigation Approaches (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized natural language processing, but their widespread use has raised significant copyright concerns. |
| Approach: | This tutorial will provide an overview of relevant copyright principles and their application to AI and examine specific copyright issues in LLM development and deployment. |
| Outcome: | The course will provide an overview of relevant copyright principles and their application to AI, followed by an examination of specific copyright issues in LLM development and deployment. |
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to jailbreak large language models have been poorly studied . a recent study showed that non-expert users can jailbreak LLMs by manipulating their prompts . |
| Approach: | They propose a formalism and a taxonomy of known (and possible) jailbreaks . they propose generating a dataset of model outputs across 3700 jailbreak prompts a 'prompt' attack is a new attack popularly categorized as "prompting injection attacks" |
| Outcome: | The proposed model exploits 3700 jailbreak prompts over 4 tasks to analyze their effectiveness . authors show that the model can learn to perform a new task on unseen examples . |
Do LLMs Know to Respect Copyright Notice? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have focused on the occurrence of copyright violations in LLM output, but a negative answer would suggest that LLMs will become the primary facilitator and accelerator of copy right infringement behavior. |
| Approach: | They propose to examine whether LLMs respect copyright information in user input . they use a set of language models, user prompts, and copyrighted materials . |
| Outcome: | The proposed model will be the primary facilitator and accelerator of copyright infringement behavior, the study finds . the study also provides a benchmark dataset serving as a test bed for evaluating infringement behaviors by LLMs . |
Vulnerability of LLMs’ Stated Belief? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly employed in question-answering tasks. |
| Approach: | They analyze how different persuasive strategies influence stated belief stability . they also examine whether verbalized confidence prompting increases vulnerability . |
| Outcome: | The proposed model exhibits extreme compliance, with 82.5% of belief changes occurring at the first persuasive turn. |
Close or Cloze? Assessing the Robustness of Large Language Models to Adversarial Perturbations via Word Recovery (2025.coling-main)
Copied to clipboard
| Challenge: | Existing models implicitly recover the original text, but it is unclear when they rely on context and when they implicitly do so. |
| Approach: | They propose to use a dictionary to recover adversarial words by using a phonetic, typo, and visual attack to study word recovery performance. |
| Outcome: | The proposed model outperforms open-source models on hateful, offensive, and toxic classification tasks. |
Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies have focused on specific domains or types of persuasion, but a general study has focused on how LLMs produce persuasive text. |
| Approach: | They construct a dataset to measure and benchmark the ability of Large Language Models (LLMs) to produce persuasive text. |
| Outcome: | The proposed model can be used to generate persuasive text across domains and domains. |