Challenge: Large language models (LLMs) often risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications.
Approach: They propose a legally-grounded framework to align LLM outputs with fair-use doctrine . LAW-LM uses a dataset containing 18,000 expert-validated examples .
Outcome: The proposed framework aligns outputs with fair-use doctrine and is validated by 18,000 experts.

Similar Papers

SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have transformed machine learning but have raised significant legal concerns due to their potential to produce text that infringes on copyrights.
Approach: They propose a lightweight, real-time defense mechanism to prevent the generation of copyrighted text by evaluating methods and testing attack strategies.
Outcome: The proposed defense significantly reduces the volume of copyrighted text generated by LLMs by effectively refusing malicious requests.
LLMs and Copyright Risks: Benchmarks and Mitigation Approaches (2025.naacl-tutorial)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized natural language processing, but their widespread use has raised significant copyright concerns.
Approach: This tutorial will provide an overview of relevant copyright principles and their application to AI and examine specific copyright issues in LLM development and deployment.
Outcome: The course will provide an overview of relevant copyright principles and their application to AI, followed by an examination of specific copyright issues in LLM development and deployment.
Do LLMs Know to Respect Copyright Notice? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on the occurrence of copyright violations in LLM output, but a negative answer would suggest that LLMs will become the primary facilitator and accelerator of copy right infringement behavior.
Approach: They propose to examine whether LLMs respect copyright information in user input . they use a set of language models, user prompts, and copyrighted materials .
Outcome: The proposed model will be the primary facilitator and accelerator of copyright infringement behavior, the study finds . the study also provides a benchmark dataset serving as a test bed for evaluating infringement behaviors by LLMs .
Copyright Violations and Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: a recent study examines the extent to which language models can memorize training data . a fair use exemption to copyright laws allows for limited use of copyrighted material .
Approach: They examine the extent to which language models can redistribute copyrighted text . they use a range of popular books and coding problems to study copyright violations .
Outcome: This study examines the extent to which language models can redistribute copyrighted text . it shows that language models may memorize entire chunks of training data .
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text (2024.eacl-long)

Copied to clipboard

Challenge: a recent study focused on detecting legal violations within unstructured textual data . a similar study focused only on associating violations with potentially affected individuals .
Approach: They constructed two datasets using Large Language Models (LLMs) they publicize the results to advance legal natural language processing research .
Outcome: The proposed datasets and the code used for the experiments have been released to advance legal natural language processing (NLP)
Certified Mitigation of Worst-Case LLM Copyright Infringement (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models are trained on vast datasets that include copyrighted material or content with usage restrictions.
Approach: They propose a "copyright takedown" method that provides certified copyright take down . they use a combination of inference-time and rewriting techniques to transform potentially infringing segments .
Outcome: The proposed method reduces infringement risk, preserves utility, and accommodates different levels of enforcement stringency with adaptive abstention.
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have been developed to deal with real-world crimes, but it remains unclear whether they internalize authentic knowledge or are forced to simulate toxic language patterns.
Approach: They construct knowledge-intensive Q&A to investigate misuse threats of Large Language Models in terms of dangerous knowledge possession, harmful task planning utility, and harmfulness judgment robustness.
Outcome: The findings raise concerns that jailbreak success is often attributable to a hallucination loop between jailbroken LLM and judger LLM .
Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: a new framework protects web content from unauthorized LLM real-time extraction and redistribution . multiple AI companies have been accused of scraping digital IP for proprietary benefit .
Approach: They propose a defense framework that empowers web content creators to safeguard their web-based IP from unauthorized LLM real-time extraction and redistribution by leveraging the semantic understanding capability of LLMs themselves.
Outcome: The proposed defense outperforms traditional defenses on LLMs and improves on black-box optimization problems.
On the Vulnerability of Safety Alignment in Open-Access LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are susceptible to malicious exploitation, but are often rejected and limited harmfulness is limited.
Approach: They propose two types of reverse alignment techniques: reverse supervised fine-tuning (RSFT) and reverse preference optimization (RPO).
Outcome: The proposed methods can significantly enhance the success rate and harmfulness of jailbreak attacks, but they face high rejection rates and limited harmfulness.
Jailbreak Open-Sourced Large Language Models via Enforced Decoding (2024.acl-long)

Copied to clipboard

Challenge: Existing studies show that Large Language Models can be misused to generate undesired content.
Approach: They propose to use large language models to manipulate the generation process to generate undesired content without heavy computations or prompt designs.
Outcome: The proposed method shows that open-sourced large language models could be misused to generate undesired content without heavy computations or prompt designs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations