Challenge: Large language models are trained on vast amounts of data, which may unintentionally or intentionally include data from commonly used benchmarks.
Approach: They propose a set of requirements that practical contamination detection methods should follow to effectively detect benchmark contamination in large language models.
Outcome: The proposed method detects whether the model is significantly more confident under the original benchmark.

Similar Papers

Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for detecting data contamination in large language models have limitations and limitations . data contamination occurs when test or evaluation data is exposed to the model during its training phases .
Approach: They evaluate five different methods for detecting data contamination in large language models . they find that current methods have non-trivial limitations in their assumptions and practical applications .
Outcome: The proposed methods have non-trivial limitations and difficulties in detecting contamination . the authors highlight the complexity of contamination detection in advanced LLMs .
Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models have demonstrated great performance across various benchmarks, but data contamination is a concern in their evaluation.
Approach: They analyze 50 papers on data contamination detection and test three of them as case studies to identify the possibility of data contamination.
Outcome: The proposed methods can detect membership inference attacks on instance-level data, and can perform similar to random guessing on LLM pretraining datasets.
Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Considering the vast size and wide-ranging sources of LLMs’ training data, it could explicitly or implicitly include test data.
Approach: They propose a Contamination Detection via output Distribution (CDD) which detects data contamination only by identifying the peakedness of LLM's output distribution.
Outcome: The proposed method improves performance by 21.8%-30.2% on humanEval and TED: trustworthy evaluation via output distribution.
Investigating Data Contamination in Modern Benchmarks for Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing evaluation benchmarks for large language models are inflated and inconsistent with actual performance.
Approach: They propose a retrieval-based system to explore potential overlaps between benchmarks and pretraining corpora and a protocol to investigate testset slot guessing.
Outcome: The proposed method exploits overlaps between evaluation benchmarks and pretraining corpora and masks a wrong answer in a multiple choice question and prompts the model to fill in the gap.
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for evaluating large language models using annotated benchmarks are in trouble . data contamination can cause wrong scientific conclusions being published .
Approach: They argue that the evaluation of NLP tasks using annotated benchmarks is in trouble . they define different levels of data contamination and propose a community effort .
Outcome: The proposed measures should detect when data from a benchmark was exposed to a model and flag papers with conclusions compromised by data contamination.
Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: In the era of evaluating large language models, data contamination is an increasingly prominent concern . static benchmarking has been used for evaluation, but there are limitations of *dynamic* benchmarks .
Approach: They propose a series of optimal design principles for *dynamic* benchmarking and analyze the limitations of existing *static* benchmarks.
Outcome: The proposed benchmarks highlight a critical gap in the evaluation of LLMs.
Data Contamination Can Cross Language Barriers (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to detect contamination of public benchmarks are too superficial to reflect deeper forms of contamination.
Approach: They propose generalization-based approaches to unmask a cross-lingual form of contamination that inflates LLMs’ performance while evading current detection methods.
Outcome: The proposed model outperforms existing detection methods while avoiding contamination of public benchmarks in the pre-training data.
Unveiling the Spectrum of Data Contamination in Language Model: A Survey from Detection to Remediation (2024.findings-acl)

Copied to clipboard

Challenge: Data contamination is a problem in Large language models due to the reliance on extensive internet-derived training corpora.
Approach: They present a survey on the topic of data contamination in large language models.
Outcome: The results of the first survey on data contamination in large language models provide a comprehensive guide for NLP researchers seeking a systematic understanding of the issue.
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models (2024.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that large language models are contaminated with data from pretraining and finetuning tasks.
Approach: They perform extensive analysis on the factors that affect model memorization and generalization, such as model size, problem difficulty, and question length.
Outcome: The results show that models perform better on the subset of the benchmarks where similar solutions are seen during training.
Data Contamination Calibration for Black-box LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Despite the rapid advancements of Large Language Models, the unchecked ultra-large-scale training sets introduce a series of potential risks like data contamination.
Approach: They propose a method to detect contaminated training data and diminish the contamination effect by using a to-be-released dataset.
Outcome: The proposed method outperforms existing methods by at least 4.5% on more 4 dataset formats, with more than 10 base LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations