Papers by Minseo Kwak
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data (2026.acl-long)
Copied to clipboard
| Challenge: | Existing state-of-the-art methods for pretraining data are largely undisclosed, resulting in ethical and copyright concerns. |
| Approach: | They propose a method that leverages the log probability gap between the top-1 predicted token and the target token, incorporating a sliding window strategy to capture local correlations and mitigate token-level fluctuations. |
| Outcome: | The proposed method outperforms baselines on WikiMIA and MIMIR benchmarks and achieves state-of-the-art performance. |