Challenge: a lack of transparency in large language models makes auditing their "digital DNA" difficult.
Approach: They propose a framework that casts DMS as an inverse problem under label-shift assumption . they propose LLMScan, a recipe-verifiable evaluation suite built from open-source LLMs .
Outcome: The proposed framework casts DMS as an inverse problem under label-shift assumption . compared with existing frameworks, it recovers domain mixtures with high fidelity .

Similar Papers

Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models have demonstrated great performance across various benchmarks, but data contamination is a concern in their evaluation.
Approach: They analyze 50 papers on data contamination detection and test three of them as case studies to identify the possibility of data contamination.
Outcome: The proposed methods can detect membership inference attacks on instance-level data, and can perform similar to random guessing on LLM pretraining datasets.
SampleMix: A Sample-wise Pre-training Data Mixing Strategy by Coordinating Data Quality and Diversity (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for pretraining data mixing for large language models neglect significant inter-domain overlaps and commonalities, failing to control the global diversity of the constructed training dataset.
Approach: They propose a sample-wise data mixture approach that performs global cross-domain sampling by systematically evaluating the quality and diversity of each sample.
Outcome: The proposed method exceeds existing domain-based methods in multiple downstream tasks and perplexity assessments.
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation.
Approach: They propose a generic workflow for LLM-driven synthetic data generation.
Outcome: The proposed workflows highlight gaps in existing research and outline avenues for future studies.
LLM×MapReduce: Simplified Long-Sequence Processing using Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on extending the context length of large language models (LLMs) due to their quadratic computational complexity and a lack of high-quality long training examples, most LLMs are trained with a limited window size.
Approach: They propose a training-free framework that enables large language models to effectively process long texts using a divide-and-conquer strategy for comprehensive document understanding.
Outcome: The proposed framework outperforms open-source and commercial long-context LLMs and is compatible with several models.
DPDLLM: A Black-box Framework for Detecting Pre-training Data from Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods to detect pretraining data from large language models are unrealistic to them.
Approach: They propose to detect pre-training data from LLM in a black-box way by using GPT-2 as reference model and feed it with sequence probabilities to detect whether it was used to train it.
Outcome: The proposed framework outperforms existing methods on the benchmark datasets and shows that it is effective on different popular LLMs.
Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for data mixture improve the generalization capability of large language models (LLMs) on downstream tasks.
Approach: They propose a fine-grained categorization of existing methods and propose three subtypes of offline and online methods.
Outcome: The proposed methods extend beyond offline and online classifications and highlight key challenges in the field of data mixture.
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMđ›„ Integration into Upcycled MoE (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are expensive and require extensive Continued Pre-Training and data-intensive alignment to expand.
Approach: They propose a method which upcycles a dense model into a Mixture-of-Experts architecture, allocating different experts to different languages.
Outcome: Experiments show that the proposed model upcycles a dense model into a Mixture-of-Experts(MoE) architecture, allocating different experts to different languages.
Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable advancements in specialized fields such as finance, law, and medicine.
Approach: They propose to provide datasets covering all major training stages including pretraining, instruction fine-tuning, and reasoning distillation with cybersecurity-specific self-reflection data.
Outcome: Extensive ablation studies show that LLMs acquire their knowledge during pretraining, while reasoning distillation leads to a 15% gain in security certification (CISSP).
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to train and test large language models that involve calls to tools and APIs are lacking.
Approach: They propose a large corpora for training and systematic testing of tool-augmented LLMs.
Outcome: The proposed datasets mimic real-world scenarios involving API-tasks and slot filling.
The Data Frontier for Large Language Models: Selection, Synthesis, and Tools (2026.acl-tutorials)

Copied to clipboard

Challenge: acquiring and curating high-quality training data remains a significant bottleneck . acquiring such high-quality data is a key challenge for researchers and practitioners .
Approach: This tutorial provides a comprehensive and practical guide to the state-of-the-art in data research directions for LLMs.
Outcome: The tutorial covers methods for curating the most valuable information from vast, noisy datasets and the synthetic data revolution.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations