Papers with MMMLU

2 papers
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation (2026.eacl-long)

Copied to clipboard

Challenge: Recent studies show that data quality can significantly boost performance and training efficiency for large language models.
Approach: They propose a German-language dataset curation pipeline that combines heuristic and model-based filtering techniques with synthetic data generation.
Outcome: The proposed pipeline can be used to create a large-scale German pre-training dataset using common Crawl web data, fineweb2 and synthetically generated data conditioned on real, organic web data.
Layer-wise Swapping for Generalizable Multilingual Safety (2026.eacl-long)

Copied to clipboard

Challenge: Existing safety datasets are predominantly English-centric, limiting progress in multilingual safety alignment.
Approach: They propose a safety-aware layer swapping method that transfers alignment from an English safety expert to low-resource language experts without additional training.
Outcome: The proposed method preserves performance on general language understanding tasks while enhancing safety in the target languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations