Papers by Matthias Bethge
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities (2025.acl-long)
Copied to clipboard
| Challenge: | ONEBench enables custom benchmarks for specific capabilities while reusing and aggregating samples. |
| Approach: | They propose a new paradigm that consolidates individual evaluation datasets into a unified, ever-expanding sample pool. |
| Outcome: | The proposed model evaluation framework is based on dynamic, sample-level evaluation. |
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve? (2024.emnlp-main)
Copied to clipboard
| Challenge: | In the last decade, the generalization and adaptation abilities of deep learning models were evaluated on fixed training and test distributions. |
| Approach: | They propose to train large language models on unlabeled text corpora and train them online. |
| Outcome: | The proposed model training on a text domain could degrade its perplexity on the test portion of the same domain. |