Papers by Adam Zahradník
Trojsten Benchmark: Evaluating LLM Problem-Solving in Slovak STEM Competition Problems (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have been used for grading open-ended responses and providing feedback beyond traditional methods. |
| Approach: | They propose a Slovak-language dataset and a rubric-based LLM grading framework . they quantify multistep reasoning performance by difficulty and show consistency under difficult items . |
| Outcome: | The proposed model outperforms existing models on Slovak-language competition problems . the model shows consistent underperformance on harder items and language sensitivity . |