Papers by Ilya Galyukshev
Processing Inconsistency Predicts Language Competence: LLM Evaluation Without Answer Labels on Turkic Languages (2026.acl-srw)
Copied to clipboard
| Challenge: | Most languages lack labeled evaluation benchmarks for large language models (LLMs). Creating labeles requires native speakers, domain expertise, and answer annotation. |
| Approach: | They hypothesize that a model's internal processing signals correlate with its actual accuracy on a language . they extract over 25 processing features per model–language pair and test them . |
| Outcome: | The proposed model outperforms the model's English/Russian benchmark score on 11 instruction-tuned LLMs across 14 language–script varieties. |