Papers by Yunsung Kim
Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments (2026.findings-acl)
Copied to clipboard
| Challenge: | Despite increasing demand for transparency and interpretability, the field has yet to develop a widely accepted solution for interpretable automated scoring to be used in large-scale real-world assessments. |
| Approach: | They propose to develop four principles of interpretability targeted at assessment stakeholder groups to address the need for transparency and interpretability in automated scoring. |
| Outcome: | The proposed framework outperforms many uninterpretable scoring methods in terms of scoring accuracy and is, on average, within 0.06 QWK of the uninterprétable SOTA. |
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact (2026.acl-long)
Copied to clipboard
| Challenge: | a recent study shows that large language models excel on benchmarks that operationalize knowledge. |
| Approach: | They compare LLM alignment on benchmarks, downstream tasks and intended impact . they find that inter-model behaviors on disparate tasks correlate higher than expert human behaviors on target tasks . |
| Outcome: | The proposed methods show that LLMs perform poorly on learning tasks . the results show that they are poorly aligned with downstream measures of teaching quality . |