Papers by Ali Abdi
Interpretability of LLM Classifiers via the Rational Inattention Theory with Application to Hate Speech Detection (2026.acl-srw)
Copied to clipboard
| Challenge: | Large language models (LLMs) perform well on text classification, but their decision strategies need to be better understood. |
| Approach: | They propose an extended rational inattention model that parameterizes linguistic noise and information processing cost and provides an interpretable behavioral framework for black-box LLM classifiers. |
| Outcome: | The proposed model provides an interpretable behavioral framework for black-box LLM classifiers. |