Papers by Nicolas Heess
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models (2024.acl-short)
Copied to clipboard
| Challenge: | In many applications of ML systems it is important to understand why the system came to a particular answer. |
| Approach: | They introduce a faithfulness metric based on counterfactual input edits that takes into account not just the binary label change, but the total shift in the model’s predicted label distribution. |
| Outcome: | The proposed explanations are more likely to mention factors when they are impactful to the model’s prediction, with the degree of association increasing with model size but varying significantly by task. |