Papers by Robie Gonzales
Long-form evaluation of model editing (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of model editing only use the ‘next few tokens’ completions after a prompt. |
| Approach: | They propose a new evaluation protocol that measures the efficacy and impact of model editing in long-form generative settings by using a machine-rated survey and a classifier which correlates well with human ratings. |
| Outcome: | The proposed evaluation protocol has little relationship with short-form metrics despite being designed to extend efficacy, generalization, locality, and portability into a long-form setting. |