Papers by Forough Poursabzi-Sangdeh
Aligning Offline Metrics and Human Judgments of Value for Code Generation Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Large language models have shown impressive capabilities on code generation tasks. |
| Approach: | They propose a metric that combines functional correctness and syntactic similarity to measure the productivity gains generated by large language models. |
| Outcome: | The proposed model achieves a 14% stronger correlation with value and better represents real-world gains when evaluating and comparing models. |