Papers by Gehao Zhang
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks emphasize correctness under limited evaluation settings . evaluation of formal specifications is time-consuming, errorprone and requires substantial expertise. |
| Approach: | They propose a multilingual benchmark for evaluating method-level postcondition generation from real-world software. |
| Outcome: | The proposed benchmarks show that evaluation remains a key bottleneck . 420 Python and Java tasks are paired with a high-quality postcondition set . |