Papers by Mingshuo Wang
ChangJuan: A Comprehensive Benchmark for Book-Length Chinese Story Evaluation (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly enhanced the capacity of Automatic Story Evaluation. |
| Approach: | They propose a method to distill raw reviews into generally agreed viewpoints across key evaluation aspects such as plot and character. |
| Outcome: | The proposed model outperforms open-source baselines and raises Qwen3’s Kendall’s tau correlation with human judgments from 24.8 to 34.1. |