Papers by Tianshuo Wei
Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom’s Taxonomy (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods suffer from cognitive dimensional simplification and methodological unreliability due to the ”LLM-as-a-Judge” approach. |
| Approach: | They propose a six-tiered benchmark that evaluates ASG systems by prioritizing deterministic algorithms and introducing a GRADE approach for abstract abilities. |
| Outcome: | The proposed method provides the ASG field with a systematic, reproducible, and theoretically grounded benchmark to guide future research. |
Feedback Is The Key for Automated Survey Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) provide a promising foundation for literature surveys, but guiding them to generate accurate, reliable content remains a fundamental challenge. |
| Approach: | They propose a feedback-driven framework that incorporates feedback across three dimensions: outline feedback for structural clarity, citation feedback for evidence validation, and content feedback for readability and analytical depth. |
| Outcome: | The proposed framework significantly improves both citation and content quality, demonstrating feedback as the critical mechanism for automatic survey generation. |