Papers by Qingchen Yu

2 papers
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Climate change adaptation requires the understanding of disruptive weather impacts on society.
Approach: They propose a large language model to evaluate the capacity of LLMs on disruptive weather impacts by using a four-stage construction pipeline.
Outcome: The proposed model is based on a four-stage well-crafted construction pipeline and requires two evaluation tasks, multi-label classification and ranking-based question answering.
GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods for large language models rely on static benchmarks and standardized evaluation protocols.
Approach: They propose an adaptive evaluation framework that integrates dynamic domain knowledge modeling with progressive reasoning assessment to improve evaluation fidelity.
Outcome: Empirical results show that the framework distinguishes LLMs in terms of domain knowledge coverage and reasoning chain completeness.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations