Papers by Yongan Yu

2 papers
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation (2026.acl-long)

Copied to clipboard

Challenge: Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components with iterative reuse of existing codes.
Approach: They propose a benchmark to evaluate LLMs' ability to perform codeflow by reusing existing functions over multiple turns.
Outcome: The proposed benchmarks show that LLMs perform significantly worse in multi-turn codeflow scenarios and that their performance inversely correlates with dependency complexity.
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Climate change adaptation requires the understanding of disruptive weather impacts on society.
Approach: They propose a large language model to evaluate the capacity of LLMs on disruptive weather impacts by using a four-stage construction pipeline.
Outcome: The proposed model is based on a four-stage well-crafted construction pipeline and requires two evaluation tasks, multi-label classification and ranking-based question answering.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations