Papers by Yuanyang Liu

2 papers
PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning.
Approach: They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios.
Outcome: The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics.
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for AI math tutoring largely overlook these skills.
Approach: They evaluate 12 leading multimodal large language models and find clear performance gaps between them.
Outcome: The proposed benchmarks show that they can solve 770 problems and provide diagnostics and guidance to students step by step.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations