Papers by Yunze Liu

5 papers
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmark datasets focus on short to moderately long videos, leaving a substantial gap in evaluating extensive, ultra-long egocentric video recordings.
Approach: X-LeBench is a benchmark dataset designed to evaluate long egocentric video recordings . it uses a life-logging pipeline to produce realistic, coherent daily plans .
Outcome: X-LeBench is a new benchmark dataset designed to evaluate long-form egocentric video understanding . the approach produces realistic, coherent daily plans aligned with real-world video data .
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics (2025.emnlp-main)

Copied to clipboard

Challenge: a study of multi-dimensional persona effects in AI-AI debates shows that personas influence moral stances and debate outcomes . political ideology and personality traits exert the strongest influence, according to our study .
Approach: They propose to use a 6-dimensional persona space to simulate structured debates . they find political ideology and personality traits exert the strongest influence .
Outcome: The study shows that personas affect moral stances and debate outcomes . political ideology and personality traits exert the strongest influence .
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language model evaluation benchmarks focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities.
Approach: They propose a comprehensive benchmark covering 29 languages, built on an English benchmark.
Outcome: The MMLU-ProX is a comprehensive benchmark covering 29 languages, built on an English benchmark.
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on single agentic capability, failing to capture long-horizon real-world scenarios.
Approach: They propose a benchmark that evaluates 6 agentic capabilities across 32 real-world scenarios.
Outcome: Experiments show that closed-source models outperform open-source model (48.4% vs 32.1%) integrating models with advanced scaffolds to form autonomous agents is a paradigm shift.
Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit anthropomorphism characteristics – human-like qualities portrayed across their outlook, language, behavior, and reasoning functions.
Approach: They propose that anthropomorphism should be treated as a design concept that can be intentionally tuned to support user goals.
Outcome: The proposed design should reflect interaction between artifact designers and interpreters, and should be based on cues embedded in the artifactor and the (cognitive) responses of interpreters to the cue.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations