Papers by Hanyu Zhou

3 papers
Measuring Psychological Depth in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Current evaluations of creative stories focus on objective properties of the text, such as its style, coherence, diversity, and creativity.
Approach: They propose a framework that measures an LLM's ability to produce authentic and narratively complex stories that provoke emotion, empathy, and engagement.
Outcome: The proposed framework shows that humans can consistently evaluate stories based on the PDS (0.72 Krippendorff’s alpha).
PRiSM: Benchmarking Phone Realization in Speech Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of phone recognition systems only measure surface-level transcription accuracy.
Approach: They propose to standardize transcription-based evaluation and assess downstream utility in clinical, educational, and multilingual settings with transcription and representation probes.
Outcome: The proposed system outperforms LALMs in clinical, educational, and multilingual settings.
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on how SAEs derive most fine-grained latent features for safety remain unexplored.
Approach: They propose a framework for interpreting SAE features in safety-critical domains . they train a suite of SAEs with human-readable explanations and systematic evaluations based on pornography, politics, violence, and terror .
Outcome: The proposed framework reduces interpretation cost by 55% and improves safety-critical features.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations