Papers by Jiayin Wang
A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks focus on specific predefined model abilities, such as world knowledge, reasoning, etc., making it difficult for users to determine which LLM best suits their particular needs. |
| Approach: | They propose to evaluate large language models from a user-centric perspective and use real-world use cases to identify their effectiveness under distinct intents. |
| Outcome: | The proposed benchmarks achieve a correlation between human preference and the user-reported scenarios and human intents. |
Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal large language models are limited to multiview diagnostics. |
| Approach: | They propose a benchmark specifically designed for medical multi-image understanding that evaluates MLLMs across four dimensions. |
| Outcome: | The proposed model performs better in multi-image contexts than open-source models . the model perform better when processing increased visual loads than closed-source ones . |
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks primarily assess static knowledge, while intelligence also entails the ability to rapidly learn from experience. |
| Approach: | They propose to use semantic games to evaluate test-time learning . they recruit eight human participants to complete the same task . |
| Outcome: | The proposed framework compares model performance under limited and cumulative experience settings and contains four forms of experience representation. |
PATIENT-π: Using Large Language Models to Simulate Patients for Training Mental Health Professionals (2024.emnlp-main)
Copied to clipboard
Ruiyi Wang, Stephanie Milani, Jamie Chiu, Jiayin Zhi, Shaun Eack, Travis Labrum, Samuel Murphy, Nev Jones, Kate Hardy, Hong Shen, Fei Fang, Zhiyu Chen
| Challenge: | Mental illness remains one of the most critical public health issues. |
| Approach: | They propose a patient simulation framework for cognitive behavior therapy training that uses large language models to act as a simulated therapy patient. |
| Outcome: | The proposed framework improves the skill acquisition and confidence of mental health trainees beyond textbooks, videos, and role-play with non-patients. |
RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation (2025.findings-emnlp)
Copied to clipboard
Tianjiao Li, Mengran Yu, Chenyu Shi, Yanjun Zhao, Xiaojing Liu, Qi Zhang, Xuanjing Huang, Qiang Zhang, Jiayin Wang
| Challenge: | Using reinforcement learning from human feedback, large language models perform poorly when applied to colloquial subtitle translation tasks. |
| Approach: | They propose an adversarial training framework that iteratively updates the offline reward model and the online LLM to improve training outcomes. |
| Outcome: | The proposed training framework significantly improves upon translation baselines. |
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Climate change adaptation requires the understanding of disruptive weather impacts on society. |
| Approach: | They propose a large language model to evaluate the capacity of LLMs on disruptive weather impacts by using a four-stage construction pipeline. |
| Outcome: | The proposed model is based on a four-stage well-crafted construction pipeline and requires two evaluation tasks, multi-label classification and ranking-based question answering. |