Papers by Lizhi Ma
From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation metrics for large language models yield numerical scores that ignore user experience. |
| Approach: | They propose a metric that suggests revision edits that mimic the human writing process . their results show that the metric offers more insightful feedback and distinguishes between texts . |
| Outcome: | The proposed metric can provide a self-explained text evaluation result in a human-understandable manner beyond the context-independent score. |
Understanding Client Reactions in Online Mental Health Counseling (2023.acl-long)
Copied to clipboard
| Challenge: | Communication success relies heavily on reading participants’ reactions, but little research is on how listeners' reactions shape trajectories and outcomes of conversations. |
| Approach: | They propose to use client reactions to predict counseling outcomes by using an annotation framework that encompasses counselors’ strategies and client reaction behaviors. |
| Outcome: | The proposed framework can predict counselors' strategies and client reaction behaviors against a large-scale text-based counseling dataset. |
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning (2026.eacl-long)
Copied to clipboard
| Challenge: | Long-context understanding is a critical capability for large language models . evaluating this capability requires extensive human annotation, which is time-consuming and costly. |
| Approach: | They propose a benchmark to assess citation-grounded long-context reasoning in academic writing. |
| Outcome: | The proposed benchmark compares state-of-the-art models with human experts on two tasks . human experts achieve 90% accuracy, but most models struggle with the cloze-style task . |
PsyGUARD: An Automated System for Suicide Detection and Risk Assessment in Psychological Counseling (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing systems for fine-grained suicide detection and risk assessment are lacking . a lack of domain-specific systems for this task poses a challenge to automated crisis intervention aimed at suicide prevention. |
| Approach: | They propose to use a fine-grained suicide detection system to assess risk in counseling . they develop a taxonomy for detecting suicide ideation and a large-scale dataset . |
| Outcome: | The proposed system detects suicidal ideation and assesses risk in counseling . it can provide safe, helpful, and tailored responses for further assessment . |
Understanding the Therapeutic Relationship between Counselors and Clients in Online Text-based Counseling using LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | In traditional face-to-face therapy, the assessment of therapeutic alliance is not directly translated to text-based settings. |
| Approach: | They propose an automatic approach to understand the development of therapeutic alliance in text-based counseling by using large language models. |
| Outcome: | The proposed approach demonstrates that the framework is effective in identifying the therapeutic alliance in text-based counseling. |