Papers by Hua Shen

20 papers
A Survey of Large Language Models in Psychotherapy: Current Landscape and Future Directions (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) can handle extensive context and multi-turn reasoning.
Approach: They propose a taxonomy dividing psychotherapy into stages of assessment, diagnosis, and treatment to examine LLM advancements and challenges.
Outcome: The proposed taxonomy reveals imbalances in current research, such as a focus on common disorders, linguistic biases, fragmented methods, and limited theoretical integration.
REALM: A Dataset of Real-World LLM Use Cases (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on LLM adoption and their social implications lack empirical grounding, weakening their validity.
Approach: They propose to integrate a dataset of over 94,000 LLM use cases collected from Reddit and news articles to provide insights into LLM adoption across different domains.
Outcome: The proposed dataset includes over 94,000 LLM use cases collected from Reddit and news articles.
MultiTurnCleanup: A Benchmark for Multi-Turn Spoken Conversational Transcript Cleanup (2023.emnlp-main)

Copied to clipboard

Challenge: Disfluency detection models focus on individual utterances, but discontinuities in spoken transcripts occur across multiple turns.
Approach: They propose a multi-turn "cleanup task" to detect discontinuities in spoken conversations . they leverage two modeling approaches for experimental evaluation as benchmarks .
Outcome: The proposed task detects "discontinuities" in spoken conversations that can be removed . the results are compared with existing methods and are expected to be validated in the future .
Gentopia.AI: A Collaborative Platform for Tool-Augmented LLMs (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing frameworks for Augmented Language Models lack flexibility, democratization, and holistic evaluation.
Approach: They propose a lightweight and extensible framework for Augmented Language Models called Gentopia.
Outcome: The proposed framework integrates language models, task formats, prompting modules, and plugins into a unified paradigm.
A Multi-Agent Framework for High-Interaction Terminal Simulation (2026.acl-long)

Copied to clipboard

Challenge: Terminal simulation is a problem of symbolic language generation in dialogue and interactive systems.
Approach: They propose a terminal command-level Turing test framework that improves realism, consistency and robustness in command-language generation.
Outcome: The proposed framework outperforms state-of-the-art benchmarks by more than 9% on multi-turn terminal simulation.
Adaptive Rank Selections for Low-Rank Approximation of Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Singular Value Decomposition (SVD) or its weighted variants has progressed in compressing language models.
Approach: They propose a binary masking mechanism for optimizing the number of ranks in a differentiable framework.
Outcome: The proposed algorithm achieves much better accuracy than previous SVD and its weighted variants.
Improving Empathetic Dialogue Generation by Dynamically Infusing Commonsense Knowledge (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on generating empathetic responses by utilizing the speaker's emotion has not been successful.
Approach: They propose an approach which incorporates an adaptive module for commonsense knowledge selection to ensure consistency between the generated empathetic responses and the speaker’s situation.
Outcome: The proposed approach outperforms baseline models in both automatic and human evaluations, exhibiting the generation of more coherent and empathetic responses.
Agentic Episodic Control (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for reinforcement learning (RL) are limited by poor data efficiency and weak generalization.
Approach: They propose a novel architecture that integrates large language models into episodic RL.
Outcome: The proposed architecture achieves 2–6 higher data efficiency than baselines and is the only method to solve complex tasks like UnlockLocal with over 90% success.
VALUE ALIGNMENT TAX: Measuring Value Trade-offs in LLM Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Existing work on value alignment characterizes value relations statically, ignoring how interventions reshape the value system.
Approach: They propose a framework that quantifies value trade-offs by measuring how alignment-induced changes propagate across interconnected values relative to achieved on-target gain.
Outcome: The proposed framework measures how value trade-offs propagate across values . it can be used to evaluate intended improvements and unintended side effects .
You Never Know a Person, You Only Know Their Defenses: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations (2026.findings-acl)

Copied to clipboard

Challenge: Psychological defenses are strategies people use to manage distress.
Approach: They propose a dialogue corpus with help seeker utterances labeled for defense level and a DMRS Co-Pilot pipeline that provides evidence-based pre-annotations.
Outcome: The proposed framework reduces annotation time by 24.0% in a counterbalanced study.
Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference (2025.emnlp-main)

Copied to clipboard

Challenge: Reasoning-focused large language models (LLMs) are rapidly evolving across various domains, yet their capabilities in handling complex legal problems remain underexplored.
Approach: They propose a large language model tailored for legal reasoning with a 7-billion parameter scale and a two-stage training strategy combining Supervised Fine-Tuning and Reinforcement Learning.
Outcome: The proposed model outperforms all models of similar scale on authoritative benchmarks and outperformed Qwen-2.5-7B-Instruct (46.6%) by an average margin of 6.6%.
Numerical Optimizations for Weighted Low-rank Estimation on Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Singular value decomposition (SVD) is one of the most popular methods for estimating a target matrix with smaller matrices.
Approach: They propose a method that approximates a target matrix with smaller matrices by two smaller . they also propose metric to predict when the SVD may introduce a significant performance drop.
Outcome: The proposed method can perform better than current SOTA methods in compressing Transformer-based language models.
Are Shortest Rationales the Best Explanations for Human Understanding? (2022.acl-short)

Copied to clipboard

Challenge: Existing models favor extracting the shortest possible rationales to explain model predictions . however, this assumption has yet to be validated .
Approach: They propose a model that extracts rationales at any target length from text inputs . they show that rationale lengths too short do not help humans predict labels better .
Outcome: The proposed model achieves compatible end-task performance and human-annotated rationale agreement compared to baseline models .
Hyperparameter-free Continuous Learning for Domain Classification in Natural Language Understanding (2021.naacl-main)

Copied to clipboard

Challenge: Existing continual learning approaches suffer from low accuracy and performance fluctuation when the distributions of old and new data are significantly different.
Approach: They propose a hyperparameter-free continual learning model for text data that can stably produce high performance under various environments.
Outcome: The proposed model outperforms the best state-of-the-art method by 20% in average accuracy and each component contributes effectively to overall performance.
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has sought to use large language models to simulate human-human and human-LLM interactions.
Approach: They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts.
Outcome: The proposed models perform similarly in simulating English, Chinese, and Russian dialogues.
SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL (2026.acl-long)

Copied to clipboard

Challenge: Recent large language models (LLMs) have significantly improved Text-to-SQL generation, but a gap remains between AI systems and human experts on challenging benchmarks such as BIRD-Sql.
Approach: They propose a multi-turn reinforcement learning agentic framework for Text-to-SQL that uses execution feedback to iteratively refine its predictions.
Outcome: The proposed framework outperforms proprietary systems on 7B and 14B models by **5% on average, underscoring the effectiveness of interactive, agentic workflows for robust Text-to-SQL generation.
Dynamic Low-rank Estimation for Transformer-based Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: RankDyna is a matrix decomposition method that can be used to compress Transformer-based language models.
Approach: They propose a matrix decomposition method that enables dynamic rank resource allocation . they say it can outperform current SOTA methods under various parameter budget levels .
Outcome: The proposed method outperforms current SOTA methods under various budget levels . the proposed method is more efficient with higher compression rates .
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing research assesses LLMs’ values by analyzing their stated inclinations . a framework to evaluate the alignment between stated values and value-informed actions is lacking .
Approach: They propose a framework to evaluate the alignment between LLMs’ stated values and their value-informed actions.
Outcome: The proposed framework shows significant misalignment between LLM-generated values and their actions . misaligned values have shown real-world risks, such as amplifying stereotypes and reinforcing bias algorithms in hiring.
IAEval: A Comprehensive Evaluation of Instance Attribution on Natural Language Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: Instance attribution (IA) aims to identify the training instances leading to the prediction of a test example.
Approach: They propose a systematic and comprehensive evaluation scheme covering four significant requirements: sufficiency, completeness, stability and plausibility.
Outcome: The proposed evaluation scheme covers four significant requirements: sufficiency, completeness, stability and plausibility.
Causally Modeling the Linguistic and Social Factors that Predict Email Response (2025.naacl-long)

Copied to clipboard

Challenge: a key intent behind many emails is to get a reply from the recipient.
Approach: They propose to model the intents, expectations, and responsiveness in email exchanges by using a dataset containing 1800 emails annotated with nuanced types of intents and expectations.
Outcome: The proposed model is based on 1800 emails annotated with nuanced types of intents and expectations . it shows that social status, argumentation, and strength of social connection influence email response rates .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations