Papers by Renfen Hu
ExpertEase: A Multi-Agent Framework for Grade-Specific Document Simplification with Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies mainly focus on sentence-level simplification, neglecting document-level and the different reading levels of target audiences. |
| Approach: | They propose a multi-agent framework for grade-specific document simplification using Large Language Models that integrates expert, teacher, and student agents that cooperate on the task and rely on external tools for calibration. |
| Outcome: | The proposed framework significantly improves the performance of large language models and compares them with human-authored texts. |
Diachronic Sense Modeling with Deep Contextualized Word Embeddings: An Ecological View (P19-1)
Copied to clipboard
| Challenge: | Existing word embeddings only assign one vector to a word for a time period, thus they face the meaning conflation deficiency. |
| Approach: | They propose a sense representation and tracking framework based on deep contextualized embeddings that can be used to answer what and when the word meaning changes. |
| Outcome: | The proposed framework is effective in representing fine-grained word senses, and brings a significant improvement in word change detection task. |
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models. |
| Approach: | They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
| Outcome: | The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies evaluate the creative capabilities of large language models (LLMs) through diverse tasks, aiming to understand their strengths and limitations. |
| Approach: | They propose to ask LLMs to generate Parallel Chains of Associations to Evaluate their creativity. |
| Outcome: | The proposed framework minimizes the risk of data contamination and offers a highly efficient evaluation. |
CMT-Eval: A Novel Chinese Multi-turn Dialogue Evaluation Dataset Addressing Real-world Conversational Challenges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation benchmarks fail to capture users’ evolving needs and how their diverse conversation styles affect the dialogue flow. |
| Approach: | They propose to use CMT-Eval to evaluate Chinese multi-turn dialogue systems. |
| Outcome: | The proposed dataset is the first dedicated dataset for fine-grained evaluation of Chinese multi-turn dialogue systems. |
Analogical Reasoning on Chinese Morphological and Semantic Relations (P18-2)
Copied to clipboard
| Challenge: | Analogical reasoning is effective in capturing linguistic regularities. |
| Approach: | They propose to use Chinese lexical knowledge to build an analogical reasoning task using a large dataset. |
| Outcome: | The proposed dataset proves to be reliable benchmark for evaluating Chinese word embeddings. |
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing Automated Essay Scoring (AES) methods focus on sentence-level features, whereas Large Language Models (LLMs) are sensitive to conventions & accuracy, language complexity, and organization. |
| Approach: | They propose to use large language models to aid in decision-making . they propose to analyze the reasoning of neural models by analyzing sentence-level features. |
| Outcome: | The proposed method improves understanding of neural approaches to Automated Essay Scoring (AES) and can also apply to other domains seeking transparency in model-driven decisions. |