Papers by Kaijie Mo
ExpertEase: A Multi-Agent Framework for Grade-Specific Document Simplification with Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies mainly focus on sentence-level simplification, neglecting document-level and the different reading levels of target audiences. |
| Approach: | They propose a multi-agent framework for grade-specific document simplification using Large Language Models that integrates expert, teacher, and student agents that cooperate on the task and rely on external tools for calibration. |
| Outcome: | The proposed framework significantly improves the performance of large language models and compares them with human-authored texts. |
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models. |
| Approach: | They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
| Outcome: | The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence (2026.findings-acl)
Copied to clipboard
Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C Wallace, Junyi Jessy Li
| Challenge: | Existing models are overwhelmingly accurate when presented with counterfactual medical evidence . prior work explored conflicts between context and LLM parametric knowledge in the general domain . |
| Approach: | They construct a counterfactual medical QA dataset that requires models to answer clinical comparison questions with evidence from randomized controlled trials. |
| Outcome: | The proposed model overemphasizes the latter, and the model overestimates the latter. |
CMT-Eval: A Novel Chinese Multi-turn Dialogue Evaluation Dataset Addressing Real-world Conversational Challenges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation benchmarks fail to capture users’ evolving needs and how their diverse conversation styles affect the dialogue flow. |
| Approach: | They propose to use CMT-Eval to evaluate Chinese multi-turn dialogue systems. |
| Outcome: | The proposed dataset is the first dedicated dataset for fine-grained evaluation of Chinese multi-turn dialogue systems. |