Papers by Kaijie Mo

4 papers
ExpertEase: A Multi-Agent Framework for Grade-Specific Document Simplification with Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies mainly focus on sentence-level simplification, neglecting document-level and the different reading levels of target audiences.
Approach: They propose a multi-agent framework for grade-specific document simplification using Large Language Models that integrates expert, teacher, and student agents that cooperate on the task and rely on external tools for calibration.
Outcome: The proposed framework significantly improves the performance of large language models and compares them with human-authored texts.
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)

Copied to clipboard

Challenge: Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models.
Approach: They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Outcome: The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence (2026.findings-acl)

Copied to clipboard

Challenge: Existing models are overwhelmingly accurate when presented with counterfactual medical evidence . prior work explored conflicts between context and LLM parametric knowledge in the general domain .
Approach: They construct a counterfactual medical QA dataset that requires models to answer clinical comparison questions with evidence from randomized controlled trials.
Outcome: The proposed model overemphasizes the latter, and the model overestimates the latter.
CMT-Eval: A Novel Chinese Multi-turn Dialogue Evaluation Dataset Addressing Real-world Conversational Challenges (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation benchmarks fail to capture users’ evolving needs and how their diverse conversation styles affect the dialogue flow.
Approach: They propose to use CMT-Eval to evaluate Chinese multi-turn dialogue systems.
Outcome: The proposed dataset is the first dedicated dataset for fine-grained evaluation of Chinese multi-turn dialogue systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations