Papers by Kehan Guo

7 papers
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark (2024.acl-short)

Copied to clipboard

Challenge: SceMQA focuses on core science subjects including Mathematics, Physics, Chemistry, and Biology.
Approach: They propose to use SceMQA to evaluate multimodal question answering at college entrance level.
Outcome: The proposed model provides specific knowledge points for each problem and detailed explanations for each answer.
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used to remove harmful knowledge and undesirable capabilities.
Approach: They propose a framework that leverages Cognitive Diagnosis Modeling to evaluate LLM unlearning.
Outcome: The proposed framework enhances evaluation and facilitates removal of harmful abilities.
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs? (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in LLMs unlearning have shown remarkable success in removing unwanted data-model influences while preserving the model’s utility for legitimate knowledge.
Approach: They propose a Selected-Expert Unlearning Framework (SEUF) that combines expert attribution and an anchor loss to ensure controlled unlearning.
Outcome: Experiments show that the proposed framework improves forget quality and model utility by 35% on MoE LLMs across benchmarks and LLM architectures compared to standard unlearning algorithms .
Compositional Mathematical Encoding for Math Word Problems (2023.findings-acl)

Copied to clipboard

Challenge: Existing MWP encoders work in a unimodal setting and map problem description to latent representation, then for decoding.
Approach: They propose a Compositional Math Word Problem Solver which maps problem description to latent representation and decodes it in an interactive way.
Outcome: Extensive experiments show that the proposed model outperforms state-of-the-art models on public benchmarks.
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly integrated into real-world decision-making, but their ability to comprehend and reason about policy-related content remains underexplored.
Approach: They propose a bilingual benchmark evaluating policy comprehension comprising 21K cases across a broad spectrum of policy areas.
Outcome: The proposed model shows stronger performance on application-oriented policy tasks than on memorization or conceptual understanding, and yields the highest accuracy on structured reasoning tasks.
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks that rely on final-answer accuracy fail to capture the quality of the reasoning process.
Approach: They propose a fine-grained evaluation framework that assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing.
Outcome: The proposed framework assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing.
Defending Jailbreak Prompts via In-Context Adversarial Game (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) demonstrate remarkable capabilities across diverse applications, but concerns regarding their security persist.
Approach: They propose an adversarial game that leverages agent learning to extend knowledge to defend against jailbreaks.
Outcome: The proposed game shows that LLMs safeguarded by ICAG exhibit significantly reduced jailbreak success rates across various attack scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations