Papers by Kai Tu

3 papers
LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software (2026.acl-long)

Copied to clipboard

Challenge: Existing automated program-repair techniques focus on repairing memory corruptions, but they struggle with logical vulnerabilities because of their limited semantic understanding of the code and its expected behavior.
Approach: They evaluated a dataset of 122 logical vulnerabilities and a framework to evaluate patches for logical weaknesses.
Outcome: The proposed framework evaluates both traditional and LLM-based approaches for addressing real-world logical vulnerabilities.
MDC-Bench: A Multidisciplinary Causal Benchmark Based on Causal Structures for Evaluating Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing causal datasets focus on the commonsense domain, but LLMs perform poorly when answering complex questions.
Approach: They propose a multidisciplinary causal evaluation benchmark to assess LLMs' knowledge and skills.
Outcome: The proposed model improves in domain specialization, structural diversity, and task complexity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations