Papers by Jason Cai
FineSurE: Fine-grained Summarization Evaluation using LLMs (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for text summarization evaluation do not correlate well with human judgments . evaluators that use Likert scale scores are limited in their ability to perform deeper analysis. |
| Approach: | They propose a fine-grained evaluator specifically tailored for the summarization task using large language models. |
| Outcome: | The proposed method improves on open-source and proprietary LLMs and shows better completeness and conciseness than existing methods. |
UniSumEval: Towards Unified, Fine-grained, Multi-dimensional Summarization Evaluation for LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks for summarization quality evaluation lack diverse input scenarios, focus on narrowly defined dimensions, and struggle with subjective and coarse-grained annotation schemes. |
| Approach: | They propose to use AI to help human annotations and identifie potentially hallucinogenic input texts. |
| Outcome: | The proposed benchmarks improve on existing benchmarks in terms of input diversity, granularity of human annotations, and evaluation dimensions. |
Active Generalized Category Discovery with Diverse LLM Feedback (2026.eacl-long)
Copied to clipboard
| Challenge: | Generalized Category Discovery (GCD) is a practical and challenging open-world task that aims to recognize both known and novel categories in unlabeled data using limited labeled data from known categories. |
| Approach: | They propose a framework for generalized category discovery that actively learns from diverse and collaborative feedback. |
| Outcome: | The proposed framework improves instance-level contrastive features, generates category descriptions, and aligns uncertain instances with LLM-selected category descriptions. |
CERET: Cost-Effective Extrinsic Refinement for Text Generation (2024.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) generate incomplete, biased or misleading outputs in their initial attempts. |
| Approach: | They propose a method for refining text generation that takes into account semantic stability, entailment and inter-sample uncertainty measures. |
| Outcome: | The proposed method outperforms self-consistency and self-rerank baselines under various task setups by 1.6% and 3.5% respectively. |
REST: Retrieval-Based Speculative Decoding (2024.naacl-long)
Copied to clipboard
| Challenge: | Retrieval-based speculative decoding (REST) is a new language model generation algorithm . it uses existing knowledge to generate draft tokens, allowing for seamless integration and acceleration of any language model. |
| Approach: | They propose a new algorithm that uses a draft language model to generate tokens from existing knowledge. |
| Outcome: | The proposed method achieves a speedup of 1.62 to 2.36 on code or text generation. |
TReMu: Towards Neuro-Symbolic Temporal Reasoning for LLM-Agents with Memory in Multi-Session Dialogues (2025.findings-acl)
Copied to clipboard
| Challenge: | Temporal reasoning in multi-session dialogues presents a significant challenge which has been under-studied in previous temporal reasoning benchmarks. |
| Approach: | They propose to augment LoCoMo dialogues and create multi-choice QAs to construct a temporal reasoning evaluation task and a framework to enhance temporal thinking capabilities of LLM-agents. |
| Outcome: | The proposed framework significantly improves temporal reasoning performance compared to baseline methods, raising from 29.83 on GPT-4o via standard prompting to 77.67 via the proposed framework. |
MemInsight: Autonomous Memory Augmentation for LLM Agents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language model (LLM) agents have evolved to intelligently process information, make decisions, and interact with users or tools. |
| Approach: | They propose an autonomous memory augmentation approach to enhance semantic data representation and retrieval mechanisms by leveraging historical interactions. |
| Outcome: | The proposed approach outperforms a baseline RAG by 34% in recall for LoCoMo retrieval on three task scenarios and boosts persuasiveness of recommendations by 14%. |
Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection (2024.naacl-long)
Copied to clipboard
| Challenge: | Semi-supervised dialogue summarization (SSDS) leverages model-generated summaries to reduce reliance on human-labeled data. |
| Approach: | They propose a scoring approach that encapsulates three primary dimensions of summarization model quality. |
| Outcome: | The proposed method reduces reliance on human-labeled data and improves the performance of summarization models. |
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training (2024.eacl-long)
Copied to clipboard
| Challenge: | End-to-end (E2E) spoken language understanding models are constrained by the cost of collecting speech-semantics pairs. |
| Approach: | They propose a model that learns E2E SLU without speech-semantics pairs . they propose cross-modal selective self-training (CMSST) to address imbalance and noise issues . |
| Outcome: | The proposed model learns E2E SLU without speech-semantics pairs . the proposed model requires the domains of speech-text and text-sensitization to match . |
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages (2025.acl-long)
Copied to clipboard
Hyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng, Nicole Hee-Yeon Kim, Taewon Yun, Hang Su, Jason Cai, Hwanjun Song
| Challenge: | Existing evaluation frameworks for text summarization lack domain-specific assessment criteria and are predominantly English-centric. |
| Approach: | They propose a multi-dimensional, multi-domain evaluation of summarization in English and Chinese that incorporates specialized assessment criteria for each domain and leverages a debate system to enhance annotation quality. |
| Outcome: | The proposed evaluation framework provides a multi-dimensional, multi-domain evaluation of summarization in English and Chinese. |
SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection (2025.emnlp-main)
Copied to clipboard
| Challenge: | Modern AI agents rely on Large Language Models (LLMs) as their reasoning engines, but they still face the challenge of generating meaningful reflections due to inadequate error analysis and a reliance on rare successful trajectories. |
| Approach: | They propose a framework for self-learning agents powered by a retrospective language model that generates reflections during inference. |
| Outcome: | The proposed framework outperforms reflection-based baselines on three challenging benchmarks. |
Learning to Summarize from LLM-generated Feedback (2025.naacl-long)
Copied to clipboard
| Challenge: | Developing effective text summarizers remains a challenge due to issues like unfaithful statements, key information omissions, and verbosity. |
| Approach: | They propose a large-scale dataset containing multi-dimensional feedback on LLM-generated summaries of varying quality across diverse domains to align them with human preferences for faithfulness, completeness, and conciseness. |
| Outcome: | The proposed model outperforms the 10x larger Llama3-70b-instruct in generating human-preferred summaries. |