Papers by Jie Cai

14 papers
HiEdit: Lifelong Model Editing with Hierarchical Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to lifelong model editing apply parameter perturbations to static and dense layers for all instances.
Approach: They propose a hierarchical reinforcement learning framework that identifies the most knowledge-relevant layers for each editing instance.
Outcome: The proposed framework boosts the performance of the competitive RLEdit by 8.48% with perturbing only half of the layers per edit.
RTADev: Intention Aligned Multi-Agent Framework for Software Development (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are efficient assistants to humans in software development tasks, but they can cause errors during the development process.
Approach: They propose an intention aligned multi-agent framework that ensures that all agents work based on a consensus.
Outcome: The proposed framework reduces errors and improves the quality of generated software code.
Deep Cognitive Reasoning Network for Multi-hop Question Answering over Knowledge Graphs (2021.findings-acl)

Copied to clipboard

Challenge: Knowledge Graphs (KGs) store structured human knowledge with nodes and edges being entities and relations between them.
Approach: They propose a deep cognitive reasoning network that uses two phases to find answers in large candidate entity sets.
Outcome: The proposed method significantly outperforms state-of-the-art methods on benchmark datasets.
A Label Informative Wide & Deep Classifier for Patents and Papers (D19-1)

Copied to clipboard

Challenge: Existing methods for classification of patents and papers are manual and limited . et al., a chinese research team has developed a label-informative classification model .
Approach: They propose a label-informative classifier based on the Wide & Deep structure . they train on millions of patents and transfer to papers by developing distant-supervised training set and domain-specific features.
Outcome: The proposed model performs comparable to the state-of-the-art model used in industry on patents and papers.
DataSciBench: An LLM Agent Benchmark for Data Science (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on single task, simple evaluation metrics, and readily available ground truth (GT) DataSciBench is built on curated, natural, and challenging prompts with complex evaluation criteria and uncertain GT.
Approach: They propose a benchmark for evaluating Large Language Models in data science that integrates LLM-based self-consistency and human verification to ensure accuracy.
Outcome: The proposed framework outperforms open-source models in all metrics and offers rigorous insights into LLM strengths and weaknesses.
DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are pre-trained on vast datasets composed of billions of tokens harvested from diverse text sources.
Approach: They propose a data engineering method to refine the pretraining corpus through data rating, tagging and editing.
Outcome: The proposed method improves the quality of the pretraining corpus by enhancing 100 billion tokens of the training corpus.
WebCPM: Interactive Web Search for Chinese Long-form Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Long-form question answering requires two procedures: information retrieval and information synthesis.
Approach: They propose a Chinese long-form question answering dataset called WebCPM . the dataset is based on a web search interface that engages with a search engine in real time .
Outcome: The proposed dataset generates answers that are no worse than human-written ones . the dataset is the first Chinese LFQA dataset .
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values (2026.findings-eacl)

Copied to clipboard

Challenge: Existing Chinese preference datasets suffer from limited scale, restricted domain coverage, and insufficiently rigorous data validation.
Approach: They propose an LLM-based data annotation pipeline with no human intervention to annotate Chinese preference datasets.
Outcome: The proposed pipeline outperforms existing Chinese preference datasets on AlignBench and Chinese Reward Benchmark.
SeDev: Structured Semantic Exploration for LLM-Driven Code Generation (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities in automating code generation, but they suffer from insufficient exploration of the vast solution space.
Approach: They propose a large-scale LLM-driven code generation framework that efficiently finds high-quality solutions in only a few iterations.
Outcome: The proposed framework outperforms baselines while maintaining reasonable time and computational costs.
Rejection-to-Acceptance Transition: Model Editing-Based Jailbreak Backdoor Injection Not Limited to Few Output Tokens (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for jailbreaking LLMs are implemented by binding backdoors to predefined phrases as first few output tokens, inducing the LLM’s next-token prediction to produce continuous responses.
Approach: They propose a model editing-based jailbreak backdoor attack that hijacks LLM representations into a acceptance domain rather than binding to a few output tokens.
Outcome: The proposed model editing method outperforms existing methods, showing stronger jailbreak capabilities across LLMs and datasets.
XDailyDialog: A Multilingual Parallel Dialogue Corpus (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets for open-domain dialogue modeling limited to a single language . absence of multilingual datasets hinders development of robust open- domain dialog systems .
Approach: They propose a multilingual parallel open-domain dialog dataset to explore multilingual and cross-lingual open- domain dialog.
Outcome: The proposed model can be used to explore multilingual and cross-lingual open-domain dialogs in other languages.
SC2: Towards Enhancing Content Preservation and Style Consistency in Long Text Style Transfer (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for short TST are difficult to implement and can cause content degradation.
Approach: They propose a method to vary the style polarity of text while preserving semantic content.
Outcome: The proposed method improves over baselines and is highly efficient.
Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning (2025.findings-emnlp)

Copied to clipboard

Challenge: Low-Confidence Gold (LCG) is a new filtering framework for Large Language Models that curates high-quality subsets while preserving data diversity.
Approach: They propose a new filtering framework that employs centroid-based clustering and confidence-guided selection for identifying valuable instruction pairs.
Outcome: The proposed framework improves performance on a subset of 6K samples while maintaining data diversity.
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on single agentic capability, failing to capture long-horizon real-world scenarios.
Approach: They propose a benchmark that evaluates 6 agentic capabilities across 32 real-world scenarios.
Outcome: Experiments show that closed-source models outperform open-source model (48.4% vs 32.1%) integrating models with advanced scaffolds to form autonomous agents is a paradigm shift.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations