Papers by Zhihua Jiang

15 papers
IM^2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Evaluation metrics for dialogue systems are expensive and time-consuming . current evaluation metrics focus on a single quality or several qualities .
Approach: They propose an interpretable, multi-faceted, and controllable framework to combine dialogue metrics which are good at measuring different qualities.
Outcome: The proposed framework integrates a large number of evaluation metrics to improve the performance of the model.
A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solution (2025.emnlp-main)

Copied to clipboard

Challenge: Low-resource language understanding is challenging for large language models (LLMs).
Approach: They propose a CompRehensive lIterary Chinese readIng comprehenSion procedure with a large dataset for CRISIS.
Outcome: The proposed procedure has the largest dataset and substantiates the effectiveness of the proposed procedure with a 7 percent hike in accuracy compared with the baseline.
LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement Feedback (2024.naacl-long)

Copied to clipboard

Challenge: Existing dialogue systems do not utilize quality dimensions specifically designed for dialogue evaluation to guide the response generation during training.
Approach: They propose a two-stage framework which generates and utilizes conversation evaluation as explicit feedback during training.
Outcome: The proposed framework generates and utilizes conversation evaluation as explicit feedback during training.
Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs (2026.acl-long)

Copied to clipboard

Challenge: Existing datasets for evaluating text-to-image generation focus mostly on real-life images, which poses challenges for assessing academic figure generation given real scientific captions.
Approach: They propose a dataset that first provides a Holistic Evaluation for Academic caption-to-Figure Generation (HE4AFG) they collect real figure captions from 8 scientific domains and generate 3,900 evaluation samples .
Outcome: The proposed model provides high-quality human ratings in terms of three aspects—scientific aesthetic (SA), topic relevance (TR), and attribute correctness (AC).
CHAE: Fine-Grained Controllable Story Generation with Characters, Actions and Emotions (2022.coling-1)

Copied to clipboard

Challenge: Existing studies on story generation focus on coarse-grained control of the story, neglecting the details of the narrative.
Approach: They propose a model for fine-grained control on the story that allows the generation of customized stories with characters, corresponding actions and emotions arbitrarily assigned.
Outcome: The proposed method has strong controllability to generate customized stories according to the fine-grained personalized guidance.
Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Knowledge-based visual reasoning (KB-VR) is a challenging task, as it requires machines not only to understand concepts and relationships of visual scenes, but also to associate them with external world knowledge to perform chain of reasoning on open-world questions.
Approach: They propose a visual knowledge card (VKC) that integrates internal visual knowledge and external world knowledge produced by a knowledge generator into an image.
Outcome: The proposed model achieves new state-of-the-art results compared to previous top-performing models on three popular KB-VR benchmarks.
ToViLaG: Your Visual-Language Generative Model is Also An Evildoer (2023.emnlp-main)

Copied to clipboard

Challenge: Recent large-scale Visual-Language Generative Models (VLGMs) generate toxic content, e.g., offensive text and pornography images, raising significant ethical risks.
Approach: They propose a bottleneck-based detoxification method to reduce toxicity while maintaining comparable generation quality.
Outcome: The proposed method could reduce toxicity while maintaining comparable generation quality.
EMPATH: An Ensemble Method for Automatic Fine-Grained Turn-Level Dialogue Empathy Evaluation with a Novel Emotional Distance Metric (2026.findings-acl)

Copied to clipboard

Challenge: Empathy evaluation metrics are lacking in the competitions, and classical dialogue evaluation metrics require further investigation.
Approach: They propose a framework which combines fine-tuned models, large language models, classical dialogue evaluation metrics, and a novel metric.
Outcome: The proposed framework improves on the WASSA 2024 benchmark and shows a statistically significant 8% improvement on the EX dataset.
DeTAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification (2025.findings-acl)

Copied to clipboard

Challenge: Existing defense methods rely on fine-tuning or input modification, which suffer from limited generalization and reduced utility.
Approach: They propose a finetuning-free approach that improves the defensive capabilities against jailbreak attacks of LLMs via targeted attention modification.
Outcome: The proposed approach outperforms baselines in jailbreak defense and exhibits robust generalization across attacks and models, maintaining its effectiveness even on in-the-wild jailbreak data.
PromptKeeper: Safeguarding System Prompts for LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: PromptKeeper is a defense mechanism designed to safeguard system prompts . adversarial and regular queries can exploit LLM vulnerabilities to expose hidden prompts.
Approach: PromptKeeper is a defense mechanism designed to safeguard system prompts . it detects both explicit and subtle leakage and regenerates responses using a dummy prompt .
Outcome: PromptKeeper detects and mitigates side-channel vulnerabilities when prompts are exposed . it regenerates responses using a dummy prompt, ensuring outputs remain indistinguishable from typical interactions .
STANKER: Stacking Network based on Level-grained Attention-masked BERT for Rumor Detection on Social Media (2021.emnlp-main)

Copied to clipboard

Challenge: Existing models for text classification are limited in performance, resulting in poor rumor detection.
Approach: They propose to use Chinese microblogs to detect rumors using pre-trained language models and auxiliary features such as comments to mask co-attention.
Outcome: The proposed model outperforms the state-of-the-art on Weibo20 and three existing social media datasets.
SEHY: A Simple yet Effective Hybrid Model for Summarization of Long Scientific Documents (2022.findings-aacl)

Copied to clipboard

Challenge: Abstractive approaches to extract salient sentences from long documents are not effective due to their size.
Approach: They propose a simple yet effective approach that exploits the discourse information of a document to select salient sections instead of sentences for summary generation.
Outcome: The proposed approach avoids full-text understanding and retains salient information given the length limit.
LLM-based Open Domain Planning by Leveraging Entity-Attribute-Level Domain Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, large language models (LLMs) based on Open domain Natural language planning have limited application potential.
Approach: They propose a dataset with a baseline for Open domain Natural language planning . the dataset provides the largest dataset for textual procedures to date .
Outcome: The proposed dataset provides the largest dataset for textual procedures to date . it leverages entity-attribute-level action models to reveal relevant physical properties .
Large-Scale and Multi-Perspective Opinion Summarization with Diverse Review Subsets (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for opinion summarization are deficient in epitomizing extensive reviews and offering opinion summaries from various angles.
Approach: They propose a supervised opinion summarization framework that takes sentiment orientation into account and trains the summarizer to learn from sub-optimal and optimal review subsets.
Outcome: The proposed framework generates pros, cons, and verdict summaries from hundreds of input reviews.
Leveraging Context-Aware Prompting for Commit Message Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for writing comprehensive commit messages focus on the changed lines or nearest context lines, but excessive contexts can lead to noise.
Approach: They propose a code model COMMIT that can generate automatic commit messages by combining a dataset with a context-aware prompt.
Outcome: The proposed model surpasses all existing models including pre-trained language models for code and large language models such as Code-LlaMa.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations