Papers by Hongzhan Lin
Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Data synthesis is a key research area in large language models (LLMs). |
| Approach: | They propose a framework that models code instruction synthesis process with a tree structure and optimization-driven evolution to alleviate constraints of unidirectional synthesis and randomness-driven generation. |
| Outcome: | The proposed framework outperforms open-weight code LLMs on five widely-used benchmarks. |
Reinforcement Tuning for Detecting Stances and Debunking Rumors Jointly with Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Social media has become a fertile ground for nurturing rumors and misinformation due to its lack of systematic moderation. |
| Approach: | They propose a framework to enhance the joint predictive capabilities of LLMs for stance detection and rumor verification tasks. |
| Outcome: | The proposed framework outperforms state-of-the-art methods and generalizes to non-LLMs accommodated as task models. |
SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on social interactions neglect hallucination while struggling with poor generalizability and implicit character fidelity judgments. |
| Approach: | They propose a generalizable and explicit paradigm for uncovering interactive patterns of Large Language Models across diverse worldviews by defining interactive hallucination through stance transfer and SHARP, a benchmark built by extracting relations from commonsense knowledge graphs. |
| Outcome: | The proposed paradigm is generalizable and explicit and demonstrates its effectiveness and stability. |
Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for harmful meme detection ignore in-depth cognition of meme text and image . authors propose a framework for learning reasonable thoughts from LLMs for better multimodal fusion . |
| Approach: | They propose to use large language models to learn reasonable thoughts from LLMs for better multimodal fusion and lightweight fine-tuning. |
| Outcome: | The proposed approach achieves superior performance than state-of-the-art methods on the harmful meme detection task. |
CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal Models (2024.acl-long)
Copied to clipboard
| Challenge: | Current methods for multimodal sarcasm target identification focus on superficial indicators in an end-to-end manner, overlooking the nuanced understanding of multimodal content. |
| Approach: | They propose a multimodal sarcasm target identification framework with a coarse-to-fine paradigm by augmenting sarcasm explainability with reasoning and pre-training knowledge. |
| Outcome: | The proposed framework outperforms state-of-the-art methods and exhibits explainability in deciphering sarcasm as well. |
Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing models struggle to balance predictive accuracy with human-understandable rationales. |
| Approach: | They propose to enhance LLMs by leveraging rationale distillation and domain knowledge injection for trustworthy multimodal rationale generation. |
| Outcome: | Experiments on real-world medical datasets show that ClinRaGen achieves state-of-the-art performance in disease diagnosis and rationale generation. |
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding? (2025.coling-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. |
| Approach: | They propose a benchmark to assess LLMs’ code understanding abilities from the perspective of code judging rather than code generation. |
| Outcome: | The proposed benchmark evaluates 12 well-known large language models to determine the correctness of provided code solutions. |
FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing fact-checking evaluation methods rely on static datasets and classification metrics, which fail to evaluate justification production and uncover the nuanced limitations of LLMs. |
| Approach: | They propose a framework that adaptively and dynamically assesses LLMs’ fact-checking capabilities by incorporating justification production alongside verdict prediction. |
| Outcome: | Experiments show that the framework differentiates among state-of-the-art LLMs, providing valuable insights into model strengths and limitations in model-centric fact-checking analysis. |
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness (2025.acl-long)
Copied to clipboard
| Challenge: | Existing models that assess mLLMs on harmful meme understanding are inaccurate and lack accuracy. |
| Approach: | They propose a framework that adaptively probes the reasoning capabilities of mLLMs . their framework systematically reveals the varying performance of different target mllms a . |
| Outcome: | The proposed framework systematically reveals the performance of different target mLLMs. |
WSDMS: Debunk Fake News via Weakly Supervised Detection of Misinforming Sentences with Contextualized Social Wisdom (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for debunking fake news rely on blending of authentic and fabricated content by creators. |
| Approach: | They propose a model that detects misinformation at sentence-level using social media conversations . they use a bag-level annotation system to train the model . |
| Outcome: | The proposed model outperforms existing state-of-the-art models on three real-world benchmarks and outperformed existing state of the art models in debunking fake news at sentence and article levels. |
Dialectical Structured Reasoning for Explainable Multimodal Fake News Detection (2026.findings-acl)
Copied to clipboard
Ruichao Yang, Yufan Bian, Wei Gao, Bo-Wen Zhang, Jing Ma, Hongzhan Lin, Ziyang Luo, Xiaobin Zhu, Xu-Cheng Yin
| Challenge: | Existing fake news detection models are opaque and lack deductive transparency . a framework for dialectical structured reasoning is proposed to address this limitation . |
| Approach: | They propose a framework that model fake news detection as an explicit dialectical process over multimodal social context. |
| Outcome: | The proposed framework achieves state-of-the-art while producing transparent explanations that mirror human reasoning process. |
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for automated fact-checking often overlook deceptive misinformation styles in generated explanations. |
| Approach: | They propose a framework that explicitly controls reasoning style by anchoring explanations to the predicted verdict. |
| Outcome: | The proposed framework achieves state-of-the-art under LLaMA-series models with 465 samples. |
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation approaches focus on mLLMs’ detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. |
| Approach: | They propose an agent-based arena-style evaluation framework that provides context-aware and unbiased assessment for mLLMs’ understanding of multimodal harmfulness. |
| Outcome: | The proposed framework reduces evaluation biases of judge agents and provides unbiased comparisons of mLLMs’ abilities to interpret multimodal harmfulness. |
A Coarse-to-fine Cascaded Evidence-Distillation Neural Network for Explainable Fake News Detection (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for fake news detection focus on fact-checked reports, resulting in limited coverage and debunking delays. |
| Approach: | They propose a Coarse-to-fine Cascaded Evidence-Distillation neural network for explainable fake news detection based on raw reports . they use hierarchical encoders and cascaded selectors to select most explainable sentences for verdicts on top of selected top-K reports based upon raw reports. |
| Outcome: | The proposed model outperforms baseline detection methods and generates high-quality explanations from diverse evaluation perspectives. |
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique (2025.findings-emnlp)
Copied to clipboard
| Challenge: | e MM-CRITIC is a holistic benchmark for evaluating the critique ability of Large Multimodal Models (LMMs) covering 8 main task types and over 500 tasks, covering 4471 samples. |
| Approach: | They introduce a holistic benchmark for evaluating the critique ability of Large Multimodal Models across multiple dimensions: basic, correction, and comparison. |
| Outcome: | The proposed benchmark covers 8 main task types and over 500 tasks and is composed of 4471 samples. |
Towards Low-Resource Harmful Meme Detection with LMM Agents (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for harmful meme detection are limited due to the dynamic nature of memes . eliciting knowledge-revising behavior within the LMM agent is a key factor in achieving this goal . |
| Approach: | They propose an agency-driven framework for low-resource harmful meme detection . they use annotated memes to leverage label information as auxiliary signals for model . |
| Outcome: | The proposed framework achieves superior performance than state-of-the-art methods on the low-resource harmful meme detection task. |
AMR-Evol: Adaptive Modular Response Evolution Elicits Better Knowledge Distillation for Large Language Models in Code Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | proprietary large language models (LLMs) have demonstrated impressive code generation performance. |
| Approach: | They propose an adaptive module-based model that refines the direct response distillation process by modular decomposition and adaptive response evolution. |
| Outcome: | The proposed framework outperforms baseline model and code generation methods on three popular benchmarks. |
Rumor Detection on Twitter with Claim-Guided Hierarchical Graph Attention Networks (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for rumor detection are limited to the strict relation of user responses or oversimplify the conversation structure. |
| Approach: | They propose a method that reinforces interaction of user opinions while reducing negative impact imposed by irrelevant posts. |
| Outcome: | The proposed method improves performance on three Twitter datasets and can detect rumors at early stages. |
DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) reasoning improves multi-step mathematical problem solving in large language models but is vulnerable to exposure bias and error accumulation. |
| Approach: | They propose a diffusion-styled CoT framework that reformulates CoT reasoning as an iterative denoising process. |
| Outcome: | The proposed framework outperforms existing methods on three multi-step CoT reasoning benchmarks. |
From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms (2026.findings-acl)
Copied to clipboard
Jinghao Luo, Yuchen Tian, Chuxue Cao, Ziyang Luo, Hongzhan Lin, Kaixin Li, Chuyi Kong, Ruichao Yang, Jing Ma
| Challenge: | Large Language Models (LLMs)-based agents have fundamentally reshaped artificial intelligence . however, the inherent statelessness of LLMs hinders their ability to maintain logical consistency across complex, multi-step tasks . |
| Approach: | They propose a framework for LLM agent memory mechanisms that formalizes the development process into three stages: storage, reflection, and experience. |
| Outcome: | The proposed framework breaks the development process into three stages . it analyzes the need for long-range consistency, challenges in dynamic environments, and the ultimate goal of continual learning. |
ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series (2025.findings-acl)
Copied to clipboard
Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, V. W., Richard Yi Da Xu, Guo Li, Xian Yang
| Challenge: | Large language models excel at processing unstructured data, but integrating time series data with text remains a challenge. |
| Approach: | They propose a self-supervised multimodal framework that uses prompt-guided learning to unify heterogeneous data types. |
| Outcome: | The proposed framework outperforms state-of-the-art approaches on disease diagnosis tasks using real-world datasets. |
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges (2025.naacl-short)
Copied to clipboard
| Challenge: | Recent advances in large multimodal models (LMMs) have demonstrated impressive code generation capabilities, primarily evaluated through image-to-code benchmarks. |
| Approach: | They propose a visual programming reasoning benchmark based on Scratch, a block-based visual programming language widely used in children’s programming education. |
| Outcome: | The proposed framework evaluates the visual programming ability of large multimodal models by integrating visual elements and embedded programming logic. |
Detect Rumors in Microblog Posts for Low-Resource Domains via Adversarial Contrastive Learning (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing rumor detection methods are poor at detecting false rumors about breaking news or trending topics due to the lack of training data and prior knowledge. |
| Approach: | They propose an adversarial contrastive learning framework to detect false rumors by adapting features learned from well-resourced rumor data to that of the low-resource. |
| Outcome: | The proposed framework improves on two low-resource datasets and shows superior performance . it overcomes restriction of domain and/or language usage and improves robustness . |