Papers by Hanwen Liu
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models (2024.emnlp-main)
Copied to clipboard
Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, Tetsuya Sakai, Tian Feng, Hayato Yamana
| Challenge: | Currently, tool-augmented large language models (LLMs) only achieve total scores of 45.3 and 37.0, respectively, on a scale of 100. |
| Approach: | They propose a multi-level diagnostic process to assess the LLM's hallucinations through two perspectives: depth and breadth. |
| Outcome: | The proposed diagnostic process assesses the hallucinations of large language models through two perspectives: depth and breadth. |
Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs? (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing large language models lack knowledge of nuanced, domain-specific details and are susceptible to hallucinations. |
| Approach: | They construct a benchmark that measures head, torso, and tail facts in terms of popularity. |
| Outcome: | The proposed model is based on 18K question-answer pairs regarding head, torso, and tail facts in terms of popularity. |
DiplomacyAgent: Do LLMs Balance Interests and Ethical Principles in International Events? (2025.emnlp-main)
Copied to clipboard
| Challenge: | a new study examines the safety implications of large language models in diplomatic positions . it identifies potential risks and ideological biases that could arise from LLMs . |
| Approach: | They propose an LLM-based multi-agent system for diplomatic position analysis . they propose ethical constraint measures to enhance the safety of LLMs . |
| Outcome: | The proposed system assesses the safety implications of large language models in diplomacy . it reveals that LLMs could exhibit a strong bias towards interests, leading to unsafe decisions . |
Global Textual Relation Embedding for Relational Understanding (P19-1)
Copied to clipboard
| Challenge: | Existing methods to learn textual relation embeddings are lacking in large open-domain corpora. |
| Approach: | They propose to learn a general-purpose embedding of textual relations using a large dataset from Freebase. |
| Outcome: | The proposed embedding can facilitate downstream tasks requiring relational understanding of the text. |
Parsing Natural Language into Propositional and First-Order Logic with Dual Reinforcement Learning (2022.coling-1)
Copied to clipboard
Xuantao Lu, Jingping Liu, Zhouhong Gu, Hanwen Tong, Chenhao Xie, Junyang Huang, Yanghua Xiao, Wenguang Wang
| Challenge: | Existing methods to parse natural language into structured logical expressions have limitations due to paucity of labeled data. |
| Approach: | They propose a scoring model to automatically learn a model-based reward . they also propose introducing a Chinese-PL/FOL dataset to compensate for paucity of labeled data . |
| Outcome: | The proposed model outperforms competitors on several datasets. |
MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information Networks (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing meta-path generation methods cannot fully exploit rich textual information in HINs. |
| Approach: | They propose a text-infilling-based approach to generate meta-paths from textual information in HINs. |
| Outcome: | The proposed approach can classify edges in the zero-shot setting, where existing methods cannot generate meta-paths. |
SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training (2024.acl-long)
Copied to clipboard
| Challenge: | Current methods focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. |
| Approach: | They propose a method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness. |
| Outcome: | The proposed method significantly improves training efficiency on deduplicated datasets and improves downstream accuracy by 1.77%. |
Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation (2025.coling-main)
Copied to clipboard
| Challenge: | Retrieval-augmented generation systems often use a fixed strategy to extract information from multiple sources. |
| Approach: | They propose a method that dynamically determines optimal granularity of a knowledge source based on input queries using a router. |
| Outcome: | The proposed method predicts optimal granularity levels and significantly improves performance in downstream tasks. |
TPS-Bench: Evaluating AI Agents’ Tool Planning & Scheduling Abilities in Compounding Tasks (2026.acl-long)
Copied to clipboard
| Challenge: | Large language model (LLM) agents have demonstrated strong problem-solving competence across domains like research and coding. |
| Approach: | They propose to use a tool repository to analyze the ability of large language model agents to solve complex problems. |
| Outcome: | The proposed model outperforms open-source and closed-source models in task completion rate and efficiency. |