Papers by Hanwen Zhang

6 papers
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Currently, tool-augmented large language models (LLMs) only achieve total scores of 45.3 and 37.0, respectively, on a scale of 100.
Approach: They propose a multi-level diagnostic process to assess the LLM's hallucinations through two perspectives: depth and breadth.
Outcome: The proposed diagnostic process assesses the hallucinations of large language models through two perspectives: depth and breadth.
DiplomacyAgent: Do LLMs Balance Interests and Ethical Principles in International Events? (2025.emnlp-main)

Copied to clipboard

Challenge: a new study examines the safety implications of large language models in diplomatic positions . it identifies potential risks and ideological biases that could arise from LLMs .
Approach: They propose an LLM-based multi-agent system for diplomatic position analysis . they propose ethical constraint measures to enhance the safety of LLMs .
Outcome: The proposed system assesses the safety implications of large language models in diplomacy . it reveals that LLMs could exhibit a strong bias towards interests, leading to unsafe decisions .
MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information Networks (2022.emnlp-main)

Copied to clipboard

Challenge: Existing meta-path generation methods cannot fully exploit rich textual information in HINs.
Approach: They propose a text-infilling-based approach to generate meta-paths from textual information in HINs.
Outcome: The proposed approach can classify edges in the zero-shot setting, where existing methods cannot generate meta-paths.
SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training (2024.acl-long)

Copied to clipboard

Challenge: Current methods focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication.
Approach: They propose a method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness.
Outcome: The proposed method significantly improves training efficiency on deduplicated datasets and improves downstream accuracy by 1.77%.
Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation (2025.coling-main)

Copied to clipboard

Challenge: Retrieval-augmented generation systems often use a fixed strategy to extract information from multiple sources.
Approach: They propose a method that dynamically determines optimal granularity of a knowledge source based on input queries using a router.
Outcome: The proposed method predicts optimal granularity levels and significantly improves performance in downstream tasks.
Logic2Text: High-Fidelity Natural Language Generation from Logical Forms (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies on Natural Language Generation (NLG) from structured data focus on surface descriptions of simple record sequences, for example, attribute-value pairs of fixed or very limited schema.
Approach: They propose to use a large-scale dataset to generate NLG from logical forms to obtain controllable and faithful generations from structured data.
Outcome: The proposed model can describe interesting facts from logical inferences across records, but it is difficult to produce such fidelity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations