Papers by Haodi Zhang

2 papers
Efficient Data Labeling by Hierarchical Crowdsourcing with Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have been gaining attention for their impressive performance in in-context dialogues.
Approach: They propose a hierarchical framework that leverages multiple LLMs for efficient data labeling under budget constraints.
Outcome: The proposed framework outperforms human labelers and GPT-4 in terms of accuracy and efficiency.
MTO: A Multi-turn Conversational Text-to-OverpassQL Dataset for Enhanced OpenStreetMap Query Generation (2026.findings-acl)

Copied to clipboard

Challenge: a framework for constructing multi-turn Text-to-OverpassQL dialogue datasets is proposed . a dataset of over 7,800 dialogues contains more than 20,000 individual utterances .
Approach: They propose a framework for constructing multi-turn Text-to-OverpassQL dialogue datasets . they convert Overpass queries into syntax trees using a custom parser based on OverpassQl .
Outcome: The proposed dataset includes over 7,800 dialogues, each containing 2 to 4 user utterances . it is the first multi-turn Text-to-OverpassQL dataset built upon the OverpassNL corpus .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations