Papers by Zi Huang

8 papers
Optimizing Deeper Transformers on Small Datasets (2021.acl-long)

Copied to clipboard

Challenge: a common belief that training deep transformers from scratch requires large datasets is wrong . however, with proper initialization and optimization, the benefits of very deep transformer can carry over to challenging tasks with small datasets.
Approach: They train 48 layers of transformers from pre-trained RoBERTa and 24 relation-aware layers from scratch.
Outcome: The proposed scheme achieves state-of-the-art performance on a text-to-sql parsing benchmark . it uses 24 fine-tuned layers from pre-trained RoBERTa and 24 relation-aware layers from scratch .
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only (2026.acl-long)

Copied to clipboard

Challenge: Recent large language model-based approaches often overlook graph context or depend on distillation from larger models, limiting generalisation.
Approach: They propose a framework for zero-shot reasoning on text-rich networks . they use a Neighbour-aware Group Relative Policy Optimisation objective .
Outcome: The proposed framework optimises base LLMs using a Neighbour-aware group relative policy optimisation objective based on a novel margin gain metric for the informativeness of neighbouring signals .
Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for out-of-distribution (OOD) detection ignore textual-structural diversity . text-rich networks (TrNs) represent complex interplay between textual content and relational structures .
Approach: They propose a framework for evaluating out-of-distribution detection in text-rich networks . they propose augmentations, structural shifts, and domain-based divisions to model interplay .
Outcome: Experiments on 11 datasets show the framework is effective in out-of-distribution detection.
Event-Content-Oriented Dialogue Generation in Short Video (2024.naacl-long)

Copied to clipboard

Challenge: Existing multi-modal dialogue models are limited to incapacity of reading visual information and multi-dimensional interactions.
Approach: They propose a novel event-oriented video-dialogue dataset called SportsVD to overcome these challenges by generating human-like response according to event contents in the video and related external knowledge.
Outcome: The proposed method outperforms existing methods on SportsVD and other baselines under several automatic metrics.
Abstract then Play: A Skill-centric Reinforcement Learning Framework for Text-based Games (2023.findings-acl)

Copied to clipboard

Challenge: Existing reinforcement learning frameworks fail to decompose the task and abstract the action autonomously.
Approach: They propose a skill-centric reinforcement learning framework capable of abstracting the action in an end-to-end manner.
Outcome: Empirical experiments on the Jericho environment validate the proposed framework against state-of-the-art baselines.
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization (2025.findings-emnlp)

Copied to clipboard

Challenge: HKVE selectively accepts gradient optimization results based on the distribution of attention scores across different layers, ensuring that every optimization step positively contributes to the attack.
Approach: They propose a framework that selectively accepts gradient optimization results based on the distribution of attention scores across different layers and selectively takes them into account when calculating the attack success rate.
Outcome: The proposed framework outperforms existing methods by achieving success rates of 75.08% on MiniGPT4, 85.84% on LLaVA and 81.00% on Qwen-VL.
Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums (P19-1)

Copied to clipboard

Challenge: Existing work on teaching machines to ask questions focused on generating fixed answers.
Approach: They propose a model to generate open-answered questions from real-world news for open discussion . they analyze how language use affects the number of answers .
Outcome: The proposed model generates questions with higher quality than most text generation methods.
LSEG: A Fine-tuning Free Method for NL2FOL via Logic-Structure and Entropy Guided Inference Controlling (2026.findings-acl)

Copied to clipboard

Challenge: Large language models struggle with natural language to first order logic (NL2FOL) translation due to logical hallucination.
Approach: They propose a fine-tuning free framework to correct hidden state deviation by leveraging logical stability across logic preserving perturbations of the input.
Outcome: The proposed framework improves logical consistency during inference and improves accuracy over baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations