Papers by Shiyu Huang

9 papers
Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing metrics for video captioning are based on text-based comparisons with ground-truth references.
Approach: They propose a reference-free benchmark that assesses video captions based on their utility . they will release the benchmark to facilitate reproducible research .
Outcome: The proposed benchmark improves on human-verified, fine-grained questions . it correlates significantly better with human judgments than existing metrics .
A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Version updates are an indispensable requirement for Large Language Models . a large learning rate in the first stage and a complete learning decay process are crucial for version updates of LLMs.
Approach: They propose a learning rate path switching training paradigm for version updates of Large Language Models.
Outcome: The proposed paradigm reduces training cost to 58% when training four versions of LLMs compared to PTFS and CPT .
Query and Extract: Refining Event Extraction as Type-oriented Binary Decoding (2022.findings-acl)

Copied to clipboard

Challenge: Existing approaches to event extraction are limited to a set of pre-defined types.
Approach: They propose a natural language query framework that uses event types and argument roles to extract candidate triggers and arguments from input text.
Outcome: The proposed framework outperforms existing methods on zero-shot event extraction.
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality (2025.emnlp-main)

Copied to clipboard

Challenge: Privacy-sensitive users require deploying large language models within their own infrastructure (on-premises) vulnerabilities in local environments can lead to unauthorized access and potential model theft.
Approach: They propose a framework that secures a few bottom layers in a secure environment . they propose metric to optimize trade-off between protection and customization flexibility .
Outcome: The proposed framework outperforms baselines on five models with 1.3B to 70B parameters.
Incremental Prompting: Episodic Memory Prompt for Lifelong Event Detection (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to improve lifelong event detection performance are limited by the limited stored examples.
Approach: They propose to use Episodic Memory Prompts to explicitly retain the learned task-specific knowledge.
Outcome: The proposed method can be used to update a model with new event types while retaining the capability on previously learned types.
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? (2025.acl-long)

Copied to clipboard

Challenge: Large Language Model (LLM) watermarking is radioactive and enables the detection of watermarks inherited by student models when trained on the outputs of watermarked teacher models.
Approach: They propose two types of watermark removal attacks that allow student models to perform untraceable knowledge distillation while avoiding watermark inheritance.
Outcome: The proposed attacks eliminate inherited watermarks while maintaining knowledge transfer efficiency and low computational overhead.
VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision (2026.acl-long)

Copied to clipboard

Challenge: Empirical evaluations demonstrate that VCORE achieves the strongest overall average performance, with especially clear gains on lower-capacity models.
Approach: They propose a framework that reformulates supervision as a constrained optimization problem.
Outcome: Empirical evaluations show that VCORE achieves the strongest overall average performance, with especially clear gains on lower-capacity models.
LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating large language models use static datasets, leading to data leakage or overlooking the complexities of multi-agent interactions.
Approach: They propose a framework that evaluates the diverse capabilities of LLM agents in multi-agent dynamic environments.
Outcome: The proposed framework assesses the diverse capabilities of LLM agents in multi-agent dynamic environments.
Pre-trained Semantic Interaction based Inductive Graph Neural Networks for Text Classification (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for text classification have vanishing or exploding gradients when dealing with long sequences, making it difficult to handle long-distance dependencies.
Approach: They propose a graph neural network based on pre-trained semantic interaction called PaSIG . they construct a text-word heterogeneity graph and use context representation capability .
Outcome: The proposed model outperforms existing methods on five datasets and achieves state-of-the-art performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations