Papers by Hai Huang

23 papers
CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling (2025.acl-long)

Copied to clipboard

Challenge: Cross-modal retrieval tasks are used to retrieve data from one modality or another based on a query from another modality.
Approach: They propose a generative cross-modal retrieval framework based on coarse-to-fine semantic modeling . they propose combining K-Means and RQ-VAE to discretize multimodal data into token sequences that support autoregressive generation.
Outcome: The proposed framework achieves excellent performance and efficiency in multimodal retrieval tasks.
Omni-Chart-600K: A Comprehensive Dataset of Chart Types for Chart Understanding (2025.findings-naacl)

Copied to clipboard

Challenge: Existing chart-related training methods lack capabilities in information extraction, mathematical reasoning, and understanding of multiple chart types.
Approach: They propose a two-stage training strategy and method for jointly training a vision encoder tailored for multi-type charts to address the deficiencies in chart types and limited scope of chart tasks in existing datasets.
Outcome: The proposed dataset includes 21 diverse chart types and tasks, including data retrieval and mathematical reasoning.
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing gaps between discrete acoustic codecs and downstream speech language models . initial channel of codebooks contains excessive information, making it difficult to generate tokens from weakly supervised signals such as text.
Approach: They propose a discrete acoustic codec for generating acustic tokens from weakly supervised signals.
Outcome: The proposed language-codec outperforms competing audio compression algorithms and validates on downstream speech language models.
Forging Multiple Training Objectives for Pre-trained Language Models via Meta-Learning (2022.findings-emnlp)

Copied to clipboard

Challenge: Empirical studies show that learning multiple training objectives in a single model makes the learned language representation barely converge to the desired optimum.
Approach: They propose a meta-learning-based adaptive sampler which learns latent sampling pattern on arbitrary pre-training objectives.
Outcome: Empirical studies show that learning multiple objectives in a single model makes it difficult to achieve the desired optimum.
Semantic and Syntactic Enhanced Aspect Sentiment Triplet Extraction (2021.findings-acl)

Copied to clipboard

Challenge: Existing approaches to extract triplets from sentences neglect the mutual information between aspects and have the problem of error propagation.
Approach: They propose a Semantic and Syntactic Enhanced aspect Sentiment triplet Extraction model to exploit the syntactical and semantic relationships between the triplet elements and jointly extract them.
Outcome: The proposed model outperforms existing methods on four benchmark datasets and significantly outperformed existing approaches.
Tracing Origins: Coreference-aware Machine Reading Comprehension (2022.acl-long)

Copied to clipboard

Challenge: a recent study has enriched pre-trained language models with syntactic, semantic and other linguistic information to improve their performance.
Approach: They use a pre-trained language model to leverage coreference information to enhance word embeddings . they use additional encoder layers to focus on coreference mentions or a relational graph convolutional network to model the coreference relations.
Outcome: The proposed model imitates the human reading process and leverages coreference information to enhance word embeddings.
Composite Backdoor Attacks Against Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated superior performance on various tasks, but untrustworthy third-party LLMs may covertly introduce vulnerabilities for downstream tasks.
Approach: They propose a composite backdoor attack that scatters multiple trigger keys in different prompt components.
Outcome: The proposed attack achieves 100% Attack Success Rate (ASR) with a False Triggered Rate (FTR) below 2.06% and negligible model accuracy degradation.
StructAM: Enhancing Address Matching through Semantic Understanding of Structure-aware Information (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to address matching rely on string-based similarity matching or manually-designed rules.
Approach: They propose a method to match unstructured addresses to standard ones in a database using pre-trained language models and graph neural networks.
Outcome: The proposed method outperforms state-of-the-art methods on real-world addresses . it incorporates spatial coordinates and contextual information from the surrounding area as auxiliary guidance.
Overcoming both Domain Shift and Label Shift for Referring Video Segmentation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods to improve the robustness of open-set domain generalization can only recognize seen objects and mark all unseen objects as “unknown” categories .
Approach: They propose a method to make the model maintain good segmentation ability for unknown objects . they propose CLIP-based Reasoning Prompt which can combine text and visual prompts .
Outcome: The proposed method can bridge the gap caused by label shift by combining text and visual prompts to improve text-object matching ability.
MELA: Multilingual Evaluation of Linguistic Acceptability (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks on linguistic acceptability have been used to evaluate language models' ability to distinguish between acceptable and unacceptable sentences.
Approach: They present the largest benchmark to date on linguistic acceptability: MELA . they establish LLM baselines on this benchmark and investigate cross-lingual transfer in acceptability judgements with XLM-R.
Outcome: The proposed model outperforms open-source models on cross-lingual transfer in acceptability judgements.
Moon IME: Neural-based Chinese Pinyin Aided Input Method with Customizable Association (P18-4)

Copied to clipboard

Challenge: a pinyin input method engine (IME) allows users to input Chinese into a computer by typing pinyan through the common keyboard.
Approach: They present a pinyin IME that integrates neural machine translation and IR to offer amusive and customizable association ability.
Outcome: The Moon IME integrates neural machine translation and IR to offer amusive association ability.
Pre-training Multi-party Dialogue Models with Latent Discourse Inference (2023.acl-long)

Copied to clipboard

Challenge: Existing studies have failed to scale up the pre-training process by putting aside unlabeled data . et al., 2019: multi-party dialogues are more difficult for models to understand since they involve multiple interlocutors resulting in interweaving reply-to relations and information flows.
Approach: They propose to treat discourse structures as latent variables and jointly infer them to pre-train a model that understands the discourse structure of multi-party dialogues.
Outcome: The proposed model outperforms baselines and achieves state-of-the-art results on multiple downstream tasks.
Faster MoE LLM Inference for Extremely Large Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing inference optimizations for coarse-grained Mixture-of-Experts models implicitly assume a fixed activation budget, which is poorly understood.
Approach: They propose a training-free policy that adapts token-level activation using router confidence and entropy while remaining within the model’s original budget.
Outcome: The proposed skipping policy can provide substantial throughput gains, but optimal static schedules vary significantly across models and routing mechanisms.
Enhancing Multimodal Unified Representations for Cross Modal Generalization (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on discrete unified representations overlook important distinctions between different dimensions of features.
Approach: They propose to use a codebook to optimize unified representations from pretraining and fine- and coarse-grained disentangling to optimize the representations.
Outcome: The proposed methods improve the interpretability of multimodal unified representations . they use training-free optimization of codebook and fine and coarse cross-modal disentangling .
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control (2025.acl-long)

Copied to clipboard

Challenge: Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS systems cannot perform speaker-specific voice generation.
Approach: They propose a style control module that captures codec representations corresponding to timbre, content, and style in a discrete decoupling codec space.
Outcome: The proposed system can fully clone the speaker's voice and perform speech-specific adjustment and control functions.
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese (2025.emnlp-main)

Copied to clipboard

Challenge: Translationese is a linguistic property that is often introduced in the translation process that is different from those of original texts.
Approach: They propose to use synthesized translations and translations in the wild to evaluate T-index's generalizability in cross-domain settings and its validity against human judgments.
Outcome: The proposed measure can generalize to unseen genres, authors, and language pairs.
Can AI Revise Research Papers with Human Review Feedback? An Empirical Study and Benchmark (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are fundamentally reshaping the scientific landscape, transitioning the role of AI from passive tools to active partners within a new paradigm of Human-AI collaboration.
Approach: They propose a benchmark to evaluate the ability of Large Language Models to improve papers with human feedback.
Outcome: The proposed benchmark tests the skills of Large Language Models (LLMs) on paper interpretation, experimental implementation, and paper formulation, using authors’ camera-ready versions as natural human baselines.
Chinese Pinyin Aided IME, Input What You Have Not Keystroked Yet (D18-1)

Copied to clipboard

Challenge: Chinese pinyin input method engine (IME) converts pinyine into character based on its core component, pinyan-to-character conversion (P2C).
Approach: They propose a sequence-to-sequence model with gated-attention mechanism for Chinese IMEs.
Outcome: The proposed model improves on existing models in benchmark datasets showing great user experience improvement compared to traditional models.
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for addressing item-level user interests are lacking in cross-domain generalization . RecBase model is domain-agnostic and can be used to enhance recommender systems' effectiveness .
Approach: They propose a domain-agnostic foundational model pretrained with a recommendation-oriented objective that leverages a large-scale, heterogeneous, cross-domain corpus with unified textual representations and feature mappings to enhance cross- domain generalization.
Outcome: The proposed model matches or surpasses baselines in zero-shot and cross-domain recommendation tasks on eight real-world datasets.
Subword-augmented Embedding for Cloze Reading Comprehension (C18-1)

Copied to clipboard

Challenge: Existing models for machine reading comprehension use word and character representations, but character is not the minimal unit.
Approach: They propose to use subword rather than character for word embedding enhancement . they also empirically explore different augmentation strategies on subword-augmented embedded embedders .
Outcome: The proposed model outperforms state-of-the-art models on public datasets.
Open Vocabulary Learning for Neural Chinese Pinyin IME (P19-1)

Copied to clipboard

Challenge: Pinyin-to-character conversion is the core component of pinyin based Chinese input method engine (IME).
Approach: They propose a neural P2C conversion model augmented by an online updated vocabulary to support open vocabulary learning during IME working.
Outcome: The proposed model outperforms commercial IMEs and state-of-the-art models on standard corpus and true inputting history dataset in terms of multiple metrics and the online updated vocabulary helps it follow user inputting behavior.
Lingke: a Fine-grained Multi-turn Chatbot for Customer Service (C18-2)

Copied to clipboard

Challenge: e-commerce chatbots usually need a mass of human dialogue data to train, but for multi-turn conversations, the performance is poor.
Approach: They propose an information retrieval augmented multi-turn chatbot which can answer questions based on unstructured documents and deal with multi-turned conversations.
Outcome: The proposed solution outperforms all other models in multi-turn conversations and can learn from conversation records.
Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization (2025.findings-naacl)

Copied to clipboard

Challenge: a new approach to timeline summarization is proposed for open-domain news content . large language models (LLMs) can be used to extract and organize news events from multiple documents .
Approach: They propose a method to integrate Large Language Models into news timeline summarization by iterating on how events are linked and posing new questions.
Outcome: The proposed system is able to generate and refresh chronological summaries based on documents retrieved in each round.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations