Papers by Zhiqi Huang
Code-Switching Can be Better Aligners: Advancing Cross-Lingual SLU through Representation-Level and Prediction-Level Alignment (2024.acl-short)
Copied to clipboard
| Challenge: | Existing code-switching-based cross-lingual spoken language understanding frameworks are limited to low-resource languages. |
| Approach: | They propose a cross-lingual spoken language understanding framework that leverages both code-switched and original sentences to achieve multi-level alignment. |
| Outcome: | The proposed framework can achieve multi-level alignment on two benchmarks across ten languages. |
Language Concept Erasure for Language-invariant Dense Retrieval (2024.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual models aim for language-invariant representations but still encode language identity. |
| Approach: | They propose a multi-task learning framework that induces language invariance in multilingual retrieval by reducing language-specific signals in the embedding space. |
| Outcome: | The proposed learning framework improves language-invariant dense retrieval over baselines on English retrieval data and general multilingual corpora. |
Knowledge-enhanced Prompt Tuning for Dialogue-based Relation Extraction with Trigger and Label Semantic (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to determine semantic relation between two arguments in dialogues are limited due to the low information density of text. |
| Approach: | They propose a Knowledge-Enhanced Prompt-Tuning method to enhance DRE model by exploiting trigger and label semantics. |
| Outcome: | The proposed method achieves state-of-the-art in F1 and F1c scores on a DialogRE dataset. |
Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for multimodal intent detection have two limitations: (i) close entanglement of multimodal semantics with modal structures; (ii) insufficient learning of causal effects of semantic and modality-specific information on the final predictions. |
| Approach: | They propose a Dual-oriented Disentangled Network with Counterfactual Intervention model that decouples semantics-oriented and modality-oriented representations and a Counterfective Intervention Module that applies causal inference to understand causal effects by injecting confounders. |
| Outcome: | The proposed model overcomes key limitations in existing systems by effectively disentangling and utilizing modality-specific and multimodal semantic information. |
GhostBERT: Generate More Features with Cheap Operations for BERT (2021.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that some parameters in pre-trained language models can be pruned away without severe accuracy degradation. |
| Approach: | They propose a method which generates more features with very cheap operations from the remaining features and can be applied to unpruned BERT models to enhance their performance. |
| Outcome: | Empirical results on the GLUE benchmark on three backbone models (i.e., BERT, RoBERTa and ELECTRA) verify the efficacy of the proposed method. |
Enhancing Code-Switching for Cross-lingual SLU: A Unified View of Semantic and Grammatical Coherence (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing models rely on annotated training data, limiting their scalability to low-resource languages. |
| Approach: | They propose a method termed SoGo for zero-shot cross-lingual SLU that uses keywords as substitution options to extract keywords and a token-level alignment strategy to ensure grammatical coherence. |
| Outcome: | The proposed method improves zero-shot cross-lingual SLU across nine languages on MultiATIS++. |
Towards Unified Spoken Language Understanding Decoding via Label-aware Compact Linguistics Representations (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for intent detection and slot filling decoders could result in misaligned predictions for both tasks. |
| Approach: | They propose a method that leverages label embeddings to jointly guide the decoding process. |
| Outcome: | The proposed method outperforms existing methods on two single- and multi-intent SLU benchmarks and can be incorporated into existing models. |
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models (2026.acl-long)
Copied to clipboard
Yang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu, Mingxuan Huang, Jingchao Wang, Zhihong Zhu, Boyan Xu, Zhiqi Huang
| Challenge: | Recent advances in vision-language models have improved performance in multi-modal learning. |
| Approach: | They propose a multi-modal benchmark that embeds a single coherent reasoning error in 1997 samples. |
| Outcome: | The proposed benchmark is based on a set of 1997 samples embedding a single coherent reasoning error. |
InfoEnh: Towards Multimodal Sentiment Analysis via Information Bottleneck Filter and Optimal Transport Alignment (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for multi-modal sentiment analysis have been developed to overcome these challenges. |
| Approach: | They propose a method that utilizes a masking technique as the bottleneck for information filtering and integrates all modalities into a common feature space via domain adaptation. |
| Outcome: | Extensive experiments on two benchmark MSA datasets show the proposed method performs better than baselines. |
TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs (2025.emnlp-demos)
Copied to clipboard
Duygu Nur Yaldiz, Yavuz Faruk Bakman, Sungmin Kang, Alperen Öziş, Hayrettin Eren Yildiz, Mitash Ashish Shah, Zhiqi Huang, Anoop Kumar, Alfy Samuel, Daben Liu, Sai Praneeth Karimireddy, Salman Avestimehr
| Challenge: | Generative Large Language Models (LLMs) produce untruthful outputs, referred to as hallucinations, which are often referred as false positives. |
| Approach: | They propose an open-source Python library with over 30 truthfulness prediction methods. |
| Outcome: | The proposed methods span diverse trade-offs in computational cost, access level, grounding document requirements, and supervision type (self-supervised or supervised). |
Federated Learning for Spoken Language Understanding (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods to improve robustness of models focus on a single dataset . but, there are few studies on how to combine merits of different datasets . |
| Approach: | They propose a federated learning framework that could unify datasets and tasks . they propose MV-Encoder as backbone of the framework to provide multi-granularity text representations . |
| Outcome: | The proposed framework improves on two SLU benchmark datasets and federated learning settings. |
PCAD: Towards ASR-Robust Spoken Language Understanding via Prototype Calibration and Asymmetric Decoupling (2024.acl-long)
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) suffers from error propagation from automatic speech recognition (ASR) in actual scenarios. |
| Approach: | They propose a framework which calibrates bias and errors and achieves adaptive-balanced decoupling training by a prototype-based loss model. |
| Outcome: | The proposed framework outperforms existing approaches and achieves state-of-the-art performance on three datasets. |
Towards Multi-modal Sarcasm Detection via Disentangled Multi-grained Multi-modal Distilling (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to sarcasm detection focus on textual and intra-modal incongruity . mainstream approaches process input of each modality in a holistic manner, resulting in redundant and unrefined information. |
| Approach: | They propose a framework for multi-modal sarcasm detection that disentangles modality representations into latent spaces and conducts multi-grained knowledge distilling. |
| Outcome: | The proposed framework overpowers existing methods on a common benchmark. |
Zero-Shot Spoken Language Understanding via Large Language Models: A Preliminary Study (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have shown promising results in zero-shot settings, which motivates us to explore prompt-based methods. |
| Approach: | They propose a two-stage framework which transforms the SLU task into a question-answering problem by directly prompting LLMs. |
| Outcome: | The proposed framework can be built by directly prompting LLMs to understand user needs without training data. |
MCLF: A Multi-grained Contrastive Learning Framework for ASR-robust Spoken Language Understanding (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Trending ASR-robust SLU systems have seen impressive improvements through global contrastive learning, but they can easily lead to severe semantic changes. |
| Approach: | They propose a two-stage multi-grained contrastive learning framework to improve ASR robustness . they first adapt pre-trained language models to downstream SLU datasets and then fine-tune it on the corresponding dataset. |
| Outcome: | The proposed framework improves on four datasets and four BERT-like backbone models. |
Syntax Matters: Towards Spoken Language Understanding via Syntax-Aware Attention (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on SLU systems have focused on integrating syntactic information into language models. |
| Approach: | They propose a model where attention scopes are constrained based on syntactic relationships. |
| Outcome: | The proposed model improves on three datasets and can be integrated into other language models to further boost their performance. |
A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiers (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Multi-modal intent recognition (MIR) requires integrating non-verbal cues from real-world contexts to enhance human intention understanding. |
| Approach: | They present a comprehensive review of multi-modal intent recognition . they provide a survey of the field covering textual, visual, and acoustic signals . |
| Outcome: | The present survey summarises the current state of multi-modal intent recognition . it includes a comprehensive taxonomy and advanced methods . |
MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts (2024.findings-acl)
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) is a crucial task in task-oriented dialogue systems. |
| Approach: | They propose an ASR-Robust SLU framework based on the mixture-of-experts technique to generate additional transcripts from clean transcripts and use it to weigh the representations of the generated transcripts, ASR transcripts . |
| Outcome: | The proposed framework achieves state-of-the-art on three benchmark SLU datasets. |
Alignment before Awareness: Towards Visual Question Localized-Answering in Robotic Surgery via Optimal Transport and Answer Semantics (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent models for visual question localized-answering (VQLA) lack the ability to relate these answers to their localization at an instance level. |
| Approach: | They propose a model which introduces optimal transport to achieve bidirectional and fine-grained alignment between images and questions, enabling more precise localization. |
| Outcome: | The proposed model outperforms state-of-the-art models on two widely-used datasets on surgical scenes and surgical instruments. |
An Automatic Method to Estimate Correctness of RAG (2025.coling-industry)
Copied to clipboard
| Challenge: | Existing methods to assess the correctness of RAG models fail to capture the model’s internal state during answer generation. |
| Approach: | They propose a method to predict the correctness of RAG models by modeling the model’s uncertainty on quantified perturbations of input. |
| Outcome: | Extensive experiments across multiple large language models show that the proposed approach quantifies RAG robustness by aligning predictions with ground truth with a MSE 0.002 while offering flexibility for diverse qualitative metrics. |