Papers by Chao Shang
Profiling-Free Mixed-Precision Quantization for MoE LLMs via Fuzzy Rule Interpolation (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models are scaling in size and capability, driving substantial computational and memory costs. |
| Approach: | They propose a mixed-precision quantization framework that uses fuzzy rule interpolation to predict quantization error from only sparse samples. |
| Outcome: | The proposed framework accelerates the profiling phase by up to 15.7 on DeepSeek-V2 while achieving comparable or slightly superior zero-shot accuracy. |
TAG: Gradient Attack on Transformer-based Language Models (2021.findings-emnlp)
Copied to clipboard
Jieren Deng, Yijue Wang, Ji Li, Chenghong Wang, Chao Shang, Hang Liu, Sanguthevar Rajasekaran, Caiwen Ding
| Challenge: | Recent studies show that publicly shared gradients in the training process can reveal the private training data to a third-party. |
| Approach: | They propose a gradient attack algorithm to reconstruct the local training data using GLUE benchmarks. |
| Outcome: | The proposed algorithm achieves 1.5x recover rate and 2.5x ROUGE-2 over previous methods without the need of ground truth label. |
Diable: Efficient Dialogue State Tracking as Operations on Tables (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing systems for dialogue state tracking use the full dialogue history as input and generate the entire state from scratch at each dialogue turn. |
| Approach: | They propose a task formalisation that represents the dialogue state as a table and formalises it as 'table manipulation task' they represent the dialogue as if it were a list with all the slots and generate the entire state from scratch at each dialogue turn. |
| Outcome: | The proposed system outperforms existing systems while maintaining competitive accuracy. |
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models (2025.findings-acl)
Copied to clipboard
Qin Liu, Chao Shang, Ling Liu, Nikolaos Pappas, Jie Ma, Neha Anna John, Srikanth Doss, Lluis Marquez, Miguel Ballesteros, Yassine Benajiba
| Challenge: | LLaVA-7B demonstrated a decline in safety alignment ability on multi-modal inputs compared to its LLM backbone. |
| Approach: | They propose a method to recover alignment ability from LLM backbone while preserving functional capabilities of VLMs. |
| Outcome: | The proposed framework recovers alignment ability that is inherent in the LLM backbone with minimal impact on fluency and linguistic capabilities of pre-trained VLMs. |
Taxonomy Construction of Unseen Domains via Graph-based Cross-Domain Knowledge Transfer (2020.acl-main)
Copied to clipboard
| Challenge: | Existing taxonomies are either entirely absent or missing. |
| Approach: | They propose a GNN-based cross-domain transfer framework for the taxonomy construction task. |
| Outcome: | The proposed framework improves on benchmark datasets from science and environment domains. |
Aligning to Constraints for Data-Efficient Language Model Customization (2025.findings-naacl)
Copied to clipboard
Fei Wang, Chao Shang, Shuai Wang, Sarthak Jain, Qiang Ning, Bonan Min, Vittorio Castelli, Yassine Benajiba, Dan Roth
| Challenge: | General-purpose language models (LMs) are aligned to diverse user intents, but fall short when it comes to specific applications. |
| Approach: | They propose a framework that uses constraints to automatically produce supervision signals for user alignment with constraints. |
| Outcome: | The proposed framework can produce supervision signals for user alignment with constraints. |
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers (2025.emnlp-main)
Copied to clipboard
| Challenge: | Autoregressive language models excel in text-to-audio generation, but lag behind diffusion models by a non-trivial margin. |
| Approach: | They propose a framework that integrates multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. |
| Outcome: | The proposed framework outperforms existing LM-based and diffusion-based systems in audio synthesis. |
Improving Time Sensitivity for Question Answering over Temporal Knowledge Graphs (2022.acl-long)
Copied to clipboard
| Challenge: | Temporal knowledge graphs record entity relations and when they occur in time . previous work fails to address time-related challenges such as time-order issues . paper proposes time-sensitive question answering framework to address these problems . |
| Approach: | They propose a time-sensitive question answering framework that uses temporal KGs to answer natural language questions. |
| Outcome: | The proposed framework outperforms the state-of-the-art on a new benchmark for question answering over temporal knowledge graphs. |