Papers by Ershad Banijamali
AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tuning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in memory-efficient zeroth-order methods have limited their widespread adoption due to performance drops and a high risk of divergence. |
| Approach: | They propose a memory-efficient zeroth-order framework to improve performance and convergence of the MeZO methods by using only forward passes. |
| Outcome: | The proposed framework improves performance and convergence of the proposed methods on Roberta-Large and Llama-2-7B models. |
QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models (2025.emnlp-main)
Copied to clipboard
Jiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu, Yequan Zhao, Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong, Zheng Zhang
| Challenge: | Large Language Models (LLMs) are quantized to lower precision to reduce memory cost and latency in inference. |
| Approach: | They propose a quantized zeroth-order framework for fine-tuning Large Language Models (LLMs) using low-precision forward passes. |
| Outcome: | The proposed method achieves comparable results to first-order methods in FP8 and superior accuracy in INT8 and INT4 training. |