Papers by Mohammad Rastegari

3 papers
Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention (2025.emnlp-main)

Copied to clipboard

Challenge: Speculative decoding is a prominent technique for accelerating LLM inference by leveraging an auxiliary draft model, but its effectiveness is limited by the autoregressive nature of draft generation.
Approach: They propose a method that integrates speculative draft generation directly within the target model using multi-stream attention.
Outcome: The proposed method improves acceptance but also latency and speculation latency, limiting overall speedup.
LLM in a flash: Efficient Large Language Model Inference with Limited Memory (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have high computational and memory requirements, especially for devices with limited memory.
Approach: They propose a method that stores model parameters in flash memory but brings them on demand to DRAM . authors propose two techniques to optimize for reading data in larger, more contiguous chunks .
Outcome: The proposed method reduces the volume of data transferred from flash and reads data in larger, more contiguous chunks.
Pyramidal Recurrent Unit for Language Modeling (D18-1)

Copied to clipboard

Challenge: Long short term memory units are powerful tools for language modeling, but their performance can be limited by the number of parameters.
Approach: They propose a pyramidal recurrent unit which enables learning representations in high dimensional space with more generalization power and fewer parameters.
Outcome: The proposed model outperforms existing models with different gating mechanisms and transformations on word-level language modeling tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations