Papers by Rishi Sharma
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference (2026.acl-industry)
Copied to clipboard
Shashank Kapadia, Deep Narayan Mishra, Sujal Reddy Alugubelli, Haoan Wang, Saipraveen Vabbilisetty, Rishi Bhatia, Anupriya Sharma
| Challenge: | Layer-aligned distillation and convergence-based early exit are dominant computational efficiency paradigms for transformer inference. |
| Approach: | They propose a training objective that aligns intermediate student layers to teacher representations and reconciles this incompatibility with standard distillation. |
| Outcome: | The proposed model achieves 1.61 measured wall-clock speedup with 91.9% of samples exiting by layer 7 and 1.80 theoretical layer reduction, where standard distilled models achieve zero effective speedup. |
Tackling the Story Ending Biases in The Story Cloze Test (P18-2)
Copied to clipboard
| Challenge: | Story Cloze Test (SCT) is a recent framework for evaluating story comprehension and script learning. |
| Approach: | They propose to use a crowdsourcing scheme to create a new SCT dataset to overcome some of the biases discovered in the original SCT. |
| Outcome: | The proposed model performs better than the baselines on the SCT dataset, despite human-authorship biases. |