Papers by Sparsh Mittal
GRIZAL: Generative Prior-guided Zero-Shot Temporal Action Localization (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to temporally localize videos without prior training examples are lacking due to the complexity of annotated videos. |
| Approach: | They propose a model that uses multimodal embeddings and dynamic motion cues to localize actions effectively. |
| Outcome: | GRIZAL outperforms state-of-the-art zero-shot temporal action localization models on ActivityNet, Thumos14 and Charades-STA datasets. |