Papers by Shiyi Qi
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending (2024.acl-long)
Copied to clipboard
| Challenge: | Existing models that use self-attention and position embedding have anomalous behavior that hinder long context window extrapolation. |
| Approach: | They propose a collinear constraint between Q and K to integrate RoPE and self-attention. |
| Outcome: | The proposed model integrates self-attention and position embedding into LLMs without fine-tuning. |
XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection (2024.findings-acl)
Copied to clipboard
| Challenge: | XMoE leverages small experts and a threshold-based router to selectively engage only essential parameters. |
| Approach: | They propose a novel MoE that leverages small experts to selectively engage only essential parameters. |
| Outcome: | The proposed model can reduce computation load at MoE layers by over 50% without sacrificing performance. |
Once is Enough: A Light-Weight Cross-Attention for Fast Sentence Pair Modeling (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies suggest that transformer-based models perform cross-attention over input pairs, leading to computational cost. |
| Approach: | They propose a lightweight cross-attention mechanism that performs query encoding only once while modeling the query-candidate interaction in parallel. |
| Outcome: | The proposed model speeds up sentence pairing by over 113x while achieving comparable performance as the more expensive models. |