#Efficient-attention
Showing 4 of 4 repositories tagged #efficient-attention, ranked by stars
thu-ml
SageAttention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Score
100
★ 3.6k
⑂ 467
+30/day
Cuda
lucidrains
CoLT5-attention
Implementation of the conditionally routed attention in the CoLT5 architecture, in Pytorch
Score
100
★ 230
⑂ 15
—
Python
jlamprou
Infini-Attention
Efficient Infinite Context Transformers with Infini-attention Pytorch Implementation + QwenMoE Implementation + Training Script + 1M context keypass retrieval
Score
0
★ 98
⑂ 7
—
Python
davidsvy
cosformer-pytorch
Unofficial PyTorch implementation of the paper "cosFormer: Rethinking Softmax In Attention".
Score
0
★ 44
⑂ 8
—
Jupyter Notebook