#Inference-acceleration
Showing 3 of 3 repositories tagged #inference-acceleration, ranked by stars
thu-ml
SageAttention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Score
100
★ 3.6k
⑂ 467
+30/day
Cuda
thu-ml
SpargeAttn
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
Score
50
★ 1.0k
⑂ 99
+1/day
Cuda
autonomi-ai
nos
⚡️ A fast and flexible PyTorch inference server that runs locally, on any cloud or AI HW.
Score
0
★ 147
⑂ 12
—
Python