#Mlsys
Showing 7 of 7 repositories tagged #mlsys, ranked by stars
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
A model compilation solution for various hardware
An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682
Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMs
Accelerating AI Training and Inference from Storage Perspective (Must-read Papers on Storage for AI)