Moonshot AI open-sources FlashKDA for Kimi Delta Attention
Moonshot AI has open-sourced FlashKDA, a high-performance CUTLASS-based implementation of Kimi Delta Attention (KDA) kernels. Designed as a drop-in backend for flash-linear-attention, FlashKDA achieves 1.72×–2.22× prefill speedups over baselines on NVIDIA H20 GPUs while offering native support for variable-length batching in production environments.
Moonshot AI is accelerating linear attention adoptability by targeting low-level GPU kernel optimizations.
- –Delivers a 1.72×–2.22× prefill speedup on NVIDIA H20 hardware compared to standard flash-linear-attention baselines.
- –Functions as a seamless drop-in backend for flash-linear-attention, minimizing integration friction for AI developers.
- –Built-in support for variable-length batching optimizes real-world LLM serving throughput and resource utilization.
DISCOVERED
2h ago
2026-07-27
PUBLISHED
2h ago
2026-07-27
RELEVANCE
AUTHOR
Kimi_Moonshot