Product
Kimi Open-Sources FlashKDA Attention Kernel Implementation
We have open-sourced FlashKDA, our high-performance Kimi Delta Attention kernel implementation based on CUTLASS. It achieves 1.72x to 2.22x prefill speedup on H20 compared to the flash-linear-attention baseline and can serve as a plug-and-play backend for flash-linear-attention. Explore on GitHub: http://github.com/MoonshotAI/FlashKDA
Read the original (opens in a new tab)
News stream data aggregated by AI HOT