EN Submit a tool
Product

Kimi Open-Sources FlashKDA Attention Kernel Implementation

Published: Source: X: Kimi.ai (@Kimi_Moonshot)

ShareXFacebookTelegramWhatsApp

We have open-sourced FlashKDA, our high-performance Kimi Delta Attention kernel implementation based on CUTLASS. It achieves 1.72x to 2.22x prefill speedup on H20 compared to the flash-linear-attention baseline and can serve as a plug-and-play backend for flash-linear-attention. Explore on GitHub: http://github.com/MoonshotAI/FlashKDA

Read the original (opens in a new tab)

News stream data aggregated by AI HOT

Related newsLatest in this category