Tag: FlashAttention
All the articles with the tag "FlashAttention".
-
LLM 算法 LeetCode(二):Prefill 与 Attention Kernel
FlashAttention 不改 attention 的计算复杂度,而是用 IO-aware 的分块 + 在线 softmax,把 N×N 中间矩阵从显存流量中拿掉,实现精确而非近似的加速。
All the articles with the tag "FlashAttention".
FlashAttention 不改 attention 的计算复杂度,而是用 IO-aware 的分块 + 在线 softmax,把 N×N 中间矩阵从显存流量中拿掉,实现精确而非近似的加速。