Flash Attentionを用いてトランスフォーマーのアテンションを最適化し、2~4倍の高速化と10~20倍のメモリ削減を実現します。長いシーケンスでトランスフォーマーを訓練・実行する際に使用します…
npx skills add https://github.com/firecrawl/ai-research-skills --skill optimizing-attention-flash