ended7월 21일· 1 sources
Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention
Why it matters
Hello, everyone. Attention becomes expensive very quickly as more text is given to an AI model. Can a Mac GPU make it faster when every token is restricted to looking only at nearby tokens? Today, I a...
1
Sources
+0
24h
—
Growth
4d
Active