ended7월 21일· 1 sources

Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention

Why it matters

Hello, everyone. Attention becomes expensive very quickly as more text is given to an AI model. Can a Mac GPU make it faster when every token is restricted to looking only at nearby tokens? Today, I a...

1
Sources
+0
24h
Growth
4d
Active

Sources

Related Issues