ended9월 13일· 1 sources

The more aggressive matmul kernel lost to the register budget

Why it matters

The WebGPU matmul sweep started with a naive kernel, then added 16 by 16 workgroup tiling and a 4 by 4 output block per thread. At a 2048 cubed matrix size, the measured time moved from 47.24 ms for t...

1
Sources
+0
24h
Growth
8d
Active

Sources

Related Issues