ended9월 13일· 1 sources
The more aggressive matmul kernel lost to the register budget
Why it matters
The WebGPU matmul sweep started with a naive kernel, then added 16 by 16 workgroup tiling and a 4 by 4 output block per thread. At a 2048 cubed matrix size, the measured time moved from 47.24 ms for t...
1
Sources
+0
24h
—
Growth
8d
Active