ended3일 전· 1 sources

They Put 7 Attention Mechanisms on a Latin Square. Then Removed Them One by One.

Why it matters

Since GPT, nearly every Transformer repeats the same attention mechanism at every layer. Forty-eight identical blocks, differing only in learned weights. Nobody tested that. It is a convention, not a ...

1
Sources
+0
24h
Growth
3d
Active

Sources

Related Issues