ended7월 2일· 1 sources
Head-level attention fusion trims compute
Why it matters
Merging full‑attention and linear‑attention at the head granularity slashes transformer FLOPs without appreciably hurting downstream quality. The trick is to keep the expensive quadratic path only whe...
1
Sources
+0
24h
—
Growth
4d
Active