ended7월 2일· 1 sources

Head-level attention fusion trims compute

Why it matters

Merging full‑attention and linear‑attention at the head granularity slashes transformer FLOPs without appreciably hurting downstream quality. The trick is to keep the expensive quadratic path only whe...

1
Sources
+0
24h
Growth
4d
Active

Sources

Related Issues