ended7월 4일· 1 sources
Dispersion loss counteracts embedding condensation in small language models
Why it matters
More severe in smaller models than in larger counterparts (Figure 2). What makes LLMs better than small LMs? Data? Parameters? Geometry might play a role! What makes LLMs better than small LMs? Data? ...
1
Sources
+0
24h
—
Growth
79d
Active