ended3월 19일· 1 sources
NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute
NanoGPT Slowrun: 무한한 컴퓨팅으로 달성한 10배 데이터 효율성
Why it matters
The NanoGPT Slowrun project achieved 10x data efficiency by training a 2.7B parameter model on only 100M tokens, vastly exceeding Chinchilla scaling laws. Key techniques include ensemble training where models pushed past individual optima improve collective performance, chain knowledge distillation that progressively improves each model in a sequence, and looped transformers that apply more compute per prediction by iterating through middle layers.
1
Sources
+0
24h
—
Growth
185d
Active
NanoGPTdata efficiencyensemblingchain distillationlooped transformersscaling laws