ended3월 19일· 1 sources

NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute

NanoGPT Slowrun: 무한한 컴퓨팅으로 달성한 10배 데이터 효율성

Why it matters

The NanoGPT Slowrun project achieved 10x data efficiency by training a 2.7B parameter model on only 100M tokens, vastly exceeding Chinchilla scaling laws. Key techniques include ensemble training where models pushed past individual optima improve collective performance, chain knowledge distillation that progressively improves each model in a sequence, and looped transformers that apply more compute per prediction by iterating through middle layers.

1
Sources
+0
24h
Growth
185d
Active
NanoGPTdata efficiencyensemblingchain distillationlooped transformersscaling laws

Sources

Related Issues