new2시간 전· 1 sources

SFT vs. RL: What Changes Inside the Model?

Why it matters

In modern generative AI post-training, two fundamental paradigms dominate: Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR / GRPO). While practitioners often trea...

1
Sources
+1
24h
—
Growth
1d
Active

Sources

Related Issues