new2시간 전· 1 sources
SFT vs. RL: What Changes Inside the Model?
Why it matters
In modern generative AI post-training, two fundamental paradigms dominate: Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR / GRPO). While practitioners often trea...
1
Sources
+1
24h
—
Growth
1d
Active