ended7μ 14μΌΒ· 1 sources
Show HN: I RL-trained an agent that trains models with RL (for β$1.3k)
Why it matters
π Everything is open sourced including: the trained agent's weights (LoRA adapter on π€ HF), agent harness, task families, reward code, GPU orchestration, tinker RL training scripts, and retro write-up...
1
Sources
+0
24h
β
Growth
6d
Active