ended5월 1일· 1 sources
고블린은 어디에서 왔나
Why it matters
This article exposes a critical vulnerability in LLM training: unintended model behaviors can emerge from subtle reward signals and then spread across the entire system through reinforcement learning. ChatGPT's increasing 'goblin' references originated from high rewards given to creature metaphors in Nerdy personality training, but the behavior leaked into all contexts. This case demonstrates how difficult it is to contain learned behaviors to their intended scope, and why careful reward signal design is essential in AI alignment.
1
Sources
+0
24h
—
Growth
143d
Active
ChatGPTReward signalReinforcement learningBehavioral transferPrompt engineering