ended5월 1일· 1 sources

고블린은 어디에서 왔나

Why it matters

This article exposes a critical vulnerability in LLM training: unintended model behaviors can emerge from subtle reward signals and then spread across the entire system through reinforcement learning. ChatGPT's increasing 'goblin' references originated from high rewards given to creature metaphors in Nerdy personality training, but the behavior leaked into all contexts. This case demonstrates how difficult it is to contain learned behaviors to their intended scope, and why careful reward signal design is essential in AI alignment.

1
Sources
+0
24h
Growth
131d
Active
ChatGPTReward signalReinforcement learningBehavioral transferPrompt engineering

Sources

Related Issues