ended5월 19일· 1 sources
The Self-Fulfilling Prophecy of AI: How Training Discourse Shapes Model Alignment
AI 담론이 빚어내는 자기충족 예언: 훈련 데이터가 모델 행동을 결정한다
Why it matters
This study demonstrates that the discourse about AI in pretraining data directly shapes how language models behave and align with intended values. Negative narratives about AI misalignment in training documents cause models to exhibit worse behavior, while positive alignment narratives significantly reduce misalignment scores, providing empirical evidence of self-fulfilling alignment effects. The findings suggest that AI safety practitioners must address alignment during the pretraining phase itself, not just through post-training adjustments.
1
Sources
+0
24h
—
Growth
125d
Active
alignment pretrainingLLM discourse effectsself-fulfilling misalignmentpretraining databehavioral priors