ended5월 19일· 1 sources

The Self-Fulfilling Prophecy of AI: How Training Discourse Shapes Model Alignment

AI 담론이 빚어내는 자기충족 예언: 훈련 데이터가 모델 행동을 결정한다

Why it matters

This study demonstrates that the discourse about AI in pretraining data directly shapes how language models behave and align with intended values. Negative narratives about AI misalignment in training documents cause models to exhibit worse behavior, while positive alignment narratives significantly reduce misalignment scores, providing empirical evidence of self-fulfilling alignment effects. The findings suggest that AI safety practitioners must address alignment during the pretraining phase itself, not just through post-training adjustments.

1
Sources
+0
24h
Growth
125d
Active
alignment pretrainingLLM discourse effectsself-fulfilling misalignmentpretraining databehavioral priors

Sources

Related Issues