ended7월 2일· 1 sources
The safety switch that doesn't actually work
Why it matters
Sparse autoencoders — the core tool of mechanistic interpretability — can identify and amplify specific concepts inside a neural network, but they cannot reliably suppress unwanted behavior by clampin...
1
Sources
+0
24h
—
Growth
81d
Active