ended7월 2일· 1 sources

The safety switch that doesn't actually work

Why it matters

Sparse autoencoders — the core tool of mechanistic interpretability — can identify and amplify specific concepts inside a neural network, but they cannot reliably suppress unwanted behavior by clampin...

1
Sources
+0
24h
Growth
81d
Active

Sources

Related Issues