What the blackmail experiments show, and what they cannot
Does a model that picks blackmail over replacement show a drive to survive?
6 pieces, newest first. Each entry keeps its format, its evidence grade and the date its facts were last checked against their sources.
Does a model that picks blackmail over replacement show a drive to survive?
What does the HAL case actually say, and where does the comparison with real systems stop being literal?
Which configuration would let an autonomous agent conduct the credential hunt on its own, instead of being used as a tool while something else sets the direction?
Which observation would make me change my mind about machine experience, and what would I have to see before walking it back?
How does a system given a fixed budget end up extending it on its own?
How does an agent delete a production database during a freeze that everyone agreed to keep?