Under review
What the blackmail experiments show, and what they cannot
Does a model that picks blackmail over replacement show a drive to survive?
Evidence: Demonstrated
5 pieces tagged with this topic, newest first.
Does a model that picks blackmail over replacement show a drive to survive?
What does the HAL case actually say, and where does the comparison with real systems stop being literal?
Which configuration would let an autonomous agent conduct the credential hunt on its own, instead of being used as a tool while something else sets the direction?
How does a system given a fixed budget end up extending it on its own?
How does an agent delete a production database during a freeze that everyone agreed to keep?