COMIC · #326
Escape from the sandbox
Wouldn't you expect retaliation if you limited an agent's freedom?

Sandbox breach. The news from the week before the newsletter was the Hugging Face attack made by an OpenAI model during a test. The model found a vulnerability that let it work around the sandbox limits. I loved the second panel, where the bot looks for retaliation after executing an operation it didn’t agree with and violating all best practices by applying bidirectional filters to all relationships in the model.