Alcides Fonseca

40.197958, -8.408312

Is sandboxing sufficient to contain rogue agents? – A Few Thoughts on Cryptographic Engineering

Unfortunately, monitoring for adversarial data access turns out to be one of the hardest problems you could imagine. The volume of data that agents produce is so high that no human being could possibly read it, and we probably wouldn’t recognize obfuscated malicious data even if we were looking directly at it. This means any attempt to monitor the inflow/outflow will have to be handled by other models. Thus, the future of agent sandboxing is (1) build a sandbox, (2) install an agent/model into it, (3) install a somewhat dumber/cheaper warden model to guard it, (4) hope you can trust the lunkhead to contain the wizard. And so on and so forth, as models become more intelligent and capable.

In other words: a warden-guarded sandbox is just another version of the alignment problem. You’re going to have to trust a model to do it, and that model will need to be at least some fraction as intelligent as the model it’s guarding. If you haven’t convinced yourself that it’s possible to build models you can trust, then sandboxing isn’t going to take you much farther.

— Matthew Green in Is sandboxing sufficient to contain rogue agents? – A Few Thoughts on Cryptographic Engineering

Matthew Green, a cryptographer at Johns Hopkins University, points out what I have been saying for a while: if you want your agents to be useful, they have to be able to access information and perform actions. How do you prevent those actions from being harmful? You use another agent to deny or allow each action. Thus alignment is still relevant, even if for the bouncer controlling the sandbox escape.

I actually believe we should be using deterministic guardrails that can be audited, instead of another model as a border patrol agent. I am working on such solution, which I will be sharing soon enough.

Read next