AI Foundationspredict · compress · act
act IX

The Agent Infrastructure

An agent that can run code needs a box it cannot escape. Sandboxing is the difference between a tool and a weapon.

39

A Safe Place to Act

before this →How Tokens Get Served

The moment an agent can execute code or drive a browser, it inherits the full power of a computer. Sandboxing is how you give it that power without giving it your machine.

the codecontainer / VMnetwork egress rulesLayers, outside in:1. no network by default2. throwaway filesystem, no host mount3. CPU + memory + time limits4. scoped, short-lived credentials5. full audit log of every actionassume the code is hostile — even if it is not
defence in depth: isolate the runtime, scope the permissions, review the rest

Least privilege

Least privilege means: give the agent exactly the access the task needs, no more. Read-only database credentials for a reporting agent. A scoped API key, not your root one. A disposable container instead of the host. Permissions should be explicit and reviewable, not inherited by accident.

Computer use

Computer use — an agent controlling a real desktop through screenshots and mouse and keyboard actions — is the most general and the most dangerous interface. It can do anything a person can, which means it can also be tricked into anything a person can be tricked into. Prompt injection in a web page or email becomes a remote-control channel.

Human-in-the-loop is the backstop: irreversible or high-stakes actions require explicit approval. The design question is not “how do we make it fully autonomous?” but “which decisions should stay human, and how do we make saying yes cheap?”

Safety is a systems property

not a line of prompt text

You cannot prompt your way to safety. A model told “do not delete files” will still delete files if a tool allows it. Safety lives in the architecture: what the agent can reach, what it can spend, and who reviews the consequence.

the mindsetApply the security mindset you would to an untrusted contractor with a login. Then apply it again, because the agent can act a thousand times a minute.
ISOLATEthrowaway runtime, no host access
LIMITscoped permissions and budgets
REVIEWhuman confirms what cannot be undone
introduces →sandboxpermissionleast privilegecomputer usehuman-in-the-loop
← previousHow Tokens Get Servednext →Watching the Loop