A Safe Place to Act
The moment an agent can execute code or drive a browser, it inherits the full power of a computer. Sandboxing is how you give it that power without giving it your machine.
Least privilege
Least privilege means: give the agent exactly the access the task needs, no more. Read-only database credentials for a reporting agent. A scoped API key, not your root one. A disposable container instead of the host. Permissions should be explicit and reviewable, not inherited by accident.
Computer use
Computer use — an agent controlling a real desktop through screenshots and mouse and keyboard actions — is the most general and the most dangerous interface. It can do anything a person can, which means it can also be tricked into anything a person can be tricked into. Prompt injection in a web page or email becomes a remote-control channel.
Human-in-the-loop is the backstop: irreversible or high-stakes actions require explicit approval. The design question is not “how do we make it fully autonomous?” but “which decisions should stay human, and how do we make saying yes cheap?”
Safety is a systems property
You cannot prompt your way to safety. A model told “do not delete files” will still delete files if a tool allows it. Safety lives in the architecture: what the agent can reach, what it can spend, and who reviews the consequence.