AWS sandbox aims to tame rogue AI agents
Amazon's new open-source Strands Box uses OS-level isolation and policy rules to stop autonomous agents from going rogue.
Autonomous AI agents are increasingly being trusted to run commands, call APIs, and push code without a human checking every move. AWS has now released an open-source tool, Strands Box, designed to put hard boundaries around that behavior at the operating system level.
The project combines a sandbox with policy enforcement, aiming to stop agents from taking destructive actions even when they are operating without human review. It is the latest in a series of open-source control tools from the cloud giant, and it arrives as developers grapple with how to safely deploy agents that can act on their own.
Why containers aren't enough
According to AWS, the usual approach of isolating agents in containers or microVMs provides strong separation but lacks contextual rule enforcement. Once an agent inside such an environment gains access to a tool, there may be nothing to stop it from deleting a production database or reaching out to the internet.
The AWS team framed the problem in its announcement:
“Agents increasingly run in ‘YOLO mode,’ approving every action without human review,” the AWS team explained in its announcement. “The usual solution to this problem is a sandbox … but access is only part of what we want to control.”
Strands Box addresses that gap by layering policy checks on top of isolation. It uses the Dogwood Local Engine to give the policy engine temporal awareness, meaning tool calls are evaluated not just on what the agent wants to do but on what it has already done.
From Slack spam to costly API calls
AWS offered concrete examples of how the policy engine can be configured. An agent could be allowed to post status updates to Slack, but no more than three times every ten minutes to prevent it from spamming its human operators, the company said.
An AWS spokesperson further explained that Box could control when an agent can perform a Git push, or cap API calls that could end up costing a small fortune.
Those controls are intended to give developers a way to set precise limits without relying on the agent to follow instructions. AWS describes the enforcement as deterministic, meaning the agent cannot talk its way around the rules.
Making agent actions visible
Strands Box also includes Strands Shell and Monty for Python, which expose shell and Python operations to the same Dogwood policy engine and event history. That makes agentic actions clearer to developers and allows policies to account for what an agent is trying to do.
Marc Brooker, AWS VP and distinguished engineer, told The Register that the interpreters are a key part of making agentic behavior more intelligible. By surfacing operations such as file deletions, API requests, and tool calls, they allow developers to write more precise policies.
“Box’s Shell and Python interpreters expose operations such as file deletions, while its gateways expose API requests and tool calls,” Brooker told us in an email. “Policies can then account for the action being attempted and earlier activity.”
Brooker added that Box enforces those rules without trusting or relying on agents to actually follow instructions, which should ideally prevent them from running roughshod over their operators’ wishes.
The limits of deterministic control
Even with properly configured permissions, Brooker cautioned that an agentic action can still produce an unwanted result. The sandbox enforces the policies developers configure, but it does not eliminate the need for human judgment.
“Box enforces the policies developers configure, deterministically, and the agent can't talk its way around these rules,” Brooker explained.
He also made clear that responsibility for access decisions remains with developers. “Developers remain responsible for deciding what access to grant and where human review is needed,” Brooker added - in other words, don’t let your YOLO mode go too YOLO. A bit of human oversight is still necessary.
Brooker said AWS sees agent safety as an area where the industry still has significant work to do, and the company is committed to continuing to invest in it both inside the AWS cloud and in open source.
Availability and platform support
Strands Box supports any agent or harness that a developer wants to confine. It is available on GitHub now, though only for macOS for the time being.
Linux support is in development, and AWS told The Register that a Windows client is “on our radar,” but neither has a planned release date. Deployment to platforms like AgentCore, ECS, and Kubernetes is also planned.
What it means for teams running agents
For organizations experimenting with autonomous agents, the arrival of a policy-aware sandbox from a major cloud provider offers a template for how to think about control. The tool does not remove the need for careful configuration, but it does provide a mechanism to enforce limits that agents cannot override.
That could matter most in environments where an agent has access to production systems, cloud APIs, or code repositories. A misconfigured or overly permissive agent can cause damage quickly, and the temporal awareness built into Strands Box is designed to catch patterns like repeated actions that might otherwise go unnoticed.
Still, the responsibility for setting sensible policies remains with the teams deploying agents. AWS has made the tool open source and available for macOS, with broader platform support in the works. For now, the message from AWS is that sandboxing alone is not enough - policy enforcement has to go hand in hand with isolation.
Sources
- The Register Original source
Continue Reading
Melius bets on ad creative after a full reset
Ex-Ramp engineers raise $25M for Melius, an AI platform for generating ad campaigns, after scrapping their first product entirely.
Ghost's $3,499 AI box sells out first run
A teen-founded startup says its on-device AI computer sold out its first batch, betting privacy beats cloud convenience.
Flai's AI books 50K dealer appointments
Flai has raised a $27M Series A after its AI software reached 50,000 monthly appointments for car dealerships.