虎嗅

OpenSandbox: Rethinking the Runtime of the Agent Era

原文:OpenSandbox:重新思考Agent 时代的Runtime

Summary of Key Points

OpenSandbox is an open-source runtime designed by Alibaba for the AI Agent era, addressing the shortcomings of traditional container systems (Docker/K8s) in agent-based scenarios. Its core design philosophy emphasizes "protocol priority," and it enhances batch delivery efficiency through pooled scheduling. It establishes a three-layer security defense system and has been implemented in various use cases such as autonomous agents, batch evaluation, and RL training. In the future, OpenSandbox will further optimize state management, persistence of working spaces, and observability, with the goal of becoming a universal, efficient, and secure foundation for AI Agents.

Why Traditional Container Systems (Docker/K8s) Are No Longer Sufficient?

AI Agents differ from traditional software: they need to operate files autonomously, expose services, and connect to networks, often requiring rapid batch initiation (for evaluation or training purposes). However, Docker/K8s are designed for "online services," not specifically for agents. The main issues with these systems include:

1. Slow Batch Delivery: Creating 100 sandboxes requires hundreds of write operations in K8s (resource creation, status updates, etc.), which becomes inefficient at larger scales.

2. Lack of a Unified Interface: When agents expose services via HTTP or WebSocket, upper-layer applications must handle different details, leading to system complexity.

3. Insufficient Fine-Grained Control: Agents need network access (e.g., to GitHub or model APIs), but traditional systems either block everything or allow everything, making it difficult to precisely control what can and cannot be accessed.

In short, while traditional containers can make agents run, they do so inefficiently and with limited control, resulting in an unsatisfactory user experience.

OpenSandbox's "Protocol Priority": Define Rules Before Choosing Tools

OpenSandbox's approach is to abstract capabilities first and then select the appropriate runtime. It defines a unified set of protocols that ensure that upper-layer applications and agents do not need to be modified, and the underlying runtime can be flexibly replaced. These protocols cover four areas:

1. Lifecycle Management: How to create, delete, and view sandboxes.

2. Command Execution: How agents execute code, manipulate files, and perform background tasks within the sandbox.

3. Network Policy: Precise control over which domain names/IP ranges agents can access (e.g., allowing access to company APIs while blocking external threats).

4. Access Control: All services exposed by agents go through a unified interface, freeing upper-layer applications from managing detailed settings.

To illustrate: Protocols act like "universal socket standards"—any runtime that meets these standards can be used, without the need for application-specific adaptations.

Batch Delivery Speed Increased by 10 Times: The Magic of Pooled Scheduling and BatchSandbox

Batch scenarios (e.g., evaluating 100 agent models) require speed. OpenSandbox addresses this issue with two techniques:

1. Pooled Preheating: Prepare sandbox resources in advance, similar to a restaurant preparing tableware, eliminating the need for initial setup.

2. BatchSandbox: Treat the creation of 100 sandboxes as a single batch operation instead of 100 individual requests. This reduces coordination overhead by more than 90%.

In tests, OpenSandbox outperforms Agent Sandboxes in the K8s community by an order of magnitude (e.g., while K8S takes 10 seconds, OpenSandbox completes the task in 1 second).

Three-Layer Security Defense: Ensuring Agents Run Properly

The more powerful agents become, the greater the potential for risks (e.g., accidental file deletion or access to sensitive data). OpenSandbox provides three layers of protection:

1. Isolation Layer: Creates independent environments for agents (using containers or lightweight VMs) to prevent them from affecting other sandboxes.

2. Network Control Layer: Dual-layer security to block dangerous accesses: first with DNS hijacking, then with IP filtering to block direct requests that bypass DNS.

3. Unified Governance Layer: All service access is through a single point, facilitating monitoring (e.g., tracking who accessed what) and auditing (e.g., extending sandbox lifetimes based on usage patterns).

These security measures are transparent to agents, allowing them to operate securely without requiring code changes.

Typical Use Cases

OpenSandbox can be utilized in the following scenarios:

1. Autonomous Agents: Integrating agents within sandboxes for both complex tasks and secure management (e.g., providing necessary authorization for projects like OpenClaw).

2. Batch Evaluation: Ensuring fairness and efficiency by assigning each evaluation task to its own sandbox.

3. RL Training: Supporting rapid creation and destruction of numerous sandboxes for trial-and-error training while maintaining security.

Future Directions

OpenSandbox is still evolving, with plans to add features such as sandbox suspension/resumption, persistence of working spaces, and enhanced observability. It has already joined the CNCF ecosystem and will continue to be open-sourced and improved.

In summary, OpenSandbox represents the "new infrastructure" for the AI Agent era, enabling agents to run faster, more securely, and with better control. If you are working on AI Agent-related projects, this tool is worth considering.