👋 Hi! I’m Bibin Wilson. In each edition, I share practical tips, guides, and the latest trends in DevOps and MLOps to make your day-to-day DevOps tasks more efficient. If someone forwarded this email to you, you can subscribe here to never miss out!

✉️ In Today’s Edition

In today’s edition, you will learn the following.

  • Why AI agents need an isolated workspace to run untrusted code

  • How Kubernetes Agent Sandbox manages these workspaces

  • How gVisor and Kata Containers isolate agent workloads

  • Hands-on: Deploy a Kubernetes troubleshooting agent and run untrusted code

  • Open Source Radar: A tool worth checking out

  • and more..

Let’s get into it.

🧱 55% OFF Linux Foundation Sale (Only Today)

If you are planning to take CKA, CKAD, CKS certifications, you can save up to 55% today during the Linux Foundation Prime Day sale.

Use code OCTPRIME26CCCT at kube.promo/devops to save a flat 40% on individual certifications.

Save up to 55% on Kubernetes certification bundles with code OCTPRIME26BCT.

In the AI Agent Basics and Building the Agent editions, I covered agent concepts and how to build and deploy one in a Kubernetes cluster.

In this edition, we will look at how to run them securely.

Use Case

An AI agent may need a workspace to run commands, install tools, and save files while completing a task.

The agent may generate some of that code or download it from external sources. We should treat it as untrusted code and run it in an isolated workspace with limited access.

This is where Kubernetes Agent Sandbox comes in. It manages a separate, single-Pod workspace environment for each agent session.

For example, a coding agent can get a sandbox, download a repository, install dependencies, and run tests. You can then pause its environment or delete it when the session ends.

Kubernetes Agent Sandbox

Agent Sandbox is a Kubernetes SIG apps project that provides isolated, stateful, singleton sandboxes (one pod) for running AI Agents or agent-generated untrusted code.

It provides,

  • Warm Pods: Pods prepared in advance so agents can start working faster.

  • Stable identity: A consistent name and address for each sandbox.

  • Scheduled deletion: Automatically deleting a sandbox when its time is up.

  • Suspend and resume: Stop a sandbox when it is idle and restart it when needed.

Next, let’s understand how the sandbox is isolated.

Agent Sandbox Isolation

To provide strong isolation for agent workloads, the Agent Sandbox uses sandboxed containers.

For example, it can use gVisor to separate the agent’s code from the host system. I covered this in our gVisor security and sandboxed containers guide.

So, the gVisor runtime provides the isolation, and the Agent sandbox orchestrates it.

❝

Note: The Agent sandbox is not tied only to gVisor. It also supports Kata Containers, which use hardware virtualization to run the sandbox in its own lightweight VM with its own kernel.

Creating a Sandbox (Example)

Once you install the Agent Sandbox CRD and controller, you can create a sandbox by applying the following custom resource.

apiVersion: agents.x-k8s.io/v1beta1
kind: Sandbox
metadata:
  name: python-sandbox
spec:
  operatingMode: Running
  podTemplate:
    spec:
      restartPolicy: Never
      containers:
        - name: workspace
          image: python:3.12-slim
          command: ["sleep", "infinity"]
          workingDir: /workspace
          volumeMounts:
            - name: workspace
              mountPath: /workspace
      volumes:
        - name: workspace
          emptyDir: {}

The custom resource above deploys an agent sandbox with the specified image and configuration and uses the underlying sandboxed runtime, such as gVisor.

❝

Note: To use a sandboxed runtime like gVisor or Kata Containers, you need to configure it in the cluster. Read this guide to learn more.

Here, you need to wait for the pod to start and define each sandbox configuration individually.

However, Kubernetes Agent Sandbox also provides a workflow with reusable templates and separation of responsibilities, loosely similar to how we manage persistent volumes and claims in Kubernetes.

Let’s look into the reusable workflow.

Agent Sandbox Workflow

The following image illustrates how the agent sandbox works

As shown in the image, the workflow starts with the platform administrator creating two resources: SandboxTemplate and SandboxWarmPool.

First, we create a SandboxTemplate with the required image and pod configuration. We can reuse this template to create multiple sandboxes.

Next, we create a SandboxWarmPool that references the template. The controller creates the number of ready sandboxes specified by replicas.

These are pre-warmed sandboxes, meaning their pods are already running and ready to use.

When a user or agent application needs a sandbox, it creates a SandboxClaim that references the warm pool. The controller notices the claim and assigns an available sandbox from the pool.

The warm pool then creates a replacement to maintain the configured number of ready sandboxes.

In the setup shown here, the agent application sends requests through the Sandbox Router (lightweight proxy). The router forwards them to the assigned sandbox pod and returns the results to the application.

This separates the responsibilities: the platform administrator manages the sandbox configuration and pool, while users and applications request sandboxes through claims.

❝

Note: Warm pools reduce startup time, but the waiting pods still use resources.

Agent Sandbox Key Features

The following two features help avoid paying for idle or forgotten sandbox workloads.

  1. Hibernation & Resume: We can pause an idle sandbox by changing its operating mode from Running to Suspended. The controller deletes the pod while keeping its persistent storage. When we switch it back to Running, the controller recreates the pod and reconnects the storage.

  2. Scheduled Deletion: We can set a shutdownTime with the Delete policy to automatically delete a sandbox at a specified time. This helps clean up forgotten sandboxes and avoid unnecessary resource usage.

❝

In cloud environments, this can reduce compute costs if the freed capacity lets the cluster scale down its nodes. Persistent storage may still incur costs while the sandbox is suspended.

Agent Sandbox Hands-on

We published a detailed hands-on guide on using the Kubernetes agent sandbox. It covers the following.

  • Setting up the Agent Sandbox Controller and Sandbox Router

  • Deploying our Kubernetes troubleshooting AI agent inside a sandbox

  • Running untrusted code inside a sandbox using the Agent Sandbox SDK

  • Testing scheduled deletion, hibernation, and resume

👉 Detailed Guide: Read the Detailed Guide Here

🛠️ DevOps Tool of the Week (Radar)

When troubleshooting Kubernetes, you may use multiple tools like kubectl, Helm, Argo CD, network dashboards, and more to understand what is happening inside the cluster.

Radar gives you a single place to see it all.

It is an open-source Kubernetes UI that runs locally as a single binary and connects directly to your cluster.

Here is what it does 👇

  • Visualize how Kubernetes resources are connected

  • Manage Helm releases, revisions, upgrades, and rollbacks

  • Check Argo CD and Flux sync/drift status

  • Visualize live service traffic using Hubble, Caretta, Istio, or Beyla

  • Connect AI tools such as Claude and Cursor through its built-in MCP server

  • and more..

You can start locally without deploying another agent into the cluster. For teams, Radar can also run inside Kubernetes with authentication.

Reply

Avatar

or to participate