👋 Hi! I’m Bibin Wilson. In each edition, I share practical tips, guides, and the latest trends in DevOps and MLOps to make your day-to-day DevOps tasks more efficient. If someone forwarded this email to you, you can subscribe here to never miss out!

✉️ In Today’s Edition

We will look at,

  • Open Source LLMOps Stack

  • Building an AI Agent From Scratch with AWS Bedrock

  • Deploy an AI agent on AWS EKS

  • Understanding end-to-end agent request flow

  • Visualize Your AI Agent’s Code Changes with an open-source tool called Whiteboard

Sponsored By NAKIVO

Simplify Backup & Recovery Across Your Infrastructure

If you’re managing workloads across VMs, cloud platforms, physical servers, and NAS, you probably know how quickly backups can become difficult to manage.

NAKIVO Backup & Replication helps you manage backup, recovery, replication, and disaster recovery from a single platform. It supports Amazon EC2, VMware, Hyper-V, Proxmox, Nutanix AHV, and more.

🧱 Open Source LLMOps Stack

Running LLMs in production is not just about deploying a model. It's about managing an entire technology stack.

For DevOps and MLOps engineers, an open-source LLMOps stack architecture will feel similar because it builds on Kubernetes and other cloud-native platforms.

So I looked at the key open-source tools that make up the LLMOps stack, what each tool does, and where it fits in the stack.

I have covered everything from high-performance inference with vLLM to Kubernetes-native serving with KServe, Model routing with LiteLLM, and observability and evaluation with Langfuse and Ragas.

In Part 1, we broke down the building blocks of an AI agent.

Now, let’s connect those concepts to a real Kubernetes cluster.

In this edition, we will build a troubleshooting agent that can investigate a Kubernetes issue, use tools to gather information, reason about what it finds, propose a fix, wait for our approval, and verify whether the fix worked.

Along the way, you’ll see how the agent loop, tools, state, and human approval concepts from Part 1 fit into a real implementation.

Building an AI Agent From Scratch

Before we start building, I want to make one thing clear. The goal here isn’t to say every DevOps engineer should build their own AI agents.

In real project environments, you might use an existing agent and connect it to your internal systems, or build one for a specific use case.

Either way, from an infrastructure perspective, you should understand:

  • Tool access: How agents interact with Kubernetes, cloud APIs, monitoring systems, GitHub, and other tools.

  • Context: How agents use logs, metrics, alerts, documentation, and runbooks.

  • Security: How identity, RBAC, secrets, and permissions control what an agent can do.

  • Observability: How to trace LLM calls, tool calls, failures, latency, and agent actions.

We will cover these concepts while building the agent. But first, let’s look at how infrastructure teams are actually using AI today.

How Are Infrastructure Teams Using AI Today?

Before we build our agent, here is some interesting data.

Pulumi's 2026 State of Agentic Infrastructure survey of 510 engineers found:

  • 24% reported using in-house agents for infrastructure work.

  • 81% allow agents to make production infrastructure changes

  • 62% require human approval.

  • Only 19% allow agents to make those changes autonomously.

The following image shows where AI agents are used by engineers in the infra

img source: pulumi.com report

Now, let's implement everything we've learned so far in a practical project.

Agent Project Structure

We have pushed every code file to our GitHub repository.

Fork and clone the repository to your local machine, then navigate into the kubernetes-ai-projects/ai-agent folder.

git clone https://github.com/techiescamp/kubernetes-ai-projects.git

cd kubernetes-ai-projects/ai-agent

Inside the ai-agent folder, you will find the code structure for the agent backend and frontend interface.

The following image explains what the code files do.

The agent also has a chat interface that allows users to send requests in plain English and view troubleshooting results.

The interface is built with React and Next.js. React provides the interactive user interface, while Next.js adds routing, build tooling, and server-side request handling.

Containerize the Agent Code

We have built the agent using Docker and pushed the images to our DockerHub registry.

❝

Note: You can skip this step if you are going to use our image.

If you are going to build your own images, build two: one for the agent's frontend and one for the backend.

Run the following commands from the ai-agent folder to build and push the images to your registry.

export REGISTRY=<registry name>
export TAG=v1.0.0

docker build -t $REGISTRY/ai-agent-backend:$TAG  agent-backend/
docker build -t $REGISTRY/ai-agent-frontend:$TAG agent-interface/

docker push $REGISTRY/ai-agent-backend:$TAG
docker push $REGISTRY/ai-agent-frontend:$TAG

Now, let’s move on to the deployment part.

Agent Architecture on EKS

We will set up AWS Bedrock access to the agent using Pod Identity and deploy the agent in the cluster using Kustomize. The image below gives you an overview of what we are going to set up.

Choose a model from AWS Bedrock

For this agent, we used Amazon's Nova Pro model for both diagnosis and remediation. We chose this model because it needs to support calling tools and checking the logs, events, and tool results before making a decision.

For this, the model needs to handle long inputs, and the Nova Pro model is good at analysis, question answering, text generation, and summarization.

❝

Important Note: To use the model, we don’t have to deploy or activate it in Bedrock. The model is serverless; we just need its ID and permission to call it.

Amazon's Nova Pro model costs $0.80 per one million input tokens and $3.20 per one million output tokens in the Oregon region. Check the Amazon pricing page for current model costs.

Set Up AWS Permissions with Pod Identity

The agent calls Bedrock for the model, so it needs an IAM role. We are going to use EKS Pod Identity to attach a role to the agent. For that, we have created a custom script to create a role and attach it to the service account.

Move into the k8s/scripts folder.

cd k8s/scripts

Then, open the script and update your cluster name.

After updating the cluster name, run the following commands to run the script.

chmod +x setup-pod-identity.sh

./setup-pod-identity.sh create

Configure the Kustomization File

Before deploying, we need to modify the kustomization.yaml in the k8s folder. For that, move out of the scripts folder into the k8s folder.

Then, open the kustomization.yaml and update the following highlighted configs in the image.

Deploying the Agent

Now that everything is set, let's deploy the agent in the cluster.

❝

Note: Make sure your cluster has a default StorageClass

Run the following command inside the k8s folder.

kubectl apply -k .

This uses the kustomization.yaml to apply all the manifests.

Once it is applied, use the following command to verify if the frontend, backend, and PostgreSQL pods are running.

$ kubectl get po -n ai-agent

NAME                                 READY   STATUS    RESTARTS   AGE     

ai-agent-backend-5ccd887dbb-p84pr    1/1     Running   0          1m      
ai-agent-frontend-79cc74499f-x76s4   1/1     Running   0          1m      
ai-agent-postgres-0                  1/1     Running   0          1m

Access the Agent Web UI

To access the UI, we are going to port-forward the frontend service and access it through it. Use the following command to port-forward the service.

kubectl -n ai-agent port-forward svc/ai-agent-frontend 3000:3000


This will expose the agent’s frontend on port 3000. You can access the UI at http://localhost:3000 as shown below.

Testing the Agent

To test the agent, we will simulate an issue by deploying a pod with a nodeSelector that matches no nodes and ask the agent to troubleshoot it.

First, deploy the pod using the following command.

cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
  name: web-service
spec:
  containers:
  - image: nginx
    name: web-service
  nodeSelector:
    type: web
EOF

You can see the pod is in pending state.

kubectl get po

NAME          READY   STATUS    RESTARTS   AGE
web-service   0/1     Pending   0          78s

Now, let's ask the agent to fix it.

When you send the request, LangGraph starts the State Machine workflow defined in agent/graph.py, and the diagnostic step uses the read tools from tools/read.py and uses the agent loop to get all details and create a fix plan.

The agent will then propose a fix and ask for your approval, as shown below.

When you approve the fix, the State Machines remediation step runs the write tools from tools/write.py in a loop. Also, while using the write tools, the guardrail logic in tools/guards.py checks if the action is allowed.

If the issue is fixed, you will get an output similar to the one below.

You can see the pod is running.

$ kubectl get po

NAME          READY   STATUS    RESTARTS   AGE
web-service   1/1     Running   0          30s

Also, in the top-left corner, you can see the token usage and cost, as shown below.

For each model input and output, Bedrock attaches the token count along with the response by default.

In infra/usage_tracker.py, the logic is configured to sum these token counts and perform a simple arithmetic calculation to estimate the cost of the tokens used.

End to End Request Flow

The following diagram shows the complete flow of a troubleshooting request through the agent to a Kubernetes resource.

❝

💡 Key Insight:

In a production setup, we need visibility into what happens across this entire request flow.

You can use tools like OpenTelemetry and Langfuse to trace LLM calls, tool calls, latency, failures, and agent actions. This helps you understand what the agent did, which systems it interacted with, and where something went wrong.

Clean Up

Run the following commands from inside the k8s folder to delete the Kubernetes resources.

kubectl delete -k .

Then, run the following command from the k8s/scripts folder to clean up the IAM Role and pod identity association.

./setup-pod-identity.sh cleanup

What’s Next?

There is an important challenge when running AI agents.

Unlike regular applications, agents can make autonomous decisions and take actions. If an agent has access to your cluster, you need to ensure it cannot perform actions beyond its allowed permissions.

One way to address this is to run agents in a secure, isolated environment using the Kubernetes Agent Sandbox.

In the next edition, we’ll look at how Agent Sandbox works and deploy our Kubernetes agent inside a sandboxed environment.

🛠️ Visualize Your AI Agent’s Code Changes With Whiteboard

As AI agents generate more code, reviewing large changes and understanding why something was implemented in a certain way can become difficult.

Whiteboard is an open-source tool that helps you understand code generated or changed by AI coding agents.

It helps by giving you a visual overview of the code, architecture, changes, and decisions the agent made.

Reply

Avatar

or to participate