👋 Hi! I’m Bibin Wilson. In each edition, I share practical tips, guides, and the latest trends in DevOps and MLOps to make your day-to-day DevOps tasks more efficient. If someone forwarded this email to you, you can subscribe here to never miss out!
✉️ In Today’s Edition
We will look at,
Open Source LLMOps Stack
Building an AI Agent From Scratch with AWS Bedrock
Deploy an AI agent on AWS EKS
Understanding end-to-end agent request flow
Visualize Your AI Agent’s Code Changes with an open-source tool called Whiteboard
Sponsored By NAKIVO
Simplify Backup & Recovery Across Your Infrastructure
If you’re managing workloads across VMs, cloud platforms, physical servers, and NAS, you probably know how quickly backups can become difficult to manage.
NAKIVO Backup & Replication helps you manage backup, recovery, replication, and disaster recovery from a single platform. It supports Amazon EC2, VMware, Hyper-V, Proxmox, Nutanix AHV, and more.
🧱 Open Source LLMOps Stack

Running LLMs in production is not just about deploying a model. It's about managing an entire technology stack.
For DevOps and MLOps engineers, an open-source LLMOps stack architecture will feel similar because it builds on Kubernetes and other cloud-native platforms.
So I looked at the key open-source tools that make up the LLMOps stack, what each tool does, and where it fits in the stack.
I have covered everything from high-performance inference with vLLM to Kubernetes-native serving with KServe, Model routing with LiteLLM, and observability and evaluation with Langfuse and Ragas.
Read it here: Best Open Source LLMOps Tools
In Part 1, we broke down the building blocks of an AI agent.
Now, let’s connect those concepts to a real Kubernetes cluster.
In this edition, we will build a troubleshooting agent that can investigate a Kubernetes issue, use tools to gather information, reason about what it finds, propose a fix, wait for our approval, and verify whether the fix worked.
Along the way, you’ll see how the agent loop, tools, state, and human approval concepts from Part 1 fit into a real implementation.
Building an AI Agent From Scratch
Before we start building, I want to make one thing clear. The goal here isn’t to say every DevOps engineer should build their own AI agents.
In real project environments, you might use an existing agent and connect it to your internal systems, or build one for a specific use case.
Either way, from an infrastructure perspective, you should understand:
Tool access: How agents interact with Kubernetes, cloud APIs, monitoring systems, GitHub, and other tools.
Context: How agents use logs, metrics, alerts, documentation, and runbooks.
Security: How identity, RBAC, secrets, and permissions control what an agent can do.
Observability: How to trace LLM calls, tool calls, failures, latency, and agent actions.
We will cover these concepts while building the agent. But first, let’s look at how infrastructure teams are actually using AI today.
How Are Infrastructure Teams Using AI Today?
Before we build our agent, here is some interesting data.
Pulumi's 2026 State of Agentic Infrastructure survey of 510 engineers found:
24% reported using in-house agents for infrastructure work.
81% allow agents to make production infrastructure changes
62% require human approval.
Only 19% allow agents to make those changes autonomously.
The following image shows where AI agents are used by engineers in the infra

img source: pulumi.com report
Now, let's implement everything we've learned so far in a practical project.
Agent Project Structure
We have pushed every code file to our GitHub repository.
Fork and clone the repository to your local machine, then navigate into the kubernetes-ai-projects/ai-agent folder.
git clone https://github.com/techiescamp/kubernetes-ai-projects.git
cd kubernetes-ai-projects/ai-agentInside the ai-agent folder, you will find the code structure for the agent backend and frontend interface.
The following image explains what the code files do.

The agent also has a chat interface that allows users to send requests in plain English and view troubleshooting results.
The interface is built with React and Next.js. React provides the interactive user interface, while Next.js adds routing, build tooling, and server-side request handling.
Containerize the Agent Code
We have built the agent using Docker and pushed the images to our DockerHub registry.
Note: You can skip this step if you are going to use our image.
If you are going to build your own images, build two: one for the agent's frontend and one for the backend.
Run the following commands from the ai-agent folder to build and push the images to your registry.
export REGISTRY=<registry name>
export TAG=v1.0.0
docker build -t $REGISTRY/ai-agent-backend:$TAG agent-backend/
docker build -t $REGISTRY/ai-agent-frontend:$TAG agent-interface/
docker push $REGISTRY/ai-agent-backend:$TAG
docker push $REGISTRY/ai-agent-frontend:$TAGNow, let’s move on to the deployment part.
Agent Architecture on EKS
We will set up AWS Bedrock access to the agent using Pod Identity and deploy the agent in the cluster using Kustomize. The image below gives you an overview of what we are going to set up.

Choose a model from AWS Bedrock
For this agent, we used Amazon's Nova Pro model for both diagnosis and remediation. We chose this model because it needs to support calling tools and checking the logs, events, and tool results before making a decision.
For this, the model needs to handle long inputs, and the Nova Pro model is good at analysis, question answering, text generation, and summarization.

Important Note: To use the model, we don’t have to deploy or activate it in Bedrock. The model is serverless; we just need its ID and permission to call it.
Amazon's Nova Pro model costs $0.80 per one million input tokens and $3.20 per one million output tokens in the Oregon region. Check the Amazon pricing page for current model costs.
Set Up AWS Permissions with Pod Identity
The agent calls Bedrock for the model, so it needs an IAM role. We are going to use EKS Pod Identity to attach a role to the agent. For that, we have created a custom script to create a role and attach it to the service account.
Move into the k8s/scripts folder.
cd k8s/scriptsThen, open the script and update your cluster name.
After updating the cluster name, run the following commands to run the script.
chmod +x setup-pod-identity.sh
./setup-pod-identity.sh createConfigure the Kustomization File
Before deploying, we need to modify the kustomization.yaml in the k8s folder. For that, move out of the scripts folder into the k8s folder.
Then, open the kustomization.yaml and update the following highlighted configs in the image.

Deploying the Agent
Now that everything is set, let's deploy the agent in the cluster.
Note: Make sure your cluster has a default StorageClass
Run the following command inside the k8s folder.
kubectl apply -k .This uses the kustomization.yaml to apply all the manifests.
Once it is applied, use the following command to verify if the frontend, backend, and PostgreSQL pods are running.
$ kubectl get po -n ai-agent
NAME READY STATUS RESTARTS AGE
ai-agent-backend-5ccd887dbb-p84pr 1/1 Running 0 1m
ai-agent-frontend-79cc74499f-x76s4 1/1 Running 0 1m
ai-agent-postgres-0 1/1 Running 0 1mAccess the Agent Web UI
To access the UI, we are going to port-forward the frontend service and access it through it. Use the following command to port-forward the service.
kubectl -n ai-agent port-forward svc/ai-agent-frontend 3000:3000
This will expose the agent’s frontend on port 3000. You can access the UI at http://localhost:3000 as shown below.

Testing the Agent
To test the agent, we will simulate an issue by deploying a pod with a nodeSelector that matches no nodes and ask the agent to troubleshoot it.
First, deploy the pod using the following command.
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: web-service
spec:
containers:
- image: nginx
name: web-service
nodeSelector:
type: web
EOFYou can see the pod is in pending state.
kubectl get po
NAME READY STATUS RESTARTS AGE
web-service 0/1 Pending 0 78sNow, let's ask the agent to fix it.

When you send the request, LangGraph starts the State Machine workflow defined in agent/graph.py, and the diagnostic step uses the read tools from tools/read.py and uses the agent loop to get all details and create a fix plan.
The agent will then propose a fix and ask for your approval, as shown below.

When you approve the fix, the State Machines remediation step runs the write tools from tools/write.py in a loop. Also, while using the write tools, the guardrail logic in tools/guards.py checks if the action is allowed.
If the issue is fixed, you will get an output similar to the one below.

You can see the pod is running.
$ kubectl get po
NAME READY STATUS RESTARTS AGE
web-service 1/1 Running 0 30sAlso, in the top-left corner, you can see the token usage and cost, as shown below.

For each model input and output, Bedrock attaches the token count along with the response by default.
In infra/usage_tracker.py, the logic is configured to sum these token counts and perform a simple arithmetic calculation to estimate the cost of the tokens used.
End to End Request Flow
The following diagram shows the complete flow of a troubleshooting request through the agent to a Kubernetes resource.

💡 Key Insight:
In a production setup, we need visibility into what happens across this entire request flow.
You can use tools like OpenTelemetry and Langfuse to trace LLM calls, tool calls, latency, failures, and agent actions. This helps you understand what the agent did, which systems it interacted with, and where something went wrong.
Clean Up
Run the following commands from inside the k8s folder to delete the Kubernetes resources.
kubectl delete -k .Then, run the following command from the k8s/scripts folder to clean up the IAM Role and pod identity association.
./setup-pod-identity.sh cleanupWhat’s Next?
There is an important challenge when running AI agents.
Unlike regular applications, agents can make autonomous decisions and take actions. If an agent has access to your cluster, you need to ensure it cannot perform actions beyond its allowed permissions.
One way to address this is to run agents in a secure, isolated environment using the Kubernetes Agent Sandbox.
In the next edition, we’ll look at how Agent Sandbox works and deploy our Kubernetes agent inside a sandboxed environment.
🛠️ Visualize Your AI Agent’s Code Changes With Whiteboard
As AI agents generate more code, reviewing large changes and understanding why something was implemented in a certain way can become difficult.
Whiteboard is an open-source tool that helps you understand code generated or changed by AI coding agents.
It helps by giving you a visual overview of the code, architecture, changes, and decisions the agent made.
👉 GitHub: https://github.com/devdotfast/whiteboard

