👋 Everything about EKS & AI Infrastructure Newsletter "#75" ☁️❤👨💻
GPU scheduling matures, agent governance gets teeth, and the boring reliability work that decides whether any of this survives contact with production.

Search for a command to run...
GPU scheduling matures, agent governance gets teeth, and the boring reliability work that decides whether any of this survives contact with production.

You've been running Kubernetes long enough to know what it's good at. You've also hit the wall that every ML team hits when they try to run serious distributed training on it. HyperPod on EKS is AWS's
You have vLLM running. You have Kubernetes. You have Karpenter. And yet, the moment you try to serve multiple models across multiple clusters, you're writing glue code that no one else can see or bene
I've spent years running containers on Kubernetes. EKS clusters, GPU workloads, Karpenter node pools — the whole stack. The mental model was stable: a container is a process with namespace isolation,
There's a specific kind of frustration that comes from reading a book that's half about what you actually needed. Most resources on running AI in production either assume you're a data scientist who p

From integer counting to structured resources — how Dynamic Resource Allocation and the AI Cluster Readiness framework finally make GPU infrastructure manageable at scale. Contents The Two Nightmares

Your vLLM cluster has a problem you probably don't know about. It's not a bug. Nothing is crashing. The metrics dashboard looks fine. But right now, every time a request hits your load balancer, there

There's a class of production incident that doesn't page anyone. No error rate spikes. No latency alert fires. The cluster health dashboard shows green. GPU nodes are online. Pods are running. And yet

Travel has been relentless lately. Back-to-back weeks, airports blurring into each other, calendar looking like a game of Tetris someone is losing badly. But AWS Community Day Pune was non-negotiable.
