Kubenatives
Subscribe
Sign in
Home
Notes
Courses
Archive
About
Triton Inference Server on Kubernetes: Multi-Model Serving
vLLM is great for LLMs. Triton is what you reach for when you need to serve 20 different models on one GPU cluster.
Jul 31
•
Sharon Sahadevan
RDMA and InfiniBand: Why GPU Networking Is Different on Kubernetes
A 70B model split across 8 GPUs moves hundreds of gigabytes per second between nodes. Standard Kubernetes pod networking handles zero of that well. Here…
Jul 27
•
Sharon Sahadevan
3
2
A/B Testing LLM Models in Production with Kubernetes
Swapping Llama 3.1 70B for your old 8B model in one deploy is how you break production. Here is how to do it safely with traffic splitting and real…
Jul 17
•
Sharon Sahadevan
3
1
2
API Server Latency: What Is Normal and What Is a Red Flag
Your API server p99 is 400ms. Is that fine or is the cluster about to fall over? Here are the numbers that matter and what to do when they go wrong.
Jul 10
•
Sharon Sahadevan
5
2
GPU Monitoring with DCGM Exporter: The Metrics That Matter
Your GPU nodes are running. nvidia-smi shows green. But are your GPUs healthy, efficient, and about to fail? DCGM tells you.
Jul 3
•
Sharon Sahadevan
6
Most Popular
View all
How I Solved a $50K Certificate Outage in 15 Minutes Using OSI Layers
Jul 22, 2025
•
Sharon Sahadevan
7
Architecture Template: vLLM Production Deployment on Kubernetes
Mar 14
•
Sharon Sahadevan
6
The OSI Model: Not Academic BS - Here's Why It Matters in Production
Jul 17, 2025
•
Sharon Sahadevan
15
3
DevOps to MLOps
Dec 16, 2025
•
Sharon Sahadevan
9
1
Latest
Top
Discussions
Resource Requests and Limits for GPU Workloads
Get requests wrong and your pods are Pending. Get limits wrong and they OOM. Here is how to size them correctly for GPU inference.
Jun 26
•
Sharon Sahadevan
2
1
Autoscaling Inference Workloads: HPA and KEDA for GPU Pods
GPU pods are expensive. Running 4 replicas at 3 AM when traffic is zero wastes thousands per month. Here is how to scale them automatically.
Jun 19
•
Sharon Sahadevan
3
Kubernetes Upgrade Strategy: kubeadm Cluster Upgrades Without Downtime
Kubernetes drops support for old versions every 12 months. Here is how to upgrade without breaking production.
Jun 12
•
Sharon Sahadevan
4
1
Network Policies in Practice: When Your Pods Cannot Talk to Each Other
You implemented network policies for security. Then DNS broke. Then inter-service communication broke. Here is how to do it without breaking everything.
Jun 5
•
Sharon Sahadevan
6
Architecture Template: GPU Node Pool Setup
Complete YAML for a multi-tier GPU cluster with taints, tolerations, affinity, quotas, and priority classes. Copy, configure, deploy.
May 29
•
Sharon Sahadevan
1
GPU Node Pools: Taints, Tolerations, and Cost Isolation
Without taints, a basic nginx pod can schedule on your $30K H100 node. Here is how to prevent that.
May 29
•
Sharon Sahadevan
3
1
LLMOps on Kubernetes: Patterns for Running LLMs in Production
Deploying the model is the easy part. Operating it in production is where most teams get stuck.
May 22
•
Sharon Sahadevan
3
Architecture Template: CoreDNS Debug ConfigMap
A production-ready CoreDNS configuration with logging, caching, and health checks for debugging DNS issues.
May 15
•
Sharon Sahadevan
1
Kubernetes DNS Troubleshooting: CoreDNS, ndots, and the 5-Second Timeout
Every DNS issue in Kubernetes traces back to one of 5 causes. Here is how to find which one in under 3 minutes.
May 15
•
Sharon Sahadevan
8
1
See all
Kubenatives
Production Kubernetes for ML/AI workloads: GPU infrastructure, control plane internals, and model serving patterns for engineers running inference at scale.
Subscribe
Recommendations
The System Design Newsletter
Neo Kim
AlgoMaster Newsletter
Ashish Pratap Singh
ByteByteGo Newsletter
Alex Xu
Kubenatives
Subscribe
About
Archive
Recommendations
Sitemap
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts