GPU utilization is probably worse than your dashboard says
A fleet dashboard showing 80% GPU utilization looks healthy. I've learned to be suspicious of exactly that number. Here's what it's actually hiding, and what to measure instead.
Lessons from production. No thought leadership, no filler, just what we learned.
Every recommendation comes from production experience. We've built and run systems at a scale most teams only read about.
DevOps, SRE, platform engineering, security, and FinOps under one roof. No handoffs between vendors, no gaps between disciplines.
We make ourselves unnecessary. Handover means documentation, runbooks, and a team that owns what it runs.
New engineering notes as they publish. Subscribe via RSS, no email required.