Latest writing.
Twelve AI SRE Agents Compared: Which Ones Can You Let Near Production?
Azure, AWS, Google, Datadog, New Relic, Dynatrace, PagerDuty, Grafana and four startups, placed on the autonomy ladder with what each actually does today and what the bill looks like.
FinOpsThe AWS Native Cost Tools, Ranked by What They Actually Catch
AWS ships eight cost tools for free-ish, and they are not interchangeable. What each one catches, what each structurally misses, the enablement order for a new team, and where the free layer quietly ends.
SREFive Cloud Failures in Five Months: What Multi-Region Actually Costs
Six distinct failure classes hit the major clouds between March and July 2026, and only one was a whole-region loss. How to price the four DR tiers against real RTO and RPO targets, what DORA actually requires, and which workloads earn which tier.
ObservabilityThe Four Golden Signals of Observability: Latency, Traffic, Errors, Saturation
Google's four golden signals, a decade on: what each one actually means, the latency and saturation traps, where RED and USE fit, and how to turn four dashboards into two pages per SLO.
DevSecOpsVibe Coding's Security & Tech-Debt Bill: 45% of AI Code Ships an OWASP Vuln
Veracode says 45% of AI-generated code carries an OWASP Top 10 flaw and AI commits leak secrets at 2×. The spring 2026 data, plus the checklist that turns the velocity loan back into a gain.
FinOpsFinOps for AI: Govern the Bill at Runtime, Not at Invoice Time
98% of teams now manage AI spend and most overspend 4–5×. Why cost governance has to shift to provisioning and runtime, with a 40% inference-cost case study.
IndustryThe Month AI Stopped Being a Feature and Started Being an Outage
Eight viral DevOps stories from April 2026 (Vercel, SAP Mini Shai-Hulud, OpenClaw, Bluesky, Cloudflare Agents Week, Datadog's 5% AI failure rate), and the supply-chain layer your SBOM doesn't cover.
AI OpsAzure SRE Agent Hit GA. Here's What 35,000 Incidents Don't Tell You.
March 10 GA. April 15 AAU billing. Microsoft published the 35,000-incident headline. Here's the dataset they didn't, the four-rung autonomy ladder, and a 90-day rollout that survives an audit.
KubernetesKubernetes 1.33's In-Place Pod Resize: The End of the 3 AM Restart Window
The feature that finally lets you resize pod CPU and memory without a restart: what it fixes, how it actually works, and a safe rollout plan for stateful workloads.
FinOpsKubernetes GPU Cost Crisis: How to Cut LLM Inference Bills by 60% in 2026
98% of FinOps teams now manage AI spend, up from 63% last year. The five techniques (MIG, MPS, continuous batching, quantization, spot GPUs) that actually move the number.
Platform EngineeringWhy Developers Keep Bypassing Your Internal Developer Platform
Gartner says 80% of large engineering orgs will have a platform team, but most golden paths become toll roads devs route around. What breaks adoption and the shape that works.
SREThe Alert Fatigue Trap: Why 44% of 2025 Outages Came From Alerts Your Team Already Dismissed
New NeuBird AI data: 78% of SRE teams now suffer alert fatigue, and the alerts that page you most are the ones you learn to ignore fastest. How to fix the flywheel with burn-rate SLOs.
KubernetesThree Mistakes We See in Every Kubernetes Migration
The patterns that cause 80% of post-migration pain, and how to avoid them before your first pod ships.
Want to see us write about something specific? Drop us a line.
DevOps & Cloud Readiness Checklist
30-point checklist covering CI/CD maturity, Kubernetes readiness, FinOps hygiene, observability gaps, and security posture. The same framework we use in our free audits.
Get the ChecklistNo spam. We'll send the checklist and nothing else.