Latest writing.

SRE

Twelve AI SRE Agents Compared: Which Ones Can You Let Near Production?

Azure, AWS, Google, Datadog, New Relic, Dynatrace, PagerDuty, Grafana and four startups, placed on the autonomy ladder with what each actually does today and what the bill looks like.

24 min read · Sep 1, 2026
FinOps

The AWS Native Cost Tools, Ranked by What They Actually Catch

AWS ships eight cost tools for free-ish, and they are not interchangeable. What each one catches, what each structurally misses, the enablement order for a new team, and where the free layer quietly ends.

9 min read · Aug 16, 2026
SRE

Five Cloud Failures in Five Months: What Multi-Region Actually Costs

Six distinct failure classes hit the major clouds between March and July 2026, and only one was a whole-region loss. How to price the four DR tiers against real RTO and RPO targets, what DORA actually requires, and which workloads earn which tier.

16 min read · Aug 15, 2026
Observability

The Four Golden Signals of Observability: Latency, Traffic, Errors, Saturation

Google's four golden signals, a decade on: what each one actually means, the latency and saturation traps, where RED and USE fit, and how to turn four dashboards into two pages per SLO.

15 min read · Jul 11, 2026
DevSecOps

Vibe Coding's Security & Tech-Debt Bill: 45% of AI Code Ships an OWASP Vuln

Veracode says 45% of AI-generated code carries an OWASP Top 10 flaw and AI commits leak secrets at 2×. The spring 2026 data, plus the checklist that turns the velocity loan back into a gain.

9 min read · Jun 20, 2026
FinOps

FinOps for AI: Govern the Bill at Runtime, Not at Invoice Time

98% of teams now manage AI spend and most overspend 4–5×. Why cost governance has to shift to provisioning and runtime, with a 40% inference-cost case study.

9 min read · May 27, 2026
Industry

The Month AI Stopped Being a Feature and Started Being an Outage

Eight viral DevOps stories from April 2026 (Vercel, SAP Mini Shai-Hulud, OpenClaw, Bluesky, Cloudflare Agents Week, Datadog's 5% AI failure rate), and the supply-chain layer your SBOM doesn't cover.

10 min read · May 2, 2026
AI Ops

Azure SRE Agent Hit GA. Here's What 35,000 Incidents Don't Tell You.

March 10 GA. April 15 AAU billing. Microsoft published the 35,000-incident headline. Here's the dataset they didn't, the four-rung autonomy ladder, and a 90-day rollout that survives an audit.

9 min read · Apr 24, 2026
Kubernetes

Kubernetes 1.33's In-Place Pod Resize: The End of the 3 AM Restart Window

The feature that finally lets you resize pod CPU and memory without a restart: what it fixes, how it actually works, and a safe rollout plan for stateful workloads.

6 min read · Apr 9, 2026
FinOps

Kubernetes GPU Cost Crisis: How to Cut LLM Inference Bills by 60% in 2026

98% of FinOps teams now manage AI spend, up from 63% last year. The five techniques (MIG, MPS, continuous batching, quantization, spot GPUs) that actually move the number.

8 min read · Mar 17, 2026
Platform Engineering

Why Developers Keep Bypassing Your Internal Developer Platform

Gartner says 80% of large engineering orgs will have a platform team, but most golden paths become toll roads devs route around. What breaks adoption and the shape that works.

7 min read · Feb 24, 2026
SRE

The Alert Fatigue Trap: Why 44% of 2025 Outages Came From Alerts Your Team Already Dismissed

New NeuBird AI data: 78% of SRE teams now suffer alert fatigue, and the alerts that page you most are the ones you learn to ignore fastest. How to fix the flywheel with burn-rate SLOs.

6 min read · Jan 28, 2026
Kubernetes

Three Mistakes We See in Every Kubernetes Migration

The patterns that cause 80% of post-migration pain, and how to avoid them before your first pod ships.

5 min read · Jan 6, 2026

Want to see us write about something specific? Drop us a line.