The Developer Experience Problem with Kubernetes at Scale

At Intuit we run thousands of production services across hundreds of Kubernetes clusters. At that scale, developers were spending 30–40% of their time on YAML, deploys, and ops (the outer loop) instead of writing product code (the inner loop).

That’s a platform problem, not a training problem. The same platform serves Intuit’s approximately 100 million customers worldwide, across products like TurboTax and QuickBooks, so we couldn’t just tell people to get better at YAML.

The first 6x: managed namespaces

Before Kubernetes, teams owned the machines: on-prem servers, or AWS accounts and EC2s they patched and scaled themselves. The first Kubernetes system at Intuit moved them onto

multi-tenant clusters the platform team ran. We owned the nodes: scaling, AMIs, security patching. The unexpected killer feature was that last part. Teams no longer had to do AMI rotations themselves; they could rely on the platform to do it, and that turned out to be the change they were happiest about.

The contract sat at the namespace. A team got a vended, managed namespace they could deploy into. They did not get a cluster, and they could not change the platform-managed parts (security policies and the rest), which RBAC and OPA admission webhooks enforced. Compared to owning VMs, that was a 6x jump in developer productivity. It worked and we ran the majority of Intuit’s production compute on it.

Kustomize in every repo

Inside that namespace, the app was still theirs to describe. Each service had a kustomize project: a shared app-base with app-specific patches plus an overlay per environment. The platform team published centrally managed remote bases (HPA, canary rollouts, metrics) which represented Intuit policies and best practices (the paved path), and services pinned them to a version tag. A typical layout looked like this:

├── app-base

├── environments

│   ├── e2e-use2

│   ├── e2e-usw2

│   ├── prd-use2

│   ├── prd-usw2

│   ├── prf-usw2

│   ├── qal-usw2

│   ├── stg-usw2

And an app-base/kustomization.yaml looked something like this: remote bases for the paved path pieces, then whatever else the team needed.

apiVersion: kustomize.config.k8s.io/v1beta1

kind: Kustomization

bases:

  – <remote>/service-hpa-base?ref=v7.0.0

  – <remote>/service-rollout-canary-base?ref=v7.0.0

  – <remote>/metrics-service-base?ref=v7.0.0

resources:

  – ConfigMap-envoy.yaml

  – PrometheusRule.yaml

patchesStrategicMerge:

  – Rollout-patch.yaml

  – Hpa-patch.yaml

  – Service-metrics-patch.yaml

From a developer’s point of view, that was power and ultimate flexibility. Inside their namespace they could patch anything, add anything, ship anything. That was also as far as the abstraction went. The cluster was ours. The application YAML was still theirs. But that left us with a big problem; we could not change Kubernetes out from under a pile of manifests in someone else’s repo.

The outer-loop trap

That leftover YAML living in the dev teams’ repos began accumulating tech debt and drift from the paved path. Kubernetes 1.16 is when leaving the app manifests in someone else’s Git repo stopped being a reasonable trade-off against the platform team’s ability to move the cluster. That was the forcing function to abstract inside the namespace as well.

API deprecations

Kubernetes 1.16 stopped serving the v1beta1 and v1beta2 workload APIs for Deployments, DaemonSets, ReplicaSets, StatefulSets, and others. While the paved path was updated so new services got the new apiVersions, existing services still had to move to apps/v1 before we could upgrade a cluster.

The upgrade to Kubernetes 1.16 was painful for the platform team. We generated PRs against onboarded repos, sent a notice asking every team to merge, test, and roll the change through every environment, then ran an Intuit-wide program to track who had and who hadn’t. Clusters had to stay on the old Kubernetes version until none of the namespaces were using deprecated APIs. Bumping the remote base was the intended path (ship a new release tag, teams adopt it, patches get updated), but without the mandate, that mostly did not happen. That work wasn’t the app teams’ job: they weren’t Kubernetes experts, and they shouldn’t have to be. They wanted to write Java. They wanted the application abstracted the same way the cluster already was.

And 1.16 was not a one-off. Kubernetes ships three new versions a year (four, back then), and the API deprecations come with them. 1.22 was Ingress, 1.25 was policy/v1beta1 PodDisruptionBudget and batch/v1beta1 CronJob, and 1.26 was autoscaling/v2beta2 HorizontalPodAutoscaler. Every one of those lived in YAML the app teams owned. Leave that YAML in their repos and the 1.16 “migration” playbook becomes a full-time job: program after program, just to clear deprecated APIs in time to upgrade the cluster before it fell below EKS’s minimum supported version. The changes we could make behind the scenes, with no app-team involvement, were the easy ones: AMI rotations already worked that way. The YAML was the part we still didn’t own.

47 commits a year

It wasn’t just the API deprecations. We audited kustomize repos and found that the typical project carried 47 infrastructure commits a year. These weren’t product-code changes. They were outer-loop YAML: HPA min/max and scale metrics, CPU and memory requests, Ingress annotations, rollout patches, remote-base version pins, the overlays that follow a service through every environment. That doesn’t sound like much until you multiply it across thousands of services. And each of those commits still had to be reviewed, tested, and rolled through every environment.

A lot of the 47 were knobs teams turned after something broke. For example, many teams left the vended HPA and resource defaults until an incident, then tuned them as part of the incident root-cause analysis. HPA is not simple, and it is easy to get wrong. Incidents where the app did not scale, or did not scale fast enough, got written up as “HPA didn’t work.” Usually HPA itself was fine. The config wasn’t.

And 47 is only what shipped. App teams were already carrying that infrastructure load on top of the product feature development, so a lot of worthwhile platform changes never made the list: newer load-balancer annotations, remote-base version bumps, best-practice improvements we’d already published. The should-have count is likely much higher.

Where the time actually went

Put 47 infrastructure commits a year on every service and you can start to see where some of the outer-loop time went. When we measured the development lifecycle, about 42% of engineering time was “inner loop”: coding, testing, debugging. Another 30–40% was that outer loop: deploying, operating, monitoring, and managing infrastructure. Industry research points the same general direction on how little of a developer’s time reaches feature work: Bain & Company’s Beyond Code Generation, from their 2024 Technology Report, finds developers spend roughly half their time writing and testing code.

Where a developer’s time actually goes: 42% inner loop (code, test, debug) vs. 30–40% outer loop (deploy, operate, monitor, manage infrastructure).

Roughly one in five support requests we fielded, in the period we sampled, were about Kubernetes basics rather than application issues, and configuration complexity was a leading source of paved-path policy drift across our fleet. We put a nominal developer-time cost on those 47 commits, plus a cost on the incidents that came from bad HPA, other Kubernetes misconfig, or shipping without progressive delivery. That number is what sent us chasing the next 6x.

The next 6x: take away the YAML

The first 6x had taken away the cluster. The next 6x had to take away the YAML. We wanted application-level changes we could make centrally and roll out on a schedule the platform team controlled.

That’s why the 2022 writeup was called unlocking the next 6X in development velocity through application abstraction. Keep the managed namespace. Take an application spec instead of raw Kubernetes inside it. That roadmap is now a production runtime, IKS AIR, and this series is the story of what we actually built.

IKS AIR, or Intuit Kubernetes Service, AI Runtime, sits on top of Kubernetes and translates application needs into platform means. A developer states intent; the platform generates the manifests, the traffic wiring, the autoscaling, the safety nets.

That breaks into three pieces, and this series takes them one at a time. This post is about the application spec that replaced the YAML. Part 2 is how the platform sizes a service from its own traffic instead of asking a developer to guess. Part 3 is what it does when you ship a new version or need to debug a running pod, and where operations go next.

From 40 YAMLs to a single intent

A new service used to mean 30 to 40 YAML files, or more. Deployments, services, ingresses, HPAs, a pile of other custom resources. And, while powerful, Kustomize’s layering of base patches and environment overrides can be quite confusing. Often a YAML would be updated correctly in preprod but not in prod.

IKS AIR takes a single application spec instead. We used an OAM-like model of components and traits, on purpose, instead of exposing raw Kubernetes resources. Components are what the app is: a webservice or a cronjob. Traits are how it behaves. A developer who declares a sizing trait does not need to know what it becomes underneath: an HPA, a pod-sizing baseline, and in some environments a VPA. They describe the outcome. The platform picks the Kubernetes-native way to deliver it. That split is what lets us change the implementation (swap an autoscaling strategy or change a default) without ever touching the spec in someone’s repo.

A spec looks something like this:

apiVersion: iks.intuit.com/v1beta1

kind: ExpressApplication

metadata:

  name: my-api

spec:

  components:

    – type: webservice

      name: my-api

      image: <registry>/my-api:v1

      traits:

        – type: sizing

          properties:

            horizontal:

              size: finetune

            vertical:

              size: finetune

  environments:

    preprod:

      – name: qal

    prod:

      – name: prd

The platform turns that into the same manifests we’d been hardening for years in production.

Before: 30–40+ hand-written YAML manifests (Deployment, Service, Ingress, HPA, NetworkPolicy, ConfigMap). After: one ExpressApplication spec with components and traits, and IKS AIR generates the rest.

With this, the 1.16 class of problem mostly goes away. Kubernetes version upgrades, observability agent swaps, other infra refreshes roll out centrally. Application teams don’t reconfigure or redeploy because the platform underneath them changed. That used to be weeks of migration work every upgrade cycle. Now the namespace is on the same footing as the managed cluster and AMI rotations: the platform moves, the app doesn’t.

Each service that moves to IKS AIR saves an estimated 25 engineering days per year in ongoing outer-loop work (modeled from the commit audit, avoided incidents, and other data). That pays back the time to adopt AIR, and then compounds for the life of the service. In practice that’s 47 infrastructure commits a year down to roughly a dozen, better than a 70% reduction. That 25 days is the other side of the 47-commit audit: the YAML they no longer have to touch, and the incidents that YAML used to cause.

Migration cost vs. annual payback per service: ~4 engineering days to migrate, one time, against ~25 engineering days saved every year afterward.

The IKS AIR application spec is what made that possible. Part 2 of this series will discuss IKS AIR Managed Autoscaling and how size: finetune on the sizing trait drives the platform to set replica counts and pod sizing from historical utilization. Developers don’t set pod resources or replica counts by hand.

If you’re curious what else we’re building, check out the Intuit engineering blog or the software engineering careers page.