Co-authored by Israel Gofman and Yaniv Weinberg
How to run N2 and N4 nodes in the same GKE cluster without your Pods quietly hanging forever
GKE’s automated disk type selection and custom ComputeClasses let a single manifest survive a mixed Persistent Disk / Hyperdisk fleet. Here’s how to wire it up — and the provisioning-time gotcha that almost nobody mentions.
Why anyone runs a mixed-generation fleet
Running Google Kubernetes Engine (GKE) at scale almost always means running more than one VM generation at a time. A typical pairing is Google Cloud’s 4th-generation N4 machine series alongside the 2nd-generation N2 or N2D series.
There are three good reasons to do this, and it’s worth being precise about which one applies to you, because they pull in different directions:
Price-performance. N4 — built on 5th-generation Intel Xeon (Emerald Rapids) with Titanium offload — generally offers better price-performance per vCPU than N2, along with custom machine types that let you size instances tightly instead of rounding up. On a pure list-price-per-unit-of-work basis, N4 is usually the cheaper place to run new work.
Existing commitments. That calculation inverts the moment you already own N2 committed use discounts. A CUD you’ve already paid for makes N2 capacity effectively cheaper at the margin than anything else in the project, and keeping those node pools full is real money. This is why most fleets are mixed rather than migrated: the new generation wins on rate, the old generation wins on sunk commitment.
Obtainability. Zonal capacity for any single machine series can run short. A fleet that can fall back to a second machine family — or to Spot capacity on either family — keeps scheduling when a single-family fleet stalls.
So: N4 as the primary target, N2 as a deliberate fallback that also soaks up your commitments. It’s a sensible architecture. And it breaks in an unobvious way the moment those workloads need a disk.
The constraint: the support matrix is asymmetric
Here is the actual compatibility picture, from the Compute Engine machine-series support tables:: N2 supports every Persistent Disk type and Local SSD, but Hyperdisk Balanced only via allowlist.
The failure: a quiet hang, not a crash
GKE’s default StorageClass, standard-rwo, provisions pd-balanced through the Compute Engine persistent disk CSI driver (pd.csi.storage.gke.io). A developer writes an ordinary PVC, doesn’t specify a storageClassName, and gets a Persistent Disk. On an N2 node this works. On an N4 node it cannot.
Diagram 1: The Compute and Storage Mismatch (Failure Mode)

Here’s the part most write-ups get wrong. It is tempting to describe this as the Pod “crashing.” It doesn’t crash. The volume never attaches, so the kubelet never starts the container, so there is nothing to crash. kubectl describe pod shows something like:
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Scheduled 2m default-scheduler Successfully assigned
default/app-0 to gke-...-n4-pool-...
Warning FailedAttachVolume 2m attachdetach-controller AttachVolume.Attach failed for
volume "pvc-8f2c..." : rpc error:
code = Internal desc = unknown
Attach error: ...
Warning FailedMount 35s kubelet Unable to attach or mount volumes:
unmounted volumes=[app-data],
unattached volumes=[app-data]:
timed out waiting for the condition
(Message text varies by driver version; the FailedAttachVolume / FailedMount pair is the signature.)
The Pod sits in ContainerCreating indefinitely. It never enters CrashLoopBackOff, never restarts, and never fires a restart-count alert. Your readiness probes have nothing to probe. This is worse than a crash, not better: a crash is loud, and this is silent.
Note also when this happens. standard-rwo already uses volumeBindingMode: WaitForFirstConsumer — GKE’s own docs spell this out — so the volume isn’t created too early and then stranded. It’s created after the scheduler picks the node, in the right zone, and it’s still the wrong type. Delayed binding was never the missing piece.
Where this bites in production
The scenario that generates support tickets is a scale-up under capacity pressure. The cluster autoscaler can’t get N2 capacity in the zone, falls back to N4, brings up healthy nodes — and every storage-attached Pod that lands on them hangs. Your fallback worked perfectly at the compute layer and failed completely at the storage layer. Nodes are Ready, Pods are Pending, and the dashboards look fine.
The fix, part 1: automated disk type selection
In GKE 1.35.3-gke.1290000 and later, a StorageClass can defer the disk type decision to GKE. You set parameters.type: dynamic, declare your preferred type on each side of the divide, and GKE picks based on the machine type of the node the scheduler chose.
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: dynamic-mixed-storage
provisioner: pd.csi.storage.gke.io
# Required: GKE can only choose a disk type once a node has been selected.
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
parameters:
# Turn on automated disk type selection.
type: dynamic
# What to use on Hyperdisk-capable nodes (N4). Default: hyperdisk-balanced.
hyperdisk-type: hyperdisk-balanced
# What to use on Persistent Disk-capable nodes (N2). Default: pd-balanced.
pd-type: pd-balanced
# On nodes that support BOTH, prefer Persistent Disk. Without this line the
# default is 'hyperdisk-type' — see the note below, this matters for N2.
disk-type-preference: pd-type
# Restrict scheduling to nodes that can actually attach the chosen type.
# Requires cluster AND node pools at 1.34.1-gke.2541000+.
use-allowed-disk-topology: "true"
# Hyperdisk-only parameters. Ignored when the PD branch is selected.
# 3,000 IOPS + 140 MiB/s is the free baseline for Hyperdisk Balanced.
provisioned-iops-on-create: "3000"
provisioned-throughput-on-create: "140Mi"
Four of those lines deserve more than a comment.
VolumeBindingMode: WaitForFirstConsumer is a prerequisite, not the mechanism
GKE cannot choose a disk type until it knows which node the Pod is going to. So delayed binding is mandatory here. But as noted above, standard-rwo already sets it — delayed binding alone changes nothing. The behavioral change comes entirely from type: dynamic.
Disk-type-preference is the line that makes the behavior deterministic
This is the parameter most guides omit, and the one your fleet’s correctness actually hinges on. It governs what GKE does on nodes that support both families. Per the docs, the values are hyperdisk-type (the default when the parameter is absent) and pd-type.
Remember that N2 sits in the Hyperdisk Balanced support matrix. If GKE treats your N2 nodes as “supports both,” then omitting this parameter means GKE prefers Hyperdisk on them — which, if you aren’t on the Hyperdisk allowlist for N2, is not what you want.
Set it explicitly. disk-type-preference: pd-type is what actually encodes the rule most people think they’re getting: Hyperdisk on N4, Persistent Disk on N2. Don’t rely on the default to mean what you assume.
Use-allowed-disk-topology is load-bearing, not decoration
Setting this to “true” makes GKE schedule Pods only onto nodes that support the volume’s disk type — it adds the topology constraint that keeps a pd-balanced volume away from N4 nodes. This is what protects you on reschedules, and I’ll come back to it in the gotchas section because it has a significant side effect.
Two operational constraints: the cluster and its node pools must be at 1.34.1-gke.2541000 or later, and volumes provisioned by this StorageClass will not schedule onto node pools running older versions. In a fleet you upgrade gradually, that’s a scheduling constraint you need to know about before you roll this out, not after.
Performance parameters apply to one branch only
When GKE creates a volume of a given type, it applies only the parameters relevant to that type — so the Hyperdisk performance settings are silently ignored on the Persistent Disk path. One StorageClass, two parameter sets, no conflict. That’s the neat part of the design.
But choose those numbers deliberately, because this is where a cost-optimization article can accidentally cost you money. Hyperdisk Balanced includes a free baseline of 3,000 IOPS and 140 MiB/s per volume; anything you provision above that is billable. The example above sits exactly on the baseline. If you’d written provisioned-throughput-on-create: “250Mi” — a very natural-looking number — every Hyperdisk volume this StorageClass creates would bill 110 MiB/s of extra throughput, per volume, for the life of the volume. On an ephemeral scratch disk, across a few hundred Pods, that adds up for no benefit.
If you do need more, the allowed ranges for Hyperdisk Balanced are:

So a 50 GiB volume at 3,000 IOPS can be provisioned anywhere from 140 to 750 MiB/s. One more limit that catches people: a volume can’t exceed the per-instance limits of the VM it’s attached to. A small n4-standard-2 will not deliver 750 MiB/s regardless of what the disk is provisioned for. Check “Performance limits when attached to an instance” before you pay for throughput you can’t reach.
The fix, part 2: a ComputeClass for the compute side
Automated disk type selection solves storage compatibility. It does nothing about getting capacity. For that, custom ComputeClasses let a platform team declare a prioritized fallback hierarchy once, as a cluster-scoped resource, instead of scattering machine-family knowledge across application repos.
Custom ComputeClasses are a ComputeClass custom resource in the cloud.google.com/v1 API group, and they’ve been available in both Autopilot and Standard since GKE 1.30.3-gke.1451000 — considerably earlier than the dynamic storage feature. If you’re not yet on 1.35, you can adopt this half today.
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
name: mixed-n4-n2-class
spec:
# Let GKE create new node pools, not just scale existing ones.
nodePoolAutoCreation:
enabled: true
priorities:
# Priority 1: N4 (4th-gen). Boot disk must be Hyperdisk — N4 can't boot from PD.
- machineFamily: n4
minCores: 4
storage:
bootDiskType: hyperdisk-balanced
bootDiskSize: 100
# Priority 2: N2 fallback, soaking up existing committed use discounts.
- machineFamily: n2
minCores: 4
storage:
bootDiskType: pd-balanced
bootDiskSize: 100
# If NEITHER rule can be satisfied, leave Pods pending rather than
# provisioning nodes from some other family. See the discussion below.
whenUnsatisfiable: DoNotScaleUp
Note that storage is set per priority rule, not once for the whole class. That’s not a stylistic choice: an N4 node pool needs a hyperdisk-balanced boot disk, and an N2 node pool wants pd-balanced. A single class-wide boot disk type would break one of the two rules. The same generational divide that motivated this entire article applies to the node’s own boot disk.
WhenUnsatisfiable deserves a real explanation
This field is usually pasted in without comment, and it’s the one place where the obvious choice and the resilient-sounding choice diverge.
whenUnsatisfiable governs only what happens when none of your priority rules can be satisfied. Fallback between your N4 and N2 rules happens regardless. The two values:
- DoNotScaleUp — Pods stay Pending. This is the default from GKE 1.33 onward.
- ScaleUpAnyway — GKE provisions nodes using cluster defaults, outside your priority list.
ScaleUpAnyway sounds like the resilient option, and in an article about surviving stockouts you’d expect me to recommend it. Don’t — not here. It permits GKE to create a node from a machine family you never listed, which is exactly how a disk-incompatible node re-enters a fleet you just spent all this effort making compatible. DoNotScaleUp keeps the guarantee: every node that runs these Pods comes from a family you vetted. A Pending Pod with a clear reason is a better failure than a running node with an unattachable disk.
Prerequisites

That second row is the quiet trip-hazard: on a pre-1.33.3 Standard cluster without node auto-provisioning turned on, nodePoolAutoCreation: enabled: true will not create node pools for you and there’s no loud error explaining why.
Also confirm the CSI driver is present — it’s enabled by default on Autopilot and on current Standard clusters, but verify rather than assume:
kubectl get csidriver pd.csi.storage.gke.io
Wiring it up: one disk per Pod, not one shared PVC
For stateless workloads with instance-lifetime scratch storage — the entire subject of this article — the correct primitive is a generic ephemeral volume. Each Pod gets its own disk, provisioned when the Pod is created and deleted when the Pod is deleted:
apiVersion: apps/v1
kind: Deployment
metadata:
name: resilient-stateless-app
spec:
replicas: 3
selector:
matchLabels:
app: resilient-stateless-app
template:
metadata:
labels:
app: resilient-stateless-app
spec:
# Route scheduling through the prioritized ComputeClass.
nodeSelector:
cloud.google.com/compute-class: mixed-n4-n2-class
containers:
- name: application-container
image: debian:12
command: ["sleep", "infinity"]
# Requests matter here: they're what GKE sizes auto-created
# node pools against.
resources:
requests:
cpu: "500m"
memory: "1Gi"
limits:
memory: "1Gi"
volumeMounts:
- name: app-data-volume
mountPath: /var/lib/data
volumes:
- name: app-data-volume
ephemeral:
volumeClaimTemplate:
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: dynamic-mixed-storage
resources:
requests:
storage: 50Gi
Diagram 2: Dynamic Storage and Compute Class Workflow (The Solution)

Three things to notice.
The disk lifecycle now matches the Pod lifecycle. Every reschedule re-runs disk type selection against the new node. This is what makes the whole design hold together, and it’s why “stateless” in the title is honest rather than aspirational.
Resource requests are not filler. With nodePoolAutoCreation enabled, Pod requests are the primary signal GKE uses to size the node pools it creates for you. A container with no requests gives the autoscaler almost nothing to work with. In a ComputeClass article, omitting them would be a meaningful bug in the example.
Yes, there’s still a nodeSelector. I want to be straight about this, because the usual claim — that this approach achieves “zero developer overhead” — isn’t quite true. Application teams still reference a compute class and a storage class. What’s changed is the kind of knowledge in the manifest: one intent-level selector owned by the platform team, instead of machine families, disk SKUs, and affinity blocks maintained by every application team. That’s a real win; it just isn’t zero.
If you do want zero, configure a cluster-level default ComputeClass — then workloads need no selector at all, and the platform team’s fallback hierarchy applies to everything by default.
Verify it actually works
Deploy the three manifests, then prove the behavior rather than trusting it. Map each Pod to its node generation and its backing disk type:
# Which generation did each Pod land on?
kubectl get pods -o wide \
-l app=resilient-stateless-app \
-o custom-columns=POD:.metadata.name,NODE:.spec.nodeName
kubectl get nodes -L node.kubernetes.io/instance-type
# Which PV backs each Pod, and what's the underlying disk?
kubectl get pv -o custom-columns=\
NAME:.metadata.name,\
CLAIM:.spec.claimRef.name,\
SC:.spec.storageClassName,\
HANDLE:.spec.csi.volumeHandle
Take a volumeHandle (it ends in the disk name) and ask Compute Engine directly:
gcloud compute disks describe DISK_NAME \
--zone ZONE \
--format='value(type)'
You’re looking for one Pod on an n4-* node backed by hyperdisk-balanced, and one Pod on an n2-* node backed by pd-balanced, from the same StorageClass. That’s the whole article in two lines of output.
Forcing the fallback so you can actually see it
You can’t order a zonal stockout for demo purposes. To exercise the N2 fallback path deliberately, either cordon the N4 nodes:
kubectl cordon -l cloud.google.com/machine-family=n4
kubectl rollout restart deployment/resilient-stateless-app
…or temporarily add a first priority rule that requests a machine type with no available capacity — an oversized shape, or a Spot rule in a constrained zone — and watch GKE fall through to the next rule. Either way you get a repeatable demonstration, which is what turns a blog post into something a team can validate in a sandbox before adopting.
The gotcha nobody mentions: selection happens once
This is the most important section in the article, and it’s the one missing from every other write-up I’ve read on this feature.
Automated disk type selection runs at provisioning time. It does not re-run.
A volume provisioned as pd-balanced on an N2 node is pd-balanced for its entire life. The PV is bound, the disk exists, and its type is fixed. If the Pod using it is later rescheduled onto an N4 node, the existing Persistent Disk cannot attach — and you are back to the exact FailedAttachVolume hang this article opened with.
That’s not a hypothetical. Reschedules happen constantly: node auto-upgrades, Spot preemptions, node repair, scale-down consolidation, and — pointedly — ComputeClass activeMigration.optimizeRulePriority, whose entire job is to move Pods back to a higher-priority rule when that capacity returns. Enable active migration on the class above and you have built a machine whose explicit purpose is to move your N2 Pods onto N4 nodes. With a bound pd-balanced volume, that’s a scheduled outage.
Two mitigations, and you want both:
1. use-allowed-disk-topology: “true”. This is why that parameter is in the StorageClass. It constrains scheduling to nodes that support the volume’s disk type, so a bound pd-balanced volume keeps its Pod on Persistent-Disk-capable nodes. Without it, nothing stops the scheduler from making the incompatible placement.
2. Generic ephemeral volumes. Because the disk dies with the Pod, there’s no stale volume to constrain the next placement, and disk type selection re-runs against whatever node the Pod lands on. For genuinely stateless workloads, this is the complete fix — and it’s why I structured the Deployment that way rather than around a standalone PVC.
The tension you should understand before rolling this out
These two mechanisms don’t fully compose, and it’s worth being clear-eyed about it.
For a Pod holding a bound, long-lived pd-balanced volume, use-allowed-disk-topology pins it to N2-capable nodes. Which means during an N2 stockout, that Pod stays Pending — and the ComputeClass’s N4 fallback, the thing you built this for, cannot rescue it. Storage compatibility and compute obtainability are in direct conflict for that Pod, and correctness wins.
The way out isn’t a cleverer parameter. It’s recognizing that only workloads whose storage can be recreated get to be generation-agnostic. If the disk must survive the Pod, it inherits the constraints of the generation that created it. That’s the real architectural boundary, and it’s usually the sentence teams are missing when they say “we’ll just make everything portable.”
What this doesn’t solve
Being explicit about the edges:
- StatefulSets and long-lived data. volumeClaimTemplates create durable per-replica volumes. Those volumes are pinned to their original generation for life, exactly as described above. Dynamic selection helps at first provision and not after.
- Existing bound PVCs. Everything here is greenfield. A fleet with thousands of bound pd-balanced volumes has a migration problem, not a configuration problem: snapshot and restore onto the new type, or leave them pinned to N2 via topology constraints and let attrition handle it. Plan that separately.
- Regional / HA disks. Hyperdisk Balanced High Availability is supported on both families here, but replicated-disk failover has its own constraints and is out of scope for a stateless-workload pattern.
- Mixed-version fleets. The use-allowed-disk-topology version floor (1.34.1-gke.2541000 on cluster and node pools) will bite before anything else does. Check it first.
- Windows node pools, which have their own storage behavior and aren’t covered by anything above.
It generalizes past GEN 2 and GEN 4 Machine Families
Nothing here is specific to these two families. The same divide — older series on Persistent Disk, newer series on Hyperdisk only — applies across the fleet: N2D → N4D, C2 → C3/C4, M1 → M4, and every future generation that ships Hyperdisk-only. Google’s own documentation example for this feature uses C2 and C3/C4 rather than N2/N4.
Write the StorageClass once, add priority rules as new generations arrive, and the application manifests never change. That’s the actual payoff — not that GEN 4 works today, but that future GEN 5 won’t require a fleet-wide PR.
Conclusion
The compatibility problem is real and its failure mode is nastier than it first appears: not a crash loop, but a silent hang on a healthy-looking node, arriving precisely when the autoscaler was supposed to be saving you.
The fix is genuinely good. One type: dynamic StorageClass, one ComputeClass with a fallback hierarchy, and ephemeral volumes that let each Pod’s disk match each Pod’s node. Application teams get to stop knowing what a machine family is.
But adopt it with the boundary in mind. Dynamic selection decides once, at provisioning. The pattern gives you a fleet where stateless workloads are genuinely generation-agnostic — and it gives you a much clearer view of which of your workloads were never really stateless at all.
Sources
All behavior and version numbers above verified against Google Cloud documentation as of September 2026:
- About Hyperdisk for GKE — automated disk type selection, disk-type-preference, use-allowed-disk-topology, supported node scheduling
- About custom ComputeClasses — CRD fields, whenUnsatisfiable, node auto-provisioning prerequisites
- Persistent volumes and dynamic provisioning in GKE — standard-rwo volume binding mode
- About Persistent Disk — machine series support
- About Hyperdisk — machine series support and restrictions
- Hyperdisk Balanced — size, IOPS and throughput limits; baseline performance
- About Local SSD — machine series support
- GKE release notes
Versions and machine-series support change. Check the docs against your cluster version before adopting.
One StorageClass, Two VM Generations was originally published in Google Cloud – Community on Medium, where people are continuing the conversation by highlighting and responding to this story.
Source Credit: https://medium.com/google-cloud/one-storageclass-two-vm-generations-0e08f9bbd6cf?source=rss—-e52cf94d98af—4
