GCP deployment guidelines

View as Markdown

As a general guideline, we recommend:

  • ARM-based CPU.
  • A 1:8 ratio of vCPU to GiB memory.
  • At least a 2:1 ratio of GiB local instance storage to GiB memory when using swap.

When operating on GCP in production, we recommend the Arm-based C4A high-memory series. Both C4A and C4 offer local SSDs only on their -lssd machine variants, which bundle a fixed number of Titanium SSD disks.

Series Examples
C4A high-memory series (recommended) c4a-highmem-16-lssd or c4a-highmem-32-lssd
C4 high-memory series c4-highmem-16-lssd or c4-highmem-32-lssd

C4A is not available in every region. Where it is unavailable, use the x86-based C4 high-memory series instead.

To maintain the recommended disk-to-RAM ratio for your machine type, see Number of local SSDs to determine the number of local SSDs to use.

See also Locally attached NVMe storage.

Number of local SSDs

Each local SSD in GCP provides 375GB of storage. Use the appropriate number of local SSDs to ensure your total disk space is at least twice the amount of RAM in your machine type for optimal Materialize performance.

C4A and C4 bundle a fixed number of Titanium SSD disks in each -lssd machine variant. The count is not configurable, but every high-memory -lssd variant satisfies the 2:1 disk-to-RAM ratio:

Machine Type RAM Bundled Local SSDs Total SSD Storage
c4a-highmem-8-lssd 64GB 2 750GB
c4a-highmem-16-lssd 128GB 4 1500GB
c4a-highmem-32-lssd 256GB 6 2250GB
c4a-highmem-64-lssd 512GB 14 5250GB
c4-highmem-8-lssd 62GB 1 375GB
c4-highmem-16-lssd 124GB 2 750GB
c4-highmem-32-lssd 248GB 5 1875GB
c4-highmem-48-lssd 372GB 8 3000GB

For other machine series, the local SSD count is configurable but may only support predefined values. To determine the valid number of local SSDs to attach for your machine type, see the GCP documentation.

Locally-attached NVMe storage

Configuring swap on nodes to use locally-attached NVMe storage allows Materialize to spill to disk when operating on datasets larger than main memory. This setup can provide significant cost savings and provides a more graceful degradation rather than OOMing. Network-attached storage (like EBS volumes) can significantly degrade performance and is not supported.

Swap support

The Materialize Terraform module supports configuring swap out of the box.

The Legacy Terraform provider, adds preliminary swap support in v0.6.1, via the swap_enabled variable. With this change, the Terraform:

  • Creates a node group for Materialize.
  • Configures NVMe instance store volumes as swap using a daemonset.
  • Enables swap at the Kubelet.

See Upgrade Notes.

NOTE: If deploying v25.2, Materialize clusters will not automatically use swap unless they are configured with a memory_request less than their memory_limit. In v26, this will be handled automatically.

CPU affinity

It is strongly recommended to enable the Kubernetes static CPU management policy. This ensures that each worker thread of Materialize is given exclusively access to a vCPU. Our benchmarks have shown this to substantially improve the performance of compute-bound workloads.

TLS

When running with TLS in production, run with certificates from an official Certificate Authority (CA) rather than self-signed certificates.

Upgrading guideline

Whe upgrading:

  • Always check the version-specific upgrade notes.

  • Always upgrade the operator first and ensure version compatibility between the operator and the Materialize instance you are upgrading to.

  • Always upgrade your Materialize instances after upgrading the operator to ensure compatibility.

Node pool resizing

The VM type of a Kubernetes node pool is immutable on EKS, AKS, and GKE, so changing it triggers a destroy + create that fails while Materialize pods are still running on the pool. The supported pattern is to add a second pool with the new VM type, roll out the Materialize instance so new pods land on it, and then drop the old pool.

For the full procedure, see Resize node pools.

Back to top ↑