Materialize, Out-of-Core: Replacing Swap with a Buffer Pool
September 29, 2026

Materialize computes and then incrementally maintains the results of SQL queries. The data structures behind this need both high-throughput scans and low-latency key lookups, making them a natural fit for main memory. For several years, we’ve designed around this assumption.
Unfortunately, physical memory is expensive. Listening to our customers, we know that many would rather spend a fixed budget on a much larger working set than on the last few microseconds of tail latency. We have chipped away at this problem for years while preserving the facade of uniform memory access.
That facade hides a profound asymmetry. A DRAM access is on the order of 100 ns; a read from local NVMe is tens of microseconds, a few hundred times slower. But bandwidth tells a different story: a local NVMe device delivers single-digit GB/s against DRAM’s tens of GB/s, so the throughput gap is closer to 10× than 1000×. Disk latency is catastrophic for a random pointer chase and merely inconvenient for a large sequential scan. That asymmetry defines our design space. An index accessed through large sequential scans can live on disk, but we need to amortize the latency of individual lookups.
Exploiting this asymmetry means controlling which bytes go to disk and the representation in which we read them back, and today we control neither. Materialize currently relies on Linux’s paging to move data between memory and disk, which gives us the illusion of a contiguous virtual address space and hands eviction policy to the kernel. The kernel evicts in 4 KiB pages, has no idea which pages belong to a cold batch and which are hot, and turns a miss into a synchronous stall on a dataflow worker. So we’re building the alternative: an out-of-core Materialize, where data lives on disk when needed and memory is an explicit cache we manage ourselves. Instead of optimizing for CPU, we’ll optimize for data transfer. Borrowing from years of buffer-pool research, we’re replacing swap with a new buffer pool. In the rest of the post, I want to outline which parts of the system this touches and the sequence of work required to get there.
From kernel paging to an application-managed buffer pool
This is our third attempt at spilling to disk. Materialize started as an in-memory database, optimized for CPU cycles and the low-latency, high-throughput access that memory affords. When memory ran out, we spilled through lgalloc, which backs allocations with memory-mapped files. That gave us real control over what went to disk, but file-backed memory comes with the kernel’s durability model attached: dirty pages are written back on the flusher’s schedule, so we paid for I/O long before any memory pressure made it necessary. Sparse files compound the problem, since disk blocks are allocated on first touch, and the space we consumed tracked everything we had ever written rather than what actually needed to live on disk. We then moved to swap, which bought a flat memory model and the ability to spill anything, but handed every policy decision back to the kernel, undoing the control lgalloc had introduced.
Each previous approach failed because it inherited a different kernel policy. lgalloc inherited a durability policy we did not want; swap inherited an eviction policy that does not know which data matter. Neither is inherently a bug; the kernel is serving a general contract while we need a specific one.
What replaces swap is a buffer pool. In this design, dataflow operators—the building blocks of Materialize—will store data in a column-first, serialized form in a managed buffer pool. The pool will decide what stays resident. Layout work and the pool work go hand in hand: a pool can only evict and re-fault a region without fixing up pointers if the region is relocatable and pointer-free to begin with. Columnar data representation enables efficient serialization and zero-copy data access.
The buffer pool will govern several important structures, most notably our indexes, which we call arrangements. Materialize uses arrangements both for information required to compute a query result and for the result itself. Their lifecycle has several stages, each with different memory-access patterns and a different tolerance for living on disk.
Memory hotspots
Let’s follow an arrangement through that lifecycle and see where memory goes.
First, we transform data into the shape the arrangement requires. We apply user functions, extract keys, and repartition data so that matching keys land on the same worker. All data needs to pass through here before we can index it, and we want to minimize the overhead of many small updates. Access is sequential and the footprint is transient, so this is primarily a throughput problem, not a major source of steady-state memory use.
Then, we hydrate an arrangement. We sort and consolidate updates into the canonical chain-chunk representation. We use log-structured chains to amortize work and keep the memory footprint bounded. Hydration ends by merging those chains into one fully consolidated chain of chunks. Merge passes over chunks are sequential by construction, making hydration the cleanest opportunity to move data out of memory: for this access pattern, disk is merely inconvenient rather than catastrophic.
info
In Materialize, we use a chain-chunk structure to amortize data insertion costs. It consists of geometrically sized chains, each segmented into roughly equal-sized chunks of consolidated data. Each chain is approximately twice the size of the preceding one. This structure allows us to incrementally build a sorted representation in O(n ⋅ log n) total merge work. We use the chain-chunk form in arrangement formation and materialized view sinks.
Once we have this chain-chunk form, we form the arrangement. We encode the chain-chunk representation into arrangement batches. The resulting arrangement is roughly a trie of key-to-value mappings and value-to-time-and-diff mappings. This step is the moment of peak memory: before the trie exists, there is no deduplication, so every (time, diff) pair carries its own copy of the key and the value. The input to formation is potentially larger than the arrangement it produces.
The data responsible for that peak is mergeable and sequentially written, making it a good candidate for the pool. Serving the finished arrangement is harder, because it must provide both high scan throughput and low lookup latency. A scan touches everything and therefore needs throughput; a lookup chases the trie from key to value range to (time, diff) and therefore needs low latency.
Turning latency into throughput
The latency of a single read is irreducible. We cannot make the device answer faster than it does. What we can do is stop paying that latency one read at a time. If an operator names many keys up front, we issue their reads together and amortize the latency across the whole set, which converts a latency problem into a throughput problem—and throughput is where disk performs well.
So operators will have to ask for data differently. Instead of performing a direct lookup that may stall, an operator will submit the keys as a batch and receive their results through a completion-based interface. This need not use Rust async; the important property is that workers no longer block on individual reads. Not everything can be batched this way: the system must first traverse the trie spine to determine which pages to request. The spine therefore stays resident, while the bulk columns can be evicted and fetched on demand.
This divides the system by access pattern fairly cleanly. Dataflow operators exchange pooled columnar data sequentially; merge batchers and materialized view sinks process chunks sequentially; arrangements require random access and therefore need the batching machinery above. Ad hoc queries and sources have hotspots of their own, which we’ll come back to in future posts.
What’s next
The goal is not to make disk behave like memory. It is to reshape Materialize’s data access patterns so that disk serves sequential and batched work, while memory remains responsible for latency-sensitive and hot data.
The work follows from that principle. First, we’ll move sequentially accessed data into a columnar representation, enabling cheap serialization and transparent compression. Then we’ll manage residency explicitly: the buffer pool will decide what to keep resident and what to evict based on application hints and system signals such as memory and I/O pressure. It will also move I/O off latency-critical threads. Finally, we’ll introduce a batched, asynchronous interface for arrangements, keeping their metadata resident while paging bulk data on demand.
Future posts will cover the pool, columnar data layouts, and API changes. They will also connect this work to the outcomes users care about: hydration time, query performance, resource utilization, and I/O efficiency.


