Customer Story

Inside Zepz's Live Context Layer: From Monolith to Microservices

"We had been debating one new fraud signal for three months and getting nowhere. Now I go to the data directly, build the feature, and test it the same day."

Anson Yu
Director of Engineering, zepz
Financial ServicesFintech

Executive summary

Zepz moves money from developed countries into developing ones through two brands, WorldRemit and Sendwave. Every transaction has to clear a fraud and trust decision in the moment it is created, and that decision is made by machine learning models that need to know features, or context about Zepz and its customers: how much a sender has moved this week, whether this is a new user, whether the device matches the last one, what happened on the payout side of the previous transfer. Fraud patterns surface and shift while a session is still open, so a signal that arrives in an offline job tens of minutes later arrives after fraud has already occurred.

Hundreds of features are used to make a decision, but unfortunately none of this context lived in one place. Over more than ten years, Zepz broke its highly normalized transactional database into dozens of microservices, each with its own database. A single transaction now fans out across tens of tables across multiple databases at write time. To make a fraud decision, the platform has to put that transaction back together as quickly as possible, joining it against user, pay-in, and payout records that live in other services entirely. The list of features is not fixed either. Fraud tactics keep changing, so the team has to be able to reach for new signals from across the business and find out quickly whether they help.

The team had been computing those features by querying the production databases directly. At some point load from feature computation landed on databases that were already at their limit, required constant index tuning, and produced serving latency that spiked to 10 to 20 seconds. More worrisome, this ML workload was destabilizing their core OLTP systems.

The engineering team had to cap the amount of features that could be included in their inference pipeline to keep things online.

Streaming infrastructure did not solve it. Streaming join engines hold up across three or four streams and do not hold up across twenty tables, and their recovery behavior was difficult to reason about.

Zepz chose Materialize to build their online feature store. Change data from eight production databases across two AWS regions flows in continuously. SQL views define the canonical transaction, user, pay-in, and payout entities, and incremental view maintenance keeps them up to date as the underlying data changes. The fraud feature store reads from that layer directly rather than from twenty tables across a dozen services.

The result is a live data platform that gives the Zepz team unprecedented flexibility when working with live data. Adding a new signal to the model used to mean asking another team to expose it, and hoping the ask made it from the backlog to the roadmap. Now the data is already in one place and a new feature can be built using standard SQL and tested in a day. That changed the cadence of model improvement from quarterly negotiation to daily experimentation.

Overview

Zepz operates two remittance brands. WorldRemit and Sendwave were separate companies that merged roughly five years ago, and for the most part they still run on separate infrastructure, written in different languages, with their own transaction systems.

The business is specifically about the harder end of the remittance market. Sending money from the US to the UK is a solved problem, handled well by large banks. Sending money from the US or the UK or Europe into a developing country is not, because the receiving banking infrastructure is thinner and every corridor requires a local partner and local settlement.

That shapes the engineering problem in two ways. Fraud and anti-money-laundering risk concentrates in exactly the corridors where identity signals are weakest. And the company has to hold working capital in each destination country to settle with local processors, which means knowing in real time how much capital is deployed where.

Both of those are data problems before they are anything else, and both start with the same question: what is true about this customer and this transaction right now?

For a company on a single database, that question can be answered with a few SQL queries, albeit expensive ones. For a company that leverages microservices, it is the central architectural difficulty.

Challenge

The monolith worked until it didn't

Zepz used to run on one large transactional database. Transactions, user records, pay-in data, and payout data all hit the same instance. There was no data consolidation problem because there was nothing to consolidate.

The problem was that the instance could not grow any further.

“We were already paying for one of the largest instances available, and it still could not keep up. Every time we broke the record for our best day, that was when the strain showed up.”

Instability that tracks record traffic days is the worst kind for a payments business. The platform came under the most pressure exactly when it was earning the most.

To continue to scale, Zepz decomposed its monolith into microservices, 20 to 30 of them, each with its own database. That removed the single point of failure and the single scaling ceiling.

Decomposition solved the write path but complicated the read path

In Zepz’s microservice architecture a single customer transaction is stored as more than twenty rows across more than twenty tables in several different databases. The fraud system needs the whole object, joined against user and pay-in and payout context, in the milliseconds before the transfer is authorized.

"At one moment a transaction record gets written into the database and breaks into ten, twenty tables. And then about ten milliseconds later we join them all back into the same original form."

Most of the work in the feature pipeline was never the feature computation. It was the denormalization, then the join across data sources, and only then the actual arithmetic.

The real workload turned out to be a 19-way join spanning eight microservice databases, one of which is still a monolith Zepz has not decomposed yet: the customer API on MySQL, two versions of the transaction data store, the trust platform, pay-in, payout, and RemitServ, which carries the majority of Sendwave’s traffic. Each of those databases contributes 20 to 30 tables, and some sit in different AWS regions from the others.

The feature store became the bottleneck

Zepz's existing approach ran feature computation as queries against the production databases themselves. The consequences compounded.

Load from analytical work landed on databases that were already maxed out. Keeping queries acceptable meant continuous index tuning. Worst of all, adding a feature made everything slower: real-time serving latency climbed to 10 to 20 seconds when new features went in, which far exceeded their latency budget for fraud decisions.

“The fraud system needs that context in real time. If we see the same pattern showing up, we need to discover it within seconds, not wait for offline processing thirty or sixty minutes later.”

The model was capped

The same database pressure set a hard ceiling on the model itself. Zepz runs tree-based models for fraud and AML detection, querying roughly 200 features per decision. Model performance was bounded by database capacity.

The organizational version of the same problem

Microservices come with a governance pattern that makes getting an up-to-date view of a business difficult: databases are siloed by design. The consequence of this is that every piece of context the fraud team needed had to be requested from the team that owned the underlying data.

"I need a new field to indicate whether this is a new user. Go talk to the user team, and that team will tell you they can give you an API endpoint in two months if you're lucky. And then the conversation just stops there."

This created a vicious cycle: the fraud team could not evaluate whether a signal was worth having without the signal. The owning team could not justify the work without evidence the signal was worth having. Iteration was measured in weeks and months, and most candidate features were never tested at all.

Why Materialize

Zepz evaluated the small number of systems that can maintain a wide join over live data, and considered building the denormalization layer itself before rejecting it.

Why not conventional stream processing

Conventional stream processing was ruled out on two specific grounds. The first is join width. Streaming join engines are genuinely good across a small number of streams. Three or four data sources joined in flight is well-trodden ground. Nineteen tables is a different regime, and the standard tools do not operate there.

"The only reason streaming joins didn't work is that they don't work when you have twenty. It just doesn't work that way."

The second is recovery, which is the failure mode nobody demos.

"One of the biggest weaknesses is outage recovery. When one of the streams is out for two hours, you have to go back and figure out what everything looked like. It's easy when you do it offline. When you have to do it online, you essentially hold the world and wait for that stream to come back. And if the outage went for twenty-four hours and it went past your memory and your buffer space, you're in a pretty doomed position."

Materialize keeps its views continuously updated and durably backed, and recovers by rehydrating from persisted state rather than by replaying history from the source. During evaluation Zepz specifically tested rehydration time, blue-green deployment of schema and feature changes, and the load a full initial snapshot places on an already stressed primary database.

The evaluation criteria went past end-to-end latency. Zepz needed one platform that could serve synchronous features on the decision path, backfill new features against historical data without leakage for model training, and hand a full change history to Databricks for long-term compliance reporting.

Materialize Team Expertise

Zepz had been down this road before. A year earlier the team had contracted for another incremental view maintenance product and spent six months trying to make it work, most of that time guessing at how much machine the workload needed and getting nothing back that helped. "It never worked, and there was no way for the product to tell us what we had to change." With Materialize the first hydration showed exactly where the memory was going, and Materialize's engineers worked through the model with Anson's team until it fit comfortably in their budget. Because of Materialize’s cloud-native architecture, they could iterate and test changes relatively quickly. Zepz landed on a configuration that was fresh enough for the inference pipeline in about a week.

Solution

Reference architecture

Diagram of Zepz's live context layer connecting eight microservice databases to fraud detection and a data lake

Eight microservice databases feed the platform. One of them is still a monolith Zepz has not decomposed yet. Each contributes 20 to 30 tables.

Inside Materialize the transformations are plain SQL, raw tables to features. The features are split across several views rather than assembled into one wide table, because a single view over the full join spikes memory during hydration. Splitting them keeps each view inside the memory envelope of the cluster it runs on.

The serving layer is a feature API doing point lookups. It joins five or six Materialize views on the fly into a single feature set for one transaction.

“Given the transaction ID, I want to know everything about the transaction and the user and the funding method.”

The fraud detection system calls that API and gets the full feature set for the decision it is about to make. Lookups are keyed on transaction ID today; a user-keyed view is a likely next step.

Rejoining the transaction

Materialize sits between the microservice databases and everything that needs to read across them.

Change data from eight microservice databases is ingested continuously and incrementally, which keeps ongoing read load off the source systems. SQL views then define the canonical business objects: the transaction as a single record rather than twenty rows, the user with current history attached, the pay-in and payout state of each transfer. Incremental view maintenance recomputes only what changed as new data arrives, so those views stay current within the tight ~1 second end-to-end latency budget the inference pipeline allows rather than being rebuilt on a schedule, which would yield stale data.

The deployment runs a three-tier cluster architecture. Source clusters handle ingestion and can be scaled up for initial snapshotting and back down afterward. Transform clusters carry the heavy join logic. Serving clusters hold the indexes that answer point lookups from the fraud service. Because compute is separated by tier, a heavy backfill or a new feature build does not compete with live serving, and neither competes with the source databases.

The important structural point is that the three layers of work are now separated and each is handled where it belongs. Denormalization happens once, in the context layer, instead of in every consuming service. The join across data sources happens in SQL over already-denormalized entities. Feature computation on top of that is comparatively simple, and it is also expressed in SQL.

Changes ship through blue-green deployment with no downtime. A second copy of the source or the view graph is built alongside the live one and swapped atomically once it has caught up, so adding a feature or absorbing an upstream schema change does not interrupt serving. That was a hard requirement rather than a nicety: during an active fraud attack, Zepz needs new features in production within the hour.

Hot path, cold path

Getting to a workable memory and local disk footprint took partnership between Zepz and Materialize teams. The first attempts at the full feature set hydrated a 19-way join across 500 million historical transactions on very large clusters, and the cost profile of holding all of that in memory was not something Zepz wanted to carry indefinitely.

Materialize's engineering team worked directly with Anson's on the model. Splitting the large join into indexed intermediate views, moving JSON unpacking earlier in the graph, and indexing narrow projections instead of wide ones brought memory down substantially. The architecture that came out of it splits hot from cold: recent activity, which is what fraud features actually depend on, is maintained live and is small enough to serve from a modest cluster, while longer history is loaded separately and refreshed on a slower cadence.

"When you're counting transactions, all you care about is what changed in the last twenty-four hours, or even the last hour. There's really no reason to keep this history of everything in memory forever."

The fraud feature store

The first and largest use case in production is the feature store behind fraud and AML detection.

Features are defined as SQL views over the canonical entities and served to the inference path where the tree-based models consume them. Because those features are maintained incrementally rather than computed on demand against twenty tables, the fraud system gets a consistent view of a transaction and its sender at the moment of authorization instead of a partially assembled one.

Consistency matters more here than in most analytical settings. A fraud feature computed from a transaction record that has half arrived is not a slightly stale feature. It is a wrong one, and it produces either a blocked good customer or an approved bad one.

History for the warehouse

Zepz is a regulated business, so the current value of a feature is not the only thing it needs. Compliance reporting requires the full history of changes.

Rather than run a second pipeline for that, Zepz syncs denormalized output from Materialize into Apache Iceberg tables, which its data team reads from Databricks, an approach sometimes called a Kappa Architecture. One set of SQL definitions produces both the live features on the decision path and the historical record in the warehouse, which removes an entire class of drift between what the model saw at decision time and what the reporting layer says happened.

One feedback loop instead of two

Feature engineering for machine learning usually runs two implementations of the same logic. One computes features online for inference. The other recomputes them offline, from raw warehouse data, to build training sets. Keeping the two in agreement is a permanent maintenance cost and a source of bugs.

Streaming the computed features into Iceberg collapses the two. The features a model trains on become the same objects, produced by the same SQL, that the model scored against in production. Training-serving skew stops being a category of bug.

Getting data without getting permission

Because the operational data is already flowing into a shared layer, a team that needs a field no longer has to negotiate for it.

"That's one of the biggest things. The data is essentially in a centralized place. If you need data for this thing, go grab it. You no longer need to negotiate to say I need this data, can you write this API endpoint for me, and then negotiate how long it takes to get there."

Results

Engineering velocity

Testing a new fraud feature or a new data signal used to take weeks or months, most of which was waiting on other teams. It now takes a day or two. Anson's team estimates the improvement at 20 to 30 times. Nearly all of the old cycle time was coordination, and the context layer removes much of it.

"Everything should be data driven. If you try one thing and it takes you three months, you're not going to get anywhere. If you can try things every day, you'll get to an optimized point within a month."

An example is a device identity signal from a new vendor, meant to establish whether a transaction is coming from a device the platform has seen before. Getting that signal into the fraud system had been under discussion for three months with no movement, because every team involved was busy and nobody could justify the work without knowing the signal's value.

"Without using that data, I can't tell you how effective it is. I actually need to get my hands on the dataset to tell you how useful it is. Now I can go to the data directly, grab it, build my own feature, try it out, and then try different combinations."

Now, instead of justifying the investment in order to get the data, the team gets the data in order to justify the investment. Signals that would have died in a backlog now get tested, and some of them turn out to matter.

That matters more in fraud than in most domains, because the adversary keeps adapting.

“Fraud is constantly evolving, and we need to be able to get our hands on any signal we can get.”

Faster iteration also changes what is discoverable at all. When the cost of testing a hypothesis drops from a quarter to a day, the team stops rationing hypotheses. Each result informs the next one, and model improvement becomes continuous rather than budgeted.

Features without database cost

Before, adding features meant adding load to a database that was already at its limit, so adding new features risked database reliability.

"We're at 100, 200 features. If someone says they have an awesome set of 100 features that would make the model better, my answer used to be: I can't, the database is already dying. We can now grow the number of features independent of our server infrastructure."

Feature count is no longer a function of transaction database capacity. The platform is being built toward 1,000 features, and adding to it is a SQL change deployed blue-green rather than a capacity conversation. That decouples model quality from operational scaling in a way that matters more over time than any single feature does.

Stability and separation

Zepz makes money when transactions complete. Feature computation is important, but it is strictly secondary to that. Under the old architecture the two shared infrastructure, which meant load from complex feature calculations could degrade or halt money movement.

"The computation is important, but it's secondary to actually making the transaction happen. Our company makes money based on the transaction happening. So if the computation becomes so heavy that it stops the transaction from happening, it hurts the business. Being able to separate it out to a separate system, so the transaction can keep processing, goes a long way."

Materialize runs the joins and the feature computation on its own compute, reading change data rather than querying the source databases. Heavy feature work, backfills, and new use cases no longer compete with authorization traffic. The failure mode where the platform goes down on its best traffic day, because success itself generated the load that broke it, is gone.

What's next

The fraud feature store is the first use case on the context layer. Zepz is looking at several more, each of which is the same problem in a different department.

Backfill for retraining. One question the POC did not answer: when a new feature is defined, how do you compute it across the last several months of transactions so a model can be retrained without waiting months for live inference data? Hydration already does this work over a 24-hour window, and extending that window is the next thing on the list to test.

“We already have it when we do the hydration process. It’s just that instead of looking at the last 24 hours of data, we may have to go back for longer.”

AI. Most of Zepz's AI work today is offline, including using language models to summarize the free-form comments operations agents leave on AML and fraud cases in order to identify the leading reasons customers get blocked. The substrate for this already exists: live, canonical, trustworthy entities that an agent could query directly rather than reassembling from twenty tables itself.

Promotions and pricing. Deciding what offer or price to show a customer requires knowing at that moment whether they are new and how much business they have done.

Capital monitoring. Knowing in real time how much capital is deployed in each destination country, and whether it needs replenishing for local settlement, is a live operational question Zepz currently answers by other means.

Unifying two brands. WorldRemit and Sendwave do nearly the same thing on separate systems written in different languages. Consolidating onto one backend is a long project. The alternative under consideration is to use Materialize as a real-time transformation layer, normalizing transactions from both brands into a single canonical dataset so downstream consumers stop caring which brand a transaction came from.

"We can do the transformation as a transaction gets created in WorldRemit and a transaction gets created in Sendwave, so that they look the same. And then whoever the consumer is, they can just consume off the same dataset."

Operational dashboards at lower cost. An operations team of about a hundred people reads dashboards that queries their data warehouse every few minutes for transaction and user detail. Their warehouse was not built for that access pattern and the bill reflects it. The data those dashboards want is already in Materialize, so Zepz has started training other teams to point at it directly.

“When you have an ops team of a hundred people looking at the dashboard and running the query every couple of minutes, it gets really pricey, because our warehouse was never designed to do that type of thing. If you want transaction information, we already have it in Materialize.”