Meta Launches ZGateway Proxy, Handles 1 Billion Ops Per Second

By Central

Meta’s infrastructure team has launched ZGateway, a proxy tier that now sits between client applications and ZippyDB, the company’s most widely used key-value store. ZGateway handles more than one billion operations per second and already carries about 40 percent of all ZippyDB traffic, a share projected to exceed 60 percent. What began as a pragmatic fix for connection sprawl across more than a million client hosts has evolved into a sophisticated platform that centralizes batching, admission control, caching, and failover. The result is a case study in how large-scale operational problems can force architectural innovation that reshapes an entire data layer.

The Origin Problem: Why ZippyDB Needed a Proxy Tier

Before ZGateway, every ZippyDB client connected directly to every database host it needed. A single client could touch tens of thousands of shards spread across hundreds of thousands of hosts. This meant that both a typical client and a typical database host carried tens of thousands of TLS connections. Each idle connection consumed memory, CPU, and a file descriptor on both ends, and inbound connection counts grew with every new client cohort.

The problems were not theoretical. Reconnection storms caused crashes from file descriptor exhaustion and out-of-memory errors. In one notable incident, a routing bug forced every client to open a connection for every shard, and the fleet fell into a reboot loop. Fixing the issue on the client side was impractical because hundreds of separate engineering teams owned different parts of the client fleet. A different approach was needed, one that could abstract away the connection complexity at a single, managed layer.

Meta’s engineering team realized that the connection problem was not merely a nuisance but a structural ceiling on scalability. Direct-access fan-in grows linearly with the number of clients, meaning that every new service or microservice that required ZippyDB data would add to the connection burden on every database host. The only sustainable solution was to interpose a stateless proxy that could absorb the connection diversity on the client side and present a stable, predictable connection profile to the database fleet.

What Is ZGateway? Architecture and Core Mechanics

ZGateway is a stateless proxy tier positioned between ZippyDB clients and the ZServer database fleet. It runs as regional tiers discovered through ServiceRouter, Meta’s service mesh, and it comes in two flavors: a pure proxy and a read-through cache. The engine is Meta’s thick C++ ZippyDB client, so ZGateway is effectively a ZippyDB client run as a managed service. This design choice is critical – it means the proxy inherits all the correctness, performance, and replica-selection logic of the native client without requiring any changes to how the database operates.

How does ZGateway process a request? A client sends a request over a sticky connection to a regional ZGateway host. The proxy terminates TLS, authorizes the request against the use case’s access control lists, applies per-tenant admission control and traffic shaping, resolves the target shard, checks the local cache on caching tiers, batches the request with other in-flight work for that shard, and forwards it to the correct replicas. Responses are demultiplexed back to the originating clients, with per-use-case metrics, traces, and quota usage recorded along the way. TLS stays within the Thrift and ServiceRouter stack, and replica selection remains embedded in the native client logic.

At roughly 6 percent computational overhead for an average use case, the proxy’s cost is remarkably low for the operational leverage it provides. Meta reports that ZGateway handles more than one billion operations per second and carries about 40 percent of all ZippyDB traffic, with plans to absorb the majority of traffic as migration proceeds.

The Fan-In and Fan-Out Math: How Connections Collapse

The quantitative impact of ZGateway on connection counts is dramatic and worth understanding in detail. Meta models the fleet using a balls-into-bins probability framework. With B shards and H hosts, the probability that a given host is hit by a client follows the formula E(H, B) = H(1 – e^(-B/H)). Using mock figures from Meta’s engineering blog – 20 regions, 500,000 database hosts, 30,000 proxy hosts, 1,000,000 clients, and 50,000 shards per client – the per-host connection counts collapse by roughly 97 to 98 percent, and total persistent connections drop about 19-fold.

The deeper structural win, however, is about scaling dynamics. In a direct-access architecture, fan-in per database host grows linearly with the size of the client population. Every new client that needs data from a given shard adds another connection to that shard’s host. Under ZGateway, fan-in per database host reduces to roughly the product of the number of regions and the shard density per host – a number that is essentially independent of both the client fleet and the database host fleet. This means that Meta can add new services, new microservices, and new use cases without increasing the connection burden on any individual database host.

For an engineering organization that operates at Meta’s scale, this shift from linear to bounded scaling is transformative. It removes a fundamental constraint on service growth and allows the database infrastructure to be provisioned based on query volume rather than connection count.

Capabilities That Followed: Batching, Caching, and Admission Control

Once the proxy layer existed, Meta’s engineering team discovered that it could serve as a platform for capabilities that were difficult or impossible to implement at the client level. These capabilities have become some of the most valuable features of the ZGateway system.

Safe Migration and Configuration Control

Configuration flags scoped per service and per shard prefix provide a percentage ramp, a region filter, and a global kill switch. This allows Meta to roll out ZGateway to new use cases incrementally, with precise control over which traffic flows through the proxy and which traffic continues with direct access. The ability to instantly disable the proxy for a specific use case or region provides a critical safety net during migration.

Discriminant Load Shedding

Under direct access, a noisy tenant flooding a database host could degrade performance for every other tenant sharing that host. ZGateway implements Discriminant Load Shedding (DLS), a mechanism that maps requests to per-tenant buckets that are drained round-robin. Each bucket has a capacity; when a bucket overflows, only that bucket’s excess requests are shed. In a controlled test at over 90 percent CPU utilization across approximately 1,350 tenant buckets, only six noisy neighbors shed load. The remaining tenants executed 99.9 percent of their requests with zero rejections, and goodput held near 97 to 98 percent. The machinery cost about 8 percent of CPU.

The key insight of DLS is that it provides tenant-level fairness without requiring any tenant to coordinate with others. Each tenant effectively has a dedicated queue, and an overflowing queue cannot starve other queues. For a multi-tenant system operating at Meta’s scale, this is a significant improvement over traditional tail-drop or random-early-detection approaches.

Read Caching

Cache tiers serve hot reads in-process, take a per-key fill lock on cache misses, and stay fresh via change-data-capture events under a bounded-staleness contract. This eliminates the need for a separate caching layer between clients and ZippyDB and simplifies the overall architecture. The bounded-staleness contract ensures that applications see data that is never older than a configurable threshold, which is critical for use cases that require relatively fresh data but can tolerate some delay.

Load Balancing Across Heterogeneous Hardware

ZGateway tiers mix hosts with widely varying capacities – roughly 26-core to 126-core machines. A control-plane balancer nudges each host’s ServiceRouter weight in the opposite direction of its recent CPU load. This heterogeneous approach allows Meta to use hardware efficiently, deploying smaller machines for lower-traffic regions and larger machines for high-traffic regions, while still keeping utilization balanced across the fleet.

Cross-Region Resilience

Global routing, mega-regions, and rings allow a saturated regional tier to fail over to healthy capacity in a nearby region. This capability provides headroom for traffic spikes and protects against regional failures without requiring clients to be aware of the underlying topology changes.

Transaction Consolidation

Client-side bookkeeping for transactions was moved into the gateway in nine phases, eventually consolidating 100 percent of transaction traffic with no reliability regression. This migration eliminated a source of complexity in client libraries and gave the operations team visibility into transaction performance that was previously scattered across hundreds of teams.

Why Client-Side Fixes Were Unworkable

A natural question is why Meta did not simply fix the connection sprawl problem by improving the client libraries. The answer lies in the structural reality of large-scale engineering organizations. Hundreds of teams own different parts of the client fleet, each with its own priorities, release schedules, and engineering capacity. A change that requires every team to update its client library, test it, and deploy it would take months or years to propagate fully – and some teams might never find the time or motivation to do so.

Furthermore, the connection problem was not a bug in the client libraries but a fundamental consequence of the direct-access architecture. Even the most efficient client library would still need to open connections to every database host it needed. The only way to escape this scaling law was to insert a proxy layer that could multiplex connections on behalf of many clients.

This decision reflects an important principle in infrastructure engineering: when a problem affects hundreds of teams, it is often more efficient to fix it once at the platform layer than to ask every team to fix it individually. The upfront investment in building a proxy is high, but the ongoing operational savings and the elimination of a structural scaling constraint can be enormous.

The Limitations: ZGateway Is Not a Product

It is important to understand what ZGateway is not. ZGateway is not a product that can be downloaded, installed, or deployed outside Meta. It is deeply integrated with Meta’s internal infrastructure – ServiceRouter, Thrift, ZippyDB itself, and the company’s monitoring and configuration systems. The proxy’s C++ engine is the same code that runs inside Meta’s ZippyDB clients, and it relies on internal libraries and abstractions that are not available externally.

The value of ZGateway for the broader engineering community lies in the patterns and principles it demonstrates, not in the code itself. The fan-in math, the discriminant load shedding mechanism, the cross-client batching approach, and the strategy of turning a client library into a managed service are all ideas that can be applied to other systems at other companies. Engineers building similar proxy layers for key-value stores, message queues, or database systems can learn from Meta’s experience.

For organizations that do not operate at Meta’s scale, the specific techniques may be overkill. A company with a few dozen database hosts and a few hundred clients does not need a proxy tier that handles one billion operations per second. But the general principle – that a proxy layer can decouple client complexity from database stability and unlock new capabilities – applies at many scales.

A Question Answered: What Makes ZGateway Different From a Standard Database Proxy?

What makes ZGateway different from a standard database proxy like a read replica or a connection pooler? The key difference is that ZGateway is not just a connection multiplexer but a full-featured platform that runs the same thick client code as the applications themselves. This gives it deep knowledge of the database’s semantics, including shard routing, replica selection, transaction handling, and consistency guarantees. A standard connection pooler or proxy can only forward bytes; ZGateway can understand and manipulate requests. It can batch requests for the same shard, coalesce requests for the same hot key, apply tenant-level admission control, and make intelligent routing decisions based on cache state and load. In essence, ZGateway is a native ZippyDB client that has been repackaged as a managed service, and this architectural choice is what enables its advanced capabilities.

The Strategic Significance for Database Infrastructure

ZGateway represents a strategic bet on decoupling two scaling dimensions: the number of clients and the number of database hosts. In traditional architectures, these two dimensions are tightly coupled. Every new client adds load to every database host it needs, and every new database host adds connection targets for every client. This coupling makes both fleets harder to manage and harder to scale independently.

With ZGateway, the client fleet and the database fleet can scale independently. New services can be added without increasing the connection burden on database hosts. Database hosts can be added or removed without requiring client-side changes. The proxy layer absorbs the coupling and provides a stable interface on both sides. This is not a minor operational convenience; it is a fundamental change in how the infrastructure can be grown and managed.

The experience also demonstrates that infrastructure teams should think about proxies not as mere intermediaries but as platforms for operational capabilities. Once a proxy exists, it becomes a natural place to implement admission control, caching, load balancing, and resilience patterns. The proxy becomes the control plane for the data layer, and its capabilities can be extended over time without touching the clients or the database.

For companies building or operating large-scale data systems, the lesson is clear: invest in proxy layers early, even if the immediate need seems manageable. The connection sprawl that ZGateway addresses is a slow-moving problem that compounds over time. By the time it becomes urgent, retrofitting a solution is much harder than designing one from the start. And once the proxy exists, it will almost certainly become the foundation for capabilities that were not even imagined when it was first built.

Share This Article