Open Source API Gateway

Membrane API Gateway Performance Benchmark

How Fast Is Membrane With Realistic API Workloads?

We tested Membrane API Gateway in seven scenarios, from simple proxying to JSON validation and legacy SOAP integration.

Membrane forwarded over 210,000 requests per second. Forwarding alone is only the starting point for an API gateway. Following the approach of other API gateway benchmarks, we added TLS encryption, Basic Authentication and rate limiting. Membrane still reached 196,028 requests per second. But Basic Authentication and rate limiting do not inspect the message body, so the high throughput is hardly surprising.

To go beyond forwarding, we combined plugins for JSON protection, JSONPath extractions and REST-to-SOAP conversion. Membrane still processed over 120,000 requests per second. A result we’re proud of.

How does an API gateway written entirely in Java achieve this throughput? Further down, we explain the architecture behind Membrane’s performance.

Benchmark Results

Here are the numbers: throughput and latency for all seven scenarios.

ScenarioWhat the gateway doesRequests/secMedian latencyp99 latency
1. Short-circuitAnswers directly, no backend539,9400.41 ms6.4 ms
2. ProxyingForwards requests to the backend216,9120.31 ms3.1 ms
3. Same as 2. plus Basic Auth and rate limitingChecks credentials, counts requests, forwards214,1250.32 ms3.0 ms
4. Same as 3. plus TLSEncrypted on both hops196,0280.34 ms3.2 ms
5. OpenAPI validationValidates every request against an OpenAPI spec167,2110.37 ms4.7 ms
6. REST to SOAPJSON protection, 5 headers from JSONPath, JSON to SOAP(XML) to JSON conversion120,4170.58 ms7.0 ms
7. Legacy ServicesXML protection, Validation of request and response against WSDL and XSD Schema97,8840.77 ms8.5 ms

Each figure is the mean of 5 runs with 10 million requests per run.

Why Plugins Matter

Many gateway benchmarks only measure proxying: the gateway receives a request and forwards it unchanged. Without any plugins or interceptors, such a test measures what a plain reverse proxy does, without any of the functionality that makes an API gateway an API gateway.

In practice, a gateway authenticates clients, validates and transforms messages. These plugins run for every request, and they affect performance more than anything else. A meaningful performance test therefore has to show how the gateway performs when plugins are engaged. That is why scenarios 5 to 7 add typical API gateway features on top of plain proxying.

Membrane is fast at forwarding requests, but it really shines in real-world scenarios. Even with demanding plugins inspecting, validating and transforming messages, Membrane maintains high throughput.

The Test Setup

Test client, gateway and backend each ran on their own virtual machine, connected by a real network:

Client
16 vCPU
Membrane
16 vCPU
Backend
16 vCPU

We used three compute-optimized Azure VMs in the same virtual network. Each vCPU corresponds to a full physical core, with no hyper-threading.

RoleAzure sizeCPUvCPUsRAMOS
GatewayStandard_F16as_v7AMD EPYC 9V45 (Turin)1664 GiBUbuntu 24.04 LTS
BackendStandard_F16as_v7AMD EPYC 9V45 (Turin)1664 GiBUbuntu 24.04 LTS
ClientStandard_F16as_v7AMD EPYC 9V45 (Turin)1664 GiBUbuntu 24.04 LTS

Here is the software setup:

ComponentDetails
GatewayMembrane API Gateway 7.6.3 with 32 GB heap and the Parallel garbage collector
JavaEclipse Temurin 21 on all machines.
BackendA minimal service that answers every request with 200 OK, or with a fixed SOAP response in the SOAP scenarios.
ClientA Java load generator based on AsyncHttpClient that keeps a fixed number of requests in flight over keep-alive connections.

How We Measured

  • 5 rounds. Each round ran all 7 scenarios once. The figures on this page are the mean of these 5 runs.
  • 10 million requests per run, after 1 million warmup requests that are not counted.
  • Fixed concurrency. 100 requests in flight, 350 in the short-circuit scenario. As soon as a response arrives, the client sends the next request.
  • Fresh start. The gateway was restarted with the scenario's configuration before every run.

Scenario 1: Short-Circuit

The gateway answers every request itself, without calling a backend. This measures the raw speed of the HTTP engine: accepting connections, parsing requests and writing responses. It is the upper limit for everything else.

api:
port: 2000
flow:
- return:
status: 200
api:
  port: 2000
  flow:
    - return:
        status: 200

A simple baseline. Now let’s give the gateway a backend to talk to.

Scenario 2: Proxying

The gateway receives each request and forwards it unchanged to the backend. This is the classic reverse-proxy benchmark and serves as the baseline for the scenarios with plugins.

api:
port: 2000
target:
host: backend
port: 2010
api:
  port: 2000
  target:
    host: backend
    port: 2010

Result: 216,912 requests per second with a ~1 KB JSON body. Let’s add authentication and rate limiting.

Scenario 3: Basic Auth + Rate Limiting

This is the minimum most APIs need. For every request, the gateway checks the user name and password (HTTP Basic Authentication), counts the request against a rate limit per client IP, and then forwards it to the backend.

api:
port: 2000
flow:
- basicAuthentication:
users:
- username: loadtest
password: "..."
- rateLimiter:
requestLimit: 100000000
requestLimitDuration: PT1H
target:
host: backend
port: 2010
api:
  port: 2000
  flow:
    - basicAuthentication:
        users:
          - username: loadtest
            password: "..."
    - rateLimiter:
        requestLimit: 100000000
        requestLimitDuration: PT1H
  target:
    host: backend
    port: 2010

We set the rate limit high enough that no request is rejected. A rejected request is cheaper than a forwarded one, so rejections would make the result look better than it is.

Result: 214,125 requests per second, only 1.3% less than plain proxying. That is within the normal variation between runs. Features that only look at the request headers, like Basic Authentication and rate limiting, have hardly any effect on performance.

What does encrypting both connections add to the cost?

Scenario 4: Basic Auth + Rate Limiting + TLS

Same as scenario 3, but with TLS on both hops: the client connects via HTTPS, and the gateway forwards to the backend over a second encrypted connection. The gateway decrypts every request, checks it, and encrypts it again. This setup is common in published API gateway benchmarks, so we included it to make the results comparable.

api:
port: 2000
ssl:
keystore:
location: gateway.p12
password: "..."
keyPassword: "..."
flow:
- basicAuthentication:
users:
- username: loadtest
password: "..."
- rateLimiter:
requestLimit: 100000000
requestLimitDuration: PT1H
target:
host: backend
port: 2011
ssl:
truststore:
location: truststore.p12
password: "..."
api:
  port: 2000
  ssl:
    keystore:
      location: gateway.p12
      password: "..."
      keyPassword: "..."
  flow:
    - basicAuthentication:
        users:
          - username: loadtest
            password: "..."
    - rateLimiter:
        requestLimit: 100000000
        requestLimitDuration: PT1H
  target:
    host: backend
    port: 2011
    ssl:
      truststore:
        location: truststore.p12
        password: "..."

Result: 196,028 requests per second, 8.5% less than the same setup without TLS.

So far, the gateway has left the message body untouched. What happens when it has to parse JSON and validate each request against an OpenAPI specification?

Scenario 5: OpenAPI Validation

This scenario shows the gateway doing real, demanding work. Validation is a typical task for an API gateway, and one a plain reverse proxy cannot do. The gateway checks every request against the Fruitshop OpenAPI specification: method, path, content type, and the JSON body against the Product schema. It only forwards valid requests to the backend. Learn more about deploying APIs from OpenAPI.

api:
port: 2000
openapi:
- location: fruitshop-v2-2-0.oas.yml
validateRequests: yes
validateResponses: no
target:
host: backend
port: 2010
api:
  port: 2000
  openapi:
    - location: fruitshop-v2-2-0.oas.yml
      validateRequests: yes
      validateResponses: no
  target:
    host: backend
    port: 2010

The request body is a valid product, {"name":"Mangos","price":2.79}. It is much smaller than the ~1 KB body of scenarios 2 to 4.

Result: Still 167,211 requests per second with a median latency of 0.37 ms and a p99 latency of 4.7 ms.

OpenAPI validation is a CPU-intensive task. Unlike authentication or rate limiting, which only look at the headers, it has to read and parse the message body and check it against the schema, for every single request. Plugins that process the body always take more time than plugins that don't. Membrane's architecture is built to process message bodies fast, and the figures show it: even with full validation, Membrane handles well over 150,000 requests per second.

Next, we connect a JSON API to a SOAP web service, converting JSON requests to XML and XML responses back to JSON.

Scenario 6: REST to SOAP

A legacy integration: the client calls a REST/JSON API, while the backend is a SOAP web service. From the WSDL alone, the gateway offers the SOAP operation as a JSON API. For every request it:

  1. checks the JSON against protection limits such as depth, size and string lengths,
  2. extracts five fields with JSONPath and sets them as HTTP headers, for example X-Name: Jane,
  3. converts the JSON into a SOAP request as described by the WSDL,
  4. calls the SOAP backend,
  5. converts the SOAP response back into JSON.
api:
port: 2000
flow:
- jsonProtection: {}
- request:
- setHeader:
name: X-Name
value: ${$.person.firstName}
language: jsonpath
- setHeader:
name: X-Last-Name
value: ${$.person.lastName}
language: jsonpath
- setHeader:
name: X-Email
value: ${$.person.email}
language: jsonpath
- setHeader:
name: X-Customer-Number
value: ${$.person.customerNumber}
language: jsonpath
- setHeader:
name: X-City
value: ${$.person.address.city}
language: jsonpath
- wsdl2openapi:
wsdl: person-service.wsdl
target:
url: http://backend:2012/person-service
api:
  port: 2000
  flow:
    - jsonProtection: {}
    - request:
        - setHeader:
            name: X-Name
            value: ${$.person.firstName}
            language: jsonpath
        - setHeader:
            name: X-Last-Name
            value: ${$.person.lastName}
            language: jsonpath
        - setHeader:
            name: X-Email
            value: ${$.person.email}
            language: jsonpath
        - setHeader:
            name: X-Customer-Number
            value: ${$.person.customerNumber}
            language: jsonpath
        - setHeader:
            name: X-City
            value: ${$.person.address.city}
            language: jsonpath
    - wsdl2openapi:
        wsdl: person-service.wsdl
  target:
    url: http://backend:2012/person-service

The request is a 414-byte JSON document: a person with 10 fields (strings, a date, numbers and a boolean), a list of 3 hobbies and an address object with 7 fields. The WSDL describes a document/literal SOAP 1.1 service with one operation, createPerson.

Result: 120,417 requests per second with a median latency of 0.58 ms and a p99 latency of 7.0 ms.

XML has a reputation for being slow. Let’s see how Membrane handles XML protection and schema validation of both SOAP requests and responses.

Scenario 7: SOAP Validation

A classic SOAP gateway: the client sends a SOAP request, and the gateway protects and validates it before it reaches the backend, and validates the backend's answer on the way back. For every request it:

  1. checks the XML against protection limits and removes any DTD, which defends against XML attacks such as entity expansion,
  2. validates the SOAP request against the WSDL and its XML Schema,
  3. forwards it to the SOAP backend,
  4. validates the SOAP response against the same WSDL.
api:
port: 2000
flow:
- xmlProtection: {}
- validator:
wsdl: person-service.wsdl
target:
host: backend
port: 2012
api:
  port: 2000
  flow:
    - xmlProtection: {}
    - validator:
        wsdl: person-service.wsdl
  target:
    host: backend
    port: 2012

The request is the same person as in scenario 6, sent as a 763-byte SOAP 1.1 message.

Result: 97,884 requests per second with a median latency of 0.77 ms and a p99 latency of 8.5 ms. This is the most demanding scenario: the gateway parses and validates XML twice per request.

Gateway Settings Used in Every Scenario

By default, Membrane keeps recent requests in memory so the admin console can show them. The forgetful exchange store turns this off, which is what you want in a high-throughput production setup. Every scenario uses it; the samples below leave it out for brevity.

components:
exchangeStore:
forgetfulExchangeStore: {}
        components:
        exchangeStore:
        forgetfulExchangeStore: {}
        

Can a Java Gateway Keep Up?

Some people expect a gateway written in Java to be slower than one written in C, Go or Rust. The results show otherwise: with over 200,000 requests per second for plain proxying and over 500,000 for answering directly, Membrane doesn't have to hide from the published figures of other gateways.

Where Java really shines is complex plugins like OpenAPI validation or SOAP processing. There are two reasons:

  • No seam between core and plugins. Membrane's HTTP core and its plugins are implemented in the same technology. Messages don't have to cross a language boundary or be handed over to a separate scripting engine; plugins work directly on the gateway's own data.
  • The Java platform. The vast Java ecosystem offers many mature, highly optimized solutions and libraries, for example for parsing and validating JSON and XML, that Membrane can build on.

That's why Membrane stays fast even when it inspects, validates or transforms every message body.

What These Numbers Mean

  • Plugins matter more than raw speed. The same gateway on the same machine goes from 540,000 requests per second when answering directly to 98,000 with full SOAP validation. Header-based features like authentication and rate limiting are almost free. Plugins that process the body cost more, but Membrane keeps them fast.
  • Low latency. The median latency is 0.3 to 0.8 ms and the p99 latency 3.0 to 8.5 ms. All but one of the 350 million requests finished within 70 ms.
  • Minimal backend. The backend does no real work. In a real deployment, your backend and its database will almost always be the bottleneck long before the gateway is.
  • Hardware matters. The numbers apply to the 16-vCPU machines described above. Cloud VMs land on different physical hosts, so expect the same order of magnitude when you repeat the tests, not the exact same figures.

Run the Tests Yourself

The whole test is scripted. Log in to Azure and run the scripts. They create the VMs, install Membrane and run the tests. Afterwards, run the teardown script to delete the VMs again, because they cost money while they are running.

Everything is in the performance-test folder of the Membrane distribution: the scripts, the complete configurations, and the detailed results of every run.

All Results

For those who want the details: the full statistics over the 5 runs of each scenario. Latency values are means over the 5 runs. The variation is the standard deviation relative to the mean. Gateway CPU per request is the gateway's CPU busy time × 16 cores ÷ requests per second; it includes the kernel's network processing.

ScenarioConcurrencyMean req/sMinMaxVariationp50 msp95 msp99 msGateway CPUCPU per request
1. Short-circuit350539,940530,480549,7761.3%0.411.246.3895%28 µs
2. Proxying100216,912212,322222,6291.7%0.311.573.0789%66 µs
3. Basic Auth + rate limiting100214,125207,392220,1422.7%0.321.493.0089%67 µs
4. Basic Auth + rate limiting + TLS100196,028186,398203,4333.1%0.341.613.2090%73 µs
5. OpenAPI validation100167,211164,438169,8681.2%0.372.174.6791%87 µs
6. REST to SOAP100120,417114,933124,1442.9%0.582.017.0495%126 µs
7. SOAP validation10097,88496,34499,5441.4%0.771.868.4896%157 µs

The gateway runs at 89 to 96% CPU, while the backend stays at 15 to 45%, so the figures show the limit of the gateway itself. Only in the short-circuit scenario does the client come close to its limit as well, at 91% CPU.