We tested Membrane API Gateway in seven scenarios, from simple proxying to JSON validation and legacy SOAP integration.
Membrane forwarded over 210,000 requests per second. Forwarding alone is only the starting point for an API gateway. Following the approach of other API gateway benchmarks, we added TLS encryption, Basic Authentication and rate limiting. Membrane still reached 196,028 requests per second. But Basic Authentication and rate limiting do not inspect the message body, so the high throughput is hardly surprising.
To go beyond forwarding, we combined plugins for JSON protection, JSONPath extractions and REST-to-SOAP conversion. Membrane still processed over 120,000 requests per second. A result we’re proud of.
How does an API gateway written entirely in Java achieve this throughput? Further down, we explain the architecture behind Membrane’s performance.
Here are the numbers: throughput and latency for all seven scenarios.
| Scenario | What the gateway does | Requests/sec | Median latency | p99 latency |
|---|---|---|---|---|
| 1. Short-circuit | Answers directly, no backend | 539,940 | 0.41 ms | 6.4 ms |
| 2. Proxying | Forwards requests to the backend | 216,912 | 0.31 ms | 3.1 ms |
| 3. Same as 2. plus Basic Auth and rate limiting | Checks credentials, counts requests, forwards | 214,125 | 0.32 ms | 3.0 ms |
| 4. Same as 3. plus TLS | Encrypted on both hops | 196,028 | 0.34 ms | 3.2 ms |
| 5. OpenAPI validation | Validates every request against an OpenAPI spec | 167,211 | 0.37 ms | 4.7 ms |
| 6. REST to SOAP | JSON protection, 5 headers from JSONPath, JSON to SOAP(XML) to JSON conversion | 120,417 | 0.58 ms | 7.0 ms |
| 7. Legacy Services | XML protection, Validation of request and response against WSDL and XSD Schema | 97,884 | 0.77 ms | 8.5 ms |
Each figure is the mean of 5 runs with 10 million requests per run.
Many gateway benchmarks only measure proxying: the gateway receives a request and forwards it unchanged. Without any plugins or interceptors, such a test measures what a plain reverse proxy does, without any of the functionality that makes an API gateway an API gateway.
In practice, a gateway authenticates clients, validates and transforms messages. These plugins run for every request, and they affect performance more than anything else. A meaningful performance test therefore has to show how the gateway performs when plugins are engaged. That is why scenarios 5 to 7 add typical API gateway features on top of plain proxying.
Membrane is fast at forwarding requests, but it really shines in real-world scenarios. Even with demanding plugins inspecting, validating and transforming messages, Membrane maintains high throughput.
Test client, gateway and backend each ran on their own virtual machine, connected by a real network:
We used three compute-optimized Azure VMs in the same virtual network. Each vCPU corresponds to a full physical core, with no hyper-threading.
| Role | Azure size | CPU | vCPUs | RAM | OS |
|---|---|---|---|---|---|
| Gateway | Standard_F16as_v7 | AMD EPYC 9V45 (Turin) | 16 | 64 GiB | Ubuntu 24.04 LTS |
| Backend | Standard_F16as_v7 | AMD EPYC 9V45 (Turin) | 16 | 64 GiB | Ubuntu 24.04 LTS |
| Client | Standard_F16as_v7 | AMD EPYC 9V45 (Turin) | 16 | 64 GiB | Ubuntu 24.04 LTS |
Here is the software setup:
| Component | Details |
|---|---|
| Gateway | Membrane API Gateway 7.6.3 with 32 GB heap and the Parallel garbage collector |
| Java | Eclipse Temurin 21 on all machines. |
| Backend | A minimal service that answers every request with 200 OK, or with a fixed SOAP response in the SOAP scenarios. |
| Client | A Java load generator based on AsyncHttpClient that keeps a fixed number of requests in flight over keep-alive connections. |
The gateway answers every request itself, without calling a backend. This measures the raw speed of the HTTP engine: accepting connections, parsing requests and writing responses. It is the upper limit for everything else.
api:
port: 2000
flow:
- return:
status: 200
api:
port: 2000
flow:
- return:
status: 200
A simple baseline. Now let’s give the gateway a backend to talk to.
The gateway receives each request and forwards it unchanged to the backend. This is the classic reverse-proxy benchmark and serves as the baseline for the scenarios with plugins.
api:
port: 2000
target:
host: backend
port: 2010
api:
port: 2000
target:
host: backend
port: 2010
Result: 216,912 requests per second with a ~1 KB JSON body. Let’s add authentication and rate limiting.
This is the minimum most APIs need. For every request, the gateway checks the user name and password (HTTP Basic Authentication), counts the request against a rate limit per client IP, and then forwards it to the backend.
api:
port: 2000
flow:
- basicAuthentication:
users:
- username: loadtest
password: "..."
- rateLimiter:
requestLimit: 100000000
requestLimitDuration: PT1H
target:
host: backend
port: 2010
api:
port: 2000
flow:
- basicAuthentication:
users:
- username: loadtest
password: "..."
- rateLimiter:
requestLimit: 100000000
requestLimitDuration: PT1H
target:
host: backend
port: 2010
We set the rate limit high enough that no request is rejected. A rejected request is cheaper than a forwarded one, so rejections would make the result look better than it is.
Result: 214,125 requests per second, only 1.3% less than plain proxying. That is within the normal variation between runs. Features that only look at the request headers, like Basic Authentication and rate limiting, have hardly any effect on performance.
What does encrypting both connections add to the cost?
Same as scenario 3, but with TLS on both hops: the client connects via HTTPS, and the gateway forwards to the backend over a second encrypted connection. The gateway decrypts every request, checks it, and encrypts it again. This setup is common in published API gateway benchmarks, so we included it to make the results comparable.
api:
port: 2000
ssl:
keystore:
location: gateway.p12
password: "..."
keyPassword: "..."
flow:
- basicAuthentication:
users:
- username: loadtest
password: "..."
- rateLimiter:
requestLimit: 100000000
requestLimitDuration: PT1H
target:
host: backend
port: 2011
ssl:
truststore:
location: truststore.p12
password: "..."
api:
port: 2000
ssl:
keystore:
location: gateway.p12
password: "..."
keyPassword: "..."
flow:
- basicAuthentication:
users:
- username: loadtest
password: "..."
- rateLimiter:
requestLimit: 100000000
requestLimitDuration: PT1H
target:
host: backend
port: 2011
ssl:
truststore:
location: truststore.p12
password: "..."
Result: 196,028 requests per second, 8.5% less than the same setup without TLS.
So far, the gateway has left the message body untouched. What happens when it has to parse JSON and validate each request against an OpenAPI specification?
This scenario shows the gateway doing real, demanding work. Validation is a typical task for an API gateway, and one a plain reverse proxy cannot do. The gateway checks every request against the Fruitshop OpenAPI specification: method, path, content type, and the JSON body against the Product schema. It only forwards valid requests to the backend. Learn more about deploying APIs from OpenAPI.
api:
port: 2000
openapi:
- location: fruitshop-v2-2-0.oas.yml
validateRequests: yes
validateResponses: no
target:
host: backend
port: 2010
api:
port: 2000
openapi:
- location: fruitshop-v2-2-0.oas.yml
validateRequests: yes
validateResponses: no
target:
host: backend
port: 2010
The request body is a valid product, {"name":"Mangos","price":2.79}. It is much smaller than the ~1 KB body of scenarios 2 to 4.
Result: Still 167,211 requests per second with a median latency of 0.37 ms and a p99 latency of 4.7 ms.
OpenAPI validation is a CPU-intensive task. Unlike authentication or rate limiting, which only look at the headers, it has to read and parse the message body and check it against the schema, for every single request. Plugins that process the body always take more time than plugins that don't. Membrane's architecture is built to process message bodies fast, and the figures show it: even with full validation, Membrane handles well over 150,000 requests per second.
Next, we connect a JSON API to a SOAP web service, converting JSON requests to XML and XML responses back to JSON.
A legacy integration: the client calls a REST/JSON API, while the backend is a SOAP web service. From the WSDL alone, the gateway offers the SOAP operation as a JSON API. For every request it:
X-Name: Jane,api:
port: 2000
flow:
- jsonProtection: {}
- request:
- setHeader:
name: X-Name
value: ${$.person.firstName}
language: jsonpath
- setHeader:
name: X-Last-Name
value: ${$.person.lastName}
language: jsonpath
- setHeader:
name: X-Email
value: ${$.person.email}
language: jsonpath
- setHeader:
name: X-Customer-Number
value: ${$.person.customerNumber}
language: jsonpath
- setHeader:
name: X-City
value: ${$.person.address.city}
language: jsonpath
- wsdl2openapi:
wsdl: person-service.wsdl
target:
url: http://backend:2012/person-service
api:
port: 2000
flow:
- jsonProtection: {}
- request:
- setHeader:
name: X-Name
value: ${$.person.firstName}
language: jsonpath
- setHeader:
name: X-Last-Name
value: ${$.person.lastName}
language: jsonpath
- setHeader:
name: X-Email
value: ${$.person.email}
language: jsonpath
- setHeader:
name: X-Customer-Number
value: ${$.person.customerNumber}
language: jsonpath
- setHeader:
name: X-City
value: ${$.person.address.city}
language: jsonpath
- wsdl2openapi:
wsdl: person-service.wsdl
target:
url: http://backend:2012/person-service
The request is a 414-byte JSON document: a person with 10 fields (strings, a date, numbers and a boolean), a list of 3 hobbies and an address object with 7 fields. The WSDL describes a document/literal SOAP 1.1 service with one operation, createPerson.
Result: 120,417 requests per second with a median latency of 0.58 ms and a p99 latency of 7.0 ms.
XML has a reputation for being slow. Let’s see how Membrane handles XML protection and schema validation of both SOAP requests and responses.
A classic SOAP gateway: the client sends a SOAP request, and the gateway protects and validates it before it reaches the backend, and validates the backend's answer on the way back. For every request it:
api:
port: 2000
flow:
- xmlProtection: {}
- validator:
wsdl: person-service.wsdl
target:
host: backend
port: 2012
api:
port: 2000
flow:
- xmlProtection: {}
- validator:
wsdl: person-service.wsdl
target:
host: backend
port: 2012
The request is the same person as in scenario 6, sent as a 763-byte SOAP 1.1 message.
Result: 97,884 requests per second with a median latency of 0.77 ms and a p99 latency of 8.5 ms. This is the most demanding scenario: the gateway parses and validates XML twice per request.
By default, Membrane keeps recent requests in memory so the admin console can show them. The forgetful exchange store turns this off, which is what you want in a high-throughput production setup. Every scenario uses it; the samples below leave it out for brevity.
components:
exchangeStore:
forgetfulExchangeStore: {}
components:
exchangeStore:
forgetfulExchangeStore: {}
Some people expect a gateway written in Java to be slower than one written in C, Go or Rust. The results show otherwise: with over 200,000 requests per second for plain proxying and over 500,000 for answering directly, Membrane doesn't have to hide from the published figures of other gateways.
Where Java really shines is complex plugins like OpenAPI validation or SOAP processing. There are two reasons:
That's why Membrane stays fast even when it inspects, validates or transforms every message body.
The whole test is scripted. Log in to Azure and run the scripts. They create the VMs, install Membrane and run the tests. Afterwards, run the teardown script to delete the VMs again, because they cost money while they are running.
Everything is in the performance-test folder of the Membrane distribution: the scripts, the complete configurations, and the detailed results of every run.
For those who want the details: the full statistics over the 5 runs of each scenario. Latency values are means over the 5 runs. The variation is the standard deviation relative to the mean. Gateway CPU per request is the gateway's CPU busy time × 16 cores ÷ requests per second; it includes the kernel's network processing.
| Scenario | Concurrency | Mean req/s | Min | Max | Variation | p50 ms | p95 ms | p99 ms | Gateway CPU | CPU per request |
|---|---|---|---|---|---|---|---|---|---|---|
| 1. Short-circuit | 350 | 539,940 | 530,480 | 549,776 | 1.3% | 0.41 | 1.24 | 6.38 | 95% | 28 µs |
| 2. Proxying | 100 | 216,912 | 212,322 | 222,629 | 1.7% | 0.31 | 1.57 | 3.07 | 89% | 66 µs |
| 3. Basic Auth + rate limiting | 100 | 214,125 | 207,392 | 220,142 | 2.7% | 0.32 | 1.49 | 3.00 | 89% | 67 µs |
| 4. Basic Auth + rate limiting + TLS | 100 | 196,028 | 186,398 | 203,433 | 3.1% | 0.34 | 1.61 | 3.20 | 90% | 73 µs |
| 5. OpenAPI validation | 100 | 167,211 | 164,438 | 169,868 | 1.2% | 0.37 | 2.17 | 4.67 | 91% | 87 µs |
| 6. REST to SOAP | 100 | 120,417 | 114,933 | 124,144 | 2.9% | 0.58 | 2.01 | 7.04 | 95% | 126 µs |
| 7. SOAP validation | 100 | 97,884 | 96,344 | 99,544 | 1.4% | 0.77 | 1.86 | 8.48 | 96% | 157 µs |
The gateway runs at 89 to 96% CPU, while the backend stays at 15 to 45%, so the figures show the limit of the gateway itself. Only in the short-circuit scenario does the client come close to its limit as well, at 91% CPU.