We use Kong as an API gateway in our cloud deployments. Our plugins authenticate and authorize requests before Kong proxies them to upstream services. Some plugins also make outbound calls to internal services before forwarding the original request.
During one investigation, several API requests took more than a minute to complete. We checked whether an upstream service was slow or whether the Kong pods were experiencing CPU throttling or memory pressure. None of those explained the delay.
The first useful clue was the difference between the total time recorded by Kong and the time recorded for the upstream service.
Here is one of kong log entry:
{
"request": "GET /api/v1/policy-api/policy-documents HTTP/1.1",
"status": "200",
"upstream_connect_time": "0.001",
"upstream_response_time": "0.002",
"request_time": "65.003"
}
The upstream service responded in milliseconds, but the total request time was about 65 seconds. We initially looked at the upstream service, but these numbers ruled it out as the primary source of the delay. The missing time had to be in Kong or in another dependency called during request processing.
Some of the internal services use an authorization mechanism that is separate from the gateway's. Before proxying a request, one of our helpers obtains and caches a token by making an HTTP request to the target service.
The helper used LuaSocket:
local http = require("socket.http")
local https = require("ssl.https")
local ltn12 = require("ltn12")
The helper returned a normal response table, so its callers did not look suspicious:
local res, err = http_client.request(uri, method, headers, body)
if not res then
-- handle err
end
if res.status < 200 or res.status >= 300 then
-- handle an unsuccessful response
end
Problem was not the response table. It was the socket implementation used to build it. LuaSocket performs blocking network operations. While it waits for DNS resolution, connection setup, TLS negotiation, server processing, or response bytes, the worker executing that Lua code cannot move on to another request.
This explained the pattern in our logs: a request could be waiting for a token from one service while unrelated requests assigned to the same worker were delayed. The effect was worse when the dependency was unhealthy, because connection and read timeouts kept the call alive longer. Retries increased the delay further.
Kong runs on the NGINX/OpenResty event-driven architecture. Each worker is generally single-threaded and handles many connections by returning to the event loop whenever a connection is waiting for network activity.
That model only works when the code running in the worker cooperates with the event loop. A blocking LuaSocket call does not yield while it waits. It stalls the worker that executes it, so other requests assigned to that worker cannot make progress.
We changed the shared HTTP helper to use lua-resty-http, which is designed for OpenResty cosockets:
local httpc = require("resty.http").new()
httpc:set_timeouts(1000, 1000, 5000)
local res, err = httpc:request_uri(uri, {
method = method,
body = post_data,
headers = headers,
})
if not res then
-- handle err
end
The Lua API still looks synchronous: the caller waits for request_uri() to return. Important difference is that the network wait uses OpenResty cosockets. The current Lua coroutine can yield while the socket is waiting, allowing the worker to process other requests.
Cosockets are not available in every OpenResty or Kong phase. They are intended for supported request-processing contexts such as rewrite, access, content, and balancer phases, but are disabled in contexts such as init_worker, log, header_filter, and body_filter. This helper runs during a request phase where the socket operation can yield cooperatively. The client request still waits for the downstream response, but the NGINX worker can process other requests while the socket is waiting.
This means:
Synchronous-looking Lua API
+ event-loop-friendly socket operations
= cooperative outbound I/O for a Kong request phase
This does not make the downstream call free or asynchronous from the client's perspective. The client request still waits for the dependency. It only prevents the network wait from blocking the entire worker.
After the fix, under the same test conditions, total request time was as follows:
{
"request": "GET /api/v1/policy-api/policy-documents HTTP/1.1",
"status": "200",
"upstream_connect_time": "0.001",
"upstream_response_time": "0.002",
"request_time": "0.005"
}
For our token request, request_uri() was a reasonable choice. It is a convenient single-shot interface for token responses, metadata, configuration, and other small payloads, but it buffers the complete response body before returning.
For large downloads or uploads, use the lower-level lua-resty-http API. It lets the caller establish the connection, send the request, and consume res.body_reader in bounded chunks. The response body must be fully consumed or the connection must be closed before it can safely be reused.
The fix addressed the worker-blocking problem. It did not remove the need to manage response size, timeouts, retries, connection cleanup, or memory usage.
lua-resty-http for network I/O in request phases.socket.http, ssl.https, and socket.tcp as review warnings in Kong plugin code.io.open and file:read in request processing.Our tests also needed to cover more than the response table. A unit test with an immediate mock verifies the caller contract, but it will not catch a blocking socket. A better integration test uses a deliberately slow downstream endpoint: send a slow request and an unrelated fast request through a controlled worker, then check whether the fast request is unnecessarily delayed.
When a Kong log shows a small upstream_response_time but a large request_time, investigate work happening inside Kong before blaming the upstream service. In our case, a convenient HTTP client had become a gateway performance problem because it blocked the worker. For OpenResty request phases, use a client that cooperates with the event loop, enforce timeouts, and verify the behavior with concurrent slow-and-fast requests.