How I made a Sidekiq job 6x faster with HTTP keep-alive
In 2017 I was working on a Rails backend with Sidekiq. We were one of the pilot teams of the Bing Ads API. Our sync was slow for a reason I didn't expect: opening connections.
One job was one client account. Each job made about 40 HTTPS requests to Microsoft, most of them uploading big XML files.
A full sync was 3,000 jobs. So 120,000 requests. Several times a day.
Fine in testing, stuck in production
We tested with batches of 10 to 20 jobs. Everything worked.
Then production. Boom. Jobs piled up in the queue and a full sync took forever.
My suspects:
- Sidekiq concurrency blocking jobs.
- A memory leak. But no memory alerts.
- One request retrying and blocking the others.
I timed the XML generation: 15 ms. I added HTTPLog: the HTTP call was 44 ms. Nothing slow.
The log that gave it away
I don't remember why, but I turned on every option in HTTPLog:
HttpLog.configure do |config|
config.log_connect = true
config.log_request = true
config.log_data = true
config.log_status = true
config.log_response = true
config.log_benchmark = true
end
And I saw this:
[2017-10-09-10:34:33.15] [httplog] Connecting: bulk.api.bingads.microsoft.com:443
[2017-10-09-10:34:33.75] [httplog] Sending: POST https://bulk.api.bingads.microsoft.com/...
[2017-10-09-10:34:33.77] [httplog] Status: 200
Did you spot it?
600 ms between "Connecting" and "Sending". Before a single byte of the request left our server. The 44 ms I measured before didn't include it.
What happens in those 600 ms
Every request opened a new connection. Over HTTPS that means two handshakes before anything useful:
- TCP: SYN, SYN-ACK. One round trip.
- TLS 1.2, what everyone ran in 2017: certificate, key exchange. Two more round trips.
Then the request itself: one more round trip.
I measured it again this week, on a server 96 ms away:
Opening the connection is 3 round trips. From a server in Europe to one on the US West Coast, a round trip is 150 to 200 ms. Three of them: 450 to 600 ms. And we didn't control the other side. It was Microsoft.
The math
We had 10 Sidekiq processes with 3 threads each. 30 jobs at a time.
Each request: 600 ms to open the connection, about 100 ms of real work. 700 ms.
- One job: 40 × 0.7 s = 28 s.
- Full sync: 3,000 jobs / 30 threads = 100 jobs per thread. 100 × 28 s = 47 min.
86% of that time, our threads were waiting for handshakes.
And testing couldn't show it. 10 jobs on 30 threads is one round: 28 s. Slow, but nothing looked broken.
The fix: keep the socket open
I switched to Excon because it has a persistent option:
connection = Excon.new("https://bulk.api.bingads.microsoft.com", persistent: true)
requests.each do |request|
connection.post(path: request.path, body: request.xml) # same socket every time
end
One connection per job. 40 requests on it. One handshake instead of 40.
- One job: 0.6 s + 40 × 0.1 s = 4.6 s instead of 28 s.
- Full sync: 100 × 4.6 s = 7.7 min instead of 47.
6x faster. From 120,000 handshakes per sync to 3,000.
And also count the slow start benefit on large XML
Those big XML uploads had a second problem. A new TCP connection starts slow.
The sender doesn't know your bandwidth, so it starts with about 14 KB in flight and doubles every round trip: 14, 28, 57, 114 KB... A fresh connection needs several round trips to move a big file. A warm one already sped up.
Here's 1 MB coming from Singapore, round trip by round trip:
And the same 1 MB from different distances:
Reusing the connection fixed that too. I didn't even know it at the time.
One catch on Linux: the window shrinks again after the connection sits idle. On the Sidekiq servers, net.ipv4.tcp_slow_start_after_idle=0 keeps it warm.
Is it still true in 2026?
I wrote a small Python script and measured 11 LLM APIs (OpenAI, Anthropic, OpenRouter, Gemini, Mistral, Groq and more) plus a few servers far away. 20 requests in a row: a new connection each time, then one connection.
Two things changed since 2017. APIs sit behind a CDN a few ms away from you, and TLS 1.3 needs one round trip instead of two. So opening a connection to OpenAI or Anthropic costs about 20 ms from my desk.
When the server takes 100+ ms to answer, that's 7% to 26% of each request. When it's fast, the handshake is most of it: OpenRouter goes from 45 ms to 14 ms per request.
And a server far away, still on TLS 1.2? 4.2× faster with one connection. Same story as 2017.
What I do now
- One HTTP client per thread, reused. Ruby:
Net::HTTP.startwith a block,net-http-persistent, HTTPX, or Excon withpersistent: true. Python:requests.Sessionorhttpx.Client. - Create the OpenAI or Anthropic client once. It holds the connection pool. A new client per request throws it away.
- Log the connect time, not only the request time. That's where my 600 ms was hiding.
- Test with production volume. 10 jobs will never show you this.