How I made a Sidekiq job 6x faster with HTTP keep-alive

In 2017 I was working on a Rails backend with Sidekiq. We were one of the pilot teams of the Bing Ads API. Our sync was slow for a reason I didn't expect: opening connections.

One job was one client account. Each job made about 40 HTTPS requests to Microsoft, most of them uploading big XML files.

A full sync was 3,000 jobs. So 120,000 requests. Several times a day.

Fine in testing, stuck in production

We tested with batches of 10 to 20 jobs. Everything worked.

Then production. Boom. Jobs piled up in the queue and a full sync took forever.

My suspects:

  • Sidekiq concurrency blocking jobs.
  • A memory leak. But no memory alerts.
  • One request retrying and blocking the others.

I timed the XML generation: 15 ms. I added HTTPLog: the HTTP call was 44 ms. Nothing slow.

The log that gave it away

I don't remember why, but I turned on every option in HTTPLog:

HttpLog.configure do |config|
  config.log_connect   = true
  config.log_request   = true
  config.log_data      = true
  config.log_status    = true
  config.log_response  = true
  config.log_benchmark = true
end

And I saw this:

[2017-10-09-10:34:33.15] [httplog] Connecting: bulk.api.bingads.microsoft.com:443
[2017-10-09-10:34:33.75] [httplog] Sending: POST https://bulk.api.bingads.microsoft.com/...
[2017-10-09-10:34:33.77] [httplog] Status: 200

Did you spot it?

600 ms between "Connecting" and "Sending". Before a single byte of the request left our server. The 44 ms I measured before didn't include it.

What happens in those 600 ms

Every request opened a new connection. Over HTTPS that means two handshakes before anything useful:

  • TCP: SYN, SYN-ACK. One round trip.
  • TLS 1.2, what everyone ran in 2017: certificate, key exchange. Two more round trips.

Then the request itself: one more round trip.

I measured it again this week, on a server 96 ms away:

Less time is betterTCP + TLS handshakesThe request itself
0100200300400msNew connection (TLS 1.2)your serverAPI serverSYNSYN-ACKClientHellocertificate, keykey exchange, FinishedFinishedPOSTresponse403 msConnection already openyour serverAPI serverPOSTresponse99 ms304 msbefore the POST
Measured this week against a server 96 ms away. Time runs down. Each arrow is one trip across the network.

Opening the connection is 3 round trips. From a server in Europe to one on the US West Coast, a round trip is 150 to 200 ms. Three of them: 450 to 600 ms. And we didn't control the other side. It was Microsoft.

The math

We had 10 Sidekiq processes with 3 threads each. 30 jobs at a time.

Each request: 600 ms to open the connection, about 100 ms of real work. 700 ms.

  • One job: 40 × 0.7 s = 28 s.
  • Full sync: 3,000 jobs / 30 threads = 100 jobs per thread. 100 × 28 s = 47 min.
Shorter is betterOpening connectionsReal work
queue3,000 jobsprocess 1process 1, thread 1: one jobjobprocess 1, thread 2: one jobjobprocess 1, thread 3: one jobjobprocess 2process 2, thread 1: one jobjobprocess 2, thread 2: one jobjobprocess 2, thread 3: one jobjobprocess 3process 3, thread 1: one jobjobprocess 3, thread 2: one jobjobprocess 3, thread 3: one jobjobprocess 4process 4, thread 1: one jobjobprocess 4, thread 2: one jobjobprocess 4, thread 3: one jobjobprocess 5process 5, thread 1: one jobjobprocess 5, thread 2: one jobjobprocess 5, thread 3: one jobjobprocess 6process 6, thread 1: one jobjobprocess 6, thread 2: one jobjobprocess 6, thread 3: one jobjobprocess 7process 7, thread 1: one jobjobprocess 7, thread 2: one jobjobprocess 7, thread 3: one jobjobprocess 8process 8, thread 1: one jobjobprocess 8, thread 2: one jobjobprocess 8, thread 3: one jobjobprocess 9process 9, thread 1: one jobjobprocess 9, thread 2: one jobjobprocess 9, thread 3: one jobjobprocess 10process 10, thread 1: one jobjobprocess 10, thread 2: one jobjobprocess 10, thread 3: one jobjob30 jobs at a time. Each thread runs 100 jobs in a row.Full sync, 3,000 jobs on 30 threads0 min10 min20 min30 min40 min50 minNew connection per requestOpening connections: 40 minReal work: 6.7 min46.7 minOne connection per jobOpening connections: 1 minReal work: 6.7 min7.7 min

86% of that time, our threads were waiting for handshakes.

And testing couldn't show it. 10 jobs on 30 threads is one round: 28 s. Slow, but nothing looked broken.

The fix: keep the socket open

I switched to Excon because it has a persistent option:

connection = Excon.new("https://bulk.api.bingads.microsoft.com", persistent: true)

requests.each do |request|
  connection.post(path: request.path, body: request.xml) # same socket every time
end

One connection per job. 40 requests on it. One handshake instead of 40.

Shorter is betterOpening the connection (TCP + TLS): 600 msRequest: 100 ms
0 s5 s10 s15 s20 s25 s30 sBefore: a new connection for every requestRequest 1: 600 ms opening the connectionRequest 1: 100 ms of real workRequest 2: 600 ms opening the connectionRequest 2: 100 ms of real workRequest 3: 600 ms opening the connectionRequest 3: 100 ms of real workRequest 4: 600 ms opening the connectionRequest 4: 100 ms of real workRequest 5: 600 ms opening the connectionRequest 5: 100 ms of real workRequest 6: 600 ms opening the connectionRequest 6: 100 ms of real workRequest 7: 600 ms opening the connectionRequest 7: 100 ms of real workRequest 8: 600 ms opening the connectionRequest 8: 100 ms of real workRequest 9: 600 ms opening the connectionRequest 9: 100 ms of real workRequest 10: 600 ms opening the connectionRequest 10: 100 ms of real workRequest 11: 600 ms opening the connectionRequest 11: 100 ms of real workRequest 12: 600 ms opening the connectionRequest 12: 100 ms of real workRequest 13: 600 ms opening the connectionRequest 13: 100 ms of real workRequest 14: 600 ms opening the connectionRequest 14: 100 ms of real workRequest 15: 600 ms opening the connectionRequest 15: 100 ms of real workRequest 16: 600 ms opening the connectionRequest 16: 100 ms of real workRequest 17: 600 ms opening the connectionRequest 17: 100 ms of real workRequest 18: 600 ms opening the connectionRequest 18: 100 ms of real workRequest 19: 600 ms opening the connectionRequest 19: 100 ms of real workRequest 20: 600 ms opening the connectionRequest 20: 100 ms of real workRequest 21: 600 ms opening the connectionRequest 21: 100 ms of real workRequest 22: 600 ms opening the connectionRequest 22: 100 ms of real workRequest 23: 600 ms opening the connectionRequest 23: 100 ms of real workRequest 24: 600 ms opening the connectionRequest 24: 100 ms of real workRequest 25: 600 ms opening the connectionRequest 25: 100 ms of real workRequest 26: 600 ms opening the connectionRequest 26: 100 ms of real workRequest 27: 600 ms opening the connectionRequest 27: 100 ms of real workRequest 28: 600 ms opening the connectionRequest 28: 100 ms of real workRequest 29: 600 ms opening the connectionRequest 29: 100 ms of real workRequest 30: 600 ms opening the connectionRequest 30: 100 ms of real workRequest 31: 600 ms opening the connectionRequest 31: 100 ms of real workRequest 32: 600 ms opening the connectionRequest 32: 100 ms of real workRequest 33: 600 ms opening the connectionRequest 33: 100 ms of real workRequest 34: 600 ms opening the connectionRequest 34: 100 ms of real workRequest 35: 600 ms opening the connectionRequest 35: 100 ms of real workRequest 36: 600 ms opening the connectionRequest 36: 100 ms of real workRequest 37: 600 ms opening the connectionRequest 37: 100 ms of real workRequest 38: 600 ms opening the connectionRequest 38: 100 ms of real workRequest 39: 600 ms opening the connectionRequest 39: 100 ms of real workRequest 40: 600 ms opening the connectionRequest 40: 100 ms of real work28 sAfter: one connection per job600 ms opening the connection, once40 requests × 100 ms of real work4.6 s
One job, 40 requests, drawn to scale. Hover a block for its duration.
  • One job: 0.6 s + 40 × 0.1 s = 4.6 s instead of 28 s.
  • Full sync: 100 × 4.6 s = 7.7 min instead of 47.

6x faster. From 120,000 handshakes per sync to 3,000.

And also count the slow start benefit on large XML

Those big XML uploads had a second problem. A new TCP connection starts slow.

The sender doesn't know your bandwidth, so it starts with about 14 KB in flight and doubles every round trip: 14, 28, 57, 114 KB... A fresh connection needs several round trips to move a big file. A warm one already sped up.

Here's 1 MB coming from Singapore, round trip by round trip:

More KB per round trip is betterFresh connection: 5 round tripsWarm connection: 1 round trip
KB02505007501,000Round trip 1, fresh connection: 82 KB82Round trip 1, warm connection: 1,000 KB1,000round trip 1Round trip 2, fresh connection: 82 KB82round trip 2Round trip 3, fresh connection: 147 KB147round trip 3Round trip 4, fresh connection: 262 KB262round trip 4Round trip 5, fresh connection: 427 KB427round trip 5
1 MB from Singapore, 290 ms per round trip, measured on 30 Sep 2026. KB received in each round trip after the first byte. Every extra round trip costs another 290 ms.

And the same 1 MB from different distances:

Shorter is betterFresh connectionWarm connection
05001,0001,5002,0002,5004 ms away4 ms away: fresh connection: 57 ms57 ms4 ms away: warm connection: 43 ms43 ms49 ms away (Germany)49 ms away (Germany): fresh connection: 329 ms329 ms49 ms away (Germany): warm connection: 147 ms147 ms · 2.2× faster106 ms away (US East)106 ms away (US East): fresh connection: 734 ms734 ms106 ms away (US East): warm connection: 114 ms114 ms · 6.5× faster290 ms away (Singapore)290 ms away (Singapore): fresh connection: 2,027 ms2,027 ms290 ms away (Singapore): warm connection: 297 ms297 ms · 6.8× faster
Downloading 1 MB after the handshake, measured on 30 Sep 2026. Same bandwidth. The fresh connection needs more round trips to get up to speed.

Reusing the connection fixed that too. I didn't even know it at the time.

One catch on Linux: the window shrinks again after the connection sits idle. On the Sidekiq servers, net.ipv4.tcp_slow_start_after_idle=0 keeps it warm.

Is it still true in 2026?

I wrote a small Python script and measured 11 LLM APIs (OpenAI, Anthropic, OpenRouter, Gemini, Mistral, Groq and more) plus a few servers far away. 20 requests in a row: a new connection each time, then one connection.

Shorter is betterNew connection for every requestOne connection for all 20
0 s2.5 s5 s7.5 s10 s12.5 sServer 107 ms away, TLS 1.2Server 107 ms away, TLS 1.2: new connection each time: 10.8 s10.8 sServer 107 ms away, TLS 1.2: one connection: 2.6 s2.6 s · 4.2× fasterServer 162 ms away, TLS 1.3Server 162 ms away, TLS 1.3: new connection each time: 12.4 s12.4 sServer 162 ms away, TLS 1.3: one connection: 5.8 s5.8 s · 2.1× fasterOpenRouterOpenRouter: new connection each time: 0.9 s0.9 sOpenRouter: one connection: 0.3 s0.3 s · 3.2× fasterFireworksFireworks: new connection each time: 1.8 s1.8 sFireworks: one connection: 0.6 s0.6 s · 3.1× fasterAnthropicAnthropic: new connection each time: 2.9 s2.9 sAnthropic: one connection: 2.5 s2.5 s · 1.1× fasterOpenAIOpenAI: new connection each time: 4.0 s4.0 sOpenAI: one connection: 3.7 s3.7 s · 1.1× faster
20 requests in a row, measured on 30 Sep 2026 from my Mac, over HTTPS. The API calls carry no key, so this is network time, not model time.

Two things changed since 2017. APIs sit behind a CDN a few ms away from you, and TLS 1.3 needs one round trip instead of two. So opening a connection to OpenAI or Anthropic costs about 20 ms from my desk.

When the server takes 100+ ms to answer, that's 7% to 26% of each request. When it's fast, the handshake is most of it: OpenRouter goes from 45 ms to 14 ms per request.

And a server far away, still on TLS 1.2? 4.2× faster with one connection. Same story as 2017.

What I do now

  • One HTTP client per thread, reused. Ruby: Net::HTTP.start with a block, net-http-persistent, HTTPX, or Excon with persistent: true. Python: requests.Session or httpx.Client.
  • Create the OpenAI or Anthropic client once. It holds the connection pool. A new client per request throws it away.
  • Log the connect time, not only the request time. That's where my 600 ms was hiding.
  • Test with production volume. 10 jobs will never show you this.