Retry fresh TLS failures before sync backoff #61

Merged
ober merged 1 commit from fix/tls-write-under-load-2.0.16 into main 2026-09-19 22:20:56 -04:00
Owner

Summary

  • retry one failed HTTPS exchange on a newly opened connection even when the failed connection was already fresh
  • report the native Rustls error for TLS write, flush, and read failures
  • add regression coverage and document the bounded transport retry
  • bump the synchronized version to 2.0.16

Cause

The prior recovery path only made an immediate fresh attempt when the failed exchange had used a pooled handle. Under a 16-worker upload, a transient failure on an already-fresh connection escaped directly to the sync layer, producing the repeated exponential upload-file-object! backoff seen by the user. The corrected path gives either source of connection exactly one fresh attempt; a failure of that second attempt still propagates to the existing bounded sync retry.

Verification

  • make test — all tests passed, including pooled and already-fresh retry coverage
  • make binary — passed on macOS arm64
  • make binary-smoke — installed-layout smoke passed
  • changed-line Jerboa security scan — no findings
  • 16-worker S3 multipart stress: 4 files, 5,015,634,768 bytes, 300 multipart parts — completed with no TLS retry or failure
  • 16-worker S3 connection-churn stress: 128 files/requests, 134,217,728 bytes — completed with no TLS retry or failure using the rebuilt 2.0.16 binary
  • both diagnostic remote prefixes were removed and verified empty
## Summary - retry one failed HTTPS exchange on a newly opened connection even when the failed connection was already fresh - report the native Rustls error for TLS write, flush, and read failures - add regression coverage and document the bounded transport retry - bump the synchronized version to 2.0.16 ## Cause The prior recovery path only made an immediate fresh attempt when the failed exchange had used a pooled handle. Under a 16-worker upload, a transient failure on an already-fresh connection escaped directly to the sync layer, producing the repeated exponential `upload-file-object!` backoff seen by the user. The corrected path gives either source of connection exactly one fresh attempt; a failure of that second attempt still propagates to the existing bounded sync retry. ## Verification - `make test` — all tests passed, including pooled and already-fresh retry coverage - `make binary` — passed on macOS arm64 - `make binary-smoke` — installed-layout smoke passed - changed-line Jerboa security scan — no findings - 16-worker S3 multipart stress: 4 files, 5,015,634,768 bytes, 300 multipart parts — completed with no TLS retry or failure - 16-worker S3 connection-churn stress: 128 files/requests, 134,217,728 bytes — completed with no TLS retry or failure using the rebuilt 2.0.16 binary - both diagnostic remote prefixes were removed and verified empty
Retry fresh TLS failures before sync backoff
All checks were successful
version-policy / required (pull_request) Successful in 3m43s
required-ci / required (pull_request) Successful in 4m37s
975f3bac5c
ober merged commit 30f91fe85e into main 2026-09-19 22:20:56 -04:00
ober referenced this pull request from a commit 2026-09-19 22:20:59 -04:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
ober/jerboa-drive!61
No description provided.