Fix manifest store hang: range-chunk flat shard buckets, parallel shard stores #37

Merged
ober merged 1 commit from fix/manifest-shard-explosion into main 2026-09-03 13:01:56 -04:00
Owner

Manifest store hang: range-chunk flat shard buckets, parallel shard stores

Problem

jd sync on a 172K-file tree (15GiB) finished all uploads, then spun in
"checkpointing manifest" for 20+ hours with zero progress, zero log output
even under -d, and 100% CPU/persistent network activity.

Root cause

manifest-expand-shard-group split oversized buckets by descending one more
path component without checking whether that component is the filename
itself
. Any flat directory with more files than the entry limit (old
default 4096) exploded into one shard per file:

  • synthetic repro of the affected tree: 172,285 entries -> 50,296 shards
    (largest 808 entries)
  • each sealed shard is a separate S3 PUT, issued sequentially by
    store-manifest-shards! with no logging, so the final manifest store
    was ~50K silent uploads — hours of work misread as a hang

Fix

  • manifest-expand-shard-group stops descending when no entry has a deeper
    path component, and range-chunks the oversized group into deterministic
    "<group-id>#k" ranges of the entry limit (sorted by path, stable ids)
  • manifest-shard-assignment emits groups in ascending shard-id order
  • store-manifest-shards! stores shards on the parallel worker pool
    (manifest-shard-fetch-width, 8) and logs the assignment plus each
    stored/unchanged shard under -d
  • default JDRIVE_MANIFEST_SHARD_ENTRIES 4096 -> 8192, matching the
    documented default; user guide documents the range-chunking

Verification

  • synthetic 172K-entry probe: 50,296 -> 305 shards (largest 5,556),
    assignment in 0.4s
  • new tests: range-chunking of flat oversized directories, component
    splitting of oversized buckets
  • make test green (live S3 test skipped)
  • version bump 1.9.1 -> 1.9.2

Upgrade note

Drives stored by an affected run (if any mid-run root landed) self-heal: the
next successful store writes the new assignment and gc-manifest-shards!
deletes the old per-file shards. Shards are content-addressed, so unchanged
shards from this fix still skip upload.

# Manifest store hang: range-chunk flat shard buckets, parallel shard stores ## Problem `jd sync` on a 172K-file tree (15GiB) finished all uploads, then spun in "checkpointing manifest" for 20+ hours with zero progress, zero log output even under `-d`, and 100% CPU/persistent network activity. ## Root cause `manifest-expand-shard-group` split oversized buckets by descending one more path component **without checking whether that component is the filename itself**. Any flat directory with more files than the entry limit (old default 4096) exploded into **one shard per file**: - synthetic repro of the affected tree: 172,285 entries -> **50,296 shards** (largest 808 entries) - each sealed shard is a separate S3 PUT, issued **sequentially** by `store-manifest-shards!` with **no logging**, so the final manifest store was ~50K silent uploads — hours of work misread as a hang ## Fix - `manifest-expand-shard-group` stops descending when no entry has a deeper path component, and **range-chunks** the oversized group into deterministic `"<group-id>#k"` ranges of the entry limit (sorted by path, stable ids) - `manifest-shard-assignment` emits groups in ascending shard-id order - `store-manifest-shards!` stores shards on the parallel worker pool (`manifest-shard-fetch-width`, 8) and logs the assignment plus each stored/unchanged shard under `-d` - default `JDRIVE_MANIFEST_SHARD_ENTRIES` 4096 -> 8192, matching the documented default; user guide documents the range-chunking ## Verification - synthetic 172K-entry probe: **50,296 -> 305 shards** (largest 5,556), assignment in 0.4s - new tests: range-chunking of flat oversized directories, component splitting of oversized buckets - `make test` green (live S3 test skipped) - version bump 1.9.1 -> 1.9.2 ## Upgrade note Drives stored by an affected run (if any mid-run root landed) self-heal: the next successful store writes the new assignment and `gc-manifest-shards!` deletes the old per-file shards. Shards are content-addressed, so unchanged shards from this fix still skip upload.
Fix manifest store hang: range-chunk flat shard buckets, parallel shard stores
Some checks failed
version-policy / required (pull_request) Failing after 3m48s
required-ci / required (pull_request) Successful in 4m36s
162c97a273
A flat directory with more files than the shard entry limit was split by
the filename component, producing one shard per file: a 172K-entry tree
yielded ~50K sealed shard objects, each uploaded sequentially with no
logging, which looked like an infinite 'checkpointing manifest' hang.

- manifest-expand-shard-group now stops descending when no entry has a
  deeper path component and range-chunks the oversized group into
  deterministic '#k' ranges of the entry limit instead
- manifest-shard-assignment appends groups in ascending shard-id order
- store-manifest-shards! stores shards on the parallel worker pool and
  logs each stored/unchanged shard under -d
- default JDRIVE_MANIFEST_SHARD_ENTRIES 4096 -> 8192, matching the
  documented default; docs note the range-chunking
- synthetic 172K-entry probe: 50296 -> 305 shards (max 5556), 0.4s

Version 1.9.2.
ober force-pushed fix/manifest-shard-explosion from 162c97a273
Some checks failed
version-policy / required (pull_request) Failing after 3m48s
required-ci / required (pull_request) Successful in 4m36s
to 8b1dcac3f1
Some checks failed
version-policy / required (pull_request) Failing after 3m46s
required-ci / required (pull_request) Successful in 4m34s
2026-09-03 12:48:35 -04:00
Compare
ober force-pushed fix/manifest-shard-explosion from 8b1dcac3f1
Some checks failed
version-policy / required (pull_request) Failing after 3m46s
required-ci / required (pull_request) Successful in 4m34s
to b5a5ff044b
All checks were successful
version-policy / required (pull_request) Successful in 3m47s
required-ci / required (pull_request) Successful in 4m33s
2026-09-03 12:53:32 -04:00
Compare
ober scheduled this pull request to auto merge when all checks succeed 2026-09-03 13:01:42 -04:00
ober merged commit 9efa05d3d4 into main 2026-09-03 13:01:56 -04:00
ober referenced this pull request from a commit 2026-09-03 13:01:57 -04:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
ober/jerboa-drive!37
No description provided.