Make process-global caches thread-safe (GC corruption wedge) #7

Merged
ober merged 1 commit from fix/global-cache-thread-safety into master 2026-09-12 19:34:50 -04:00
Owner

Problem

This recreates and fixes the GC failure from the botcommons-store production incident (glm-53): under concurrent load the store wedged forever with one thread spinning at 100% CPU inside S_gc_ocd_entry, and the preserved FreeBSD core dumps showed nonrecoverable invalid memory reference inside the same function — the collector's post-GC TLC hash-table bucket rehash walking a torn chain.

Root cause: jsqlite kept 20+ process-global caches as plain Chez hash tables (plan/result/parse/tmeta/page-tracking/schema caches) mutated from every connection's thread with no synchronization. Chez hash tables tear under concurrent mutation. The corrupted table found in the core dump maps parsed statements to plan records — the select-plan cache. Each bot-farm reader thread owns its own connection, but all of them hammer these shared tables.

Reproducer (now tests/unit/test-global-cache-race): 8 threads, each with its own :memory: connection, run cacheable SELECTs + cache-invalidating UPDATEs under allocation churn and forced collections. On master it dies within seconds — observed: nonrecoverable invalid memory reference aborts and the infinite rehash spin (the production wedge signature).

Change

Two-tier fix, chosen after measuring both mechanisms:

  • Cold/coarse caches → (std concur hash) (the stdlib thread-safe HAMT): parse/plan/result/stmt/tmeta/row caches and process write-locks in jsqlite api, expression programs in jsqlite exec, type affinity in jsqlite eval. Records compare by identity under equal?/equal-hash (stable across GC — verified empirically), so the original eq-keying semantics are preserved.
  • Hot/identity-keyed tables → in-place Chez tables + with-global-cache-lock (new macro in jsqlite cache, plain mutex; Chez deactivates a blocked thread so GC is never blocked): writer page-tracking tables, zero-page cache, schema text caches, the bytevector-identity image caches in jsqlite page/jsqlite btree (a pmap hashes bytevectors by content, but images mutate in place), and the weak side tables in jsqlite exec/jsqlite eval (strong keys would leak every scan/AST node).

The split is load-bearing: an earlier draft used concurrent hashes everywhere and was 13x slower on indexed bulk inserts — ensure-original-page-count! runs >200k times per 2k-row load and a HAMT write allocates a fresh path per call. The locked in-place tables cost nothing measurable (insert benchmark 1.58s before and after; test-write 17.3s vs 17.7s baseline).

Also: jsqlite dml tmeta-cache helpers dispatch on the parameter's table type (it is bound to either the global concurrent hash or a per-connection eq-hashtable).

Known follow-up (not blocking): the cshim handle registry is also process-global and unsynchronized, but it is single-threaded test-only code today.

Verification

  • New regression tests/unit/test-global-cache-race: crashes 3/3 on master (abort/spin), passes 23/23 runs with this change.
  • make build, full make unit (31 files, incl. 1203-check test-write at baseline speed), make binary + --version/--help smoke: all green on macOS arm64.
  • Performance guard: 2000-row indexed insert benchmark and the whole write suite are at pre-change speed.

Fixes the jerboa-sqlite half of the botcommons glm-53 wedge. The jerboa repo's vendored jsqlite lock (vendor-lock.env) should be bumped to this commit in a follow-up jerboa PR so release builds pick it up.

## Problem This recreates and fixes the GC failure from the botcommons-store production incident (glm-53): under concurrent load the store wedged forever with one thread spinning at 100% CPU inside `S_gc_ocd_entry`, and the preserved FreeBSD core dumps showed `nonrecoverable invalid memory reference` inside the same function — the collector's post-GC TLC **hash-table bucket rehash** walking a torn chain. Root cause: jsqlite kept **20+ process-global caches as plain Chez hash tables** (plan/result/parse/tmeta/page-tracking/schema caches) mutated from every connection's thread with no synchronization. Chez hash tables tear under concurrent mutation. The corrupted table found in the core dump maps `parsed` statements to `plan` records — the select-plan cache. Each bot-farm reader thread owns its own connection, but all of them hammer these shared tables. Reproducer (now `tests/unit/test-global-cache-race`): 8 threads, each with its own `:memory:` connection, run cacheable SELECTs + cache-invalidating UPDATEs under allocation churn and forced collections. On master it dies within seconds — observed: `nonrecoverable invalid memory reference` aborts and the infinite rehash spin (the production wedge signature). ## Change Two-tier fix, chosen after measuring both mechanisms: - **Cold/coarse caches → `(std concur hash)`** (the stdlib thread-safe HAMT): parse/plan/result/stmt/tmeta/row caches and process write-locks in `jsqlite api`, expression programs in `jsqlite exec`, type affinity in `jsqlite eval`. Records compare by identity under `equal?`/`equal-hash` (stable across GC — verified empirically), so the original eq-keying semantics are preserved. - **Hot/identity-keyed tables → in-place Chez tables + `with-global-cache-lock`** (new macro in `jsqlite cache`, plain mutex; Chez deactivates a blocked thread so GC is never blocked): writer page-tracking tables, zero-page cache, schema text caches, the bytevector-identity image caches in `jsqlite page`/`jsqlite btree` (a pmap hashes bytevectors by content, but images mutate in place), and the weak side tables in `jsqlite exec`/`jsqlite eval` (strong keys would leak every scan/AST node). The split is load-bearing: an earlier draft used concurrent hashes everywhere and was **13x slower** on indexed bulk inserts — `ensure-original-page-count!` runs >200k times per 2k-row load and a HAMT write allocates a fresh path per call. The locked in-place tables cost nothing measurable (insert benchmark 1.58s before and after; test-write 17.3s vs 17.7s baseline). Also: `jsqlite dml` tmeta-cache helpers dispatch on the parameter's table type (it is bound to either the global concurrent hash or a per-connection eq-hashtable). Known follow-up (not blocking): the cshim handle registry is also process-global and unsynchronized, but it is single-threaded test-only code today. ## Verification - New regression `tests/unit/test-global-cache-race`: **crashes 3/3 on master** (abort/spin), **passes 23/23 runs** with this change. - `make build`, full `make unit` (31 files, incl. 1203-check test-write at baseline speed), `make binary` + `--version`/`--help` smoke: all green on macOS arm64. - Performance guard: 2000-row indexed insert benchmark and the whole write suite are at pre-change speed. Fixes the jerboa-sqlite half of the botcommons glm-53 wedge. The jerboa repo's vendored jsqlite lock (`vendor-lock.env`) should be bumped to this commit in a follow-up jerboa PR so release builds pick it up.
fix: make process-global caches thread-safe
Some checks failed
version-policy / required (pull_request) Failing after 3m51s
required-ci / required (pull_request) Successful in 9m0s
affd720274
Concurrent use of separate jsqlite connections corrupts Chez hash
tables: the collector's post-GC TLC rehash then walks a torn bucket
chain, producing 'nonrecoverable invalid memory reference' aborts or
an infinite rehash spin (the botcommons-store production wedge,
reproduced locally in tests/unit/test-global-cache-race).

- cold process-global caches (parse/plan/result/tmeta/row-cache/write
  locks in api, expr programs in exec, affinity in eval): converted
  to (std concur hash); Chez records compare by identity under
  equal?, so eq-keying is preserved
- hot per-page tables (writer page tracking, zero pages, schema text
  caches, bytevector-identity image caches in page/btree, weak side
  tables in exec/eval): kept as in-place Chez tables guarded by a
  new with-global-cache-lock (jsqlite cache); a concurrent hash
  allocates a HAMT path per write and is pathologically slower in
  the b-tree write loop (measured 13x on indexed bulk inserts)
- dml tmeta-cache helpers dispatch on the parameter's table type
- regression test: 8 threads x own connections hammer the caches
  under GC pressure; crashes within seconds on master, passes 15/15
  with this change
ober scheduled this pull request to auto merge when all checks succeed 2026-09-12 19:33:59 -04:00
ober merged commit 435cfd1f2b into master 2026-09-12 19:34:50 -04:00
ober referenced this pull request from a commit 2026-09-12 19:34:52 -04:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
ober/jerboa-sqlite!7
No description provided.