Make process-global caches thread-safe (GC corruption wedge) #7
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "fix/global-cache-thread-safety"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
This recreates and fixes the GC failure from the botcommons-store production incident (glm-53): under concurrent load the store wedged forever with one thread spinning at 100% CPU inside
S_gc_ocd_entry, and the preserved FreeBSD core dumps showednonrecoverable invalid memory referenceinside the same function — the collector's post-GC TLC hash-table bucket rehash walking a torn chain.Root cause: jsqlite kept 20+ process-global caches as plain Chez hash tables (plan/result/parse/tmeta/page-tracking/schema caches) mutated from every connection's thread with no synchronization. Chez hash tables tear under concurrent mutation. The corrupted table found in the core dump maps
parsedstatements toplanrecords — the select-plan cache. Each bot-farm reader thread owns its own connection, but all of them hammer these shared tables.Reproducer (now
tests/unit/test-global-cache-race): 8 threads, each with its own:memory:connection, run cacheable SELECTs + cache-invalidating UPDATEs under allocation churn and forced collections. On master it dies within seconds — observed:nonrecoverable invalid memory referenceaborts and the infinite rehash spin (the production wedge signature).Change
Two-tier fix, chosen after measuring both mechanisms:
(std concur hash)(the stdlib thread-safe HAMT): parse/plan/result/stmt/tmeta/row caches and process write-locks injsqlite api, expression programs injsqlite exec, type affinity injsqlite eval. Records compare by identity underequal?/equal-hash(stable across GC — verified empirically), so the original eq-keying semantics are preserved.with-global-cache-lock(new macro injsqlite cache, plain mutex; Chez deactivates a blocked thread so GC is never blocked): writer page-tracking tables, zero-page cache, schema text caches, the bytevector-identity image caches injsqlite page/jsqlite btree(a pmap hashes bytevectors by content, but images mutate in place), and the weak side tables injsqlite exec/jsqlite eval(strong keys would leak every scan/AST node).The split is load-bearing: an earlier draft used concurrent hashes everywhere and was 13x slower on indexed bulk inserts —
ensure-original-page-count!runs >200k times per 2k-row load and a HAMT write allocates a fresh path per call. The locked in-place tables cost nothing measurable (insert benchmark 1.58s before and after; test-write 17.3s vs 17.7s baseline).Also:
jsqlite dmltmeta-cache helpers dispatch on the parameter's table type (it is bound to either the global concurrent hash or a per-connection eq-hashtable).Known follow-up (not blocking): the cshim handle registry is also process-global and unsynchronized, but it is single-threaded test-only code today.
Verification
tests/unit/test-global-cache-race: crashes 3/3 on master (abort/spin), passes 23/23 runs with this change.make build, fullmake unit(31 files, incl. 1203-check test-write at baseline speed),make binary+--version/--helpsmoke: all green on macOS arm64.Fixes the jerboa-sqlite half of the botcommons glm-53 wedge. The jerboa repo's vendored jsqlite lock (
vendor-lock.env) should be bumped to this commit in a follow-up jerboa PR so release builds pick it up.