No description
  • Scheme 75.5%
  • Common Lisp 23.7%
  • Makefile 0.4%
  • Shell 0.3%
Find a file
ober f9e705f7d0
All checks were successful
required-ci / required (push) Successful in 7m5s
Merge pull request #54
2026-09-19 18:43:44 -04:00
.forgejo Bootstrap pinned Jerboa toolchain in CI 2026-09-19 15:08:52 -06:00
benchmark feat: fair multi-mode benchmark harness and fixed-slot execution metrics 2026-09-14 16:31:34 -06:00
cmd Stack CSV scanner fixes after statistical aggregates 2026-09-19 15:01:15 -06:00
contracts Stack statistical aggregates after MAP support 2026-09-19 14:55:39 -06:00
data Stack parallel set operations after decimal arithmetic 2026-09-19 14:43:04 -06:00
docs Stack parallel set operations after decimal arithmetic 2026-09-19 14:43:04 -06:00
lib/jerboa Merge remote-tracking branch 'origin/feat/statistical-aggregates' into feat/csv-scanner-opts 2026-09-19 16:19:13 -06:00
support Bootstrap pinned Jerboa toolchain in CI 2026-09-19 15:08:52 -06:00
test Merge remote-tracking branch 'origin/feat/statistical-aggregates' into feat/csv-scanner-opts 2026-09-19 16:33:03 -06:00
.gitignore Bootstrap pinned Jerboa toolchain in CI 2026-09-19 15:08:52 -06:00
.gitsafe.json Initialize Jerboa DuckDB Phase 0 contracts 2026-07-19 22:04:39 -06:00
.gitsafeignore Set up Forgejo CI/CD policy 2026-08-03 12:55:30 -06:00
AGENTS.md docs(agents): add checkout hygiene policy (work dirs under ~/work, cleanup when done) 2026-09-12 19:10:46 -06:00
astra-complete.md Document complete DuckDB parity plan and repair audit verification gate 2026-09-07 12:39:51 -06:00
glm-53-handoff.md Fix unit tests for jerboa 0.8.15 taint/prelude bugs; refresh handoff 2026-08-25 11:44:11 -06:00
jpkg.sexp Stack CSV scanner fixes after statistical aggregates 2026-09-19 15:01:15 -06:00
Makefile Build complete FreeBSD standalone binaries 2026-09-19 15:45:47 -06:00
Project.md feat: add DuckDB JSON aggregate parity 2026-08-26 17:23:06 -06:00
README.md feat: add typed hash join build and probe 2026-09-14 23:41:49 -06:00
VERSION Stack CSV scanner fixes after statistical aggregates 2026-09-19 15:01:15 -06:00

jerboa-duckdb

jerboa-duckdb is an independent DuckDB-compatible analytical database being implemented in Jerboa Scheme. DuckDB is used as a pinned development oracle and is not linked into the production engine.

The governing scope and delivery gates are in Project.md. The current code implements a partial Phase 1 in-memory scalar slice: catalog tables, columnar chunks, tokenizer/parser, binder, vectorized expression kernels, scan/filter/project/aggregate/ distinct/order/limit execution, GROUP BY validation with post-aggregate, aggregate DISTINCT arguments, aggregate recursive LIST, fixed ARRAY, MAP, STRUCT, and UNION SQL types with constructors, subscripts/field extraction, typed SHREDDED child vectors, grouping/order/distinct, DML/CTAS integration, and persistent reopen, deterministic logical filter pushdown, adjacent projection composition/identity removal, and basic cardinality-guided INNER/CROSS join side selection with conservative window/subquery/lateral/outer-join barriers, ORDER BY expressions, first-slice aggregate FILTER (WHERE predicate) for count/sum/avg/min/max, count()/count(*), NULL-literal and empty-set aggregate semantics, aggregate input validation, basic EXPLAIN result rendering for supported query plans, SQL-level named PREPARE/EXECUTE for the current statement subset, connection-local SET/RESET for scheduler, memory, insertion-order, immediate-transaction, and transaction-invalidation settings plus typed current_setting(...), explicit BEGIN/COMMIT/ROLLBACK with optimistic snapshot isolation, transaction-local inserts, updates, deletes, indexes, and catalog data, commit-conflict detection, and invalidation, typed commit-version history and per-object scan-conflict watermarks, schema-aware physical row versions with update/delete chains and snapshot-safe GC, statement before-images for failure atomicity plus explicit undo-record state transitions, serial commit validation for primary, unique, foreign-key, check, and not-null constraints, atomic multi-table catalog/data publication with pre-publication failure-path coverage, first-slice ANALYZE/VACUUM statement handling with optional table validation, first-slice PRAGMA table_info, PRAGMA version, and PRAGMA database_list metadata results, first-slice metadata table-function sources for pragma_version(), pragma_table_info(...), and a two-setting duckdb_settings() subset, numeric BIGINT range(...) and generate_series(...) table-function sources, first-slice repeat(value, count) table-function sources, first-slice glob(pattern) table-function sources, first-slice SUMMARIZE table/query metadata plans, first-slice duckdb_databases()/duckdb_schemas()/duckdb_views()/ duckdb_indexes()/duckdb_constraints()/duckdb_types()/ duckdb_sequences()/duckdb_dependencies()/duckdb_functions()/ duckdb_extensions()/duckdb_optimizers()/duckdb_memory()/duckdb_logs()/ duckdb_prepared_statements()/duckdb_variables()/duckdb_secrets()/ duckdb_temporary_files()/duckdb_external_file_cache()/duckdb_secret_types()/ duckdb_approx_database_count()/duckdb_connection_count()/ duckdb_coordinate_systems()/duckdb_keywords()/duckdb_log_contexts()/ duckdb_table_sample(...)/ duckdb_tables()/duckdb_columns() system catalog table functions, first-slice SHOW TABLES and SHOW ALL TABLES catalog metadata results for tables and views, first-slice DESCRIBE table, DESCRIBE SELECT/VALUES ..., and DESCRIBE table-source metadata results, first-slice INNER JOIN ... ON ..., LEFT [OUTER] JOIN ... ON ..., RIGHT [OUTER] JOIN ... ON ..., FULL [OUTER] JOIN ... ON ..., SEMI JOIN ... ON ..., and ANTI JOIN ... ON ... execution plus typed same-storage equi-key hash joins with residual predicates and first-slice JOIN ... USING(...), NATURAL inner/outer/semi/anti joins, POSITIONAL JOIN, and ASOF inner/left/right/full joins, materialized global window execution with named windows, independent orderings, ROWS/numeric and row-dependent numeric or temporal-interval RANGE/GROUPS frames, ranking/distribution functions, lag()/lead() with signed row-dependent offsets, first_value()/last_value() and nth_value() with row-dependent indexes, ntile() with row-dependent bucket counts, numeric and DATE/TIME/TIMESTAMP/fixed-offset TIMESTAMPTZ fill() interpolation/extrapolation, window aggregates, and bounded QUALIFY filtering over projected windows, including right/full coalesced key output, CROSS JOIN, and comma-style cross joins with qualified column binding, implicit SELECT aliases, table aliases and qualified column references, schema-qualified table names and CREATE [TEMP] TABLE [IF NOT EXISTS] for the in-memory main/temp catalog, first-slice CREATE [TEMP] TABLE [IF NOT EXISTS] name [(aliases...)] AS SELECT/VALUES ..., first-slice CREATE [OR REPLACE] [TEMP] VIEW [IF NOT EXISTS] ... AS SELECT/VALUES ... with rebound view scans, DESCRIBE view, and DROP VIEW [IF EXISTS], DuckDB identifier behavior for keyword-like column names such as unknown, DROP TABLE [IF EXISTS], basic DELETE FROM ... [WHERE ...], basic UPDATE ... SET ... [WHERE ...], first-slice ALTER TABLE name RENAME TO new_name, first-slice ALTER TABLE name RENAME [COLUMN] old_column TO new_column, first-slice ALTER TABLE name ADD [COLUMN] [IF NOT EXISTS] column type [DEFAULT expr], first-slice ALTER TABLE name DROP [COLUMN] [IF EXISTS] column, first-slice ALTER TABLE name ALTER [COLUMN] column SET/DROP DEFAULT, first-slice ALTER TABLE name ALTER [COLUMN] column TYPE/SET DATA TYPE type [USING expr], first-slice ALTER TABLE name ALTER [COLUMN] column SET/DROP NOT NULL, INSERT with full-row, column-list, DEFAULT row items, or DEFAULT VALUES sources plus DEFAULT/NOT NULL column handling, expression-bearing ON CONFLICT [(columns...)] DO NOTHING and bounded DO UPDATE SET ... [WHERE ...] handling with excluded values, atomic unique-index validation, and RETURNING, VALUES statements and VALUES table sources, derived SELECT and DESCRIBE table sources, first-slice uncorrelated scalar subqueries, first-slice uncorrelated EXISTS/NOT EXISTS and IN (SELECT/VALUES/WITH ...) subqueries, plus bounded correlated EXISTS/NOT EXISTS, scalar, and IN/NOT IN subsets with correlation stacks spanning nested query levels; explicit comma-style and inner/left LATERAL derived-table joins are also implemented. Aggregate outer-shape limitations, implicit lateralization, right/full lateral joins, and broader correlated subquery forms are not included, first-slice non-recursive WITH CTEs over SELECT/VALUES queries and INSERT ... SELECT sources, including AS [NOT] MATERIALIZED modifiers, basic UNION/EXCEPT/INTERSECT set operations including ALL over SELECT and VALUES with final ORDER BY/LIMIT, DuckDB-style literal coercion and typeof including DuckDB's quoted NULL type name in the covered scalar subset, parameterized DECIMAL type metadata plus DECIMAL casts/rendering, DuckDB SQL type-name aliases in DDL for unsigned integer spellings, bare DEC/DECIMAL/NUMERIC, TIMETZ/TIME WITH TIME ZONE, JSON, and GEOMETRY, DATE/TIME/TIMESTAMP casts/rendering, fixed-offset TIMETZ casts with offset-preserving DuckDB ordering, fixed-offset TIMESTAMPTZ casts normalized to UTC +00, INTERVAL casts/rendering, DATE-to-TIMESTAMP coercion, direct temporal extraction aliases (year, month, day, dayofmonth, quarter, century, decade, millennium, era, dayofweek, weekday, dayofyear, week, weekofyear, isoyear, isodow, yearweek, hour, minute, second), generic date_part/datepart plus extract(... FROM ...) temporal extraction, date_diff/datediff plus date_sub/datesub temporal difference functions, date_trunc/datetrunc, last_day, scalar make_date/make_time, julian, epoch, and epoch_ms/epoch_us/epoch_ns, IN list comparison literal coercion, LIKE/ILIKE including ESCAPE and DuckDB ~~/!~~/~~*/!~~* operator aliases, BLOB casts/rendering, boolean/numeric comparison coercion, typed string equality coercion against numeric/boolean operands, low-precedence JSON extraction ->/->> with chained bare-key, JSONPath/JSONPointer, constant JSONPath object/array wildcards and [#-N] back indexes, integer and negative-index paths, VARCHAR path columns, LIST-of-path extraction, json_extract_path aliases, JSON/VARCHAR result typing, and parenthesized JSON-to-numeric equality, JSON '...' prefix literals, chained explicit JSON casts, strict JSON validation, and recursive JSON conversion for scalar, temporal, LIST/ARRAY, MAP, STRUCT, and UNION values, json_structure inference with ordered duplicate-aware object merging, and bind-time typed json_transform/from_json plus strict aliases for scalar, LIST, ARRAY, MAP, and STRUCT targets, ordered json_group_array, duplicate-preserving json_group_object, and cross-row json_group_structure aggregation, mixed numeric ///divide and %/mod zero-divisor semantics, DuckDB scalar abs including DECIMAL preservation and signed-min overflow checks, unary DECIMAL negation preserving width/scale, DuckDB half-away-from-zero scalar round including scale arguments, truncation toward zero with trunc, numeric sign, signbit, isfinite, isinf, and isnan, pi, radians/degrees, trigonometric and hyperbolic scalars, pow/power, factorial, even, gcd/lcm, bit_count, integer xor, cbrt, exp, ln, log2, log10, log, gamma, lgamma, and nextafter, scalar function arity/type checks and NULL overload defaults, version, current_schema, current_database, ceiling, ucase/lcase, and substr alias support, format_bytes/formatReadableDecimalSize, to_base, bin/to_binary, hex/to_hex, unhex/unbin, md5, sha1, sha256, concat_ws, length aliases plus strlen/string bit_length, substring including FROM/FOR syntax, ascii/unicode/ord/chr, hamming/mismatches, levenshtein/editdist3, damerau_levenshtein, jaccard, left/right/repeat/lpad/rpad/replace/translate/reverse, url_encode/url_decode, position(... IN ...), instr/strpos, string predicates contains/prefix/suffix plus starts_with/ends_with, trim/ltrim/rtrim including standard FROM syntax, constant_or_null, can_cast_implicitly, scalar error, and mixed numeric coalesce/ifnull/greatest/least, lambda x, i: ... LIST expressions plus deprecated x -> ... and (x, i) -> ... forms with DuckDB-compatible global lambda_syntax DEFAULT/ENABLE_SINGLE_ARROW/DISABLE_SINGLE_ARROW enforcement, transform/filter/reduce aliases, 1-based indexes, captured row values, nested captures, and reduce initial values, nullif comparison coercion including typed string equality, CASE string-literal result normalization, ternary if as CASE-compatible sugar, unary integer typing, boolean string cast aliases, integer cast range checks, and BIGINT overflow checks, wide signed/unsigned integer boundary storage, unsigned 128-bit literal typing, and arithmetic overflow checks, assignment casts for inserted values, and a small public API with cached prepared plans for unparameterized SELECT/VALUES. Prepared parameter substitution covers scalar expressions, VALUES table sources, INSERT ... VALUES, INSERT ... SELECT, DELETE ... WHERE, and UPDATE ... SET/WHERE, plus SUMMARIZE SELECT/VALUES payloads, including mixed ? and $n placeholders, with count validation for missing or extra arguments. LIMIT/OFFSET, including standalone OFFSET, support DuckDB constant coercion for numeric, string, boolean, NULL, LIMIT ALL, and prepared-parameter inputs in the current subset. SQL PREPARE name AS ... and EXECUTE name(...) are connection-local and reuse the same prepared execution path as the public Scheme API. Each execution now receives a monotonically increasing connection-local query ID. duckdb-query-progress reports its running/completed/failed state, emitted chunk and row counts, timestamps, and failure message; duckdb-current-query-id reports an active query. duckdb-set-query-timeout! installs a per-connection deadline, and duckdb-interrupt! cooperatively cancels at pull and join-loop checkpoints. Admission starts before tokenization and also checks recursive parsing and optimization, scans, lateral/ASOF and unmatched-side joins, aggregates, sorting, CSV/storage I/O boundaries, and vector kernels. Connections reject overlapping top-level executions, and duckdb-close is idempotent and prevents further prepare/query operations. Across connections, database read owners may coexist while autocommit writers remain exclusive. Explicit transactions use private snapshots, expose only committed state to other connections, and detect concurrent-write conflicts at commit. Closing a connection rejects new work, cancels and drains active synchronous or pending queries, rolls back an open transaction, releases database ownership, and joins its workers before returning. duckdb-pending-query reserves a query on a bounded connection-local CPU pool. Each duckdb-pending-execute-task call performs one prepare or chunk-pull slice; callers can poll DuckDB-style pending states, wait on the terminal notifier, cancel execution, or materialize and cache the final result with duckdb-execute-pending. duckdb-execute-tasks lets a caller thread drain queued external work. CPU and blocking-I/O pools are distinct, honor threads and external_threads, and task groups preserve the first worker failure while cancelling and draining siblings. SQL SET name=value and RESET name currently cover connection-local threads, external_threads, memory_limit, and preserve_insertion_order; current_setting('threads') returns BIGINT, current_setting('memory_limit') returns VARCHAR, and current_setting('preserve_insertion_order') and current_setting('immediate_transaction_mode') return BOOLEAN. default_transaction_invalidation_policy selects all-error invalidation or DuckDB's syntactic-error exemption; an active transaction can override it with current_transaction_invalidation_policy. JOIN ... USING(...) exposes DuckDB-style visible key output and keeps qualified references to the original left/right key columns available. SQL BEGIN/COMMIT/ROLLBACK cover one active transaction per connection. Transactions read a stable private snapshot, see their own catalog/data changes, publish them atomically on a revision-checked commit, and become invalid after an execution or commit error until ROLLBACK. This is whole-database optimistic snapshot isolation. Committed table mutations also retain schema-aware physical row IDs, begin/end commit visibility, update/delete chain boundaries, and active-snapshot-aware version garbage collection. Query execution still reads the transaction's private catalog snapshot; write-set merging, attached-database meta-transactions, and row-chain-driven DuckDB MVCC execution remain later work. SQL ANALYZE [table] and VACUUM [table] currently validate optional target tables and return an empty result; file-backed database snapshots can be reopened through duckdb-open/duckdb-close using a checksummed 4 KiB binary header with DuckDB's DUCK magic at byte 8, an explicit Jerboa format marker, and a versioned payload envelope. Legacy v1/v2 text and v3 binary snapshots remain readable. New v4 writes encode table data into validated 2,048-row groups with per-column segments that deterministically select constant, RLE, dictionary, integer bitpacking, all-valid payload elision, or uncompressed encoding and persist exact total/NULL/non-NULL counts plus conservative min/max zone maps, then publish duplicate recoverable metadata headers followed by checksummed 256 KiB blocks at DuckDB's byte-12,288 block start, with atomic replacement, highest-valid-generation temporary recovery, and synchronized stale-handle write rejection. Column definitions accept USING COMPRESSION; forcing PFOR matches DuckDB through V2_0_0 by returning its not-available-yet binder error. The storage-version contract inventories all 67 pinned StorageVersion enum entries and 65 accepted names, while explicitly recording native DuckDB reads and writes as unsupported until bidirectional file-format fixtures exist. Index snapshots include deterministic versioned radix key-to-row entry lists; the live ART compresses shared byte prefixes and grows child storage through 4, 16, 48, and 256-way representations. Reopen rebuilds entries from table rows and rejects mismatches, while INSERT, UPDATE, DELETE, and TRUNCATE keep live index materializations synchronized. Plain INSERT and CSV COPY tail-appends stage radix-entry updates without rescanning table rows; stable-row UPDATEs stage old-key deletion followed by new-key insertion. ON CONFLICT combines that stable-prefix update with appended-tail insertion in the same publication. DELETE remaps stored row ordinals through the ordered kept-row subsequence, and TRUNCATE replaces each tree with an empty index. VACUUM [table] deterministically rebuilds either the target table's indexes or every index in the database. The optimizer uses complete indexes for conjunctive column = literal filters. It also traverses validated one-column scalar ART entries for <, <=, >, and >= predicates over BOOLEAN, NUMBER, and VARCHAR values; reversed literal comparisons are normalized before lookup. Composite indexes support the same range operators when direct equalities cover every index column before the ranged column; the optimizer prefers the longest eligible equality prefix. BOOLEAN, NUMBER, and VARCHAR indexes use an invalidation-aware ordered cache plus bounded compressed-prefix ART seeks. Finite numeric keys preserve exact integer and IEEE-double order through canonical binary magnitude encoding; infinities have explicit endpoints, while NaN predicates remain table scans. The original filter remains above every index scan as a correctness guard, and row identifiers are range- and uniqueness-checked before indexed rows are materialized. Saves emit typed, checksummed state-delta WAL records bound to base and target generations. The committed WAL is fsynced before the temporary image, the image is fsynced before rename, and the WAL is removed only after publication; complete transactions replay idempotently while torn tails fall back to the last valid snapshot. Legacy full-snapshot journals remain recoverable. SQL CHECKPOINT and FORCE CHECKPOINT publish file-backed committed snapshots, including from active transactions without leaking their private changes; configurable relaxed fsync modes, row-chain-driven execution, native DuckDB catalog/data blocks, and full native DuckDB storage-file compatibility remain later phases. Recovery and publication are serialized across processes with a persistent POSIX advisory-lock sidecar; contention fails without touching the database generation. SQL PRAGMA table_info, PRAGMA version, and PRAGMA database_list expose the first in-memory metadata subset; CALL pragma_version(), CALL pragma_table_info(...), and a two-setting CALL duckdb_settings() slice are implemented as the first table-function-style CALL metadata coverage. The same pragma_version(), pragma_table_info(...), and duckdb_settings() metadata slice can also be used as table functions in FROM. Numeric BIGINT range(...) and generate_series(...) table-function sources cover DuckDB's one/two/three argument forms, exclusive versus inclusive stop semantics, negative steps, NULL-empty results, and prepared arguments. repeat(value, count) table-function sources cover constant scalar value expressions, typed value output, zero-row counts, NULL values, prepared counts, and DuckDB-shaped count errors. glob(pattern) table-function sources cover non-recursive regular-file patterns with */?, empty matches, and prepared string patterns. SUMMARIZE table/query returns DuckDB-shaped column summary metadata for the current in-memory table/query subset, including direct statement and parenthesized table-source forms. duckdb_databases(), duckdb_schemas(), duckdb_views(), duckdb_indexes(), duckdb_constraints(), duckdb_types(), duckdb_sequences(), duckdb_dependencies(), duckdb_functions(), duckdb_extensions(), duckdb_optimizers(), duckdb_memory(), duckdb_logs(), duckdb_prepared_statements(), duckdb_variables(), duckdb_secrets(), duckdb_temporary_files(), duckdb_external_file_cache(), duckdb_secret_types(), duckdb_approx_database_count(), duckdb_connection_count(), duckdb_coordinate_systems(), duckdb_keywords(), duckdb_log_contexts(), duckdb_table_sample(...), duckdb_tables(), and duckdb_columns() expose the current in-memory catalog as first-slice system catalog table functions, including direct CALL and FROM forms for stable database/schema/view/index/constraint/type/sequence/dependency/function/ extension/optimizer/memory/log/prepared/variable/secret/temp/cache/count/ coordinate/keyword/sample/table/column metadata fields. duckdb_schemas() includes case-preserving user schemas with internal = false alongside the built-in system schemas. duckdb_constraints() populates NOT NULL, autogenerated UNIQUE, and autogenerated PRIMARY KEY rows with global zero-based constraint indexes; duckdb_types() exposes built-in type and alias metadata, including the geometry alias rows and system-only JSON alias row from the pinned oracle; duckdb_functions() exposes first-slice metadata for supported scalar and table functions; duckdb_extensions() exposes the v1.5.1 built-in/available extension list; duckdb_optimizers() exposes the v1.5.1 optimizer-name list; duckdb_memory() exposes v1.5.1 memory categories with first-slice zero counters; duckdb_secret_types() exposes the built-in http secret type; duckdb_approx_database_count() and duckdb_connection_count() expose fresh in-memory count metadata; duckdb_coordinate_systems() exposes the stable v1.5.1 OGC:CRS83 and OGC:CRS84 coordinate-system rows with first-slice empty projjson/WKT fields; duckdb_keywords() exposes the full v1.5.1 keyword/category list; duckdb_table_sample('table') returns the target table schema with first-slice empty sample rows; CREATE [TEMP] SEQUENCE, DROP SEQUENCE, nextval(), and currval() are implemented for the in-memory catalog; duckdb_dependencies() remains a schema-correct empty result set; duckdb_logs(), duckdb_log_contexts(), duckdb_prepared_statements(), duckdb_variables(), duckdb_secrets(), duckdb_temporary_files(), and duckdb_external_file_cache() currently expose schema-correct empty result sets until those subsystems have backing state; duckdb_views() exposes first-slice user view metadata from CREATE VIEW, and CREATE [UNIQUE] INDEX [IF NOT EXISTS] name ON table(cols...) and DROP INDEX [IF EXISTS] store first-slice in-memory index metadata; duckdb_indexes() lists those rows and duckdb_tables().index_count reflects them. Unique indexes reject duplicate non-NULL keys on create, insert, and update. Indexes are not used for scan planning yet. Full DuckDB pragma and table-function coverage remains later work. SQL SHOW TABLES and SHOW ALL TABLES expose the current in-memory table and view catalog with DuckDB-compatible display ordering. SQL DESCRIBE table returns DuckDB-shaped column metadata from the catalog; DESCRIBE SELECT/VALUES ... returns bound output names/types with conservative nullability metadata; relation forms such as SELECT * FROM (DESCRIBE t) are available in the current table-source subset. EXPLAIN returns explain_key/explain_value rows for supported query plans; EXPLAIN ANALYZE executes supported query plans and returns a first-slice analyzed_plan row with elapsed time, row count, and plan summary. The SQLLogicTest runner currently covers basic statement ok/error/maybe and query blocks, expected error fragments, fresh per-file databases, tab-delimited upstream expected rows, type-directed integral rendering, emptyset blocks, persistent mode skip/unskip, skipif/onlyif conditionals, halt, reconnect, restart, require/require-env-driven skips, test-env defaults, labels, named connections, hashed expected-output blocks, and rowsort/valuesort result normalization. loop/foreach expansion covers the current sequential subset, including continue, nested blocks, and {name}/$name substitution. test/sqllogic/run.ss --events PATH ... emits JSONL run, file, statement, query, and summary events for tooling.

Requirements

  • jerbuild on PATH
  • Pinned DuckDB source is fetched into vendor/duckdb; set DUCKDB_SOURCE to use an explicit checkout
  • git for upstream revision verification

Commands

make build            # compile all DuckDB Jerboa library modules
make binary           # build dist/jduckdb and its relocatable runtime bundle
make binary-smoke     # build and test DuckDB-compatible CLI behavior
make install          # build the package and install locked jpkg dependencies
make contracts        # regenerate the checked-in upstream inventory
make contracts-check  # verify revision and deterministic generated output
make unit             # run Jerboa unit tests
make diff             # run the current DuckDB CLI differential tests
make bench            # run the current scan/aggregate/group/order benchmark
make sqllogic         # run the current SQLLogicTest subset
make sqllogic-upstream # run the selected unmodified pinned DuckDB fixtures
make sqllogic-events  # smoke-test SQLLogicTest JSONL event output
make concurrency-stress # replay deterministic connection/query interleavings
make native-race-check # run native TSan or the platform race-analysis equivalent
make concurrency-check # run both concurrency gates
make check-docs       # check Markdown hygiene
make verify           # run all current gates

make diff and make bench resolve the DuckDB oracle as duckdb from PATH by default. Set DUCKDB_CLI to use a specific compatible executable. Both targets report a skip when no external DuckDB CLI is available.

concurrency-stress repeatedly forces reader/writer ownership conflicts and close/query races with a fixed schedule. native-race-check runs ThreadSanitizer on platforms with a working runtime; on macOS, where Apple TSan crashes before the launcher can execute its core process, it uses Clang thread-safety warnings and the Clang static analyzer over the repository's native launcher instead.

The native command is named jduckdb and uses DuckDB's jduckdb [OPTIONS] FILENAME [SQL] shape. It supports -c/-s, -cmd, -f, -init, stdin and interactive input, common header/separator/null controls, box/list/CSV/JSON/JSON-lines/line/Markdown/HTML/ quote modes, and the .mode, .headers, .separator, .nullvalue, .read, .open, .tables, .bail, .echo, .help, and .quit commands. The relocatable dist/ bundle contains the public jduckdb launcher, a legacy duckdb launcher alias, WPO-compiled jduckdb-core, and its native runtime library; keep those files together when installing it. Full readline history, completion, terminal-width rendering, safe-mode policy, import/export helpers, and the UI command remain later shell work.

The compatibility ledger begins at contracts/compatibility.sexp. The dependency-ordered roadmap for removing the remaining DuckDB gaps is in Project.md (Workstreams 0-18, Phases 0-7); the consolidated working handoff is in glm-53-handoff.md. SQL execution support is still partial; unsupported work includes remaining advanced join forms, correlated subquery shapes beyond the implemented nested EXISTS/NOT EXISTS, scalar-subquery, and IN/NOT IN subsets, implicit and right/full lateral joins, aggregate outer-shape limitations for correlated subqueries, and broader subquery forms, true materialization guarantees, advanced set-operation clauses, remaining advanced window forms such as timezone-specific temporal fill and streaming execution, aggregate or advanced QUALIFY shapes, persistent storage, and external file formats.