Make DTrace performance profiles attributable and actionable #75

Closed
opened 2026-09-14 22:44:14 -04:00 by ober · 0 comments
Owner

Summary

Jerboa needs a trustworthy, attributable DTrace/performance-profiling workflow
before downstream projects can use traces to choose optimizations. In a recent
Jerboa-SQLite investigation, the available historical trace data was useful for
finding measurement contamination, but not for identifying a defensible engine
hotspot. We ended up building project-local counters and a controlled workload
to find one optimization that DTrace could not prove.

This issue is about making the Jerboa runtime/toolchain produce actionable
profiles, not about optimizing one downstream SQL implementation.

Observed problems

  • Allocation and GC samples were globally scoped, so unrelated host processes
    could be included in totals.
  • CPU, syscall, and I/O totals were not reliably attributable to a specific
    Jerboa worker or process incarnation.
  • A repeated four-ioctl startup pattern looked suspicious, but without
    request/FD/path attribution it could not distinguish bridge startup from
    workload behavior.
  • Historical syscall latency totals were quantized and overlapping; they were
    not suitable for assigning an optimization budget.
  • The old trace did not prove that repeated scalar-subquery planning was the
    dominant cost. Controlled counters later showed that it was a real local
    hotspot, but the wall-clock improvement was small and noisy.
  • Live attachment is difficult to diagnose: on macOS, DTrace may be present
    while SIP/provider/target privileges prevent attachment. The failure needs
    to be represented as an explicit capability result, not as an ambiguous
    empty trace.
  • A successful -dtrace-named artifact is not sufficient evidence that probes
    were emitted, attached, or observed.

Requested runtime/toolchain capabilities

  1. Add an explicit tracing capability probe and mode manifest covering OS,
    architecture, DTrace path/version, SIP/privilege state, provider
    availability, compiler flags, bridge mode, and whether probes were actually
    attached.
  2. Provide a cohort API that registers target PIDs and process incarnations
    before workload activity. Scope CPU, allocation, GC, syscall, and I/O
    collectors to that cohort rather than the host globally.
  3. Give every runtime/process incarnation a stable runtime ID and every span a
    unique ID, with parent IDs preserved across threads. Use a monotonic,
    high-resolution clock and declare units in the event schema.
  4. Make collector lifecycle observable and fail closed: record startup,
    attachment, drop counts, detach/exit status, malformed records, and signal
    handling. A missing or failed collector must invalidate the result bundle.
  5. Keep startup/bridge/DOF costs separate from steady-state workload costs.
    Provide a standard warmup boundary and an untraced baseline with identical
    workload settings.
  6. Add standard no-consumer, light-sampling, and full-tracing A/B modes. Keep
    seed, workload, dataset, process count, and settings identical; record A/B
    order and quantify observer overhead before interpreting traced timings.
  7. Add documented allocation and GC sampling controls. Report sample rate,
    observed samples, estimated bytes, dropped samples, GC pause units, and
    target attribution separately; never present host-wide values as worker
    values.
  8. Add syscall/I/O attribution for open/read/write/fsync/flock/ioctl, including
    bounded request type, file descriptor, path/file identity, and process
    incarnation. Preserve raw records alongside normalized summaries.
  9. Expose stable probe names/field layouts and a versioned ABI contract for
    the Jerboa trace bridge. Include compile-time and runtime examples that
    verify the actual providers/events before collection.
  10. Ensure standalone/native builds preserve the trace probes and emit useful
    diagnostics when the requested tracing mode is unavailable. The default
    build must remain unchanged when tracing is disabled.

Requested artifact/validation workflow

  • Retain exact binary, debug information, native libraries, symbol maps,
    source/toolchain revisions, and configuration hashes with every trace.
  • Validate raw headers, units, monotonicity, overflow markers, event pairing,
    parent/child spans, and collector completion using checked-in fixtures.
  • Bound run length and output volume; use atomic result finalization and create
    the checksum manifest only after all validation and logs are complete.
  • Provide a deliberately failing collector fixture and ensure the top-level
    validator rejects it. Archive and re-validate bundles after copying them.
  • Separate “not available”, “not requested”, “no events observed”, and
    “collector failed”; do not collapse these into pass or SKIP.

Downstream workflow improvements enabled by this

Once the above exists, the recommended optimization workflow is:

  1. Establish a clean untraced baseline before changing engine code.
  2. Run deterministic correctness, fault-injection, rollback/savepoint/WAL, and
    independent multiprocess ledgers first.
  3. Run cold versus persistent connections and dataset cells below/above known
    retention thresholds.
  4. Keep writer, reader, and mixed lanes explicit; do not label a writer-only
    sweep as mixed concurrency. Measure worker counts 1/2/4/8/16/32 within
    resource limits.
  5. Pair each proposed optimization with a focused counter and differential or
    independent-oracle correctness evidence. Preserve negative results and do
    not invent regression budgets from contaminated historical data.

For query engines, subquery attribution should distinguish correlated from
uncorrelated plans, and any result cache must be statement-scoped, parameter
aware, conservative around user/nondeterministic functions, and tested for
NULLs, multirow scalar behavior, nested CTEs, errors, reset, and reexecution.
Benchmark matrix manifests should report which axes were actually measured
versus merely planned/delegated.

Acceptance criteria

  • A bounded sample workload produces a bundle whose events are attributable to
    only the registered target processes and whose collector status is verified.
  • No-consumer, light-sampling, and full-trace runs have comparable settings and
    a measured observer-overhead report.
  • A synthetic unrelated background process does not change target allocation,
    GC, CPU, or syscall totals.
  • Missing attachment, dropped events, malformed records, and collector exit
    produce a failing validation result.
  • The same workflow works on supported macOS/Linux configurations, or records
    a precise capability failure with remediation instructions.
  • The example report identifies a dominant cost with enough attribution to
    justify a code change, rather than only presenting host-wide syscall totals.
## Summary Jerboa needs a trustworthy, attributable DTrace/performance-profiling workflow before downstream projects can use traces to choose optimizations. In a recent Jerboa-SQLite investigation, the available historical trace data was useful for finding measurement contamination, but not for identifying a defensible engine hotspot. We ended up building project-local counters and a controlled workload to find one optimization that DTrace could not prove. This issue is about making the Jerboa runtime/toolchain produce actionable profiles, not about optimizing one downstream SQL implementation. ## Observed problems - Allocation and GC samples were globally scoped, so unrelated host processes could be included in totals. - CPU, syscall, and I/O totals were not reliably attributable to a specific Jerboa worker or process incarnation. - A repeated four-`ioctl` startup pattern looked suspicious, but without request/FD/path attribution it could not distinguish bridge startup from workload behavior. - Historical syscall latency totals were quantized and overlapping; they were not suitable for assigning an optimization budget. - The old trace did not prove that repeated scalar-subquery planning was the dominant cost. Controlled counters later showed that it was a real local hotspot, but the wall-clock improvement was small and noisy. - Live attachment is difficult to diagnose: on macOS, DTrace may be present while SIP/provider/target privileges prevent attachment. The failure needs to be represented as an explicit capability result, not as an ambiguous empty trace. - A successful `-dtrace`-named artifact is not sufficient evidence that probes were emitted, attached, or observed. ## Requested runtime/toolchain capabilities 1. Add an explicit tracing capability probe and mode manifest covering OS, architecture, DTrace path/version, SIP/privilege state, provider availability, compiler flags, bridge mode, and whether probes were actually attached. 2. Provide a cohort API that registers target PIDs *and process incarnations* before workload activity. Scope CPU, allocation, GC, syscall, and I/O collectors to that cohort rather than the host globally. 3. Give every runtime/process incarnation a stable runtime ID and every span a unique ID, with parent IDs preserved across threads. Use a monotonic, high-resolution clock and declare units in the event schema. 4. Make collector lifecycle observable and fail closed: record startup, attachment, drop counts, detach/exit status, malformed records, and signal handling. A missing or failed collector must invalidate the result bundle. 5. Keep startup/bridge/DOF costs separate from steady-state workload costs. Provide a standard warmup boundary and an untraced baseline with identical workload settings. 6. Add standard no-consumer, light-sampling, and full-tracing A/B modes. Keep seed, workload, dataset, process count, and settings identical; record A/B order and quantify observer overhead before interpreting traced timings. 7. Add documented allocation and GC sampling controls. Report sample rate, observed samples, estimated bytes, dropped samples, GC pause units, and target attribution separately; never present host-wide values as worker values. 8. Add syscall/I/O attribution for open/read/write/fsync/flock/ioctl, including bounded request type, file descriptor, path/file identity, and process incarnation. Preserve raw records alongside normalized summaries. 9. Expose stable probe names/field layouts and a versioned ABI contract for the Jerboa trace bridge. Include compile-time and runtime examples that verify the actual providers/events before collection. 10. Ensure standalone/native builds preserve the trace probes and emit useful diagnostics when the requested tracing mode is unavailable. The default build must remain unchanged when tracing is disabled. ## Requested artifact/validation workflow - Retain exact binary, debug information, native libraries, symbol maps, source/toolchain revisions, and configuration hashes with every trace. - Validate raw headers, units, monotonicity, overflow markers, event pairing, parent/child spans, and collector completion using checked-in fixtures. - Bound run length and output volume; use atomic result finalization and create the checksum manifest only after all validation and logs are complete. - Provide a deliberately failing collector fixture and ensure the top-level validator rejects it. Archive and re-validate bundles after copying them. - Separate “not available”, “not requested”, “no events observed”, and “collector failed”; do not collapse these into `pass` or `SKIP`. ## Downstream workflow improvements enabled by this Once the above exists, the recommended optimization workflow is: 1. Establish a clean untraced baseline before changing engine code. 2. Run deterministic correctness, fault-injection, rollback/savepoint/WAL, and independent multiprocess ledgers first. 3. Run cold versus persistent connections and dataset cells below/above known retention thresholds. 4. Keep writer, reader, and mixed lanes explicit; do not label a writer-only sweep as mixed concurrency. Measure worker counts 1/2/4/8/16/32 within resource limits. 5. Pair each proposed optimization with a focused counter and differential or independent-oracle correctness evidence. Preserve negative results and do not invent regression budgets from contaminated historical data. For query engines, subquery attribution should distinguish correlated from uncorrelated plans, and any result cache must be statement-scoped, parameter aware, conservative around user/nondeterministic functions, and tested for NULLs, multirow scalar behavior, nested CTEs, errors, reset, and reexecution. Benchmark matrix manifests should report which axes were actually measured versus merely planned/delegated. ## Acceptance criteria - A bounded sample workload produces a bundle whose events are attributable to only the registered target processes and whose collector status is verified. - No-consumer, light-sampling, and full-trace runs have comparable settings and a measured observer-overhead report. - A synthetic unrelated background process does not change target allocation, GC, CPU, or syscall totals. - Missing attachment, dropped events, malformed records, and collector exit produce a failing validation result. - The same workflow works on supported macOS/Linux configurations, or records a precise capability failure with remediation instructions. - The example report identifies a dominant cost with enough attribution to justify a code change, rather than only presenting host-wide syscall totals.
ober closed this issue 2026-09-22 00:52:33 -04:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
ober/jerboa#75
No description provided.