Make DTrace performance profiles attributable and actionable #75
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Jerboa needs a trustworthy, attributable DTrace/performance-profiling workflow
before downstream projects can use traces to choose optimizations. In a recent
Jerboa-SQLite investigation, the available historical trace data was useful for
finding measurement contamination, but not for identifying a defensible engine
hotspot. We ended up building project-local counters and a controlled workload
to find one optimization that DTrace could not prove.
This issue is about making the Jerboa runtime/toolchain produce actionable
profiles, not about optimizing one downstream SQL implementation.
Observed problems
could be included in totals.
Jerboa worker or process incarnation.
ioctlstartup pattern looked suspicious, but withoutrequest/FD/path attribution it could not distinguish bridge startup from
workload behavior.
not suitable for assigning an optimization budget.
dominant cost. Controlled counters later showed that it was a real local
hotspot, but the wall-clock improvement was small and noisy.
while SIP/provider/target privileges prevent attachment. The failure needs
to be represented as an explicit capability result, not as an ambiguous
empty trace.
-dtrace-named artifact is not sufficient evidence that probeswere emitted, attached, or observed.
Requested runtime/toolchain capabilities
architecture, DTrace path/version, SIP/privilege state, provider
availability, compiler flags, bridge mode, and whether probes were actually
attached.
before workload activity. Scope CPU, allocation, GC, syscall, and I/O
collectors to that cohort rather than the host globally.
unique ID, with parent IDs preserved across threads. Use a monotonic,
high-resolution clock and declare units in the event schema.
attachment, drop counts, detach/exit status, malformed records, and signal
handling. A missing or failed collector must invalidate the result bundle.
Provide a standard warmup boundary and an untraced baseline with identical
workload settings.
seed, workload, dataset, process count, and settings identical; record A/B
order and quantify observer overhead before interpreting traced timings.
observed samples, estimated bytes, dropped samples, GC pause units, and
target attribution separately; never present host-wide values as worker
values.
bounded request type, file descriptor, path/file identity, and process
incarnation. Preserve raw records alongside normalized summaries.
the Jerboa trace bridge. Include compile-time and runtime examples that
verify the actual providers/events before collection.
diagnostics when the requested tracing mode is unavailable. The default
build must remain unchanged when tracing is disabled.
Requested artifact/validation workflow
source/toolchain revisions, and configuration hashes with every trace.
parent/child spans, and collector completion using checked-in fixtures.
the checksum manifest only after all validation and logs are complete.
validator rejects it. Archive and re-validate bundles after copying them.
“collector failed”; do not collapse these into
passorSKIP.Downstream workflow improvements enabled by this
Once the above exists, the recommended optimization workflow is:
independent multiprocess ledgers first.
retention thresholds.
sweep as mixed concurrency. Measure worker counts 1/2/4/8/16/32 within
resource limits.
independent-oracle correctness evidence. Preserve negative results and do
not invent regression budgets from contaminated historical data.
For query engines, subquery attribution should distinguish correlated from
uncorrelated plans, and any result cache must be statement-scoped, parameter
aware, conservative around user/nondeterministic functions, and tested for
NULLs, multirow scalar behavior, nested CTEs, errors, reset, and reexecution.
Benchmark matrix manifests should report which axes were actually measured
versus merely planned/delegated.
Acceptance criteria
only the registered target processes and whose collector status is verified.
a measured observer-overhead report.
GC, CPU, or syscall totals.
produce a failing validation result.
a precise capability failure with remediation instructions.
justify a code change, rather than only presenting host-wide syscall totals.