How to benchmark collection updates¶
Use Python 3.11 or newer for benched history and Sphinx reports. Install the project’s development dependencies and Chromium, then build the browser runtime:
pnpm --dir js install --frozen-lockfile
pnpm --dir js build
python -m playwright install chromium
Run a small wire benchmark first:
python -m benched run benchmarks/test_collection_wire.py -k '1k and json'
For browser-only comparisons, run benchmarks/test_each.py. Use -k '100k' for the largest
collections, or omit -k for the complete matrix. The wire matrix covers 1k, 10k, and 100k records,
JSON/MessagePack/CBOR, positional/keyed diffs, and insert, remove, update, move, move-with-update,
reorder, streamed-burst, and batched-burst workloads. Large eager mounts make the full matrix slow.
Use plain python -m pytest with --benchmark-disable for a single correctness pass without recording
history. make benchmark-history builds the runtime and records the full suite; use
BENCHMARK_ARGS="-k '1k and json'" to restrict it.
Benched saves immutable records in benchmarks/results. Keep the records you want published alongside
the benchmark changes. Build the Sphinx reports from those records:
yardang build
Open docs/html/docs/src/benchmarks.html through a local HTTP server to inspect the interactive reports.
The docs build reads stored results; it does not rerun benchmarks. For exploratory measurements that
should not enter the published history, pass --results-dir .benchmarks/benched to benched run.
Keyed cases skip when Python transports lacks Store.host(list_keys=...). To test an unreleased
transports checkout, build its Python extension first and select it explicitly:
PYTHONPATH=/absolute/path/to/transports python -m benched run benchmarks/test_collection_wire.py \
-k 'keyed and 1k' --subject-label transports=local-keyed-diffs
The browser uses the transports package installed by Spaday’s lockfile. No local JS replacement is needed: keyed Python diffs emit existing wire operations. Each round checks values against the workload, agreement between the transports mirror and Spaday Store, retained element identity, and unsaved input, focus, and selection. Batched bursts must write the changed item’s label once.
Label local builds explicitly; their installed distribution version need not describe the checkout. Benched also retains Spaday’s revision and dirty-worktree flag. Do not present local-development runs as released-package baselines.
If installed Spaday metadata differs from the browser checkout, pass --subject-version with the
version in js/package.json; make benchmark-history supplies it automatically. Use a clean Python
environment for published-dependency measurements, and check extra_info.transports_python_distribution_version
against the intended transports version before publishing a record.
Compare like-for-like runs on idle hardware. Inspect each wire benchmark’s extra_info.rounds for all
rounds’ counters and timings (top-level fields describe the final round). Pytest-benchmark’s distribution covers the complete Python-to-browser
test call, including correctness checks. roundtrip_render_ms excludes setup, initial snapshot/mount,
and final correctness checks, and ends after two animation frames. Server bridge, diff, and encode
timings are separate. client_receive_apply_ms combines codec decoding, client reduction, and
synchronous store propagation; deferred item-scope rendering is included in the roundtrip time.
Payload sizes count encoded message bodies, excluding WebSocket/TCP framing. Compression is disabled. The streamed burst sends ten revisions as they are generated; the batched burst sends those revisions in one transports batch. Long tasks come from Chromium’s observer; estimated missed frames use the median idle animation-frame interval. They are diagnostics, not hardware-independent release budgets. This single-connection harness does not measure Hub fan-out, reconnect, or slow-consumer policy.