Reduce repeated grouping allocations for committed game inputs
perfloop-oss/spacetimedb · ALLOCATION HOT LOOP
https://perfloop.ai/t/oss/case_b6yy3j6z2n
Verdict
VALIDATED BY A DIFFERENT CHANGE · settled 2026-10-06 · pull request opened as clockworklabs/SpacetimeDB#6136
What happened: The paired measurements met the required improvement.
Hypothesis
The v2 delivery worker rebuilds a disposable grouping index for every committed update that affects subscriptions. Blackholio steering inputs can repeat this storage work for hundreds of clients even though their encoded row data is shared. Its contribution to full server allocation cost, and the saving from reuse, remain unmeasured.
At revision 257a1916ca8cdc20367b321e28e4593d31e2cd5d, the benefiting consumer operation is the Blackholio browser client's `connection.reducers.updatePlayerInput({ direction })` with its default `subscriptionBuilder().subscribeToAllTables()`. The demo documents hundreds of synchronized players (`demo/Blackholio/README.md:3-8,55-58`). Its browser README identifies the existing Rust module and repository-linked TypeScript SDK (`demo/Blackholio/client-ts/README.md:3-4,16-21`). `GameManager.connect` installs the all-table subscription (`client-ts/src/game/GameManager.ts:43-57`). The SDK generates one SELECT query per table and registers them under one query-set ID (`crates/bindings-typescript/src/sdk/subscription_builder_impl.ts:152-158`; `sdk/db_connection_impl.ts:555-576`). `PlayerController.update` sends steering input at most every 50 ms while the local player owns circles (`client-ts/src/game/PlayerController.ts:8,78-104`). Entering the game creates one initial circle; before splitting, a changed steering input updates that circle only (`server-rust/src/lib.rs:198-239,273-286`). The required benefiting replay is 100 healthy subscribed connections with an entered, unsplit caller and changed steering directions. This uses the documented game scale and source-defined initial player state, not an observed production distribution.
This is an ordinary update route, not a retry or administrative path. The SDK prefers v3 transport with v2 fallback (`crates/bindings-typescript/src/sdk/websocket_protocols.ts:3-5`). Both carry the same logical messages: the v3 handler dispatches to the v2 handler, which enqueues the reducer with its caller and request ID (`crates/core/src/client/message_handlers_v3.rs:18-20`; `message_handlers_v2.rs:50-56`; `client_connection.rs:1102-1121`). Successful reducer execution reaches `commit_and_broadcast_event` (`crates/core/src/host/wasm_common/module_host_actor.rs:1189-1207`). The subscription manager emits row-list references per client/query-set membership, then transfers the computed updates to the broadcast queue (`module_subscription_manager.rs:1500-1514,1374-1391`). One replica queue binds the manager and actor to the send worker (`crates/core/src/host/host_controller.rs:977-982`). The worker waits for the transaction offset before calling the v2 sender (`module_subscription_manager.rs:1879-1884,2158-2164`).
The sender creates a fresh `BTreeMap<(ClientId, ClientQuerySetId, TableName), Vec<TableUpdateRows>>`, appends each successful membership's rows, and consumes the map to construct client messages (`:2050-2103`). For the required one-table change and one installed set per healthy client, there are 100 distinct grouping keys per affected successful broadcast. This bound does not apply to failed, unchanged or unsubscribed inputs. The row vectors become boxed slices in messages that escape into client send queues (`:2090-2098,2177-2186`; `client_connection.rs:434-488`). Those owned message arrays cannot simply be retained as scratch storage. The separate ordered index is disposable bookkeeping: its nodes are consumed and discarded after each broadcast. The row payload itself already uses shared `Bytes` and `Arc` storage (`crates/client-api-messages/src/websocket/common.rs:67-101`), so payload cloning is not the proposed waste. The v1 sender demonstrates worker-owned, drained aggregation capacity (`module_subscription_manager.rs:1905-1954,1998-2000`), but is only a storage-lifecycle comparison, not an equivalent protocol implementation.
The cost and priority are inferred from source. Index storage grows with the number of distinct grouping keys and is reconstructed on each eligible input commit; at this documented scale that is storage for 100 composite-key records each time. The expected removed delta is recurring allocation of this index storage after capacity warm-up, not removal of the required message boxes or a claimed percentage of CPU time. For a changed unsplit-player input, database writes and row encoding concern one circle, while the input membership vector, grouping index and owned message descriptors grow with client count. The index therefore has the same per-client growth order as the main message-construction allocation population, rather than constant bookkeeping beside a large per-client row batch. Serialization buffers also have an existing bounded reuse pool (`crates/client-api/src/routes/subscribe.rs:1563-1606`), though transport backlog can prevent immediate reclamation. This makes an allocation benefit plausible, not established. The 20-Hz client setting is an upper send cadence, not a measured changed-input or successful-commit rate. Actual steering mix, queue backlog and deployed fan-out remain outside the claim.
The migration guide identifies v2 as the new 2.0 protocol (`docs/docs/00300-resources/00100-how-to/00600-migrating-to-2.0.md:15-21`), and v3 explicitly retains its server-message schema (`crates/core/src/client/messages.rs:208-218`). These passages support work on the shared delivery representation; they do not establish a future support timeline. The proposed storage change is private to the worker and adds no consumer operations, configuration or protocol types. It is independent of the known duplicate-plan evaluation and borrowed-search-key repairs.
A Case can publish the checked-in Blackholio Rust module on this revision, connect 100 clients through the linked SDK with the default subscription, enter an unsplit caller, and replay changed steering inputs with ordinary timers and real client delivery enabled. After warming the worker, compare server allocated bytes and allocation requests per successful changed-input delivery, attributing grouping separately from delta evaluation, encoding, serialization and queue wait. The smallest falsifier is that index allocations do not disappear after warm-up, or that full-operation allocation savings are not material. Report the retained scratch-memory tradeoff rather than hiding it. Required regression controls cover ordered offsets, repeated broadcasts without stale groups, multiple query sets and fragments, per-set row multiplicities, event-table rows, failed-set exclusion, cancelled clients, and caller results with both populated and empty updates. The existing `subscribe_batch_applies_sets_and_reports_errors` test (`module_subscription_actor.rs:2857-2944`) is a delivery regression control, not fan-out cost evidence. No runtime measurement or benchmark result is claimed here.
Change to test: Replace the disposable v2 grouping tree with bounded, worker-owned reusable aggregation storage, retaining the current composite-key emission order and fragment order within each key. Preserve transaction-offset waiting, per-set bag updates, event rows, failed-set exclusion, cancellation and both caller ReducerResult forms without changing wire types or consumer APIs.
Where it lives
perfloop-oss/spacetimedb · crates/core/src/subscription/module_subscription_actor.rs
Evidence
100-client Blackholio all-table steering delivery · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
rust_process_alloc_requests_per_input_replay400 |
2402 |
1462 |
−39.1% (−939.7) |
−946.4 to −937.6 |
< −120.1 |
PASSED |
rust_process_requested_alloc_bytes_per_input_replay400 |
536629 |
302145 |
−43.7% (−234485) |
−242020 to −231442 |
< −26831 |
PASSED |
v2_sender_aggregation_enqueue_alloc_requests_per_input_replay400 |
613.1 |
300.3 |
−51% (−313) |
−314.5 to −312 |
< −30.66 |
PASSED |
v2_sender_aggregation_enqueue_requested_alloc_bytes_per_input_replay400 |
95999 |
17128 |
−81.7% (−78387) |
−81224 to −77043 |
< −4800 |
PASSED |
delivery_latency_ns_median_replay400 |
4514447 |
4539972 |
+1.4% (+64210) |
−19366 to +100107 |
≤ 225722 |
PASSED |
Checks: 5 of 5 passed. Verification: no defect found.
Timeline
2026-10-04· Case opened2026-10-09· PR opened