Do not credit capacity for an already-absent sandbox

perfloop-oss/runtime · UNCATALOGUED MECHANISM

https://perfloop.ai/t/oss/case_apass4xmx1

Verdict

VERIFIED · settled 2026-10-07

What happened: The assertion is violated on the comparison and satisfied with this change.

Hypothesis

Cleanup of an already-absent sandbox still reduces the API's estimate of allocated CPU and memory. If a newer node report already excluded that sandbox, the subtraction makes surviving allocations look smaller. This is a bookkeeping correctness proposal, not a claim of permanent leakage, exact live totals, production oversubscription, or performance gain.

The scope is delayed failed-publication rollback after successful placement at revision 92197909dce5a1bef33e764ae4af76f0732fd7a8. Placement optimistically adds the new sandbox's resources after node Create succeeds (placement/placement.go:159–176). CreateSandbox subsequently launches copied, uncancelled node rollback on Add failure, returns a 500 error, and completes the reservation without waiting for cleanup (create_instance.go:287–295, 549–577). The failed execution may independently retire before that cleanup reaches the node. Node metrics can then refresh the allocated CPU and memory view without it (nodemanager/sync.go:72–82; metrics.go:37–78).

Server.Delete returns NotFound when its live registry has no sandbox under the supplied ID (server/sandboxes.go:723–728). killSandboxOnNode treats that absence as idempotent success, but falls through to the same OptimisticRemove used for acknowledged deletion (delete_instance.go:347–378). OptimisticRemove subtracts the old sandbox's CPU and memory quantities whenever the aggregate counters are large enough (nodemanager/node.go:233–265). Its underflow checks prevent unsigned wrap; they do not distinguish those quantities from surviving allocations in a refreshed view. Other delete errors return before subtraction.

The triggering ordering is: E1 is placed and its Add fails; original and joined requests settle as errors; rollback remains delayed; E1 retires; a newer authoritative report contains only remaining allocations; then the stale E1 delete receives NotFound and subtracts E1 again from that report. A private authored TestPerfloopPublicationLifecycle exercised this through actual CreateSandbox, Store, Redis reservation scripts on miniredis, actual node Map retirement, Node.UpdateMetricsFromServiceInfoResponse, and killSandboxOnNode. A refreshed 4-CPU/1024-MiB view became 2 CPUs/512 MiB after stale NotFound for E1. The original parent was cancelled, while original and joined errors had returned and pending was zero. An Unavailable-delete control did not subtract. Reproduction used go test ./packages/api/internal/orchestrator -run '^TestPerfloopPublicationLifecycle$' -count=1 -v with authored publication and deletion barriers. It did not create physical surviving VMs or demonstrate a production placement error.

The estimates are explicitly optimistic and overwritten by the next real node report (placement/placement.go:171–174). The proposed change is conservative: without an intervening refresh, skipping NotFound subtraction may retain the old optimistic addition until the next report. That tradeoff avoids inventing free capacity from absence alone; it does not promise exact counters between reports or immediate correction of every stale estimate. This proposal does not change API errors, reservation ownership, Redis record cleanup, or node deletion identity.

The required Case check faults a whole start or resume after successful node creation, holds rollback, independently retires E1, and obtains a real node report that excludes E1 while retaining known surviving sandboxes. Release deletion and verify that NotFound leaves that refreshed allocation view unchanged. Observe the original API and joined errors, reservation result, records, node executions, and resource estimates together. Preserve optimistic credit for successful matching deletion, no credit for other errors, and cancellation-detached cleanup. A no-refresh NotFound control may conservatively retain E1's old estimate, but the next authoritative report must reconcile it. Successful publication and synchronous filesystem-boot rejection are immediate comparison controls. Subtracting E1's quantities from the refreshed surviving view, or failing to accept the subsequent authoritative report, falsifies this limited property.

README.md:31–41 describes continuing API/node responsibilities, and CONTRIBUTING.md:12–23 welcomes clear bug fixes and integration tests; neither is approval of this proposal. This is an internal branch correction. It adds no consumer operation, configuration, state type, or wire protocol.

Change to test: Keep NotFound an idempotent deletion success, but do not subtract the stale sandbox's resource quantities when the node reports no live target. Let authoritative node metrics reconcile absence, retaining optimistic subtraction for acknowledged deletions and existing handling for other errors.

Where it lives

perfloop-oss/runtime · packages/api/internal/handlers/sandbox.go

Evidence

The assertion is violated on the comparison and satisfied with this change: `In CreateSandbox start and resume calls where node Create succeeds but Store.Add fails, once E1 is retired from the real node live map, delayed NotFound rollback leaves a newer service-info CPU, RAM, and huge-page view for surviving sandboxes unchanged. A report after a no-refresh NotFound reconciles the allocation view. The original and joined calls return the same 500 and settle their reservation; successful publication still stores E1, and synchronous filesystem-boot rejection waits for acknowledged Delete before returning 503.`

Checks: 6 of 6 passed. Verification: no defect found.

Timeline