Prove 15-way cloud computer fan-out and operator readiness #11

Open AtlantisPleb opened this 6h ago

Project

Cloud computer platform

Source: Cloud computer scale architecture audit

Outcome

Create the exact-candidate qualification and observability gate required before one chat can activate 15 OpenAgents-managed computers. Qualify logical inventory and active concurrency separately so 30 cold workspaces do not imply 30 reserved runtimes.

Scope

Build the reproducible load, fault, security, metering, and operator-readiness gate for logical inventory and active runtime concurrency.

Required scenarios

  • Create 30 logical computers in one chat with no compute allocated while all are cold.
  • Activate 15 standard runtimes through bounded admission without a provider API storm.
  • Run competing multi-chat and multi-tenant requests and verify weighted fairness.
  • Restart the controller before dispatch, after dispatch, during streaming, and during checkpoint commit.
  • Crash one runtime and lose one complete host or sandbox node.
  • Reattach without duplicate commands or cross-generation events.
  • Deny unauthorized egress and metadata access while admitted broker traffic continues.
  • Exercise CPU, memory, process, disk, time, network, and provider-operation exhaustion.
  • Saturate quota while cleanup, host replacement, and the production reserve remain available.
  • Cancel and time out work at every lifecycle stage.
  • Verify zero residue after success, failure, cancellation, timeout, runtime crash, and host drain.
  • Compare text and voice requests and prove identical admission, authority, budget, placement, usage, and terminal outcomes.

Measurements

  • Queue and cold-start time by class and provider.
  • Active, cold, queued, failed, leaked, and quarantined computers.
  • Host or cluster allocatable and reserved CPU, memory, scratch, and warm capacity.
  • Quota age, safety headroom, denied admission, and cleanup reserve.
  • Command ambiguity, duplicate-dispatch prevention, event gaps, and reattachment.
  • Checkpoint duration, bytes, restore time, integrity failure, and garbage collection.
  • Cleanup duration and zero-residue failures.
  • Runtime cost per active minute and per completed job.

Deliverables

  • Reproducible workload and fault corpus.
  • Exact image, kernel, controller, policy, provider, and configuration digests.
  • Machine-readable qualification report and human operator summary.
  • Capacity recommendation for four, eight, fifteen, and thirty active standard runtimes.
  • Go, hold, or rollback decision for each provider and runtime class.

Acceptance criteria

  • All required scenarios pass on one exact candidate without hidden manual repair.
  • Usage totals match provider observations within a documented tolerance.
  • The report proves that application, cleanup, and replacement headroom remain intact.
  • Any leak, cross-tenant access, duplicate dispatch, secret exposure, or unproven cleanup blocks promotion.
  • The operator can identify capacity, queue, failure, quarantine, and quota state without raw cloud administration access.

Dependencies

Depends on every Foundation and Runtime providers issue in this project, OpenAgentsInc/openagents.com#37, and OpenAgentsInc/openagents.com#38.

  1. AtlantisPleb opened this issue 6h ago
Sign in with GitHub to comment on this issue.