Skip to main content
Marketplace recovery is an instance-admin operation. It diagnoses every company, applies only the existing idempotent repair paths, and returns safe diagnostic codes. It does not guess at a new repair.

CLI recovery flow

Authenticate once, then use the same API base for every command:
Before the POST on a Compose host, verify the exact reviewed release and storage contract:
The CLI prints the operation UUID to stderr before the POST, keeping --json stdout as one parseable response document. If the local request times out, do not retry. Inspect that UUID first. A second operation may be safe only when inspection returns safeToRetry: true. Retry with a new UUID linked to the inspected predecessor:
The server re-runs the inspection and rejects an omitted, stale, mismatched, or unsafe predecessor reference. A failed-before-mutation attempt does not clear the retry barrier.

Reading the result

skips[] means a safety gate deliberately left a company unchanged. Every skip counter has exactly one matching entry. failures[] means a company stage threw or could not persist its result. Both use fixed messages and recovery instructions; raw exception text and paths are not returned.
Join the HTTP response, durable operation ledger, company activity rows, and server logs using operationId. The instance-scoped marketplace_reconciliation_operations row is authoritative for operation-ID ownership, terminal state, and the cross-replica lease. Company activity rows remain the per-company audit detail; each completion row contains only that company’s skips and failures. This also keeps a zero-company or pre-audit operation inspectable after a restart.

Diagnostic recovery table

Failure codes are marketplace_update_failed, crew_repair_failed, legacy_steward_failed, crew_update_failed, team_reconcile_failed, and unknown_internal_failure. Correct or inspect the named company stage before retrying.

Endpoint error recovery table

Outcome unknown after mutation

This state means domain writes may have committed but the completion audit did not. The inspection endpoint re-runs read-only diagnosis and fails closed for active writers, customized rows, unaccounted crew, or query ambiguity. Never retry merely because the POST returned 500. Only one fleet reconciliation lease can be active in the database. The lease is heartbeated while the operation waits for local maintenance locks and while it runs, so another app replica reports operation_in_flight rather than starting overlapping writes. After a crashed worker’s lease expires, inspection reports outcome_unknown_after_mutation; a later POST atomically fences that stale owner and requires the stale operation ID as retryOfOperationId before it can claim a new operation. Follow the inspection result and never bypass the ledger with manual activity rows.

Advanced board-key request

The supported workflow is the CLI. For a controlled diagnostic, an explicitly supplied board key is non-ambient authority and does not require a fabricated Origin or Referer:
For a verified-safe retry, add "retryOfOperationId":"<inspected-operation-id>". The referenced operation must be the server’s current outcome-unknown retry barrier and must still inspect with safeToRetry: true. Never place the key in the URL, command history, incident bundle, or committed fixture.

Managed bundle storage

/aoa is the persistent data mount.