process.env reads in server/src/, cli/src/, and packages/.
Use the AOA_* names shown below.
Server Configuration
For horizontally scaled deployments, local-file run logs require sticky routing or shared durable storage. Object-storage-backed run logs are a future backend; do not rely on per-container ephemeral disk for production run history.
Authentication
CLI subscription authentication
Execution targets & gVisor pool egress (multi-tenant cloud, Phase 5)
Phase 5 adds a tenant-scopedexecution_targets registry (fleet inventory) on
top of the AOA_EXECUTION_TARGET_ID identity above. Runs route to a target by
credential kind — execution-target-resolver.ts: a business (company) API key
routes to the shared pooled_gvisor target, a personal subscription routes to
the dedicated target whose slug matches its bound AOA_EXECUTION_TARGET_ID. No
new environment variable governs this routing; it reads the execution_targets
table and the P4 credential-kind seam. Self-hosted single-tenant installs are
unaffected — they never populate execution_targets beyond the seeded
control-plane row, and resolveExecutionTargetForRun falls back to the local
driver when no pool target exists.
Pool egress allowlist policy. A pooled gVisor run’s Docker hardening
(--runtime=runsc, dropped capabilities, read-only rootfs, etc.) is applied by
the app layer via opt-in buildDockerRunArgs isolation flags — see
docs/aoa/guides/gvisor-worker-image.md
for the exact flag set. Network egress filtering is NOT an app-layer
concern: --network none is the safe default (no egress at all), and a
pooled run that needs the provider API must run on a filtered bridge
network — filtering is a worker-image deliverable (a DOCKER-USER iptables/
nftables policy or an egress proxy) that denies RFC1918, 169.254.0.0/16
(cloud metadata, incl. 169.254.169.254), and the control-plane CIDR, while
allowing only the provider API hosts and package registries. There is no
environment variable for this allowlist yet — it is configured on the worker
image/host, not via AoA server env vars. As of this writing that firewall has
not been validated on real hardware (Task 0’s spike is a pending Gate-B step,
not yet run) — see the guide’s status banner before deploying a pool on
bridge.
Unsandboxed multi-tenant execution gate
AOA_RUNTIME_PROCESS_OWNER_ID prevents one replica from interpreting another
machine’s numeric PID as local. It does not turn the process-local runtime
maps, desired-state restart, or control APIs into a distributed scheduler.
Run at most one owner of local_process services for a shared deployment;
horizontal cloud_auth deployments are production-safe only with the
unsandboxed override disabled until the worker/gVisor runtime lands.
Agent JWT (signing for AOA_API_KEY)
Secrets
Storage
Database backups
Migrations / startup
Legacy Steward reconciliation backout
DisableAOA_STEWARD_RECONCILE_ENABLED before investigating or running a
backout so the next 24-hour pass cannot reapply the change.
template_version = '0.0.0-legacy' is shared by every pointer-only crew
adoption, so it is not a safe rollback selector by itself. Every successful
background reconciliation writes an activity_log row with action
marketplace.legacy_steward_adopted in the same transaction as the adoption.
The application log named legacy Steward adopted in place includes that
row’s auditId for lookup, but the durable database row is the source of truth.
After taking a database backup, inspect the audit rows, choose the exact audit
IDs from the affected deployment window, and create a transaction-local target
table from only those IDs. Never select the whole fleet by version:
memberInserted=true, the pass also created the exact
team_members(teamId, agentId) link in that transaction. Delete that link only
after confirming no later team operation now relies on it; a false value
means the link predated reconciliation and must be kept. The delete uses only
rows that the guarded pointer update above actually reverted, so a Steward that
has since advanced cannot be detached: