receipt factory commands.
Two names recur throughout. A computer is a throwaway container Factory rents for one task and releases afterwards; OpenSandbox is the controller that hands those containers out, and it is the only one Receipt supports. Inside each container the thing doing the work is the Codex CLI, OpenAI’s command-line coding agent, run once per task against the files Factory staged for it.
This page is the engine’s internals. For what a run looks like from the outside — what you see while it works and how to read the result — start with background runs and objectives and tasks.
The application’s Tasks page is titled Beetle Tasks. Beetle is the assistant’s name in the interface; the engine described here is the same one behind chat, that page and the CLI.
From a goal to a finished run
A delivery objective moves through this sequence. Every state change is a receipt on the streamfactory/objectives/<objectiveId>, so the whole run is reconstructable afterwards; the boxes below name those receipts and the jobs that write them.
An investigation objective takes a shorter path: it never consumes a repository execution slot, moves from collecting_evidence to evidence_ready, dispatches a synthesis task, and ends when that task writes its investigation.reported receipt — no checks, no promotion gate. Investigation is the default mode in a stock checkout, so it is the path you meet first.
The objective, task, candidate, job, check and promotion vocabulary — the twelve objective statuses, the display states, the profile and the promotion gate’s exact refusal messages — is defined once in Factory concepts and configuration.
Planning is a supervisor decision at a receipt boundary
Factory does not plan continuously. At explicit points the runtime runs an objective supervisor: one model call whose structured answer is validated and then applied deterministically. The decision carries an assessment (on_track, needs_more_evidence, needs_replan, sufficient or blocked), a confidence, a rationale, and exactly one action — continue, apply_plan, synthesize or block.
apply_plan is what creates the task graph: it appends task.added receipts, each task carrying its dependsOn node ids. A reconcile pass then activates every ready task and dispatches worker jobs up to the objective’s policy limits. Mutations are compare-and-append, so a mutation that loses the race is retried against the new head rather than overwriting it.
The supervisor model comes from RECEIPT_FACTORY_OBJECTIVE_SUPERVISOR_MODEL (falling back to RECEIPT_FACTORY_SUPERVISOR_MODEL); the task worker model comes from RECEIPT_FACTORY_TASK_MODEL.
Dispatch, reaction and recovery
Factory has no scheduler of its own. It enqueues jobs on the receipt-backed queue and lets durable execution carry them through Resonate: a driver begins the worker RPC, the worker leases the job, heartbeats under a fence, and settles it with a terminal receipt. Which process picks a Factory job up is decided by its agent id first and its kind second.
Anything whose kind starts with
factory.integration. lands on worker-codex whatever agent id it carries. The queue’s kind contract records a workerGroup as well, and for factory.dispatch and factory.task.monitor it names control — but the dispatcher routes on the agent id, so both of those run on worker-chat.
Factory jobs also carry longer leases than the 300-second default: objective control and the computer-path task monitor both default to 900 000 ms, and a codex job takes the larger of CODEX_JOB_LEASE_MS and its own payload timeout plus five minutes, capped at one hour.
Three behaviours are specific to Factory:
- Every codex-lane job reacts the objective when it finishes.
factory.task.run,factory.integration.validateandfactory.integration.publisheach end by calling the objective’s reconcile pass, unless the result wasskipped_terminal_state. That is how one finished task activates the next. - A watchdog reconciles objectives nobody is driving. It runs on a cron, scans objective summaries and enqueues an objective-control job with the reason
reconcilefor one of three decisions:no_active_objective_work,phase_supersession_has_active_stale_joboractive_objective_work_stalled. - Repeated identical failures stop the objective without another model call. When two consecutive sandbox-backed tasks are blocked by the same failure family, a deterministic circuit blocks the objective instead of replanning around it.
The computer lane
There is exactly one execution path,computer, and exactly one computer target, opensandbox — the sandbox controller that starts, tracks and stops the containers tasks run in. There is no local or host-managed execution lane, and neither value is a per-objective choice.
Capacity is a queue, not a pool
Computer capacity is a receipt-backed ledger with two limits, both defaulting to 1: a global one (OPEN_SANDBOX_GLOBAL_MAX_ACTIVE) and a per-organization one (OPEN_SANDBOX_ORG_MAX_ACTIVE, or OPEN_SANDBOX_ORG_<ORG>_MAX_ACTIVE for a named organization). Capacity leases last 240 seconds and are renewed every 160 seconds while the task runs; a waiter polls every second, is told it is waiting after 15 seconds, and a claim abandoned by a dead worker expires after 3 minutes.
While a task waits and runs, the progress summaries it writes are the ones you see in chat and in replay: Requesting a computer., Computer acquired; resolving the workspace., Computer acquired., Preparing the computer workspace. (detail: Syncing task files, selected skills, credentials, and execution metadata.), Computer workspace is ready; starting the agent., Starting the agent in the computer. and Released the computer.
Idle sandbox hosts are not stopped automatically. Host shutdown is an explicit operator action, so a host left running after a burst of work stays up until someone stops it.
The task packet
Everything a worker reads and writes for one task lives underreceipt/current/ in its workspace:
output/<candidateId>.integration.json and the matching .integration.stdout.log and .integration.stderr.log.
What sits around the packet depends on the objective’s mode. An investigation task that carries a connected-system contract gets a receipt-only workspace — AGENTS.md, the receipt tree (the packet plus the organization skill registry) and the selected skill roots, and nothing of the application source. Every investigation task also skips the dependency bootstrap. Delivery workers get the repository worktree, because they may implement, check or publish.
receipt/current/receipt-cli.md is generated per task and tells the worker which commands it may run. It is where the credential rule is stated to the agent directly: Receipt applies provider credentials server-side; never request or print them. The task prompt adds the matching discipline: Never print or persist raw secret, token, password, API key, or credential values in stdout, stderr, artifacts, or the final JSON.
What runs inside
The agent in the sandbox is the Codex CLI, invoked once per task inexec mode against the workspace path, reading prompt.md on stdin and writing its last message and a structured result back into the packet. Web search (--search) is switched on only for tasks whose execution class is broad.
Codex’s own kernel sandbox is bypassed deliberately — the container does not allow the nested namespace operations it needs — so the OpenSandbox computer is the isolation boundary, not Codex’s sandbox flag.
Three timers supervise the run: a 60-second startup timeout (RECEIPT_CODEX_STARTUP_TIMEOUT_MS), a 300-second stall timeout (RECEIPT_CODEX_STALL_TIMEOUT_MS) and a 500 ms abort poll; both timeouts are additionally capped by the job’s own execution window. A steer or abort command posted to the job is picked up within about half a second — the command stream is aborted, the computer lease is released, and the run is recorded as a controlled abort rather than an infrastructure failure.
Credentials inside the sandbox
Connected-system credentials are not copied into the computer. For AWS, GCP, kubectl and Jira, Factory installs a credential helper: a small shell script that reads a token file, callsPOST <gateway>/connect/credential/<provider> on the Receipt Connect gateway, and hands fresh material to the CLI at the moment it is used. Generic HTTP integrations get no helper at all — the worker calls receipt connect call … and the credential never leaves the server. Generated AWS profiles pin region = us-east-1, so name a region explicitly when your resources live elsewhere.
The token in that file is a job-scoped Receipt Connect JWT, minted with connect:read plus, when the task’s contract needs credentials, connect:credential and the contract’s write scopes. It is written with mode 0600 inside the workspace, and the gateway endpoint re-checks the scope and current workspace membership on every call.
Credential setup writes its own receipts: computer.credential_setup.started (Installing and validating remote credentials before command startup.), then either computer.credential_setup.succeeded — Remote credential setup completed with <n> helper file(s)., or Remote credential setup completed; no Receipt Connect helpers were required. — or computer.credential_setup.failed carrying the error.
When the contract cannot be satisfied, the task reports a readiness failure with one of these codes: workspace_missing, codex_missing, codex_auth_missing, receipt_cli_missing, receipt_connect_unavailable, credential_helper_missing, provider_cli_missing.
Results, checks and promotion
After Codex exits, the runtime collects artifacts and patches — up to three attempts, refreshing the computer lease between them, when the sandbox command transport is interrupted, rather than rerunning the model — then writesstdout.log and stderr.log and reads token usage out of the Codex event stream. The structured result is output/result.json when it exists, otherwise the last message parsed as JSON. If neither parses, the job fails with missing structured factory task result from codex.
That result is a contract, not free text. A delivery result carries an outcome of approved, changes_requested, blocked or partial; a completion object listing what changed, the proof for it and any remaining work; and an alignment verdict of aligned, uncertain or drifted. An investigation result carries a status of answered, partial or blocked and findings each marked confirmed, inferred or uncertain.
The outcome decides what happens to the task. approved approves the candidate and the task, which is what lets the work move toward integration. changes_requested and partial both write a changes_requested review and return the task to ready for another pass — unless the rework cap or the alignment gate fires, which blocks it instead; a partial whose checks all passed and whose only remaining notes the controller can resolve itself is upgraded to approved. blocked writes a task.blocked receipt carrying the worker’s handoff as the reason, and no candidate is produced. Those completion fields are what the promotion gate checks later, which is why a task that records no proof, or still reports remaining work, cannot be promoted.
Delivery work is merged in a dedicated integration worktree, and its checks run inside the computer, not on the controller: each check command goes through the sandbox lease with a 60-minute timeout, with one repair-and-retry when a check fails because a tracked file is missing. A passing run emits validated events with the summary Integration checks passed for <candidateId>. and the handoff Integration checks passed for <candidateId>. Controller may continue toward promotion., commits an integration memory entry, and reacts the objective. Promotion happens behind the promotion gate: when the source checkout is clean it is a fast-forward merge of the integrated commit into the source branch, and when the source checkout has uncommitted changes Factory commits only the promoted paths — or refuses with a conflict when those changes overlap the promoted files.
Publishing a pull request
A delivery objective can end with one more Codex run,factory.integration.publish, driven by a checked-in publisher skill. That run reads the objective’s history through the receipt CLI, pushes the branch, checks for an existing pull request with gh pr view, creates one with gh pr create if there is none, retries transient GitHub failures at most twice more, and returns a strict JSON object: summary, prUrl, prNumber, headRefName and baseRefName, with null allowed for the last three when GitHub does not return them. The runtime also accepts an optional handoff and falls back to the summary when it is absent. The publisher is explicitly forbidden from running builds or tests and from changing code.
If the result carries no valid http or https prUrl, the job fails — with the worker’s own blocker summary when it reported one, otherwise factory publish result missing valid prUrl.
Memory
Memory is receipt-backed like everything else: a scope maps to its own stream, and reads and writes appendmemory.accessed and memory.committed events. Each packet mounts six scopes, reached through memory.cjs.
Publishing has a scope of its own,
factory/objectives/<objectiveId>/publish.
Memory search is keyword matching in this release. The memory layer can rank entries by embedding similarity, but no call site supplies it with an embedding function — not the runtime, not the in-repo CLI, not the Factory service — so every search a worker or a runtime memory route performs matches on the terms in the stored text.
summarize is a character-capped join of the matched entries, not a model-written summary.Audits after a run
A terminal objective enqueues afactory.objective.audit job. The audit reconstructs the run from its receipts, writes objective.audit.json and objective.audit.md into the objective’s artifacts, and commits summaries to the scopes factory/audits/objectives/<objectiveId> and factory/audits/repo. A repository-wide system-improvement report is produced only when FACTORY_OBJECTIVE_AUDIT_SYSTEM_IMPROVEMENT is set to true.
If newer receipts moved the objective’s head while the audit was running, the audit completes as superseded rather than failing — findings are never published against a stale snapshot.
Concurrency and limits
The board calls the first row “the repo execution slot”, which reads like one objective at a time. The runtime admits queued delivery objectives up to
repoSlotConcurrency minus the number already holding a slot, so the shipped default is twenty concurrent delivery objectives per repository, not one. Investigation objectives take no slot at all.