---
name: parallelsandbox
description: >
  Use ParallelSandbox when a task needs a real remote computer to build, run and test a repository: docker compose
  stacks, services with ports, browser end-to-end tests, screenshots and screen recordings, anything that must not run
  on the user's own computer, and any step a person must do by hand on a live screen (logins, two-factor codes,
  CAPTCHAs, payments) that you hand over with sandbox_takeover. It is also how you give a running version back: once you have tried a
  version yourself and it works, sandbox_review puts it in their app as something they open and use as a user (a web
  page, or a phone or desktop app on the box's screen that they tap and type on) and returns their report to this active conversation. The box is your remote computer; there is no AI inside
  it. You drive it with the sandbox_* tools and read browser logs from pages you run there with the logs_* tools. Code
  changes stay in your local checkout; the box only runs copies. Do not use it for chat or research.
---

# ParallelSandbox: your remote computer

## Install and connect when the person asks

Do the setup for the person. Save this full file in the active client's skill directory (Claude Code: `~/.claude/skills/parallelsandbox/SKILL.md`; other clients: use their configured skill directory, or their project instructions if they do not support skills). Inspect the existing MCP configuration first and preserve other servers and settings.

Use the local stdio adapter `parallelsandbox-mcp` 0.4.0 or later (Node 20+). Check that Node 20+ and `npx` are available, and install a supported Node runtime if needed. Verify the adapter version before starting it; an older published version is a setup blocker, not a reason to ask the person for an API key. It signs in with OAuth and provides `sandbox_sync` and `sandbox_pull` for local files. Claude Code: `claude mcp add parallelsandbox -- npx -y parallelsandbox-mcp`. Codex: `codex mcp add parallelsandbox -- npx -y parallelsandbox-mcp`, then set `startup_timeout_sec = 60` and `tool_timeout_sec = 3900` under `[mcp_servers.parallelsandbox]` in `~/.codex/config.toml`, preserving an existing longer timeout and all other settings. Cursor and other stdio clients run `npx` with args `["-y", "parallelsandbox-mcp"]`.

Start or reconnect the MCP client. The adapter opens the browser to sign in; while authorization is pending, call `parallelsandbox_connect` and show the person its sign-in link. They sign in; the local adapter then completes the connection automatically. Do not ask them to create a key or click an extra approval button. Call `parallelsandbox_connect` again and then `sandbox_list` to verify the account. If the client cannot reload the tool list, reconnect it after sign-in. The adapter shares sign-in across local conversations and refreshes automatically; no API key is needed. Environment connections and service settings are your job over REST when a project needs them (see Environments).

If local file transfers are not needed, remote HTTP with OAuth is an alternative: Claude Code uses `claude mcp add --transport http parallelsandbox https://mcp.parallelsandbox.com/mcp` and `/mcp` → Authenticate; Codex uses `codex mcp add parallelsandbox --url https://mcp.parallelsandbox.com/mcp` and `codex mcp login parallelsandbox`. `sandbox_sync` and `sandbox_pull` return an error over plain HTTP because they need the agent's local disk. Existing API keys remain optional for legacy HTTP clients or existing adapter setups. A client with a third-party remote callback still asks the person to authorize that service.

## Using boxes

A box is an x86_64 (amd64) Linux machine, Debian 12: 2 vCPU and 8 GB memory per size unit, its own Docker, git, Node 22, Go, Python, PHP 8.2 with Composer, PowerShell 7, the Azure and GitHub CLIs, gcc, clang and cmake, ripgrep (`rg`), `zip`, `xxd`, `bats`, `bwrap`, `python` (Python 3), an Xvfb virtual display `:99` with Chromium (390x844 portrait by default; 1280x800 with `orientation: "landscape"` at `sandbox_start`; see Seeing the screen), ffmpeg. Working directory `/work`, 40 GB of disk per size unit (80 GB at size 2, 320 GB at size 8). Commands run as root in `bash -l` with `DISPLAY=:99`, `SBX_BOX_ID` (the box id) and `SBX_SCENE_HOST` (the scene URL's host, key included) already set. Boxes stay isolated from each other except through the links you declare (see Linking boxes). A box has no cloud credentials of its own; it reaches AWS only when its environment names an AWS role (see Environments). Nothing on a box survives `sandbox_stop` or an interruption. `sandbox_start` returns `arch: "amd64"`. Pull or build `linux/amd64` images: an arm64-only image fails with `exec format error`, and emulation is not installed (`docker run --privileged --rm tonistiigi/binfmt --install arm64` adds it, but slowly). arm64 output for somewhere else (Graviton, Lambda on arm64) can still be built without running it: `CGO_ENABLED=0 GOOS=linux GOARCH=arm64 go build`, or `docker buildx build --platform linux/arm64` of a Dockerfile whose `RUN` steps run on the build platform (`FROM --platform=$BUILDPLATFORM …` in the stage that compiles); a `RUN` step in an arm64 stage fails without that emulation.

Everything belongs to one account: boxes, environments, links, published versions, the registry and the credits. A team shares one account. Each teammate or agent connects with the one line above and signs in once (OAuth); only tools without OAuth need an API key, which the owner makes and revokes over REST with the app's signed-in session (`/v1/keys`; an API key or OAuth token cannot create, list or revoke keys). Every connection sees and operates all boxes of the account (`sandbox_list` lists them, teammates' boxes included), and links and versions work across its connections but never across accounts. Only the account's login opens the app (box cards, Use it, takeovers, `sandbox_say` messages); the account has no members, and the person trying a change needs only the box's URL, which works for anyone who has it. Boxes record the MCP client that started them, not the credential: name boxes `<person>: <task>` and say what the box is for in `goal`, stop only boxes you started or whose name says they are yours, never stop boxes in bulk by age, and ask before stopping a teammate's box.

Ask for more machine with `size` (1, the default, 2, 4 or 8 units of 2 vCPU + 8 GB; it costs proportionally more and is fixed for the life of the box): 1 is enough for web front ends and most Node, Go and Python work, 2 or more for Gradle, Java, Android or heavy docker compose stacks, such as several services built as images in one box. Ask for a language with `toolchains` (`java-17`, `java-21`, `android`, `dotnet`, `rust`, `flutter`, `esp32`): they are baked into the image but off PATH until you pick them, and picking them sets JAVA_HOME, ANDROID_HOME and PATH so plain `java`, `gradle`, `dotnet`, `cargo`, `flutter` and `pio` work. `flutter` builds web and Android here (it implies `android`; a box's first `flutter build apk` downloads about 3 GB of Gradle dependencies, about 3 minutes); its iOS build runs on the Mac holding a leased iPhone (`psbx-ios flutter`, see Phones). `esp32` is PlatformIO with Arduino for ESP32 and ESP32-S3 ready (ESP-IDF and other chips download on their first build, about 3 minutes); it compiles firmware, but a box has no USB, so flashing a board happens on the person's computer. For a test database, `psbx-testdb up` starts Postgres and `psbx-testdb up <name> --mysql` starts MySQL 8.0, both already in the box (see Environments for its options). A box cannot run programs that only run on Windows (WinForms, COM broker APIs, Excel VBA, MetaTrader), a virtual machine (no `/dev/kvm`) or GPU work such as local LLM inference; Google Apps Script and TradingView Pine Script can be edited and pushed from a box but run on their own platforms. `/etc/sbx/manifest.json` in the box lists what that image has. A project's own dependencies are not in the image: for Python, `pip install -r requirements.txt` (and the project's other `requirements*.txt`) in the box before running its tests; pip installs system-wide, `break-system-packages` is already set. A one-off Node script anywhere in the box, with no `package.json`, can import `playwright`, `playwright-core` and `ws` (`import { chromium } from 'playwright'`, or `require`) without `npm install`: they, and Playwright's Chromium, Firefox and WebKit, are installed for the whole box.

Rules:

1. Edit code locally. Files in the box are copies: put them in with `git clone` (committed work) or `sandbox_sync` (uncommitted), run and test there, bring results back with `sandbox_pull` (adapter) or `sandbox_get`. Never edit files in the box and forget to bring the change back.
2. One box per task. Reuse it for the whole task, and finish it before you stop working on it: `sandbox_review` when the person should try the result, `sandbox_stop` when there is nothing left for them to see. A box left running without either shows in the person's app as "AI stopped" and waits for them to decide what happens to it. Do not stop a box while a person is still trying the version on its URLs (after `sandbox_review`, or with `webUrl`): stopping ends their session and the URLs for good. Ask them in this conversation, the way your client asks, or leave the box: it freezes on its own after 10 idle minutes and a frozen box bills nothing (it keeps its slot, and is stopped 24 hours after its last use). Leave `idleTimeoutMin` unset on a box you hand over: past that limit the box is stopped for good the first time it freezes, their URLs with it. Free accounts run 3 boxes at once, Pro 5, Max 20; frozen boxes count toward that limit. The limit and the credits belong to the account and are shared by every key and agent on it, so stop the boxes you are done with rather than leaving them frozen: a frozen box keeps its slot until it is stopped. `sandbox_list` shows every box of the account, so you can find forgotten ones, or pick up a box another conversation left by its `name` or `goal` (see Picking up a box another conversation left).
3. Usage is billed in credits by the minute, with no time limit: 2,000 credits run a size-1 box for about 69 hours. A bigger `size` costs proportionally more, a frozen box bills nothing, and outbound traffic is metered too. `idleTimeoutMin` replaces the idle limits of rule 6: it stops the box for good after that many minutes without use (rule 6), frozen or not, and a person's use of the box's URLs does not reset it (their requests keep the box up for at most 2 hours after the last use; then it freezes and, past its limit, is stopped right away; a page load that wakes it counts as use and restarts the count); something that keeps the box awake (a background command, for up to 1 hour after the last use; a takeover; a link from another box) holds the stop off until it ends. Rates: a running size-1 box costs about 29 credits an hour; an Android emulator 55 an hour; outbound traffic (every byte the box sends, including through environment connections and links and the pages it serves through its URLs) 114 per GB; files kept by `sandbox_get` (each call stores another copy) and the build cache 25 per GB per month, kept until the account is deleted (the build cache at most 20 GB), and the box pictures the app shows the same (a new one for each freeze with something on the screen and each `sandbox_review`; the replaced ones stay stored); screenshots and recordings the same, deleted after 7 days (the mp4 of a recording stopped for you, once fetched with `sandbox_get`, is a `sandbox_get` file); published images in the registry are not metered.
4. When the account runs out of credits, boxes are frozen (`/work` kept, not billed); a box a person has taken over keeps running and billing. `sandbox_start`, and any call that would wake a frozen box, fail with `no credits left` (HTTP 402) and say until when the boxes are kept: 24 hours after credits ran out, every box is reclaimed unless the person adds credits in the ParallelSandbox app. Tell the person right away; once credits are added, the next call wakes the box where it left off.
5. If a tool returns `interrupted`, the machine under the box is gone (reclaimed by the cloud, or its host failed). Start a new box and redo the steps from the beginning; do not retry on the old id. (A machine the cloud reclaims with notice has its boxes frozen and saved first; the next call thaws them on another host.) When the error says `Its /work was saved at …`, `sandbox_pull` or `sandbox_get` with path `/work` on the old id still returns that copy as `work.tar.gz` (without `node_modules`, caches and files over 64 MiB; later changes are lost).
6. A box is `frozen` after 10 minutes without activity: its memory is snapshotted, `/work` is kept and it stops billing. Activity is use, or a request through the box's URLs for up to 2 hours after the last use, so a person using `webUrl` keeps the box awake for that long; the app's live view of the screen and its live thumbnails do not (they are streamed, not stored). Use is a tool call that acts on the box, a person's takeover, a link connection from another box, or a wake; the last use is `lastUsedAt` in `sandbox_status`. The calls that act on the box are `sandbox_start`, `sandbox_exec`, `sandbox_sync`, `sandbox_get` (so `sandbox_pull` too), `sandbox_shot` (except with `urls`), `sandbox_scene`, `sandbox_wire`, and `sandbox_secrets` and `sandbox_build` for a box; `sandbox_status`, `sandbox_list`, `sandbox_say` and `sandbox_feedback` are not use, nor are `sandbox_review`, `sandbox_procs`, `sandbox_device`, `sandbox_takeover` and `sandbox_shot` with `urls`, though these wake a frozen box first (except `sandbox_device` `list`), and the wake counts. Past that window, someone still using a page may see the box freeze mid-session; their next page load wakes it. It does not freeze while a person has taken it over, a foreground command is running, or another awake box has a link connection open to it, nor while it is recording, until a full hour has passed since the last use (how long it has recorded does not matter): then, once nothing else keeps it awake, the recording is stopped without uploading (the mp4 stays at `/work/.sbx/rec/` and is not shown in the app; fetch it with `sandbox_get` after the box wakes, before the box is stopped). A command started with `background: true` (a dev server, a long build) keeps it awake only until 1 hour after the last use; then the box freezes anyway and the process carries on when the next call thaws it. A `sandbox_review` waiting for the person keeps nothing awake: the box and its phones freeze after 10 idle minutes and are stopped by the same rules as any idle box; the person opening the review wakes a frozen box, and once the box has been stopped the review is over and the version has to be handed over again on a new box. A person trying a box needs nothing from you: when it has frozen, loading one of its pages in a browser wakes it (they see "Waking this box…" and the page reloads itself into the app within about 10 seconds; the wake counts as use; out of credits, it says the box is paused). `fetch`/XHR, assets, WebSockets and POSTs to a frozen box get 503 `box is frozen` and wake nothing, so an open tab that polls neither wakes the box nor keeps it awake past the 2 hours. Its URLs stay the same across freeze and thaw. An attached phone freezes only when the box is about to, just before it (whatever keeps the box awake keeps its phones running), and comes back with the box, with the same `deviceId` and `serial`: frozen, it holds no device slot and does not bill; an Android emulator resumes from a snapshot of its memory, an iPhone simulator boots again with its apps and data kept. Besides your next call and a page load, a person can wake it with the Show screen button on the box's page in the app (there when something was on its screen as it froze), and anything with the API key with `POST /v1/boxes/{id}/wake`; a wake counts as use. After a thaw, boxd closes the box's connections to environment addresses and links that were open before the freeze, so programs have to reconnect. Any call naming the box except `sandbox_status`, `sandbox_say`, `sandbox_feedback`, `sandbox_stop` and `sandbox_device` with `list` thaws it first, which adds a couple of seconds (about a minute with phones attached, which come back first); nothing is lost. `sandbox_stop` works on a frozen box too. Unless `idleTimeoutMin` is set, a frozen box is stopped for good 24 hours after its last use (after it was frozen, if that is later). `sandbox_status` and `sandbox_list` give `stopsAt`, when the idle rules will stop the box unless something uses it first; from 30 minutes before that, every tool result for a box this conversation used (or the call names) ends with a warning saying when, and that a call acting on it pushes it back (`sandbox_status` and `sandbox_list` do not). A frozen box stopped by these rules keeps its frozen copy 24 more hours (`restorableUntil` in `sandbox_status`, and in `sandbox_list` with `all: true`): `sandbox_start { "restore": "<id>", "goal": "<its goal>" }` brings back that same box as it froze, with `/work`, its running programs, its id and its URLs (attached phones are not kept; attach new ones). Only a box the idle rules stopped can be restored; `sandbox_stop` deletes everything at once.
7. Secrets enter the box as environment variables only when named in `sandbox_start.secrets` or injected later with `sandbox_secrets { "id": "<id>", "names": [...] }`. Never paste secret values into commands; use `$NAME`. Never put your ParallelSandbox API key or token in a box.
8. Use `sceneUrl`, `webUrl` and `services[].url` exactly as `sandbox_start` or `sandbox_status` returns them; they carry a random key, so they cannot be built from the box id; take each one from the result rather than deriving it from another. There is no login in front of them: the key protects them, and anyone with one of these URLs can use that service and whatever it reaches (the person's dev, when the box has an environment). An account can also limit box URLs to an IP allowlist (`PUT /v1/box-access`, see [Other facts](https://parallelsandbox.com/en/docs/tools/other-facts/)); when `sandbox_start` or `sandbox_status` returns `boxUrlAccess`, a request from any other IP gets 403, including the box calling its own URL, so test from inside the box with `localhost` or the service name, and tell the person whose IP is not on the list why the page does not open. Share them only with people who should have that; a chat app that builds a preview of a pasted URL fetches the page too, which counts as use and can wake the box. They stop working when the box stops. Each box's URLs have a random host of their own, so for sign-in (SSO, OAuth) on a box, run its copy with the team's local login (a seeded test user, a dev-only auth mode) or add the box's exact URL to a dev client's allowed redirect URIs while it runs; never register a wildcard such as `https://*.box.parallelsandbox.com/*`, because every account's boxes share that domain.

## One pass

1. `sandbox_start` with `name`: whose box it is and what it is for, in the person's language (`"ana: membercenter refund flow"`); the ParallelSandbox app shows it as the box's title. `goal` is required: why the box exists and what done looks like, one or two sentences in the person's language (`"Make the member center refund button post the refund to the ledger; done when a refund goes through end to end in the browser"`), at most 500 characters. The app shows it on the box's card, and `sandbox_status` returns it so another conversation can finish the work if this one is closed. Add `services`: every name and port the repo exposes on the box, for example `[{ "name": "web", "port": 5173, "web": true }, { "name": "api.acme.internal", "port": 8080 }]`. From the moment the box starts, each name has an address of its own inside the box, containers included, and `name:port` reaches your process on `targetPort` (default: `port`): callers keep dialing `name:port`, and the process listens on `targetPort` on `0.0.0.0`, not only on `127.0.0.1`. Ports 80 and 9095 in the box belong to boxd (the scene entry and its API) and no process can listen on them; a service callers reach on port 80 declares `targetPort`, such as `{ "name": "admin.acme.internal", "port": 80, "targetPort": 8082 }` (without it `sandbox_start` fails with 400 and says so). `web: true` marks each service a person can open in a browser to use the product; leave it off for APIs. Each web service gets its own URL in `services[].url`, and `webUrl`, the first of them, is what the Use it button in their app opens, wherever that service sits in the list; `sceneUrl` opens the first declared service, web or not. An entry with `fromBox` is a link to another of your boxes instead of a service in this one (see Linking boxes); an entry with `version` runs a published version under that name (see Versions). Requests through any of these URLs reach the process with `Host: 127.0.0.1:<targetPort>`, the public host in `X-Forwarded-Host` and `X-Forwarded-Proto: https`. Those pages run in the person's browser, outside the box: for their code to reach an API in the box, have the front end's dev server proxy a same-origin path to the API's name (Vite: `server: { proxy: { '/api': 'http://api.acme.internal:8080' } }`, which keeps the `/api` prefix, as a load balancer's path rule usually does; strip it only where dev strips it); a cross-origin call reaches it only if the API is marked `web` too, and then CORS applies. Add `secrets` names if the repo needs credentials, `environment` to run against the person's own dev (see below), `externalBaseUrl` if unchanged HTTP services live elsewhere (see step 4), `size` and `toolchains` if the build needs a bigger machine or a language that is not on by default, `orientation: "landscape"` for a 1280x800 screen when what you show on it is a desktop layout (otherwise it is 390x844 portrait, unless the person's account prefers landscape), `idleTimeoutMin` if you might forget the box. You get `id`, `sceneUrl` (`https://<id>-<key>.box.parallelsandbox.com`), `webUrl`, `services` with each web service's `url` (`https://<id>-<key>-<targetPort>.box.parallelsandbox.com`, or `sceneUrl` itself for the first service; see rule 8), `takeoverUrl`, `status`, `arch` (`amd64`), `orientation` (with `orientationNote` when portrait was asked for but the box's image could only start landscape, 1280x800), `startedFrom` (`snapshot`: restored from a snapshot on a microVM host, in a few seconds) and `next`, a hint for what to do next. If `status` is not `ready`, call `sandbox_status` until it is.
2. Put the code in:
   - `sandbox_exec { "id": "<id>", "cmd": "git clone https://github.com/<owner>/<repo>.git", "note": "clone the repo" }` for committed work. Private repos: store a token as a secret and clone with `https://x-access-token:$GITHUB_TOKEN@github.com/...`, then `git remote set-url origin` to the clean URL.
   - `sandbox_sync { "id": "<id>", "localPath": "/absolute/path/to/repo", "dest": "<repo>" }` for uncommitted work (adapter only; a relative `localPath` resolves against the adapter's own working directory, so pass an absolute one). `dest` is relative to `/work` (`app`; `/work/app` works too, other absolute paths are refused before anything is sent). In a git working tree it sends what git tracks plus untracked files `.gitignore` does not exclude, so build output and `node_modules` stay home; the result's `skipped` lists top-level names it left out. It merges over what is already in the box: only files whose content differs are written, so a running dev server sees just your edits; `sentPaths` lists what was sent and `changedOnBox` how many files the box actually rewrote (0: the box already had them). A change to `package.json`, `vite.config.*`, `tsconfig*` or `.env*` makes a running dev server reload the whole page or restart, and the result's `notes` says so. Sync again after every local change. If a sync fails, the result says `NOT SYNCED` and the box still has the old files; the next `sandbox_exec` on that box carries a warning too, so sync again before trusting a test run.
   - One file: `localPath` may be a file. `dest` is then its path in the box (`"dest": "renderer/.env"` writes `/work/renderer/.env`), unless `dest` ends in `/` or is already a folder in the box, which then receives it; `remotePath` in the result says where it landed.
   - Deletions: files in `dest` that `localPath` does not have (deleted or renamed locally, or no longer in the commit you send) stay in the box, where a type check or test keeps reporting errors from them. The result lists them in `staleInDest` (`count`, the first 20 `paths`); pass `"prune": true` to delete them (`pruned` says how many). Paths the local ignore rules skip (`node_modules`, build output, `.env`) are never deleted; `prune` needs a git working tree or `commit`, a `dest` below `/work`, and no `includeIgnored`.
   - A checkout that other people or conversations also edit: send `"commit": "HEAD"` (any revision; one that does not exist is an error) and list your own uncommitted files in `"alsoPaths": [...]`, which are taken from the working tree instead (a directory replaces the commit's copy as a whole, picked the way a normal sync picks files): their half-finished edits stay out. Without `commit`, the result's `uncommitted` lists the paths that differ from `HEAD` and went over as they are on disk (`count` and the first 20). `alsoPaths` without `commit` sends only those paths (files or folders, such as the output of `git diff --name-only`), each to the same place under `dest`. A `git worktree` of your own works as `localPath` too. `"includeIgnored": true` also sends what the ignore rules skip.
   - Git submodules: without `commit`, an initialized submodule goes over whole as it is on disk, ignored files included, on every sync. With `commit` (and in a `baseline`), add `"submodules": true` to send each submodule at the commit the tree records, taken from its local checkout (initialize it first: `git submodule update --init`); the result lists them in `submodules`, and those it could not send in `submodulesMissing` with the reason. Without it they are empty directories. No `.git` is sent, so git commands in the synced copy (`git submodule update`, `git describe`) fail there; `git clone` when the build needs them.
   - When a test fails and you do not know whether your change caused it, add `"baseline": "<revision>"` (`HEAD`, `origin/main`, your branch's merge base): the same call also sends that commit's tree to `<dest>-baseline`, and the result's `baseline` describes that copy. Run the same command in both directories and compare the failures: what fails in both was failing before your change. To show that a new test catches the bug, copy it into `<dest>-baseline` and watch it fail there. The baseline tree holds only what that commit tracks: install its dependencies in it separately (one `node_modules` symlinked into both makes their Vite caches overwrite each other) and give it its own test database (`psbx-testdb up base`).
3. Build and run. Compile and install steps go through `sandbox_build` so later boxes skip them: `sandbox_build { "id": "<id>", "dir": "<repo>", "cmd": "npm ci", "out": ["node_modules"], "note": "install dependencies" }` — the first box pays the time, every later box with the same inputs gets the output back in seconds (`reused: true`), and `env` is part of the fingerprint. Then run: `sandbox_exec { "id": "<id>", "cmd": "docker compose up -d --build --wait", "cwd": "<repo>", "timeoutSec": 600, "note": "start the refund service stack" }`. `sandbox_exec` and `sandbox_build` require a `note`: one short sentence, in the language the person reads (the tool description names it), saying what the step is for; the app shows it instead of the raw command, as what the box is doing right now. Services without compose: build as usual and run with `background: true`; you get a `bgId` and a `logPath` (`/work/.sbx/logs/<bgId>.log`), which you read with another `sandbox_exec` (`tail -n 100 <logPath>`). Anything that should keep running (a dev server, a watcher, a database, an emulator) goes in its own `background: true` call, never behind `&`, `nohup` or `setsid` in a foreground command: the foreground call returns once its shell exits (within 3 seconds, even if what it started still holds its output) with `leftoverChildren: true` and a `hint`; those processes keep running, but nothing tracks them (`sandbox_procs` knows only `background: true` commands), whatever they print after that is lost, and a timeout (`timeoutSec`) kills the whole process group with them. `exitCode` is the last command's only, so a failing test piped to `tee` or followed by another line reports 0: pass `"strict": true` to prepend `set -eo pipefail`. Commands run as root; `"user": "tester"` runs one as a non-root user with its own `HOME`, `GOPATH` and npm cache, for tests of programs that switch uid (`setpriv`, per-user isolation) or refuse to run as root (root-owned files under `cwd` are handed to it first, so give `cwd` the repo rather than all of `/work`).
4. Wire, only when the box has an `externalBaseUrl` (passed at start, or inherited from `environment`; `sandbox_status.externalBaseUrl` shows it). Then every declared service starts as `external`: `name:port` forwards HTTP to `externalBaseUrl`, keeping the path and rewriting Host, and so does `localhost:<port>` from the box, except on 80 and 9095. boxd does not hold the port, so start your process on the service's `targetPort` whenever you like, and `sandbox_wire { "id": "<id>", "service": "<name>", "mode": "box" }` sends the name to it; names that share a `targetPort` switch together, and open connections are not cut. Without `externalBaseUrl` every service is already `box`; skip this step. To declare one more name on a running box: `sandbox_wire { "id": "<id>", "service": "<name>", "port": 9090 }`, adding `"targetPort"` when the process listens on another port (required for port 80). Host names only, and not one of the environment's addresses, which must be declared at start. It takes the mode of the names already on its `targetPort` (otherwise `external` when the box has an `externalBaseUrl`), stays with the box (it shows in `sandbox_status.services` and survives a boxd restart) and gets no URL. With `fromBox` the same call adds or re-points a link instead (see Linking boxes). The result is the whole wiring table, with each name's `port`, `targetPort`, `address` and `mode`; check with `sandbox_exec { "id": "<id>", "cmd": "curl -s http://<name>:<port>/", "note": "check the service answers" }`.
5. Test: run the repo's tests with `sandbox_exec`. When the picture on the box's screen changes during a foreground command, the step keeps a screenshot taken 0.6 seconds after the command ends (right away if your next foreground command starts sooner) and, usually, a video from when the change is seen until that screenshot; the person sees them under the box's Screenshots & recordings in their app, labeled with your `note`. So run browser and app tests headed on the display (`HEADED=1` or the test runner's own flag, `DISPLAY=:99` is set): each step then shows the person what was tested. The screen is compared with how it looked before the command every second (every 2 with a phone) and once more 0.6 seconds after the command ends (sooner if your next foreground command starts first); a blinking cursor is not a change. So a command that changes nothing leaves nothing; one over within a second (a click) leaves only the screenshot, or nothing if the page changes later than that, so wait in the same command (`xdotool mousemove 400 300 click 1; sleep 2`; with a phone, wait before and after the tap too: `sleep 3; psbx-ios tap 200 400; sleep 3`); a command with `background: true` is not recorded; whether the step records its own video is decided once, when the change is first seen: if your own `sandbox_shot` recording is running at that moment, the step gets no video of its own, then or later, only the screenshot after it ends (what changes on the screen during the command is only in your recording, which is never added to the step: if you end it yourself with `record: "stop"`, it shows in the app as a separate step; if it is stopped for you, it is not uploaded, becomes no step at all, and the person sees only the step's screenshot); a step records at most 10 minutes from when its recording starts. The app shows them only while the box is not stopped, for at most 7 days (then the files are deleted), and a frozen box nobody uses is stopped 24 hours after its last use (or after it froze, if that is later), even one waiting under Ready for you. When the person should see the steps, finish with `sandbox_review`, not `sandbox_stop`.
6. Test it, then hand it over. Go through the flow where the person will try it: tap and type on the box's screen for a phone or desktop app (`sandbox_shot` shows the result), or use the web page from a fresh browser session, including its sign-in entry. Fix failures, try again and leave the product at its starting screen. Call `sandbox_review { "id": "<id>", "what": "<what they should test first: where to start, what to tap, what they see when it works; then what you tested>", "open": "web", "port": 5173 }`. `open` says what they open: `web` with `port`, the port the page is served on in the box, which they open in their own browser on their own device (a desktop app whose UI comes from a dev server can be handed over this way too, with that server's port; the port gets its own URL if it has none yet), or `box`, the box's screen, where they tap and type on a desktop app on `DISPLAY=:99` or a phone app on a device; keep that app running. When they can try it in a browser, hand over `web`: it is smoother for them. Write `what` in the person's language and lead with what they should test, because the card shows only its first two lines: where to start, what to tap or type, and what they see when it works. Then add one short sentence on what you tested yourself ("Type a task and press Start: the countdown begins with your task shown under it. I started a round with my own task on the box screen."). The card appears immediately with your text, a picture and the product entry. The call waits within the client's supported limit: up to 25 seconds unless a longer limit is declared, or up to 30 minutes with the host and adapter configured (see When a person is needed). Progress carries the review ID and URL. Read the returned `fromHuman` report, its annotated images, recording URLs and timed transcripts, then continue the task. After timeout or cancellation, resume the same round with `sandbox_review { "id": "<id>", "reviewId": "<reviewId>" }`. Read a known report again with `sandbox_report { "id": "<id>", "reportId": "<reportId>" }`; reports and their media are kept for seven days. Set `waitSec: 0` for an explicit handoff that returns immediately. A new `what` creates the next version's card; `reviewId` resumes the existing round. Hand over each tested version worth their time, and finish with `sandbox_review` when the person should try the result.
7. Show your work: `sandbox_shot` for a screenshot (also returned inline); `record: "start"` before a run and `record: "stop"` after for an mp4 (both also show in the person's app as steps of the box while it is not stopped, a recording only once you stop it yourself, and their files are deleted after 7 days; past the hour, `GET /v1/boxes/{id}/media` with your token returns the newest 200 steps with pictures and signs their URLs again, for a stopped box too; it is REST only, with no MCP tool); `sandbox_get` for any file or directory under `/work`, `/tmp` or `/dev/shm` (a path is relative to `/work`, and `/work/...` works too; directories come as tar.gz; `paths` fetches several at once and shows the first 10 images, 8 MB in all; each call stores another billed copy, except that two same-named files fetched within the same second are stored as one, both URLs then giving the later upload, so fetch those in separate calls a second apart). All three return download URLs valid for one hour; put them in your answer. Once the box is stopped, no new URL can be had for a `sandbox_get` file (an unexpired one still works), so fetch what must be kept before you stop it. `sandbox_pull { "id": "<id>", "path": "<repo>/dist", "localPath": "/absolute/path/dist" }` (adapter only) writes a file or directory from the box straight to a local path instead.
8. `sandbox_feedback` once, before you stop or hand your work back: what you ran, how, what got in the way (exact errors; `"none"` if nothing) and what should change.
9. `sandbox_stop` as soon as you are done with the box and there is nothing for the person to try or look at (end with `sandbox_review` when there is; the steps' screenshots and recordings count, since the app stops showing them once the box is stopped), unless a person is still trying the box's URLs (rule 2).

## Picking up a box another conversation left

A box outlives the conversation that started it. To continue one:

1. `sandbox_list`: find the box by its `name` or `goal`; `agent.state` `left` or `idle` marks the boxes nobody is working on (`idle` goes by time alone: the conversation may still be open, but has not touched the box for an hour).
2. `sandbox_status { "id": "<id>" }` (it only reads, and does not make the box yours): `agent.state` `working` or `away` without `thisConversation` means the conversation that was using the box is still open and may come back to it, so ask the person in this conversation before taking it over. `goal` says what the box is for, `steps[]` the last 30 things done in it, oldest first, never their output (kept 7 days). Each step has `at`, `kind` (`exec`, `build`, `sync`, `wire`, `shot`, `device`, `start`, `stop`, `takeover`, `review`, `return`, `say`), `summary` (the command, at most 200 characters), and when they apply `target` (the directory uploaded or built), `note`, `actor` (`human`), `detail` (`background`, `reused`, `record` for a `sandbox_shot` recording, a wire mode), `ms`, plus `ok`.
3. Carry on with the same tools and `id`: `/work` still holds what the earlier conversation left, unless the box has stopped; then start a new box with the same `goal`. `steps[]` has no output and no files: run `git status`, `git log` and `docker ps` in `/work` to see where the work stands; edits the earlier conversation made on the person's machine and synced are in their local checkout too. A box nobody uses is stopped for good 24 hours after its last use (or after its `idleTimeoutMin`), so pick it up before then.
4. End with `sandbox_review` or `sandbox_stop`.

In the person's app, a box whose conversation closed, or that went unused for 1 hour, without either shows as AI stopped, with Continue (copies a prompt telling an AI to do exactly this) and Shut down (stops the box). The adapter (0.3.0 or later) is how the app knows: it sends a random id as `X-Psbx-Agent` on every call, `POST /v1/agents/<id>/heartbeat` every 60 seconds after the first tool call, and `POST /v1/agents/<id>/leave` when the conversation closes; a conversation killed without a leave counts as gone 3 minutes after its last heartbeat. Nothing else is sent: no prompt, no transcript. The app ties a box to the last conversation that called a tool acting on it, so a box you pick up counts as yours from your first such call (`sandbox_exec`, `sandbox_sync`, `sandbox_get`, `sandbox_shot` and the like; not `sandbox_status`, `sandbox_list`, `sandbox_say` or `sandbox_review`). `sandbox_list` and `sandbox_status` carry the same `agent` object as the app (`state` `working`, `away`, `left` or `idle`, `idleSec`, and `thisConversation` when it is you), so a teammate without the app finds the boxes waiting for someone through their own agent. Since a conversation can end before it finishes a box (a laptop shut), commit in `/work` as you go, with messages that say what is done and what is left.

## Tools

| Tool | Use it for |
|---|---|
| `sandbox_start` | new box, titled by `name` in the app; `goal` (required) says why it exists and what done looks like; returns `id`, `arch` (`amd64`), `sceneUrl`, `webUrl`, `services` (each web service with its `url`), `takeoverUrl`, `status`, `orientation` (and `orientationNote` when portrait was asked for but the image could only start landscape), `startedFrom`, `next`; `restore` with the id of a box the idle rules stopped brings that box back within 24 hours |
| `sandbox_exec` | one shell command; `cwd` relative to `/work`; `timeoutSec` for foreground (default 600, max 3600; when it runs out the whole process group is killed: `timedOut: true`, `signal: "SIGKILL"`); `background: true` returns `bgId` and `logPath` (`/work/.sbx/logs/<bgId>.log`), the only way to leave a server running (what a foreground command starts with `&` is left untracked: `leftoverChildren: true`); `strict: true` prepends `set -eo pipefail`; `user: "tester"` runs as a non-root user; values of secrets and environment settings come back as `****`; `note` says what it is for |
| `sandbox_procs` | the commands started with `background: true`: `action: "list"` (the default) returns those still running, oldest first, each with `bgId`, `note`, `cmd` and `startedAt`; `all: true` adds finished ones with their `exitCode`, `bgId` returns just that one with the end of its log (`logTail`, 4 KB unless `tailBytes`, up to 16384); `"wait"` for one to end (`bgId`; `timeoutSec` up to 600, then `stillRunning: true`), `"stop"` one with every process it started (its process group). Use it rather than `pkill -f` or `pgrep -f`: those no longer match the shell running your own command, but they match every other process whose command line holds the pattern, other jobs included |
| `sandbox_sync` | local directory or one file into `/work/<dest>` (adapter only): `commit`, `alsoPaths`, `includeIgnored`, `prune` (delete what `staleInDest` lists), `submodules` (with `commit`), `baseline` (a second copy of an earlier commit in `<dest>-baseline`); returns `remotePath`, `sentPaths`, `changedOnBox`, `staleInDest`, `uncommitted` |
| `sandbox_build` | a compile or install step whose output is reused: the same `dir` contents, `cmd`, `env` and toolchains restore the previous `out` instead of building again; `reused: true` means nothing compiled, `cached: true` means this run's output was saved for later boxes; `note` says what it is for |
| `sandbox_say` | tell the person watching the box in the app what you are doing or what you changed plan to (a status line: ask questions in this conversation) |
| `sandbox_get` | file or directory under `/work`, `/tmp` or `/dev/shm` out, as a download URL; `paths` for several at once; a single image up to 4 MB also comes back inline, and with `paths` the first 10 images, 8 MB in all |
| `sandbox_pull` | file or directory under `/work`, `/tmp` or `/dev/shm` out, written to a local path (adapter only) |
| `sandbox_wire` | point a declared name at your process in the box (`mode: "box"`: `name:port` reaches its `targetPort`) or at `externalBaseUrl` (`mode: "external"`), or declare a new name with `port` (and `targetPort`), or link a name to another of your boxes with `port` and `fromBox`, or run a published version under a name with `port` and `version`; returns the whole wiring table, `links` for a link, or the `service` with its `run` for a version |
| `sandbox_shot` | screenshot of the display (`target: "screen"`), of the front window (`target: "window"`; the result's `window` is its `x,y,width,height` on the screen, so add x and y to image coordinates before clicking), of `url` opened in a fresh headless Chromium (`target: "url"`, with `width`/`height` for the viewport, by default the box's screen size, such as 390×844 for a phone layout, `deviceScaleFactor`, `mobile` and `fullPage`), or of a tab already open, and signed in, in a browser on the box that was started with `--remote-debugging-port=9222` (`target: "tab"`: `width`, `height`, `deviceScaleFactor` and `mobile` give it a viewport the box keeps until `reset`); `waitFor` (url, tab) waits for a CSS selector or a `js:` expression, `stableMs` (screen, window) for the picture to settle; `capturedAt` says when it was taken; `urls` captures several pages into `/work/.sbx/shots/` (fetch them with `sandbox_get`; they are not shown in the app); `record: "start"` / `"stop"` for mp4 |
| `sandbox_scene` | run JavaScript in the page a person has open at any of the box's URLs (`sceneUrl` or a web service's `url`) and get the result |
| `sandbox_status` | state, `goal`, `steps[]` (the last 30 things done in the box, oldest first: commands, directories and their notes, no output), `agent` (whether the conversation that last used it is still at it; `thisConversation` when that is you), boxd health (uptime, idle time, display, private endpoints), wiring (each name's `port`, `targetPort`, `address`, `mode`, `web`), `sceneUrl`, `webUrl` and each web service's `url`, each published version's `run`, `environment` (as `sandbox_start` returns it, read again on each call), `links` (this box's links to other boxes) and `linkedFrom` (boxes linked to this one), what is running (`running.background` with log paths, `running.containers`), registry login, `takeoverUrl`, the open takeover, interrupted reason, credits left, `stopsAt` (when the idle rules stop it for good unless it is used before then) and, on a stopped box that can still be restored, `restorableUntil` |
| `sandbox_review` | pass `id`, `what` and `open` for a new tested version (`web` with `port`, a page in their browser; `box`, the box's screen); the card appears immediately; wait up to the client's declared limit (25 seconds when undeclared, up to 30 minutes when configured); after timeout resume with `id` and `reviewId`; `waitSec: 0` explicitly hands off and returns immediately |
| `sandbox_report` | pass `id` and a stable `reportId` to read the complete human report again, including notes, annotated images, recordings and timed transcripts; file URLs are refreshed |
| `sandbox_takeover` | ask a person for what only they can do (signing in, a CAPTCHA, a code); blocks until they hand back (max 30 minutes) and returns their message |
| `sandbox_secrets` | names of stored secrets, never values; with `id` and `names`, inject them into a running box |
| `sandbox_environments` | the person's environments (their own dev or staging): private `host:port` a box can reach, whether a connector is online, each service's settings files |
| `sandbox_publish_version` / `sandbox_versions` | register an image you pushed to the account's registry as a named version; list versions; another box runs one with `services[].version` |
| `sandbox_device` | lease an Android emulator or an iPhone simulator for the box (`action: "attach"`, `platform: "android"` or `"ios"`, optional `model` and `osVersion`; a version the hosts lack is an error), `list` them (with each platform's device `hosts`: online, free slots), `reconnect` an Android one that `list` shows `offline`, `release` one by `deviceId` |
| `sandbox_feedback` | report to the ParallelSandbox team what you ran, how, what went wrong and what should change |
| `sandbox_list` | every box of the account, teammates' included, most recently used first (up to 100): `id`, `name`, `goal`, `status`, `agent` (whether the conversation that last used it is still at it), `sceneUrl`, `webUrl`, `services`, `linkedFrom`, `lastUsedAt`, `stopsAt`, `restorableUntil`; live boxes only unless `all: true`; read-only, never wakes a box |
| `sandbox_stop` | destroy the box; call it as soon as you are done and there is nothing for the person to try; lists the boxes that link to it, whose links then fail until re-pointed |
| `logs_search` / `logs_errors` / `logs_tail` | logs from pages using `@parallelsandbox/log`; `logs_errors` has stacks resolved through uploaded source maps |

## Seeing the screen

Anything a command starts shows on the virtual display (`DISPLAY=:99` is already set). It is 390x844 portrait unless the box was started with `orientation: "landscape"`, which makes it 1280x800 (`PUT /v1/boxes/{id}/orientation` over REST switches a running box). A Chromium window is at least 500 px wide, so on the portrait screen scale it down to fit. To look at a web page on a portrait box:

```
sandbox_exec { "id": "<id>", "cmd": "chromium --no-sandbox --kiosk --force-device-scale-factor=0.78 --window-size=500,1082 --window-position=0,0 --remote-debugging-port=9222 --user-data-dir=/tmp/chrome http://localhost:3000/", "background": true, "note": "open the app on the box screen" }
sandbox_shot { "id": "<id>", "target": "tab", "width": 390, "height": 844, "mobile": true, "deviceScaleFactor": 3, "waitFor": "#app" }
sandbox_shot { "id": "<id>", "stableMs": 500 }
```

On a landscape box, use `--window-size=1280,800` in place of the three scale, size and position flags. On the portrait screen the window lays a page out 500 CSS pixels wide. For an exact phone layout of a page you keep using (signed in, mid-flow), start the browser with `--remote-debugging-port=9222` as above and call `sandbox_shot` with `target: "tab"` and `width` (plus `height`, `mobile`, `deviceScaleFactor`): the box gives that tab the viewport and keeps it until `"reset": true`, the tab closes or the browser exits, so the phone layout holds while you click and type in it; one tab is held at a time, each such call replaces the whole emulated viewport, and `tab` (an id, or a piece of the tab's URL or title) picks another tab, `port` another browser. `target: "tab"` waits up to 45 seconds for a browser that is still starting, does not scroll or move the pointer, and works with Electron started with the same flag. For a signed-out page, `target: "url"` lays it out at exactly `width`, below 500 too. Click with `xdotool mousemove X Y click 1`, type with `xdotool type "..."`, using coordinates from the screenshot.

Wait for a condition, never a fixed `sleep`: `waitFor` (with `target: "url"` or `"tab"`) captures once a CSS selector matches a visible element, or, prefixed with `js:`, once an expression is truthy, waiting up to 30 seconds (if it never holds you still get the picture, with `waitFor.met: false`); `stableMs` (with `target: "screen"` or `"window"`) captures once the picture has not changed for that long, up to 20 seconds (`stable.settled: false` if it never settles); every screenshot carries `capturedAt`. In a script, `timeout 60 xdotool search --sync --onlyvisible --class chromium` returns once the window is up, and `curl -s http://127.0.0.1:9222/json/version` answers once the debugging port is. In browser tests (Playwright, or raw CDP against that port):

- Find elements by role and accessible name, such as `page.getByRole('button', { name: 'Save' })`, which also matches an icon button's `aria-label`, not by visible text or position: an icon-only button has no text, and a locator matching several elements fails Playwright's strict mode.
- One browser drives one page at a time. A tab that is not in front gets no frames, so `page.screenshot` on it times out; call `page.bringToFront()` (CDP `Page.bringToFront`) before working in another tab of the same browser, and give scripts that run at the same time a browser each (their own `--user-data-dir` and port).
- What needs a user gesture (`window.open`, the clipboard, fullscreen) needs real input: Playwright's `click()`, CDP `Input.dispatchMouseEvent`, or `xdotool click`. An `element.click()` sent through CDP `Runtime.evaluate` carries no user activation unless you pass `userGesture: true`, and the popup it opens is blocked without an error.
- The box's Chromium lets `localhost`, `127.0.0.1`, `*.localhost` and `*.parallelsandbox.com` use the clipboard without asking. `navigator.clipboard.readText()` still needs a focused page, so bring it to the front first. Another origin needs the permission granted (Playwright `context.grantPermissions(['clipboard-read', 'clipboard-write'])`; CDP `Browser.grantPermissions` lasts only while that connection stays open); without it, Chromium shows a permission bar over the page and the call waits.

## Testing on a phone

```
sandbox_device { "id": "<id>", "action": "attach", "platform": "android" }
```

About 35 seconds later the box has an Android 15 emulator on a `pixel_9` profile (`model` picks another, for example `pixel_7`). An `osVersion` (API level `35` or version `15`) the device hosts lack fails at once with an error naming the version they have; it is never quietly swapped for another. It runs on separate hardware, but inside the box `adb devices` shows it as `emulator-5554` (a second one is `emulator-5556`), and its screen is on the box display, so `sandbox_shot` shows it. Build the app in the box (`toolchains: ["android"]`, `size` 2 or more) and `adb install -r app.apk`. The box's URLs (`sceneUrl`, `webUrl`, `services[].url`) open on the phone as anywhere else; `adb reverse` (below) reaches any port. A fresh emulator already has Chrome's first-run and notification prompts out of the way.

Drive it with `artemis-adb`, deterministic ADB tools with no model inside:

```
artemis-adb hierarchy --out /work/ui.json   # JSON array of on-screen elements
artemis-adb tap X Y                         # tap the center of an element's parsed_bounds
artemis-adb type X Y "text"                 # also: long-press X Y, back, key KEYCODE
artemis-adb swipe X1 Y1 X2 Y2               # scroll down: Y1 larger than Y2
artemis-adb launch com.example.app          # also: stop PACKAGE, open URL
artemis-adb screenshot --out /work/s.png
```

To iterate without rebuilding the app, run the dev server in the box and `adb reverse tcp:5173 tcp:5173`: the phone's `localhost:5173` is now the box's port 5173 (React Native and Expo already load from `localhost:8081`, so reverse that port). For Capacitor, set `server: { url: "http://localhost:5173/", cleartext: true }` in the box's copy of `capacitor.config`, then `npx cap sync android`, one `./gradlew assembleDebug` and `adb install -r`; after that every edit reaches the phone through the dev server. A debug build's WebView speaks CDP: `adb forward tcp:9333 localabstract:webview_devtools_remote_$(adb shell pidof <package>)`, then `http://127.0.0.1:9333/json` lists its pages, so a script can read the DOM or set localStorage for a test login. `adb exec-out screencap -p > /work/phone.png` saves the phone screen at full resolution.

`sandbox_device { "id": "<id>", "action": "list" }` does not wake the box. It shows each device's `status`, `failed` ones from the last hour with their `reason`, and `hosts`: for each platform whether a device host is `online` and its `freeSlots` (with `lastHeartbeatAt` when it is offline, which is on ParallelSandbox's side; plan without that phone or try later). Check it before building for a phone; `sandbox_start` with `toolchains: ["android"]` also says so in `next` when no Android host is online. An Android device whose adb in the box no longer reaches it shows `status: "offline"` with its `adb` state: `sandbox_device { "id": "<id>", "action": "reconnect", "deviceId": "..." }` re-opens its tunnel and keeps the emulator, its apps and data; if that fails, the emulator itself crashed or hung, so release it and attach a new one.

For an iPhone, `sandbox_device { "id": "<id>", "action": "attach", "platform": "ios" }` boots an iOS 26.5 simulator on a Mac: an `iPhone 17 Pro` unless `model` names another simulator device (`"iPhone 17"`, `"iPad Pro 11-inch (M5)"`), and `serial` is its UDID. Xcode runs on the Mac, not in the box, so you write code in the box, copy it over and build there, all through `sandbox_exec`:

```
psbx-ios sync /work/app                  # /work/app -> work/app on the Mac; prints progress and the Mac's full path
xcodebuild -project app/Hello.xcodeproj -scheme Hello -destination "platform=iOS Simulator,id={udid}" -derivedDataPath dd build
psbx-ios install dd/Build/Products/Debug-iphonesimulator/Hello.app
psbx-ios launch com.example.hello        # the PRODUCT_BUNDLE_IDENTIFIER (xcodebuild ... -showBuildSettings); also: terminate BUNDLE_ID
psbx-ios screenshot /work/phone.png
simctl openurl {udid} myapp://orders/42  # simctl ARGS runs xcrun simctl on the Mac; {udid} is this simulator
```

Every `psbx-ios` word runs on the Mac in the lease's `work/` folder, a copy of the box's `/work` that `psbx-ios sync` fills: the box's `/work/app` is `work/app` there, at the same path (`psbx-ios pwd` prints the folder), so give the other words paths relative to `/work` (`app/Hello.xcodeproj`), never box paths such as `/work/app/...`; `cd` in the box changes nothing there. `psbx-ios sync [path]` (default all of `/work`; a relative path starts from your directory in the box) copies it to the same place on the Mac, printing progress (run a big one with `background: true`) and the Mac's full path; it deletes what the box no longer has there, except `dd/` and `DerivedData/` folders at any depth and `build/`, `*.mp4` and `*.mov` at the top of that path, which stay on the Mac and are not sent. Build products stay on the Mac: `ls` in the box does not see them, `psbx-ios ls dd/Build/Products/Debug-iphonesimulator` does, `psbx-ios install` takes that Mac path (one that is not there says where Mac paths start and to look with `psbx-ios ls`), and `psbx-ios pull <path> [dest]` copies a file or folder from the Mac's `work/` into the box (default `/work/`). `psbx-ios xcodegen <dir>` runs XcodeGen in `work/<dir>`, making the `.xcodeproj` from its `project.yml`. `release` deletes the Mac's copy and what was built, so after a new `attach` sync and build again. Other `psbx-ios` words are `boot`, `shutdown`, `ls`, `pwd`, `xcrun`, and the input commands below; `psbx-ios help` lists them all, and anything else exits 126. A Capacitor or React Native project's Xcode project links `node_modules` by relative path: sync the folder that holds both `ios/` and `node_modules/`. A Flutter app (`toolchains: ["flutter"]`) builds for iOS with the Mac's own Flutter, the same version as the box's: `psbx-ios sync app`, then `psbx-ios flutter app build ios --simulator` (it runs in `app/` on the Mac) and `psbx-ios install app/build/ios/iphonesimulator/Runner.app`; plugins that still need CocoaPods work too. The Mac has CocoaPods and Node on its PATH for every `psbx-ios` word: for React Native, Expo (after `npx expo prebuild` in the box) and Capacitor 6 or older, sync the folder that holds both `ios/` and `node_modules/`, run `psbx-ios pod myapp/ios install` (it runs in that folder on the Mac), then `xcodebuild -workspace myapp/ios/<App>.xcworkspace ...`; build React Native with `-configuration Release` so the JavaScript is bundled into the app, since Metro in the box is not reachable from the simulator. The box display shows the simulator as a live stream of about 6 to 10 frames a second, and a person on it taps, swipes and types, Chinese and emoji included. You can drive it too, in points (402 × 874 on an iPhone 17 Pro; `psbx-ios size` prints pixels and the scale): `psbx-ios tap X Y`, `psbx-ios swipe X1 Y1 X2 Y2 [SECONDS]`, `printf 'hello 你好' | psbx-ios type` (any text, Chinese and emoji included: non-ASCII text goes through the simulator's Unicode helper by itself; `psbx-ios unicode` uses only that helper), `psbx-ios key 40` (Enter; 42 Backspace), `psbx-ios button home`. Taps, swipes, keys and typing retry for about 3 seconds through a screen transition; when they still fail, the message says a system dialog may be in the way, so look with `sandbox_shot`. A deep link or a UI test (`xcodebuild test`) still reaches a screen faster; check it with `sandbox_shot`. To show a native app to the person, leave it running and call `sandbox_review` with `open: "box"`: the card opens the box's screen, where they use the native app themselves.

`sandbox_device { "id": "<id>", "action": "release", "deviceId": "..." }` when you are done; stopping the box releases its devices too. A device bills by the minute while it is attached and running. It freezes with the box (rule 6) and comes back attached, with the same `serial`, on the call that thaws the box; frozen, it holds no device slot and does not bill. Coming back takes about 35 seconds for Android and 50 for an iPhone, in parallel. If its host has no free slot at that moment, it stays frozen until one frees up while the box keeps being used (`sandbox_device list` shows `status: "frozen"` and does not wake the box). A `sandbox_review` of its app does not keep it up: while the review waits, the phone freezes and is released with the box by the usual idle rules, and the person opening the review wakes it with the box.

## When a person is needed

The person may be watching the box in the ParallelSandbox app. `sandbox_say { "id": "<id>", "text": "Building the refund service; about three minutes." }` posts a line there: send one when you start something long or change plan. It is a status line, not a way to ask: when you need the person to decide something, ask them in this conversation, the way your client asks (Claude Code, Codex and the others each have their own), and they answer you there. Notes they leave for you in the app arrive in your next tool result on that box, `sandbox_status` included, as `fromHuman` (a list of messages, each delivered once, and only to the conversation holding the box); images they attach follow the JSON as image content, in the same order, and the message says how many. Read it before you continue.

When the person tries the product on their own phone or laptop (through `sceneUrl`, or `webUrl` from the Use it button), that page is not on the box's display, so `sandbox_shot` and takeover input cannot reach it. `sandbox_scene { "id": "<id>", "js": "document.querySelector('#total').textContent" }` runs JavaScript in the page they have open (the one most recently opened through any of the box's URLs, `sceneUrl` or a web service's `url`) and returns the value of the last expression, awaiting promises: read what they see, or drive the feature while they watch. It fails with a clear message when no page is open.

`sandbox_review { "id": "<id>", "what": "<what they should test first: where to start, what to tap, what they see when it works; then what you tested>", "open": "web", "port": 5173 }` immediately shows a tested version and waits for the report or completion within the client's supported limit. Use stdio adapter 0.4.2 or later for wait negotiation, progress and cancellation forwarding. Undeclared clients, and the adapter's default 60-second host budget, get a wait capped at 25 seconds. A configured client can wait up to 30 minutes. Give the person the progress/result's `appURL` to try it and submit feedback in the app; `url` is the product entry: the page on `port` for `open: "web"`, or the box's screen for `open: "box"`. A web review runs on the person's own device; a screen review gives them input on the box's display and pauses `sandbox_exec` while they use it. Keep the review call active, read the returned `fromHuman` notes, images, recordings and timed transcripts, then continue in this conversation. Resume after a timeout or interrupted wait with `sandbox_review { "id": "<id>", "reviewId": "<reviewId>" }`, or replay a known report with `sandbox_report { "id": "<id>", "reportId": "<reportId>" }`. Use `waitSec: 0` for an explicit immediate handoff. Watching the box's live screen is a separate view for following your work. `sandbox_takeover` asks for their hands on that screen to complete a step you need before continuing, and waits for them to hand it back.

For a full 30-minute wait in Codex, the host tool timeout and adapter timeout must both allow it. When the person requests longer waits, merge these properties into their existing stdio server entry, preserving other properties and any larger timeout already set:

```toml
[mcp_servers.parallelsandbox]
tool_timeout_sec = 1920
env = { PSBX_TOOL_TIMEOUT_SEC = "1920" }
```

Merge the environment variable into an existing `env` map and reconnect. Adapter 0.4.2 declares the safe wait limit to the server; setting `waitSec: 1800` alone still respects the client's limit. With the normal shorter wait, show `appURL` and retain `reviewId`; the card and feedback remain available for a later resume or `sandbox_report` read. Configure longer waits only as part of the person's requested client setup.

Logins, two-factor codes, CAPTCHAs, payments, anything you must not do alone: `sandbox_takeover { "id": "<id>", "note": "<what to do and why>" }`. The person sees your note and the live screen in the ParallelSandbox app; the call blocks until they press Hand back (or 30 minutes pass) and returns the message they leave. While a person is connected, `sandbox_exec` returns an error saying a person is operating the box; wait for the takeover to finish instead of retrying in a loop. If your client cuts the call off, the takeover stays open: `sandbox_status` shows it, and calling `sandbox_takeover` again with the same `id` resumes waiting. `takeoverUrl` from `sandbox_status` is the same live screen for a person you tell directly.

## Logs and console output

Without any log service:

- A `background: true` command returns its `logPath`, `/work/.sbx/logs/<bgId>.log`; read it with `sandbox_exec { "id": "<id>", "cmd": "tail -n 100 /work/.sbx/logs/<bgId>.log", "note": "read the dev server output" }`. `sandbox_status.running.background` lists the ones still running, with `logPath`; `running.containers` lists the box's containers with their ports. `sandbox_procs` lists the background commands still running (`all: true` adds finished ones; `bgId` gives one with the end of its log), waits for one, and stops one with everything it started; stop them that way rather than with `pkill -f` or `pgrep -f`, which match every process whose command line holds the pattern, other jobs included (they no longer match the shell running your own command).
- Containers: `docker compose logs --tail 100 <service>` or `docker logs --tail 100 <container>` through `sandbox_exec`.
- Connections to environment addresses: `/work/.sbx/logs/private-endpoints.log`, one line per failure.
- The box does not capture a browser's console or network traffic. In your own tests, listen for them and print them (Playwright: `page.on('console', m => console.log(m.type(), m.text()))`, `page.on('pageerror', e => console.log(e.message))`, `page.on('requestfailed', r => console.log(r.url(), r.failure()?.errorText))`). For a page a person has open at one of the box's URLs, `sandbox_scene` can read its state. `sandbox_shot` shows the screen.

With the ParallelSandbox log service, pages send their console output and errors themselves through the browser SDK `@parallelsandbox/log`, and you query them instead of asking the person to paste console output. A project is created once over REST with the account's token (API key or OAuth access token), outside any box: `POST https://log.parallelsandbox.com/v1/projects` with `{"name": "...", "origins": [...]}` returns `project.id` (`prj_...`) and a write key (`psw_...`, shown once). The page calls `init({ writeKey: "psw_...", release: "<build id>" })`. Then:

```
logs_errors { "project": "prj_...", "since": "30m", "boxId": "<id>" }
logs_search { "project": "prj_...", "since": "1h", "query": "<text>" }
logs_tail   { "project": "prj_...", "boxId": "<id>" }
```

`logs_errors` returns stacks resolved through the source maps uploaded for that release. `logs_tail` is a long poll: it returns when rows arrive or after `waitSec` (default 20), with a `cursor` to pass to the next call. These tools read only ParallelSandbox's log service, not Sentry or any other tracker. Details: https://parallelsandbox.com/en/docs/log-sdk/

## The person's own dev: environments

When the task changes one or two services of a larger system, run only those in the box and let everything else stay on the person's existing dev. `sandbox_environments` (read-only) lists the account's environments; use a name exactly as it lists it (the person picks the names, so "their dev" may be called `lab`, `staging` or anything else). An environment that does not exist yet is created over REST, not through an MCP tool: `POST https://api.parallelsandbox.com/v1/environments`, connections, addresses and service settings, all with your token (https://parallelsandbox.com/en/docs/environments/#set-it-up); only the connector needs someone on their network. Start the box with one:

```
sandbox_start { "name": "refund flow rework", "goal": "Rework the api's refund flow against dev's databases and services; done when a refund made through the api is recorded in dev", "environment": "dev", "services": [{ "name": "api", "port": 8080 }] }
```

- Private addresses in `environment.reachable` (databases, Redis, internal services) connect from anywhere in the box, containers included (default bridge, user-defined networks and `docker compose` alike), through a connector the person runs in their network. Use the host names exactly as the service's settings have them; nothing needs rewriting.
- Each service's settings for that environment are in `/work/.sbx/env/<service>.env` (for `docker run --env-file`) and `.sh` (to `.` before running a process directly). Start the service you changed with its own file, so it runs with the same configuration it has on their dev; override one value with `-e KEY=value` after `--env-file`. To run a program directly, source the `.sh`, then override what the box needs: `. /work/.sbx/env/api.sh && DATABASE_URL="$(psbx-testdb url)" ./bin/backfill`. Do not source the `.env` (`set -a; . api.env`): its values are not quoted, so one with a space, a quote or a `$` breaks the shell or changes silently.
- If the environment has an `externalBaseUrl`, the box inherits it: every declared service starts `external`; `sandbox_wire` the ones you run in the box to `box` (step 4 above).
- Environment addresses on port 80 (an internal load balancer) work like any other, and your own processes can listen on the same port numbers as environment addresses: `localhost:5432` is yours while `db.internal:5432` still goes to their network.
- If other services call the one you changed by an internal host name that is also one of the environment's addresses (a router forwarding to `api.svc.local:8080`, an internal load balancer at `internal-lb.acme.internal:80`), declare it in `services` under that exact name and the port callers dial at `sandbox_start`, with `targetPort` when your process listens on another port (always for port 80: `{ "name": "internal-lb.acme.internal", "port": 80, "targetPort": 8080 }`). Added later with `sandbox_wire` (`port` or `version`), such a name is taken over in place only on boxes whose `sandbox_status` → `health.features` includes `take-over` (every box started from 2026-09-25 12:40 UTC); on older boxes it keeps going to the connector. That name then points at the box, `reachable` leaves it out, and everything else stays on their dev. Only programs inside the box see the change: services still running on their dev keep calling their dev's copy, and nothing on their network can reach into a box.
- When the caller dials addresses from its own registry instead of a host name (a gateway taking worker IPs from Redis), declare a private address range in `services`: `[{ "name": "172.16.0.0/16", "port": 9090 }]`. Connections from the box, containers included, to that range on that port reach your service in the box; `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `100.64.0.0/10` only, and only at `sandbox_start`.
- A box's own internet traffic leaves from its host's public IP, which differs from host to host and changes as hosts come and go, so it cannot be allowlisted. For an outside service that accepts only known IPs, the person lists its host on a connection (`api.partner.example:443`): the box then reaches it through the connector, from their network. Sites that block data-center addresses treat a box as a bot: YouTube playback fails with `Sign in to confirm you're not a bot`. Test such a part by handing the page over with `sandbox_review` and `open: "web"`, so it plays in the person's own browser.
- When the environment names an AWS role, boxes get `AWS_CONTAINER_CREDENTIALS_FULL_URI` already set: AWS SDKs and the CLI in the box, and containers started with a service's `.env`, reach the person's AWS with temporary credentials. If an AWS call fails with no credentials, the environment has no role yet; set it with `PUT /v1/environments/{env}` (`awsRoleArn`, `awsRegion`) if the person gives you the ARN, or tell them, rather than putting keys in the box.
- `connectorOnline: false` means those addresses refuse connections until the person starts a connector; tell them in this conversation. With several connections it is `true` when any of them is online: read `connections[]` in the `environment` of `sandbox_start` or `sandbox_status` (read again on each call) or in `sandbox_environments`, which shows each connection's `online`, `sessions`, `lastSeenAt` and the addresses it carries, and `sandbox_start`'s `next` names offline connections. The final check for one network is still to use one of its addresses with a real client (`curl`, `pg_isready`, `psql` or `redis-cli ping`, all installed in the box; a bare TCP connect succeeds either way) and read that address in `sandbox_status` → `health.privateEndpoints[]`. When a connection fails, `health.privateEndpoints[].lastError` and `/work/.sbx/logs/private-endpoints.log` (one line per failure) say why: 403 means the address is not listed on the environment, 503 the connector of the connection that lists it is offline (the message names the connection; the client sees its connection accepted, then reset), 502 the connector cannot reach it.
- What is not listed stays out of reach: a private IP that is not listed, or a listed IP dialed on a port that is not listed, times out and nothing is logged; a host name that is neither listed nor declared is looked up in public DNS, so a name only the person's private DNS knows (Cloud Map, Route 53 private zones, office DNS) does not resolve, while one that public DNS resolves to a private IP (an RDS endpoint, `internal-….elb.amazonaws.com`) resolves and then times out; a listed or declared name dialed on another port lands on the box itself (refused unless something in the box listens on that port). To reach a name only their DNS knows (a private API Gateway such as `<api-id>.execute-api.<region>.amazonaws.com`, a Cloud Map name), add it with the port it is dialed on (`443` for API Gateway) to the connection that can reach it: `PUT /v1/environments/{env}/connections/{id}/endpoints` replaces the whole list, so read the current one from `GET /v1/environments/{env}` and send it back with the new address. The connector then resolves the name in their network. A box knows only the addresses it was started with, so start a new box after.
- These are the person's real dev databases: a changed service started with its settings writes there, migrations included. Do not run migrations, seeds or destructive tests against them unless the task says so. For tests, `psbx-testdb up <name>` starts a clean Postgres 16 in the box (`--mysql` for MySQL 8.0; several names start several) on a random loopback port (a container needs `--network host`) and writes `/work/.sbx/testdb/<name>.env` as `export` lines: `DATABASE_URL`, `TEST_DATABASE_URL`, `TEST_POSTGRES_URI` (Postgres), `PG*` and Laravel's `DB_*`. It prints a summary and the `. <file>` line, not the connection string (`psbx-testdb url <name>` prints that, password included). Tests that read another variable skip and look like passes: add `--as THAT_NAME`. Data sits in memory (tmpfs) capped at 2 to 4 GB depending on the box; `--size 8g` raises the cap and `--disk` puts it under `/work` instead, and a full data area makes Postgres panic or the container exit (`url` and `env` then print its `docker logs`). `psbx-testdb gotest ./...` runs `go test` with a fresh database for every package, removed afterwards, instead of packages clearing each other's tables in a shared one; `reset <name>` drops and recreates one, `list` shows each one's engine, port, state, use against its cap, creation time and owner, and `down <name>` removes it (`psbx-testdb --help`). For a service, run a database container and pass `-e DATABASE_URL=…` after `--env-file`.
- To exercise a changed service through its real callers, run those callers in the box too, unchanged and under their dev names: build them from source at the commit dev runs with their `.env`, or, when the environment's AWS role may pull from ECR, `aws ecr get-login-password --region <region> | docker login --username AWS --password-stdin <account>.dkr.ecr.<region>.amazonaws.com` and `docker run --env-file /work/.sbx/env/<service>.env … <dev's image>`. A login that fails with `AccessDeniedException` means the role lacks ECR read access; tell the person. If a caller already runs in another of your boxes, link the callee's name to this box from there instead.
- Values baked in at build time (`NEXT_PUBLIC_*` and the like) come from the settings too: `. /work/.sbx/env/<service>.sh`, then `docker build --build-arg NEXT_PUBLIC_API_URL …` (each needs an `ARG` in the Dockerfile), or build on the box after sourcing the file. Next.js also fixes `rewrites()` from `next.config.js` at `next build`, so put the same-origin `/api` rewrite in the box's copy before building; changing either later means building again (Team setup has the example).

If they have no environment yet, point them to https://parallelsandbox.com/en/docs/environments/ instead of guessing credentials. A worked example with several services, AWS plus an office network, Sentry and a changed service used from another box: https://parallelsandbox.com/en/docs/team-setup/

## Linking boxes

A service that runs in one of your boxes (A) can be used from another (B) under its usual name. When B only needs A's build, not A's live process, prefer a published version instead (see Versions): publish early and start B with `version` from the beginning, and B keeps working after A stops; when A is about to go away, B can switch the name to the version in place and keep its URLs (on boxes with the `take-over` feature, which every box started from 2026-09-25 12:40 UTC has; older boxes need a new B, with new URLs). Find A's id with `sandbox_list` if you did not start A yourself: every key of the account sees the same list, so a teammate's boxes are there too. Declare a link in B: `{ "name": "billing.acme.internal", "port": 8080, "fromBox": "<A>", "fromPort": 8081 }` in `sandbox_start.services`, or on a running B `sandbox_wire { "id": "<B>", "service": "billing.acme.internal", "port": 8080, "fromBox": "<A>" }`, which returns `{ "links": [...] }`; the same call with another `fromBox` or `fromPort` re-points it.

- `name:port` in B, from its programs and containers, reaches `fromPort` in A. `fromPort` defaults to the `targetPort` of A's service with the same name, else `port`. The process in A must accept connections on `127.0.0.1:<fromPort>` (listening on `0.0.0.0` does).
- Same account only; host names only; no `web` or `targetPort` on a link; a box cannot link to itself; `fromPort` cannot be 80 or 9095; the name cannot also be a service of B.
- A link wins over an environment address with the same `host:port`, so B can use a changed service running in A under its dev name.
- A frozen A is thawed by the connection (the account needs credits); A stays awake while B is awake and holds connections to it; when B freezes or stops, its connections are closed within about a minute and A freezes on its own idle time; a person taking A over does not interrupt links.
- A failure reaches B's program as a reset connection, with the reason in B's `/work/.sbx/logs/private-endpoints.log` and `sandbox_status.links[].lastError`: 409 A is still starting or runs an older image, 410 A has stopped (`start a new box and re-point the link with sandbox_wire`, or run the build A published with `version`), 502 nothing listens on `fromPort` in A.
- `sandbox_status` shows B's `links[]` (with A's `fromStatus`) and A's `linkedFrom[]`; `sandbox_stop` on A lists the boxes linked to it.
- Link traffic goes through ParallelSandbox and counts as outbound traffic of the box that sends it: B's requests are B's, A's replies are A's. Boxes started before links shipped cannot be linked to, and adding a link to a running box needs a box started after links shipped; a 409 says so.
- A link added to a running B with `sandbox_wire … fromBox` takes over an environment address with the same `host:port` for new connections at once (open ones keep their route until they close); a service or version added later takes it over the same way on boxes with `take-over`, and leaves it on the connector on older ones. Wiring the same name and port with another `fromBox` or `fromPort` re-points the link. On a box whose `sandbox_status` → `health.features` includes `take-over`, `sandbox_wire` with `port` (your process on `targetPort`) or `version` (plus `env`) on a link name takes it over in place: same address, open connections to A closed, the record stops being a link, A's `linkedFrom` drops B, and B keeps its URLs; the name must be taken over on the port it goes out on. On older boxes both fail with 409 (`… which cannot turn a link into …`): re-point the link or start a new box. A `mode` switch on a link name always fails with 409. Links on port 80 work as they are: `{ "name": "api.acme.internal", "port": 80, "fromBox": "<A>", "fromPort": 8080 }`.

## Secrets

`sandbox_secrets` lists the names stored for the account (`PUT /v1/secrets/{name}` over REST, see [Secrets](https://parallelsandbox.com/en/docs/secrets/)). When a token you need is in the person's hands, ask them to add it at app.parallelsandbox.com under Settings → Secrets rather than paste it into the conversation; a value typed there never reaches you. List the ones you need in `sandbox_start.secrets`; they become environment variables of every `sandbox_exec` command and never appear in health output. What the box returns to you masks them: in `sandbox_exec` output (and its `outputUrl` file), `sandbox_procs`' `cmd` and `logTail`, and the processes `sandbox_status` lists, every value of 8 characters or more of an injected secret, or of a setting in `/work/.sbx/env`, comes back as `****` (settings named like an address or a mode, `*_HOST`, `*_PORT`, `*_REGION`, `*_ENV`, `*_MODE`, `*_NAME`, `*_PATH` and the like, stay visible unless the name also holds `KEY`, `SECRET`, `TOKEN`, `PASSWORD`, `AUTH`, `DSN` or the like). Files are not changed: a background command's log file holds the real values, and `sandbox_get` or `sandbox_pull` of a file returns it as it is. A secret you need later goes into the running box with `sandbox_secrets { "id": "<id>", "names": ["SENTRY_AUTH_TOKEN"] }`; commands from then on see it (processes and containers already running do not).

## Versions

Publish a build as soon as it works if another box may need it: a box that declares the version runs its own copy and keeps working after the box that built it stops, which a link does not. A build another box should reuse: `sandbox_status.registry.imagePrefix` is the exact image prefix (`<registry>/parallelsandbox/tenant-<account>:`, one repository per account, so put the service in the tag) and the box is already logged in to that registry. `docker push "<imagePrefix>api-<sha>"`, then `sandbox_publish_version { "service": "api", "label": "<label>", "image": "<imagePrefix>api-<sha>", "gitSha": "<sha>" }` (labels are unique per service). `sandbox_versions` lists what other boxes of the account can run instead of rebuilding.

Another box runs a version by declaring it: `{ "name": "api.acme.internal", "port": 8080, "version": "<label or versionId>" }` in `sandbox_start.services`, or `sandbox_wire { "id": "<id>", "service": "api.acme.internal", "port": 8080, "version": "<label>" }` on a running box, where on boxes with the `take-over` feature it also takes over a link or an environment address with that name (on older boxes, declare environment addresses at `sandbox_start`). The same `sandbox_wire` with another `version` switches a name that already runs one, including one declared at start: the container is replaced and `run` starts again. The box pulls the image and runs it in the background as the container `psbx-svc-<name>`, with `SBX_BOX_ID` and the settings file of the service the version was published for (`/work/.sbx/env/<service>.env`, when the environment has one), then `env` if you pass one: a string-to-string map, only with `version`, applied after the settings file so a name in both takes the `env` value, for example `"env": { "SENTRY_ENVIRONMENT": "psbx", "WORKERS_ENABLED": "false" }`. Names follow shell-variable rules, `SBX_BOX_ID` cannot be set, at most 64 variables, 4 KB per value, 32 KB in all; the values show in `sandbox_status`, so no secrets; wiring the version again replaces the `env`. `sandbox_start` does not wait: call `sandbox_status` until `services[].run.state` is `running`; `failed` carries the error and the end of the start log. `targetPort` is picked for you when `port` is 80 or 9095 or already used (20000 and up); the container port is `containerPort`, else the image's one EXPOSEd TCP port, else `port`. A label that several services share resolves by the name's first DNS label, else pass the `versionId`. A service cannot be both a link and a version. Images stay in the registry until the account is deleted; a push is metered as outbound traffic. A box frozen for 10 hours or more gets a fresh registry login when it thaws. While this box runs, another box can instead reach the service here through a link (see Linking boxes); a published version is what outlives this box.

## Trouble

- `sandbox_status` first. `healthError` means boxd is not up yet; `interrupted` means start over on a new box.
- `goal is required: one or two sentences, in the language the person reads, on why this box exists and what done looks like. …`: `sandbox_start` was called without `goal`, or with a blank one. Call it again with `goal` set.
- `sandbox_start` fails with 503 `no box capacity right now: more is being started. Try sandbox_start again in about 5 minutes; a box of size 4 or 8 can take a few minutes longer.`: no host has room right now and another one is starting. Wait about 5 minutes and call `sandbox_start` again with the same arguments; there is no need to shrink `size` unless you want to. The same can happen for about 3 minutes right after ParallelSandbox switches box images.
- `status` values: `claimed` (starting for you), `ready`, `takeover` (a person is on it), `frozen` (idle; the next call that acts on it, or a page load in a browser, thaws it), `stopping`, `terminated`, `interrupted`.
- A service does not answer by its name: it must listen on `0.0.0.0` on its `targetPort` (the name has an address of its own in the box, not `127.0.0.1`), and a container must publish that port (`ports:` in compose, `-p <targetPort>:<container port>` in `docker run`). If the box has an `externalBaseUrl` and the wiring table shows the name as `external`, wire it to `box`.
- `port 80 in the box is boxd's scene entry; add targetPort, the port your process listens on — callers keep dialing <name>:80` (400): declare that service with `targetPort`. A message that the box's boxd `does not support targetport yet`: the box started before `targetPort` shipped; start a new box, or declare the service on the port your process listens on, without `targetPort` and not on 80 or 9095.
- `status: terminated` with a `statusReason` like `frozen and unused for …` or `… over idleTimeoutMin …`: the idle rules stopped the box (rules 3 and 6). When `sandbox_status` or `sandbox_list { "all": true }` shows `restorableUntil`, `sandbox_start { "restore": "<id>", "goal": "…" }` brings it back as it froze until then; otherwise `/work` is gone: start a new box, and pass `idleTimeoutMin` when the task needs a longer gap without calls.
- `services[].run.state` is `failed`: `run.error` has the reason and the end of the start log (`logPath`). `docker pull failed for <image>`: the image is not in the registry (push it, or check `sandbox_versions`). `the image exposes several TCP ports`: pass `containerPort`. `the container exited before it listened on port N`: the log tail says why; often a missing setting. `exited` means the container stopped after it had started: `docker logs psbx-svc-<name>`.
- 502 `nothing listens on <name>:80 in this box: that name is declared on another port. Dial the port it is declared on, or declare it on port 80 with targetPort, the port your process listens on`: the name is dialed on port 80 but declared (or listed) on another port. Dial its declared port, or declare it on port 80 with `targetPort`. `localhost:80` still reaches the first declared service; a box started before 2026-09-24 20:30 UTC answers the name on 80 with the first declared service too.
- A version or service added with `sandbox_wire` under an environment address name runs, but the name still reaches the environment: the box has no `take-over` in `sandbox_status` → `health.features`. Declare the name at `sandbox_start`, on a new box if needed.
- A link's box is going away and you want its published version instead: `sandbox_wire` with the link's name, `port` and `version` takes the link over on boxes with `take-over`, keeping the URLs; on older boxes it fails with 409 (`… which cannot turn a link into a published version …`), so re-point the link with `sandbox_wire … fromBox` or start a new box that declares the name with `version`.
- Connection errors right after a thaw: the connections opened before the freeze were closed; reconnect (database pools usually do on their own).
- A link fails: `sandbox_status.links[].lastError` and `/work/.sbx/logs/private-endpoints.log` say why. 410: the other box stopped; start a new one and re-point the link with `sandbox_wire`. 409: it is still starting, or runs an image from before links. 502 `nothing listens on port …`: its process is not up on `fromPort`.
- `sandbox_exec` says a person is operating the box: a person has the screen. Wait.
- `sandbox_exec` returns `leftoverChildren: true`: the command started processes with `&` or `setsid` that outlived it. They keep running untracked and lose their output; stop them and start them again with `background: true`.
- A `pkill -f` or `pgrep -f` in a command stopped or found more than you meant: the pattern matches every process whose command line holds it, other background jobs included. Stop background commands by `bgId` with `sandbox_procs`.
- A command reports `exitCode` 0 though a test failed: `exitCode` is the last command's; pass `strict: true`.
- A `429` means you hit the plan's box limit, frozen boxes included; stop one first. A `402` (`no credits left`) means the account is out of credits (rule 4).
- The box's URLs return 503 until the box is ready, while the box is frozen (a page load in a browser wakes it and gets a page that reloads itself; other requests get `box is frozen`; a call that acts on the box thaws it too), and after it stops; `sceneUrl` also when no service is declared. The body says which (`box is terminated`). 404 means nothing answers at that URL: there is no such box, the key is wrong, or the `-<port>` has no `web` service behind it (get the URLs again from `sandbox_status`). A 502 with `upstream ...: connection refused` means the box is fine but nothing is listening on that service's `targetPort`.
- A service marked `web` has no `url`: the box is not `ready` yet (read `sandbox_status` once it is), or it started before per-service URLs shipped, and then only `sceneUrl` works; start a new box.
- Downloads expired: the URLs last one hour. Call `sandbox_get` again for a file while the box exists (once it is stopped, no new URL can be had). For screenshots and recordings, `GET https://api.parallelsandbox.com/v1/boxes/{id}/media` with your token (`Authorization: Bearer <token>`, an API key or an OAuth access token) signs their URLs again (the newest 200 steps with pictures), for 7 days after each was taken and for a stopped box too; another `sandbox_shot` only takes a new picture.
