sandbox_exec

Runs one shell command in the box, in the foreground with a timeout or in the background with a log file. While the screen changes, a foreground step keeps a screenshot and a recording for the person.

Input Output
id, cmd, note, cwd, timeoutSec, background, strict, user (root or tester), readOnly, tailBytes, readyPort, readyLog, baseline stdout, stderr, exitCode, truncated, durationMs, timedOut, signal, boxTime, note, execId; outputUrl, outputBytes and outputNote when the output was cut (or was binary); endReason when a signal other than the timeout ended it; leftoverChildren and hint when the command left processes running; privateEndpointErrors and privateEndpointHint when its connections to environment addresses or links failed; bgId and logPath for background; with baseline, the comparison below
  • note (required): one short sentence, in the language the person reads (the tool description names it), saying what the command is for. The app shows it instead of the raw command, as what the box is doing right now.
  • cwd is relative to /work, the box's working directory, and defaults to it.
  • timeoutSec bounds foreground commands: default 600, at most 3600. When it runs out, the command is killed with SIGKILL, along with the processes it started in its process group (those started with & too), and the call returns what it printed so far, with timedOut: true, signal: "SIGKILL" and exitCode -1.
  • Every foreground call has an execId (the local adapter picks it before sending the call). If the call is cut off before its result comes back (the connection drops, or the client gives up waiting), the command is not stopped: it runs until it ends or timeoutSec runs out, and sandbox_procs wait with bgId set to the execId returns what the call would have returned (result: stdout, stderr, exitCode…), so do not run it again first. The adapter's error names the execId and that call. The last 200 foreground results are kept. stop with the execId ends it early.
  • Output: stdout and stderr each keep their last 16 KB (tailBytes sets fewer, 1 to 16384). When more was printed, the beginning is cut: truncated: true, outputUrl links the complete output (both streams, in the order they were printed; valid one hour), outputBytes is its size and outputNote says how much each stream kept. A stream that is mostly binary (NUL bytes, invalid UTF-8) is not returned: it reads [binary output: N bytes, not shown; …] and the complete output is behind outputUrl; a few NUL bytes in text come back as ␀ or [N NUL bytes]. Output is not a terminal, so programs print their non-interactive format: node --test prints TAP (# fail 2, not ok 3 - name), not the ℹ fail lines of a terminal.
  • A signal that is not the timeout comes with endReason: stopped by sandbox_procs stop, the out-of-memory killer (the box ran out of memory: run fewer heavy jobs at once, such as go build -p 2, or start a larger box), or another process (kill, pkill).
  • cmd is one JSON string. To write a file, use a quoted heredoc, cat > path <<'EOF' … EOF: nothing inside it is expanded and backslashes stay as written (an unquoted <<EOF expands $ and backquotes). Send long files, and many files, with sandbox_sync instead of putting them in cmd: a cmd over 256 KB is refused, and a very long one can make the client send broken arguments, which come back as arguments: …. command is accepted as another name for cmd.
  • readOnly: true is for a foreground command that only reads (logs, files, status): it runs without DISPLAY, nothing is recorded, and it runs even while a person has taken the box over through takeoverUrl, when other commands wait. It is not for looking at or driving the screen. It cannot be used with background.
  • A foreground call returns once its shell exits, waiting at most 3 more seconds for what still holds its output. Processes it started with &, nohup or setsid that are still running then are left alone: the result has leftoverChildren: true and a hint, they keep running, nothing tracks them, and whatever they print after that is lost. When timeoutSec runs out first, they are killed with the rest of the process group. Anything that should keep running (a dev server, a watcher, a database, an emulator) goes in its own background: true call; sandbox_procs knows only those.
  • exitCode is that of the last command, so a failing test piped to tee, or followed by another line, reports 0. strict: true puts set -eo pipefail in front of the command: exitCode is then that of the first command or pipeline stage that fails, and the rest does not run.
  • Long test output: keep what matters in the result instead of scrolling past unrelated log lines, for example go test ./... 2>&1 | grep -E '^(--- FAIL|FAIL|ok|panic)' (with strict: true, or set -o pipefail, so a failure still shows in exitCode), or go test -json ./... | jq -r 'select(.Action=="fail") | .Package + " " + (.Test // "")'. Output over 16 KB is cut from the start, and outputUrl has all of it.
  • A throwaway test that only the box needs (a fixture that builds fake data and checks a rollup, say) can be written straight into the box's copy and removed after the run, so the local repo never sees it: cat > /work/<repo>/internal/rollup/rollup_sbx_test.go <<'EOF' … EOF, then go test ./internal/rollup -run Sbx and rm the file in the same command. Bring it home with sandbox_pull only if it should become a real test.
  • user: "tester" runs the command as a non-root user (uid 1000 when it is free, in the docker group) with its own HOME, TMPDIR (/tmp/u<uid>, 0700), GOPATH and npm cache, created the first time. Use it for tests of programs that switch uid (setpriv, a uid per member) or refuse to run as root: as root, a child that switches uid cannot enter the 0700 temporary directories root made, and such suites fail in bulk for that reason alone. Before each such command, root-owned files under cwd (except /work/.sbx) are handed to tester so it can write there, so give cwd the repo rather than all of /work. The default is root. /tmp itself is shared by every uid: a test that runs a CLI under another uid can create that uid's /tmp/claude-<uid> (or similar) first, and the CLI then refuses it (Temp directory /tmp/claude-1001 is owned by uid 97884). Give each uid its own temporary directory (TMPDIR; for Claude Code CLAUDE_CODE_TMPDIR), and remove a stale one to recover.
  • background: true returns a bgId and a logPath at once; the process keeps running after the call and its output goes to that file, /work/.sbx/logs/<bgId>.log, which you read with another sandbox_exec. sandbox_procs lists these commands, finished ones included, waits for one to end, and stops one together with every process it started; sandbox_status.running.background lists the ones still running with their logPath. Stop them by bgId rather than with pkill -f or pgrep -f. The command runs from a script file, so a pattern in it no longer matches the shell running your own sandbox_exec, but it matches every other process whose command line holds it, other background jobs included. For a process you did not start in the background, match its exact name with pkill -x <name> or use kill <pid>.
    • Programs buffer output that goes to a file or a pipe instead of a terminal, so a log can stay empty for minutes and fill all at once. PYTHONUNBUFFERED=1 is set for every command (unless you set it yourself); in a pipeline, use sed -u, grep --line-buffered or stdbuf -oL <command> so each line reaches the log as it is printed. Node's console.log to a file is written at once.
    • readyPort (a port that accepts connections on 127.0.0.1 or ::1 once the server is up) and readyLog (a regular expression, RE2 syntax, matched against each line of the log, such as "Listening on|ready in") say when the command is ready. sandbox_procs wait with until: "ready" returns as soon as either holds, instead of a fixed sleep; it can also set them when you wait.
    • psbx-step <name> -- <command> [args...] runs one step of a longer script and records its start, exit code and duration; sandbox_procs with that bgId (or a foreground execId) lists them as steps, so you see which step is running and which one failed even when the log says nothing. The exit code passes through (psbx-step build -- make && psbx-step test -- make test stops at the first failure); put a pipeline inside bash -c '…'.
  • baseline: true runs the same command twice, one run after the other: in cwd, then in the same place in <dest>-baseline, the earlier tree that sandbox_sync baseline put next to dest (cwd must be inside a dest synced that way). It picks the failing tests out of both outputs (go test --- FAIL and FAIL <package> lines, node --test TAP not ok, the ✖, ✕ and × lines of node, jest and vitest, pytest FAILED and ERROR), drops timings and test numbers, and returns onlyMine (fails only with your changes), alreadyFailing (fails in both), fixedByMine (failed only before them), counts, exitCode and baselineExitCode, and mineLog and baselineLog with both full outputs under /work/.sbx/logs. No names recognized but a run that exited non-zero is said in summary: read the logs then. It cannot be used with background.
  • A project's own dependencies are not in the image. For Python, run pip install -r requirements.txt (and the project's other requirements*.txt) in the box before its tests; pip installs system-wide, with break-system-packages already set. A one-off Node script anywhere in the box, with no package.json, can import playwright, playwright-core and ws (import { chromium } from 'playwright', or require) without npm install: they, and Playwright's Chromium, Firefox and WebKit, are installed for the whole box. Those browsers belong to Playwright 1.56.0: a project pinned to another @playwright/test fails with Executable doesn't exist at /usr/local/share/playwright/…, so run npx playwright install chromium in it first (about 20 seconds; it lands in the shared browser directory), or launch with executablePath: '/usr/bin/chromium'. Another Node major version runs with npx -y node@20 script.js (or node@24, about 5 seconds the first time). The image also has ripgrep (rg), zip, xxd, bats, bwrap, xclip, python (Python 3.11, with Pillow, numpy, boto3, fonttools and brotli; 3.12 and 3.13 through toolchains), ImageMagick (convert, montage), the en_US.UTF-8 locale and a CJK font. git has an identity (ParallelSandbox Box, in /etc/gitconfig and root's ~/.gitconfig) and safe.directory *, so tests that commit need no setup. nproc and free -g show the box's own CPUs and memory; there is no cgroup quota inside (/sys/fs/cgroup/cpu.max does not exist). The box's IBus input method is set for every command through IBUS_ADDRESS, so a test that starts its own ibus-daemon needs env -u IBUS_ADDRESS. /etc/sbx/manifest.json lists what the image has.
  • Commands run as root in bash -l. Their environment already has DISPLAY=:99 (the virtual display, so anything you start shows in sandbox_shot and through takeoverUrl), SBX_BOX_ID (the box id), SBX_SCENE_HOST (the scene URL's host, key included; treat it like the URL), the secrets given to the box, and AWS_CONTAINER_CREDENTIALS_FULL_URI and AWS_REGION when the box's environment names an AWS role. Containers get none of these unless you pass them (docker run -e SBX_BOX_ID ...). Each command also gets PYTHONUNBUFFERED=1, PSBX_JOB_ID (its bgId or execId) and PSBX_STEPS (where psbx-step records).
  • Secret values are masked in what comes back. In stdout, stderr and the full output behind outputUrl (as in sandbox_procs' cmd and logTail and the processes sandbox_status lists), each value of 8 characters or more of a secret injected into the box, or of a setting in /work/.sbx/env, is replaced by ****; a multi-line value, such as a PEM key, line by line. Settings whose name says they are an address or a mode stay visible so you can debug: HOST, PORT, ENV, TZ, LANG, STAGE and names ending in _HOST, _HOSTNAME, _PORT, _REGION, _ENV, _ENVIRONMENT, _STAGE, _MODE, _LEVEL, _PATH, _DIR, _NAME, _TIMEZONE or _LOCALE, unless the name also contains SECRET, TOKEN, PASSWORD, PASSWD, PWD, KEY, CREDENTIAL, AUTH, PRIVATE, SIGNATURE, SALT, COOKIE, SESSION or DSN. Values shorter than 8 characters (true, 3000, dev) are never masked, nor is anything else, such as a psbx-testdb connection string. Files are not changed: a background command's log file holds the real values (reading it through sandbox_exec masks them again), and sandbox_get and sandbox_pull return files as they are.
  • While a foreground command runs (including the 0.6 seconds after it ends, before the screenshot is taken), if the picture on the box's screen (DISPLAY=:99, or an attached phone) changes, the step automatically keeps a screenshot and, usually, a recording (exceptions below: sometimes only the screenshot, sometimes nothing): the recording runs from the moment the change is seen (the screen is compared only every second, every 2 seconds with a phone attached, so this is a little after it actually changes) until the screenshot is taken, and the screenshot is taken 0.6 seconds after the command ends (right away, before the next command starts, if you run another foreground command within those 0.6 seconds). If the screen does not change, nothing is kept. The person sees them under the box's Screenshots & recordings in their app, each step labeled with its time and your note. So run browser and app tests headed on :99 (Playwright's --headed, HEADED=1), and each step shows the person what was tested. How long each kind of picture is kept, and who can still see it once the box is stopped, is summed up in one table under Finishing a box. Details and exceptions:
    • What counts as a change: the screen is noted before the command starts, then compared with that every second (every 2 seconds with a phone attached); a blinking cursor is not a change. After the command ends, capture waits 0.6 seconds (less if the next foreground command starts sooner); if no change has been seen by then, it compares one last time and, if the screen changed, takes the screenshot. If the screen does not change during the command (a plain npm test), nothing is kept, not even a screenshot.
    • A command over in under a second (a click with xdotool) ends before a video can start: if the screen changes within 0.6 seconds after it ends, the step gets only that screenshot; if it changes later (a page that takes seconds to load after the click), nothing is kept. To keep what the click led to, wait a few seconds in the same command, for example xdotool mousemove 400 300 click 1; sleep 2. With a phone attached, wait longer: how the screen looked before the command is only grabbed after the command starts (an iPhone takes a second or two per grab), then compared every 2 seconds after that grab, so recording starts no sooner than 2 seconds into the command (later still with an iPhone, whose grabs are slow), and what a tap (artemis-adb tap, psbx-ios tap) changes may be taken for the screen before the command, and nothing is kept. Wait a few seconds before and after the tap, for example sleep 3; psbx-ios tap 200 400; sleep 3.
    • Whether the step itself records a video is decided once, when the change is first seen, and never revisited. If your own sandbox_shot recording is running when the change is first seen, what changes on the screen during the command is recorded only in the recording you started with sandbox_shot; the step itself has no video (your recording does not count as the step's either) and keeps only the screenshot taken after the command ends, and no video is added to it later either. If yours is not running when the change is first seen, the step records its own video as usual, and takes the screenshot too. The recording you started with sandbox_shot is never added to the command's step, and it ends one of two ways: if you end it yourself with record: "stop", it is uploaded and shows in the app as a separate step (one with kind shot and detail record); if it is stopped for you, it is not uploaded, becomes no step at all and never shows in the app, and for the foreground commands run while it was recording, the person sees only each step's own screenshot. ParallelSandbox stops your recording for you a full hour after the box's last use (the end of a tool call that acts on it, or a wake; see Box states) at the earliest, and only once nothing else keeps the box awake; how long it has been recording does not matter, and it never happens while a foreground command runs. Where the mp4 of a recording stopped for you is kept, how to fetch it and how long the fetched copy is kept: see the record item under sandbox_shot.
    • A step records at most 10 minutes, counted from when its recording starts (when the change is first seen), which only a command with timeoutSec over 600 can reach (recording starts no sooner than a second into the command, later with a phone attached, so one that ends within 600 seconds never records a full 10 minutes): then the command keeps running and the recording stops; the 10 minutes recorded are kept, and the screenshot is still taken after the command ends.
    • Commands with background: true are not recorded.
    • These screenshots and recordings are kept for seven days after capture. After a box stops, the app lists it under All sandboxes → Stopped (last 7 days), where its unexpired media remains accessible. The box and its /work cannot be restored. idleTimeoutMin has a maximum of 10080 minutes (seven days); see Box states.
    • After the box is stopped you can still get them, through the REST API only, with no MCP tool for it: GET https://api.parallelsandbox.com/v1/boxes/{id}/media with the API key returns each step until 7 days after it was taken, the newest 200 steps at most (see REST). When the person should see what each step did, hand the box over with sandbox_review instead of stopping it with sandbox_stop (Finishing a box).
  • Boxes are x86_64 (amd64): Debian 12 on Intel Xeon hosts. Pull or build linux/amd64 images. An arm64-only image, such as one built for Graviton or on an Apple silicon Mac without --platform linux/amd64, fails with exec format error. Emulation is not installed: docker run --privileged --rm tonistiigi/binfmt --install arm64 adds it to that box, but emulated containers are slow (a CPU loop took about 6 times as long on a box we tried), so rebuild for amd64 instead, or push a multi-architecture image (docker buildx build --platform linux/amd64,linux/arm64). arm64 output for somewhere else (Graviton, Lambda on arm64) can be built in a box without emulation as long as nothing arm64 runs there: CGO_ENABLED=0 GOOS=linux GOARCH=arm64 go build, or docker buildx build --platform linux/arm64 of a Dockerfile whose RUN steps run on the build platform (FROM --platform=$BUILDPLATFORM … in the stage that compiles, then COPY --from that stage). A RUN step in an arm64 stage fails without emulation, and the image it makes cannot be run in the box (exec format error). The box is Debian 12 with glibc 2.36: a Go binary built in it with cgo (the default when a C compiler is present, as here) needs that glibc, so copied into an older runtime image it fails with GLIBC_2.34 not found; build with CGO_ENABLED=0, or inside the same base image the program will run on.
  • Browsers on the box's screen:
    • The screen is 390 × 844 portrait unless the box was started with orientation: "landscape" (1280 × 800; sandbox_start). A Chromium window is at least 500 pixels wide, so on the portrait screen start it with --force-device-scale-factor=0.78 --window-size=500,1082 --window-position=0,0 to fill the screen: chromium --no-sandbox --kiosk --force-device-scale-factor=0.78 --window-size=500,1082 --window-position=0,0 --remote-debugging-port=9222 --user-data-dir=/tmp/chrome http://localhost:3000/ with background: true. The page then lays out 500 CSS pixels wide. For an exact phone layout of that page, signed in, call sandbox_shot with target: "tab" and width (it needs the --remote-debugging-port): the box keeps that viewport on the tab while you work in it. For a signed-out page, target: "url". On a landscape screen, --window-size=1280,800 replaces those three flags. Chromium needs --no-sandbox because commands run as root.
    • Wait for a condition, not a fixed sleep: timeout 60 xdotool search --sync --onlyvisible --class chromium returns once the window is up, a Chromium started with --remote-debugging-port=9222 is ready for CDP when curl -s http://127.0.0.1:9222/json/version answers, and sandbox_shot can wait itself (waitFor, stableMs; target: "tab" waits up to 45 seconds for a browser that is still starting).
    • Openbox manages the windows: xdotool windowactivate and getactivewindow work, a new window takes the keyboard focus, and dialogs (Electron's dialog.showMessageBox, GTK) open centred with a title bar. Other windows get no frame, so a window opened at 0,0 at the screen's size still fills it exactly, and openbox has no shortcuts of its own: every key reaches the app. The pointer rests in the bottom-right corner, out of the picture; move it yourself to test hover. DBUS_SESSION_BUS_ADDRESS points at the display's session bus, so desktop notifications (Electron's Notification, notify-send) show at the top right.
    • The screen comes in those two sizes only. For a page at another size, use sandbox_shot with target: "url" or "tab" and width and height. For a desktop app at, say, 1440 × 900, start a display of your own (Xvfb :100 -screen 0 1440x900x24 with background: true), run the app with DISPLAY=:100, and capture it with ffmpeg -f x11grab -video_size 1440x900 -i :100 -frames:v 1 shot.png and sandbox_get: sandbox_shot and the person's view only show :99.
    • Find elements by role and accessible name (page.getByRole('button', { name: 'Save' }) also matches an icon button's aria-label), not by visible text or position: an icon-only button has no text, and a locator that matches several elements fails Playwright's strict mode.
    • One browser drives one page at a time. A tab that is not in front gets no frames, so page.screenshot on it times out; call page.bringToFront() (CDP Page.bringToFront) before working in another tab of the same browser, and give scripts that run at the same time a browser each, with their own --user-data-dir and debugging port.
    • What needs a user gesture (window.open, the clipboard, fullscreen) needs real input: Playwright's click(), CDP Input.dispatchMouseEvent or xdotool click. An element.click() sent through CDP Runtime.evaluate carries no user activation unless the call passes userGesture: true, and the popup it opens is blocked without an error.
    • A minimal headed Chromium for your own CDP script: start chromium --no-sandbox --remote-debugging-port=9222 --user-data-dir=/tmp/chrome <url> with background: true (as root it exits at once without --no-sandbox), wait with until curl -sf http://127.0.0.1:9222/json/version >/dev/null; do sleep 0.2; done, then connect: Playwright's chromium.connectOverCDP('http://127.0.0.1:9222'), or the webSocketDebuggerUrl that /json/version returns.
    • A window stays until its program ends: start it with background: true and end it with sandbox_procs stop; sandbox_shot with target: "window" lists what is still on the screen in windows (xdotool windowkill <id> closes any one).
    • With --force-device-scale-factor, --window-size is in unscaled points and rounds down: at 2, a 1280 × 800 screen needs --window-size=641,401; 640,400 leaves a 1-pixel black edge on the right and bottom, and 1280,800 asks for a window twice the size of the screen.
    • The box has no GPU, so WebGL is unavailable by default (Chromium logs WebGL1 blocklisted and pages see no WebGL). For a page that needs it, add --use-angle=swiftshader --enable-unsafe-swiftshader: software rendering, correct but slow.
    • Voice input: /opt/psbx/fake-mic/en.wav and /opt/psbx/fake-mic/zh.wav hold a spoken English and a spoken Mandarin sentence. Start Chromium with --use-fake-ui-for-media-stream --use-fake-device-for-media-stream --use-file-for-fake-audio-capture=/opt/psbx/fake-mic/en.wav and getUserMedia gets that recording as the microphone, looped, with no permission prompt. For other words, espeak-ng -v en-us -w /tmp/say.wav "..." (Mandarin: -v cmn-latn-pinyin; the plain cmn voice reads the tone numbers in English).
    • To swipe in a page emulating a phone, send CDP Input.dispatchTouchEvent: touchStart, several touchMoves, then touchEnd. Input.synthesizeScrollGesture with gestureSourceType: "touch" does not scroll on the box.
    • The box's Chromium lets localhost, 127.0.0.1, *.localhost and *.parallelsandbox.com use the clipboard without asking, and shows no translate prompt over pages in another language, so it stays out of screenshots and recordings (a managed policy, /etc/chromium/policies/managed/psbx.json). navigator.clipboard.readText() still needs a focused page, so bring it to the front first. Any other origin needs the permission granted (Playwright: context.grantPermissions(['clipboard-read', 'clipboard-write']); CDP Browser.grantPermissions lasts only while that connection stays open); without it, Chromium shows a permission bar over the page and the call waits.
    • To check which font a page actually renders with, ask Chromium: CDP CSS.getPlatformFontsForNode on the element lists the fonts used and how many glyphs each drew. document.fonts.check() is unreliable for web fonts split by unicode-range (most CJK and Google Fonts): it can return false for text already drawn with the web font.
    • A Chrome extension (Manifest V3) loads with --load-extension=/work/<ext> --disable-extensions-except=/work/<ext> and a --user-data-dir of its own; Playwright needs chromium.launchPersistentContext with those args and headless: false (on DISPLAY=:99). chrome.runtime.reload() and other developer actions need developer mode on in that profile (chrome://extensions), or the extension is disabled with reason 16777216. Its service worker shows up as a CDP target of type service_worker; to restart it, reload the extension from chrome://extensions or stop the target. Leave downloads alone: Playwright's acceptDownloads and CDP Browser.setDownloadBehavior with allow handle downloads before the extension's chrome.downloads.onDeterminingFilename sees them, so test those with the default behavior.
  • When a program in the box cannot reach an environment address or a link, it only sees the connection reset or closed. The result then carries privateEndpointErrors ([{target, error, count}], every failure during this command) and privateEndpointHint, with what the connector or ParallelSandbox said, such as the connector could not reach chatbot-dev.svc.local:8000: lookup chatbot-dev.svc.local: no such host (the service has no running task in their network) or the connector "office" of this environment is offline. These are on the far side of the box, so changing the code does not fix them (Environments). A name that is neither an environment address nor declared is looked up in public DNS: ERR_NAME_NOT_RESOLVED or NXDOMAIN for a public-looking name means public DNS has no record of it, so check it with dig <name> in the box before suspecting the box.
  • While a person is connected through takeoverUrl (or is trying a box review on its screen), sandbox_exec returns an error (HTTP 423) instead of running, unless the command is readOnly: it says since when a person has the box, why, and the latest time it comes back. Wait and retry; sandbox_get, sandbox_procs and sandbox_shot of the screen keep working meanwhile, and to keep working on the code you can start another box and sync it there.

Every tool and topic is listed in the tool reference.