Private pilot — invite only

Real-world examples

Four examples of firebox in action.

Each one is a problem, the exact API calls that solve it, and why the isolation model matters. Every call below is against the documented API — create a sandbox, do the work, destroy it.

Example 01

The AI data analyst

A user uploads sales_q3.csv and asks in plain English: “which region grew fastest?” The agent writes analysis code — but nobody vetted it, and the dataset can’t leak. So the code runs sealed.

1

Boot a sandbox

POST /v1/sandboxes. The dataset is about to live only inside this VM.

2

Upload the data

POST /v1/sandboxes/{id}/files with the CSV, base64-encoded. Files up to 10 MiB.

3

Ship unvetted code

The LLM generates analyze.py. Same files API. At this point the code could contain anything — it goes straight into the sealed box.

4

Run it, watch it fail safely

POST /v1/sandboxes/{id}/exec → KeyError: 'Region' — the model guessed a column name. A wrong guess, contained. If it had been malicious instead of merely wrong, the outcome is identical: one dead VM.

5

Fix, re-run, stream the long pass

Corrected script, same warm sandbox. For the full-dataset run, POST /v1/sandboxes/{id}/exec/stream delivers stdout over SSE as it happens.

6

Collect and destroy

GET /v1/sandboxes/{id}/files?path=summary.json returns the answer; then DELETE the sandbox. Metering logged the CPU-seconds.

Why firebox fits: LLM-generated code is untrusted, so it runs sealed — the data never leaves the VM, a bad run can’t touch anything else, and destroying the sandbox is the cleanup.

Try it yourself

# pip install firebox
from firebox import Firebox

fb = Firebox(api_key="fb_...")
sb = fb.create_sandbox(
    template="python",
    secrets=["ANTHROPIC_API_KEY"],   # injected, never in code
    egress=["api.anthropic.com:443"], # locked down
)

sb.write("sales_q3.csv", csv_data)
code = llm.generate("which region grew fastest?", columns)
sb.write("analyze.py", code)

result = sb.exec("python3 analyze.py")
print(result.stdout)   # "West, +23% QoQ"
sb.destroy()

Example 02

The coding agent that proves its work

An AI engineer is fixing a checkout bug. Its code is untrusted by construction — written by a language model, reviewed by no one — and “trust me, it works” isn’t proof. So it tests itself, inside a box that gets thrown away.

1

Boot a clean room

POST /v1/sandboxes returns sb_9c2e…. The agent is about to run code it wrote itself. That runs here — not on your infrastructure.

2

Drop in the suspect code

POST /v1/sandboxes/{id}/files uploads the checkout module and its test file. Kilobytes, not megabytes.

3

Reproduce the bug

POST /v1/sandboxes/{id}/exec runs pytest. It fails: assert 95.0 == 90.0. The agent finally has a failing test it can see — the feedback loop that makes agents work.

4

Iterate warm

Fix via the files API, re-run. Still red — a rounding error this time. Fix again, re-run. Green. Same sandbox, no reboot; a runaway iteration hits timeout_s and gets killed.

5

Ship with proof

The agent shows the user the diff plus the passing test output it captured. Fixed — and here’s the run proving it.

6

Destroy it

DELETE /v1/sandboxes/{id}. Nothing persists — no stale state poisoning the next task.

Why firebox fits: tiny payloads, short runs, code that is hostile by default. The blast radius of one bad iteration is a disposable VM, and the tight fix-and-observe loop is exactly where fast boots pay off.

Try it yourself

# pip install firebox
from firebox import Firebox

fb = Firebox(api_key="fb_...")
sb = fb.create_sandbox(template="python")

for attempt in range(5):
    code = llm.generate(task)          # your LLM call
    for path, content in code.files.items():
        sb.write(path, content)        # upload to microVM

    result = sb.exec("pytest -x -q")
    if result.exit_code == 0:
        break                          # tests pass
    task = fix_prompt(result.stdout)  # feed failure back

sb.exec("zip -qr /tmp/app.zip .")
artifact = sb.files.read("/tmp/app.zip")
sb.destroy()

Example 03

Team scripts, without the shared Jenkins

Forty product teams, hundreds of automation scripts, one shared cluster. Noisy neighbors, credential sprawl, zero isolation. The platform team replaces it with an execution API: every team gets its own key, its own VMs, its own bill.

1

One key per team

POST /v1/keys (admin) mints team-a’s key. The secret is shown once. Per-key concurrency caps and rate limits apply.

2

Submit and isolate

Team A’s nightly sync runs POST /v1/sandboxes with their key. Their jobs land in their own microVMs — invisible to every other team.

3

Run with the blast radius contained

POST /v1/sandboxes/{id}/exec with a timeout_s. A script that loops forever gets killed; a script that tries to wander the network hits the egress allowlist.

4

Meter per team

GET /v1/usage per key gives the platform team chargeback numbers: executions, CPU-seconds, memory-seconds, active sandboxes.

5

Revoke when someone leaves

DELETE /v1/keys/{id}. The key dies immediately; in-flight sandboxes are reaped. No shared credentials to rotate.

Why firebox fits — and why microVMs, not containers: one kernel per tenant. A kernel exploit can’t cross team boundaries, which is the property you need when teams run each other’s untrusted code. Containers share the one thing you most need isolated.

Example 04

The agent that brings its own sandbox

Your coding assistant already writes code — but it runs that code on your laptop. Give it firebox tools over MCP and the untrusted code runs in a microVM instead. The agent never touches your filesystem.

1

Register the tools

One command — claude mcp add firebox — gives your agent ten tools: create, exec, files, pause, resume, delete, usage.

2

Ask in plain English

“Clone this repo, run the test suite, and tell me what fails.” The agent calls firebox_sandbox_create — the work is about to happen in a microVM, not on your machine.

3

Run it sealed

firebox_exec runs git clone and pytest inside the sandbox. stdout/stderr come back as text, capped at 64 KB so a runaway cat can’t flood the agent’s context.

4

Iterate warm

12 failures, 9 from a rounding bug. The agent patches the file with firebox_files_write and re-runs in the same sandbox — the feedback loop that makes agents work.

5

Pause between turns

firebox_sandbox_pause snapshots the VM when the user walks away; firebox_sandbox_resume restores it exactly. Metering stops while paused.

6

Check spend, clean up

firebox_usage reports the key’s totals any time. When the task is done, firebox_sandbox_delete — or let it expire.

Why MCP matters: every AI coding tool is becoming an agent that runs code. Without sandboxing, that code runs with the user’s privileges on the user’s machine. MCP tools move the blast radius into a disposable microVM — the agent keeps its full workflow, and your laptop stays out of it.

See your workload here

firebox is in private pilot — access is invite-only for now. If you’re building agents that run code, we’d like to hear what you’re running.

Try the quickstart