About sandboxes
This page explains what a sandbox is, why Fountain does not let you create one, and what changes when you move to a different provider. For the executable contract a backend must meet, read the sandbox contract. To choose and configure one, read Self-host Fountain. For the endpoints, read the Sandboxes section of the API reference.
What a sandbox is
A sandbox is the isolated machine that a conversation runs in. Several conversations can run on one sandbox at the same time, each with its own transcript.
You never create one on its own. Fountain provisions a sandbox when a
conversation starts. Fountain reclaims it when every conversation on it goes
quiet, or when it runs too long. You can name one, with sandbox_id, to put
a second conversation on it. GET /api/sandboxes lists them, and
fountain sandbox does the same from the CLI.
That is deliberate. A sandbox you could create on its own would be a resource you could leak. The machine that nobody remembered to stop is what makes agent infrastructure expensive.
One conversation builds one sandbox. A conversation can end up with two processes behind it — a rolling deploy, a lost node — and only one of them builds the machine; the other finds it already being built and stands down. A build that stops part-way is torn down and started again rather than continued, because the steps it runs cannot be repeated on top of themselves. Building has an absolute ceiling: after about half an hour Fountain stops it, records the sandbox as failed, and the next prompt starts a fresh one. If Fountain cannot write that record — because something else is working on the sandbox at that moment — it asks again for up to a quarter of an hour and then stops the build anyway, leaving the sandbox for the hourly cleanup pass.
Two modes
An agent chooses a default, and a launch can name the other.
With ephemeral, the default, each conversation gets a sandbox of its own.
When the conversation ends, the sandbox ends. This is the mode for a
one-shot task, for a fan-out, and for input you do not trust.
With persistent, the agent has one machine of its own. Each conversation
with the same environment and vault lands on it. Notes, clones and installed
tools accrue across conversations. What a bad turn leaves behind also
accrues, because all of them use the one disk. The machine survives a
conversation that ends. Nothing stops it while it runs. When you delete the
agent, Fountain destroys the machine. A reset (DELETE /api/sandboxes/:id)
also destroys the machine, and keeps the conversations. The next prompt builds
a clean one.
The machine belongs to one identity. Three changes move that identity. You
move the agent to a different environment. You delete the environment. You
delete the vault. Fountain retires the machine for each of them, and keeps the
conversations, so the next prompt builds a machine on the identity that exists
now. Fountain refuses each change with 409 sandbox_mid_turn while a
conversation on that machine runs a turn. Let the turn end, or stop it, then
send the request again.
Only a ready or suspended machine resets. A machine in the pending or
starting status answers 422 sandbox_not_resettable with that status. It
has no disk to replace yet.
A reset blocks new turns before it calls the provider. It releases capacity
only after the provider confirms deletion. A timeout or lost request keeps
the reset fence and quota reservation. Another reset returns
409 sandbox_reset_pending, and so does an attach or a wake. None of them
send a second delete.
409 sandbox_reset_pending is not only about resets. A machine is fenced the
same way when the last conversation on it ends and it is being torn down, so
an attach that arrives after that point answers the same code, and the message
says the computer is being torn down or reset. The answer is the same in both
cases: the machine is going away, and a new one is what the request needs.
One operation at a time acts on a machine. If a different teardown of the same
machine is already running, the reset returns 503 sandbox_unavailable with a
Retry-After header. The fence is already in place, so the reset is accepted
and queued: Fountain completes it, and a repeat of the request answers
409 sandbox_reset_pending.
A prompt that wakes a machine, an attach that opens a conversation on one, a
prompt to a conversation whose process is up on one, and a request to end a
conversation on one, answer 503 sandbox_unavailable with a Retry-After
header for the same reason: Fountain is in the middle of an operation on that
machine. A prompt to a conversation whose process is up, and a request to end
one, wait a few seconds for the operation first; the other two answer at once.
Deleting a machine, resetting one and parking an idle one each take the
machine for the length of one provider round trip. Send the request again.
When a machine is deleted, every turn still running on it is marked
interrupted: nothing on that machine can finish it. Each SDK reports
this as a not-ready error, the launch queue and a team schedule wait and try
again on their own, and the boot sweep leaves the machine to the operation that
holds it.
Every five minutes, Fountain looks for resets and deletions that were asked for and then abandoned. A reset whose delete the provider did not confirm is tried again on that run, and on each run after it until the provider confirms. A provider without credentials waits until it is enabled again. A deletion is finished when it is fifteen minutes old and no server holds the machine. A machine that another operation is working on is left to that operation. The fence and capacity reservation stay in place until deletion is confirmed.
For a persistent machine with a pending reset in ready or suspended, an
administrator can choose Retry reset on the admin sandbox list. This
checks the provider first. A confirmed missing machine retires immediately;
an existing machine gets another delete attempt. A failed probe or delete
keeps the fence and capacity reserved. The admin audit records the attempt
and its outcome. The action checks the administrator's current access again,
so an old browser tab cannot retain revoked authority.
For a manual override, an operator reaps the sandbox from the admin
sandbox list. That terminates the row and releases the quota slot. The
machine at the provider is then the operator's to check, because Fountain
has no evidence that the delete completed. Do not use a reap as a routine
retry. The audit trail shows sandbox.reset_requested when the fence commits
and sandbox.reset only when the provider confirms the deletion.
When a home parks, Fountain can take a checkpoint of its disk. The operator
turns this on with CHECKPOINT_CREATION_ENABLED, and only a provider with
checkpoints (Sprites) does it. The checkpoint belongs to that one machine.
It can roll the machine back to its last park. It cannot rebuild a machine
that the provider lost. GET /api/sandboxes/:id shows the checkpoint as
checkpoint.
A second agent on a machine
A conversation you attach with sandbox_id usually belongs to the agent the
machine was built for. A conversation of a different agent can also attach, as
a guest, to that agent's persistent machine. A claude agent and a codex agent
can then work on one disk. All of these must hold:
- The launch names the machine's environment and vault.
- The machine and the guest make a claude and codex pair: a codex agent on a
claude agent's machine, or a claude agent on a codex agent's. No agent of any
other runtime can have run there, and no other agent of the guest's runtime.
Two agents of the same runtime would share one set of instructions and
skills, so Fountain refuses them. Gemini and
opencode agents don't take part, and an
acpagent runs a command whose files Fountain can't place. - Fountain also checks that the two runtimes keep their settings,
instructions and skills in separate directories:
~/.claudefor claude,~/.codexfor codex. - The machine has a recorded runtime.
- The request is made with a full-scope API key, such as one you created
yourself. A sandbox's own token (
FOUNTAIN_TOKENinside a sandbox), or any other key below full scope, gets403 guest_attach_requires_full_scopewithreason: "insufficient_scope". Without this rule, an agent could put another agent onto a machine without you, and leave files there that the machine's own agent loads when it next starts.
If any of the first four doesn't hold, the attach answers
422 sandbox_identity_mismatch, whatever key made it. A 403 therefore means
the attach would have passed the pairing checks with a full-scope key.
A guest teammate keeps its machine when you start a new conversation for it:
the team ends the old conversation and opens the new one in its place. Any
other request of a sandbox token that names another agent is refused. That
includes a second conversation of a guest already there, and a channel_id
request with fresh, which leaves the old conversation running beside the
new one.
Every conversation that has attached keeps counting until the machine is reset, because its files stay on the disk. That includes conversations that ended, conversations you deleted, a conversation whose first prompt was refused, and a guest that moved to a machine of its own.
A guest stays on the environment and vault it attached with, even if its agent later moves to another environment. If a guest's environment or vault changes anyway, such as when you rebind a teammate, the guest does not write its new values to the shared machine. Its next prompt, or a new conversation for the teammate, moves it to a machine of its own. The machine it leaves stays with its agent. The guest's working files stay on that machine, its runtime session starts fresh on the new one, and the new machine counts against your concurrency limit alongside the one it left.
The machine stays its own agent's. Deleting that agent destroys the machine,
and the guest's conversations end with it. Deleting the guest's agent ends
only the guest's conversations. A guest cannot reapply its conversation. It
answers 409 rebuild_required, with field set to shared_sandbox while
another conversation is on the machine, and to guest once the guest is
alone there. Each conversation registers the inference credentials of every
other runtime that has run on the machine, including a guest that has since
moved or been deleted, so no transcript there shows another runtime's key.
Skills you write inline are installed into each runtime's own directory. Skills from GitHub are different: every runtime on a machine installs them through skills.sh with the same home directory, so two agents that pick the same GitHub skill may share one installed copy. Reinstalling it at another version then changes it for both agents.
A claude machine created before guests were admitted cannot take a codex
guest. Its codex credential file may already exist on the disk, so the attach
answers 409 codex_inference_conflict. Reset the machine, and the rebuilt one
takes the guest.
Why the lifecycle belongs to the conversations on it
Tie the machine to the runs on it, and three problems solve themselves.
Reclamation gets an owner. When every conversation on a machine is idle, the machine is idle, and no separate record can go wrong.
Cost gets a shape that a user recognises. "This conversation cost something" reads clearly, and "you have 14 sandboxes" does not. Two conversations that each run for an hour on one machine spend two turn hours, on a machine that was busy for one.
Memory gets somewhere to live. The runtime keeps its session on the sandbox's disk. To suspend and wake the machine is therefore the same act as to pause and resume the conversations on it. Read About conversations.
A prompt to any conversation on a parked machine wakes it, and a prompt that
arrives while the machine is being parked is answered 503 sandbox_unavailable: Fountain finishes the park and the retry wakes it. The
wake serves that conversation only. The other conversations on the machine start again
on their own next prompt, and that prompt lands on the machine that is
already awake. The idle clock counts the activity of every conversation on
the machine, so one active conversation keeps the machine up for all of them.
Two prompts that arrive together wake the machine once. The second waits
for the first, and then finds a machine that is already running. If the first
takes long enough that the second gives up waiting, the second is answered
503 sandbox_unavailable with a Retry-After, and the retry lands on the
running machine. Waking a machine is compute again, so it passes the same
account limits as starting one: the balance, the account's concurrent-machine
cap and the deployment's fleet ceiling. A machine being woken counts against
those from the moment Fountain allows the wake, not from the moment the
provider answers.
A wake the provider refuses answers 503 sandbox_resume_failed, and the
machine stays parked with its disk exactly as it was. Retrying is safe.
One seam, four backends
Each backend plugs into the same seam, Managoat.Sandbox. The contract is
executable, and not merely described. The behaviour specifies callbacks,
capabilities and an error taxonomy. A conformance suite that each adapter must
pass covers four things. It covers replay-from-start attach, total
write_stdin, exactly one terminal frame for each command, and allow: [] as
deny-all.
| Provider | What it is | Suspend |
|---|---|---|
| Sprites | The first backend, and the instance default. | Parks on its own, scales to zero, keeps the disk. |
| E2B | Hosted. | An explicit pause, with a snapshot of filesystem and memory. |
| Daytona | Hosted, and you can self-host it. | An explicit stop, keeps the disk, archives a long park. |
| Self-hosted runner | fountain runner on a machine the user owns. |
Stops the processes. The directory stays. |
Capabilities are honest, and that has consequences
An adapter advertises only what it can do, from :suspend,
:network_policy, :attach, :tty, :checkpoint and :public_url. The
lifecycle then degrades to match, and it pretends nothing.
A provider without :suspend destroys on idle. It does not fake a park. A
resume with a fresh disk would lose the agent's memory without a sound, which
is worse than an honest teardown.
So the idle timeout does not mean the same thing everywhere. On Sprites it costs nothing. On a provider with no suspend it costs the memory the agent works from. Read Change sandbox lifetimes.
Three more consequences matter.
A runner gives you no isolation. It runs in trusted mode on a machine the user owns, with no egress policy. It is a way to use hardware you already have. It is not a security boundary.
Only Sprites gives you a TTY.
A sandbox never moves. A parked disk stays where Fountain wrote it, so a wake never crosses providers. Change the instance default and only new sandboxes follow it.
Anything an agent serves can be public
A provider that gives a sandbox its own HTTP endpoint advertises
:public_url. Fountain sets the URL in the sandbox as SANDBOX_URL, so an
agent can answer when you ask where it runs.
On Sprites that endpoint asks for a platform credential by default. Fountain opens it, because an agent that serves a page serves it to a person, and that person holds no such credential.
So anyone with the URL can reach what an agent serves. Set
config :fountain, :sprites_public_urls, false to keep the default instead.
A sandbox then keeps its URL, and only a token holder can open it.
Not a container you can hand us
Fountain builds no image and stores no image. A provider applies your environment as a recipe at provision time. Two conversations from the same environment install their packages one by one, on their own machines.
That costs provision time. It buys you no registry, no build pipeline, and no garbage-collection problem.
A Sprites machine that drops quiet connections
This is a known fault in the Sprites platform. Fountain cannot fix it or work around it.
Some Sprites machines drop a network connection that is idle for about one second. A connection that moves data without a pause is not affected. The machine is not told the connection is gone, and neither is the other end. The program on the machine waits for data that never comes.
We measured this on 2026-09-26 on a hosted Codex machine. We ran the same requests through the same credential broker session on that machine, on two other Sprites machines and on a computer outside Sprites.
| Request | Affected machine | Other machines |
|---|---|---|
| A 20 MB download | Complete | Complete |
| One byte every 0.3 seconds | Complete | Complete |
| A reply that starts after 1 second | Nothing arrives | Complete |
| One byte every 2 seconds | The first byte, then nothing | Complete |
The affected machine got worse over three hours of use. Early on, about half of its model replies finished. By the end, none did. Nothing inside the machine differed from the healthy ones. The network settings were the same and it had no firewall. The drop happens outside the machine, in the Sprites network. We do not know what starts it. The affected machine had opened several thousand short connections through the broker when it failed. That may be related, but we have not shown it.
What you see. A model streams its reply and pauses while it reasons. On an affected machine the stream stops at the first pause. Codex prints:
stream disconnected before completion: Transport error: timeout
stream disconnected before completion: Transport error: network error: error decoding response body
Codex retries five times, each after about 30 seconds, and then ends the turn. The reply stops partway. A later retry can repeat text from an earlier attempt. Other runtimes are also affected if their replies pause for a second or more.
What to do. Reset the machine: fountain sandbox reset <id>, or
DELETE /api/sandboxes/:id. The conversations stay, and the next prompt
builds a new machine. A reset also deletes everything on the machine's disk,
as described in Two modes. On 2026-09-26 new machines did not
have the fault.
Fountain does not reset an affected machine by itself. The operator of a
Fountain deployment can see affected machines on the admin page
/admin/broker, under Sandboxes cutting streams, as described in
Observability.
Where to go next
- The sandbox contract, the executable version of this page.
- About conversations, whose lifecycle this one follows.
- Change sandbox lifetimes.
- Add a provider, if you want a fifth.