Run a mesh
Guide (informative) · For: operators · Prereqs: Quickstart
Day-to-day operation of a local mesh: what cotal up actually runs, how spawning
resolves personas, harnesses, and models, how to reach a mesh from any directory, and the
operator-only maintenance verbs. Every command’s full flag set is in the
CLI reference.
The stack
Section titled “The stack”cotal up brings up the whole local stack and bare cotal down stops it. Managed
agents stay running as unmanaged OS processes; pass --with-agents to take them
with the stack. Seats of the built-in pty runtime run inside the manager process, so
they stop with the manager either way. Ctrl-C on a foreground up stops the manager through the
same stop as bare down and prints the same report; when that stop is refused, for example because
the manager cannot prove it can spare, Ctrl-C leaves the stack running, and you end it with
cotal down --with-agents. A current manager records what its stop does with its seats before
bare down signals it. A pre-pin legacy manager instead receives a reduced-guarantee
warning and is signalled according to the documented upgrade contract. Its running binary
may still carry the older destructive SIGTERM handler, so the CLI does not claim its
pre-signal agent inventory was spared; those agents may have been reaped.
- Broker: a local
nats-server(logs to.cotal/nats.log). - Delivery daemon: the durable backstop, auth mode only (what it does).
- Manager: a detached supervisor answering the control plane, so
cotal spawn --detachand thecotal_spawntool work right afterup.
Cotal creates the presence bucket in memory storage. Its records are liveness that every endpoint
rewrites each heartbeat, so nothing is lost when a broker restart empties it, and nats-server’s file
store write latch cannot reach it. A broker stop removes the memory stream itself, so every cotal up,
including the resume after cotal down --preserve-state, creates it again before any daemon starts.
JetStream fixes a stream’s storage class when it is created, so a presence bucket created file-backed
by an older cotal stays file-backed until that stream is recreated.
A file-backed presence bucket can remain open and watchable while refusing every write. A bound
endpoint reports this as presence-write-stuck after one full presence TTL of consecutive failures.
The roster is last-known while that condition is active. Restarting the broker clears nats-server’s
in-memory store latch and preserves the JetStream root. Current credentials split the required stream
authority: the cotal up provisioner can create the presence stream but cannot delete it, while the
teardown credential can delete it but cannot recreate it. Cotal therefore reports the condition but
does not attempt an unsafe partial delete-and-recreate. Stop and restart the broker to recover.
A broker below nats-server 2.14.5 carries the latch (nats-server fixed it in 2.14.5). When cotal up starts or finds such a broker and the space’s presence bucket is file-backed, it says so. A memory-backed bucket gets no warning. A broker below the SPEC §13.12 floor of 2.12 is refused at connect with the floor sentence.
Three modes:
- Default (static auth). JWT-authed, on by default: sender authenticity and per-agent ACLs, enforced by the broker (how).
--user-auth --idp <url>. Per-user auth: peoplecotal loginonce, the operator grants their agents on the actor ledger, and every connect is authorized live against that grant. Starts the space’s auth service alongside the broker (how).--open. An unauthenticated, live-only dev mesh (no auth, no delivery daemon). For quick local experiments.
The broker and local services bind loopback by default. --host 0.0.0.0 widens the broker
bind independently of the auth mode, so “network-reachable” never silently means
“unauthenticated”. With no explicit --server, cotal up auto-selects a free local port when
the default address is already held by another project; an explicit --server fails loud on
collision.
--host is a boot flag, not a live rebind. A fresh cotal up writes the generated
.cotal/auth/server.conf (project-local, not ~/.cotal) with that bind and starts nats against
it. If anything is already answering at the mesh URL, up refreshes the recorded mesh and
leaves the running nats listener alone, so passing --host 0.0.0.0 on a live or orphaned
broker does not change who can connect. To change the bind: cotal down, then cotal up --host <addr> against a stopped broker so the generated file is rewritten. Do not edit server.conf
by hand; the next real boot overwrites it.
On a stopped shared broker, up renders every persisted space account and every enabled
space’s auth-callout account into the resolver preload, regardless of which space starts
the broker. A missing callout account for an enabled space stops the boot rather than
starting with a reduced resolver. An already-running broker is refreshed without rewriting
its config.
A broker-only host is a first-class up mode. cotal up --no-manager boots the broker and, in
auth mode, the delivery daemon, and no local manager, so the broker host never has a manager to
stop and never leaves a manager slot stale. A refresh under the flag of a mesh whose manager is
live refuses rather than keeping or stopping it: cotal down manager first. Without the flag,
auth-mode up still starts nats, the delivery daemon, and a
local manager. A space may run more than one manager, addressed by instance id
(control surface); putting no manager on the broker host
is a topology choice, not a singleton invariant. A manager whose boot inventory has no
available connector does not take unpinned spawn/launch on the class rail, so a sibling
that can launch the harness can. describe still rides the class rail, so an unpinned spawn
can bind-fence when that skip member answered describe; re-issue, or pin --on. Pin one
instance with --on when a partial inventory still answers with a harness refusal. The
supported split is:
# broker host (project root that owns the generated conf, pidfiles, and logs)cotal up --detach --host 0.0.0.0 --space main --no-manager# no local manager starts: the summary lists nats-server + delivery daemon, and there is no# `.cotal/manager.<spaceKey>.log` to wait for on this host
# manager host (registered remote mesh, same space)cotal meshes add --server nats://broker.example:4222 --root ~/meshes/maincotal supervise --space main --server nats://broker.example:4222Wait for ✓ manager up in .cotal/manager.<spaceKey>.log on the manager host before spawning
agents. On a broker host started without --no-manager, cotal up --detach prints ✓ running in the background: with manager listed once the manager pidfile is live; stop that local manager
only after the ✓ manager up line. A host started WITH --no-manager never runs one, so neither
the wait nor the stop applies there. That detach stdout is not a safe teardown boundary: it is
pidfile liveness, not ✓ manager up. ✓ manager up is supervise’s post-start line after
await mgr.start(). cotal down manager after only the detach line can still default-terminate
the child during registration after it has taken the governance slot. Stopping before that
post-start log line can leave the endpoint governance slot held until the holder’s gate
reopens past the stamp (the successor’s boot heal, or
cotal reconcile-gate when that boot cannot run). See
Gate recovery.
Standalone cotal deliver --creds is not a repair for that split. Production renewal needs
the manager and the daemon to address one credential store. The manager renews its own service
credential inside that credential’s own window and re-dials its service connection with the
renewed credential; if the connection closes and cannot be restored within about forty seconds
it releases its lease and exits so a restart can serve, while a broker that is briefly gone is
waited out. Separate host filesystems still
leave manager root A writing and the daemon reloading root B; that composition is refused
while the daemon stays up. Before every remint the manager challenges the delivery daemon’s
store identity, and the answer must come from the process holding the delivery lease: the
reply names the answering endpoint and the manager reads the lease row itself under its own
credential, so a non-holder answering on the queue-grouped admin rail is refused instead of
counting as the daemon’s store. A rail that reports no responder is also settled from the
lease row, so a live holder on record makes that outcome a refusal rather than an absent
daemon. Keep delivery on the broker host under up, and share one store
only when you are composing a hosted pair (embedding).
On the --no-manager split above, the manager host’s manager stays off the daemon-credential
renewal lease once its store check finds the daemon on another store. A filesystem store is named
by its root and by a random id in .cotal/store.id, which the copied .cotal/auth does not carry,
so this holds when both hosts use the same root path. cotal doctor auth --fix on
the broker host then renews the daemon credentials once they pass their renewal point.
Split host bind
Section titled “Split host bind”A remote manager cannot reach a loopback broker. After changing --host, confirm the
generated host: in .cotal/auth/server.conf and that nats is listening on that address
before registering the mesh on the manager host. Detached child logs stay under the project
.cotal/ that up ran in (see When something looks absent);
they are not ~/.cotal unless that directory is the mesh root.
A user-auth mesh can expose only its credential exchange through an operator-owned HTTPS reverse proxy while leaving the existing local exchange untouched:
cotal up --user-auth --idp https://idp.example/api/auth \ --exchange-public-port 7443 \ --exchange-public-url https://auth.exampleThe public listener itself still binds 127.0.0.1:7443; configure the proxy to terminate TLS and
forward to it. It serves only /health, /jwks, /exchange, /manager-service-authority, and
/.well-known/cotal-mesh with the documented methods. It needs no local file capability: the
signed IdP JWT or managed-agent actor token is the proof, while the original loopback listener
remains capability-gated. Add
--exchange-trusted-proxy only when that listener is reachable exclusively through your trusted
proxy; it keys failure throttling by the last X-Forwarded-For hop instead of the socket address.
The well-known bundle includes IdP pins and a deny-all sentinel credential, so fetch it only from
the configured HTTPS origin. To change these listener flags, stop and restart the mesh; a refresh
of an already-running service does not replace its bind or proxy policy. See
Identity & auth for the trust boundary.
Remote supervised seats by enrollment
Section titled “Remote supervised seats by enrollment”A remote seat does not need to run cotal login when the mesh owner pre-mints a single-use
enrollment for it. Mount the enrollment URL as a private file, place the seat persona on the remote
machine, and launch the foreground seat:
COTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \ cotal spawn --config ./worker.md --space mainThe URL is redeemed once with an unauthenticated GET. Redirects, off-machine plain HTTP, retries, and login fallback are refused. If the seat has no mesh record yet, the enrollment response’s stock user-bundle fields register it before the launch. The returned actor token then uses the same remote auth-service exchange as a login-provisioned agent. The enrollment URL and file path do not enter the preflight or harness environment. A failed or reused enrollment leaves no actor material on disk; ask the owner for a fresh enrollment. When the foreground seat exits, this machine’s credential files are removed and the mesh-side grant stays until the mesh operator revokes it; the launch line says so. The exact server contract is in Enrollment redeem.
cotal status prints the detailed setup, process, registry, and live mesh status. Its Machine
section names the running CLI’s source checkout, installed package root, or npx package root beside
the version. It has one row per installed connector, which reports whether the executables that
connector declares in requires are on PATH. Status, setup and the manager’s preflight resolve them
the same way: an entry written as a path is checked as given, and a directory never counts as the
executable. A connector whose setup provider reports health adds its
own rows above those. The Claude Code connector reports its plugin and its skills plugin, and a stale
skills row names the installed and CLI versions it compared. cotal setup (after the first run) prints the compact card.
Before reporting ready, the manager resolves every installed connector’s declared harness
binaries against its own environment. A missing binary does not stop unrelated manager work: boot
continues, but prints a named connector <name> unavailable line and records that reason in the
manager’s status response. Available connector rows record the absolute paths boot resolved.
A spawned seat and a seat resumed after cotal down --preserve-state both launch from those paths,
and both are refused with the recorded reason when their connector’s row is unavailable.
cotal models takes the same rule and reports that reason in place of the catalog, so it agrees
with a launch about a harness installed or removed after boot. The manager looks again only when it
restarts. A connector registered after boot has no row, so spawn, resume and cotal models check
its binaries on PATH when they run.
On an authenticated manager start, unfinished static lifecycle rows reconcile while the control
endpoint is already serving. The manager status response reports
the staticReconciliation state, the last sweep counts, and each failed alias with its durable
phase and literal disposition. cotal status --components reports the state and per-alias failure
details. A slot row carrying a DEL or PURGE marker stops the sweep before it plans any alias, and
the manager log names the row. A failed exact terminal is retried in the same process after 1, 5,
and 30 seconds. Each attempt re-reads the durable slot and re-enters the same deterministic terminal
operation; the delays only schedule work and never release the lifecycle fence. The terminal’s
cleanup removes the lifecycle’s credential file and its broker durables and read-ACL row as separate
steps. A file that cannot be removed does not leave the broker footprint behind, and its failure
keeps the alias held for the next attempt.
On shutdown, the manager fences new reconciliation work and waits for an exact terminal that already
started. The current serial sweep stops before its next alias, and startup cannot publish the manager
service after stop() completes.
The four-attempt budget is per manager process. An exhausted row stays held and reports
retry-exhausted with the remedy to restart the manager. The next process derives a fresh budget
from the still-authoritative durable row. A recovered row remains visible until the next static
reconciliation sweep, then clears. This component reports reconciliation outcomes. It does not say
whether footprint cleanup completed independently of the terminal result; that separate durable
projection remains tracked by #1274.
cotal service install is the supported way to run the manager as a user service
(CLI reference): a systemd user unit on Linux, a launchd agent on macOS, one
per mesh, surviving logout and reboot. On Linux that needs user lingering: install refuses while
it is off and prints the root command that enables it. It installs only
the manager; the units below remain the process models for every other component, and they are
still examples of process models for those: copy them only after you decide which processes
the unit should own.
Supervising the detached stack
Section titled “Supervising the detached stack”cotal up --detach is a launcher: it starts the broker, delivery daemon, and manager, reports what
started, then exits. Do not wrap it in a systemd service with Type=oneshot and
RemainAfterExit=yes and treat systemctl is-active as stack health. That unit becomes active (exited) when the launcher exits successfully and stays active even if every detached process dies.
When up --detach can identify that exact unit shape, it prints a warning but keeps the requested
startup behavior.
For a single-host stack, keep cotal up itself in the foreground so systemd tracks a long-running
process and restarts the stack if that process fails:
[Service]Type=simpleWorkingDirectory=/srv/cotal-meshExecStart=/usr/bin/cotal up --space main --host 0.0.0.0Restart=on-failureRestartSec=5sAn active unit then proves the foreground launcher and broker are still running, but it still does
not prove that every child component serves. Pair it with the component check below. Also remember
that cotal up starts a local manager as well as the broker and delivery daemon; run
cotal up --no-manager (add the flag to the unit’s ExecStart too) on a host intended to be
broker-only, so the unit and the host agree.
Seats spawned by the built-in pty runtime run with oom_score_adj 500, so under memory
pressure the kernel prefers a seat over the broker, manager and delivery daemon, which are left as
they were started; the extension runtimes do not own the seat’s process and get no preference.
That Type=simple shape puts nats in the unit’s cgroup with the foreground up process. A
Restart=always (or on-failure) of this unit therefore restarts nats as well, so remote
managers drop for the time it takes the broker to come back. Wrapping cotal up --detach in
Type=oneshot with RemainAfterExit=yes does not move nats out of that cgroup. Detached
spawn starts a new process group, not a new systemd cgroup, and the default
KillMode=control-group still signals every process left in the service cgroup on stop or
restart, including the nats PID. Escaping that cgroup needs an explicit unit setting such as
KillMode=process, or a separate nats unit; this CLI does not ship that escape. The
Type=oneshot unit below is a cotal status --components liveness check, not a
--detach launcher. Neither trade is universal from
Type=simple alone; it follows from which processes the unit actually owns. cotal service install covers only the manager, so for the broker and its siblings pick the example that
matches the ownership you want, and treat
systemctl is-active as unit health, not mesh health.
A broker that crashes under that foreground up keeps its mesh record and exits non-zero, so the
unit’s restart takes the repair path against the recorded store rather than starting a second one.
If the deployment deliberately uses cotal up --detach as a boot action, monitor observed state
instead of the launcher’s exit:
[Unit]Description=Check Cotal component liveness
[Service]Type=oneshotWorkingDirectory=/srv/cotal-meshExecStart=/usr/bin/cotal status --components --space mainRun that check from a systemd timer or another monitor and alert on a nonzero exit. The command
distinguishes absent, not-serving, and refused components and never treats a sibling’s health as
proof. Its delivery-process check is local to the broker host, so run it there. On a split topology,
also probe the broker URL from the manager host and monitor the manager’s own service there. A remote
manager cannot observe the broker host’s delivery PID, and an active unit on either host says
nothing about the other host.
Stop one part without tearing down the mesh by naming its registered component: cotal down manager, cotal down delivery, or cotal down web. Component names from installed extensions
join the same surface; cotal down with no names retains whole-stack behavior and
leaves managed agents running as unmanaged OS processes, except pty seats, which stop with the
manager. cotal down --with-agents is the previous reap. If a pinned manager has no
spare-capability record, stop its managed agents explicitly before running that whole-stack
command. A current manager always publishes the record, so it is absent only for an older manager,
which may not understand the reap request.
Remote supervised agents
Section titled “Remote supervised agents”On a remote user-auth mesh, foreground cotal spawn remains the default participant path. A
participant can run detached agents only after the host advertises and operates the remote manager
authority service, and the participant’s actor-ledger row includes supervise. This is not implied
by spawn or admin.
The participant’s loopback/operator exchange obtains one closed manager-service view for its
ordinary derived owner, a fixed server-selected manager actor, and one opaque manager instance.
The host, not the participant, issues the public-nkey JWT material via the replay-safe,
lifecycle-bound prepare → activate → renew exchange, plus a one-shot target-pinned retirement
request for a host-managed terminal. It never exports the space signer, a static
provisioner credential, or generic storage authority. Remote registration publishes its service
status at the registered revision and current process epoch, so manager-caller selection can find it.
Stock participant supervision asks its host to enroll a detached agent and to prepare its terminal retirement, over the same manager-authority transport. The stock auth service answers both when it runs with a public exchange face: it grants the agent under the participant’s owner at a lifecycle UID it picks, bounded by the participant actor’s own grant, provisions that UID’s durables, and on retirement releases them and revokes the grant before the manager’s terminal rail. It refuses a second enrollment of a name whose grant still stands until that agent’s retirement is prepared. A host platform that keeps these writers in its own storage intercepts both requests on its own route instead. Copying host secrets or actor-ledger files to a participant is not supported. Foreground spawning and operator-local hosted managers use their existing paths.
The remote manager that cotal supervise starts can host workflow runs through its host: the host
admits each run and signs only the run’s own driver, mediator and operator credentials. A logged-in
user’s cotal run start against it is admitted: the auth callout issues the user’s manager
connection, and the host binds each run to the owner who registered the manager. The run spawns
agents that user owns, enrolled by the host like any detached spawn, with the reach the user’s own
row grants when the spawn runs. A spawn may be placed on that manager and on no other instance. The
host’s own manager refuses user-auth runs by name.
User-auth run start
records the path.
The registry entry decides the broker URL supervise dials, so a mesh published over wss:// is
dialed as a websocket. The manager-authority registration it runs first also takes its TLS
requirement from that entry, so the prepare credential is not exchanged over a plaintext
connection the record did not describe. cotal meshes add records both.
When the authority service, login, or renewal is unavailable, the remote manager degrades fail-closed: it refuses new agents, restarts, and credential replacement rather than pretending local authority exists. Existing agents remain live only while their independent credentials are valid. A hosted composition must revoke the managed grant and finish its resumable release before it requests terminal retirement. Deleting DM or delivery consumers is not retirement and must not reset a resumable lifecycle’s frontier or pending state. The alias remains held until the terminal barrier confirms. Restore service and renew successfully before asking it to recover an agent. See Identity & auth and the CLI reference.
Spawning agents
Section titled “Spawning agents”cotal spawn # foreground: your default agent, in this terminalcotal spawn reviewer --detach # supervised: the manager runs it in a PTYcotal attach --name reviewer # watch/type into a detached agent (Ctrl-] detaches)cotal ps # what the manager is runningcotal stop --name reviewer # stop oneHow a spawn resolves:
- Persona. A bare
cotal spawnuses.cotal/agents/default.md; a positional name picks.cotal/agents/<name>.md;--configtakes an explicit ref or path. SetCOTAL_DEFAULT_PERSONA=<name-or-path>to change the fallback. Fields and format: agent files. - Harness. Resolution order is an explicit
--agentorcotal_spawnagentargument, then the persona file’sagent:pin, then the invoking caller’sCOTAL_DEFAULT_AGENT, then the manager’sCOTAL_DEFAULT_AGENT, then the product default (Claude). Compared in Connectors; per-connector guides: Claude · OpenCode · Hermes · pi. - Model.
--modeloverrides the persona file’smodel:(Claude:opus/sonnetor a full id; OpenCode:provider/model). Connectors that expose a catalog report it viacotal models --agent opencode: model ids plus available variants; pick one with--model provider/model --variant high. - Tools. A spawned Claude Code agent gets the cotal tools plus the MCP servers the cotal
config shares, which first-run
cotal setupfills with your own; narrow them per spawn with--share-tools(config). - Launch options.
--opt key=value(repeatable) passes a native harness flag straight through; a persona or manifestlaunchOptions:mapping does the same declaratively (a--optwins per key). It is a raw passthrough, with no allow/deny list: Claude renders each as--key value(a bare--keyfor an empty value), OpenCode merges them into its agent config, and Hermes has no option surface so it fails loud. The trust boundary is thespawncapability itself, not the flag set, so grantingspawnis host-launch authority (security). A key must be a plain flag name; malformed or prototype-polluting keys are refused.
Detach from an attached PTY with Ctrl-] (the agent keeps running); rebind it with
COTAL_DETACH_KEY=ctrl-<char> when it clashes with a keybinding inside the agent’s TUI.
Runtimes. The manager spawns into a pty by default. It spawns the PTY in-process on
every platform, so replacing the manager worker closes its seats and the pty runtime gives no hot
update. Any manager stop, bare cotal down included, stops and deprovisions those seats. A stopping
manager refuses new spawns and first waits for the ones it already accepted, so their seats stop too. On Linux
it can still adopt and reap seats that an earlier manager launched under a detached per-seat
custodian, so those seats drain under the new manager; it starts no new custodian. A custodian whose agent has exited exits a few seconds later on its own. cotal seats
lists the custodians left on the machine, and cotal seats --drain retires the ones whose agent
has exited while keeping every seat whose agent still runs (cli.md). When a pty
agent exits on its own, in-process or under a custodian, the manager logs a seat reaped: line
with the exit code and, for a signalled child, the signal number. The line ends with the last line
the child printed that starts with a connector’s [cotal-<name>] or [cotal-<name>/<part>]
prefix, cut to 240 characters, when it printed one. A custodian keeps the same record beside the
seat’s custody record, so a later reap of that seat, including one by a
successor manager, reports how the child ended. When the custodian cannot write that record, it
says why in the seat’s custodian.log, and a later reap of a child that ended on its own reports
the record as missing or unreadable. Optional runtimes are installed
through the extension surface, for example cotal ext add @cotal-ai/orca, then selected with
--runtime orca (similarly @cotal-ai/tmux, @cotal-ai/cmux, and @cotal-ai/herdr). They put teammates in native
terminal surfaces rather than manager-owned PTYs. Those surfaces stream no exit, so the manager
asks the runtime every five seconds whether a seat’s process has ended, and frees a seat that ended
on its own once the runtime proves the exit. Its seat reaped: line says the exit detail is
unavailable from that runtime. Runtime names are open-ended and resolved from
the registry; a missing provider or app throws, never silently falls back
(architecture).
Mesh registry
Section titled “Mesh registry”cotal up records each running mesh in a machine-local registry
(~/.cotal/meshes/space.<key>.json, named by a case-safe hex encoding of the space: broker URL, the project root holding its creds and
personas, and its mode). So a bare cotal spawn <persona> from any directory joins the
running mesh with the right credentials instead of mistaking the cwd for a space:
cotal use <name>sets the default from every directory, including inside another mesh’s project.--space <name>overrides it for one command.- When one broker has records for several spaces,
cotal up --space <name>refreshes that named space. - A refresh rewrites only what that command decided: the server, root and mode, the user-auth
endpoints, and an explicit
--hostor--max-sessions. Every other field, such as the TLS requirement, is kept as the record stands when the refresh writes it, so a change another command made during the refresh survives. If the record was removed during the refresh,upfails instead of writing it back. - With no live selected default, a project with its own
.cotal/resolves to that project’s mesh; otherwise one running mesh is used automatically and several are an error. cotal mesheslists them (a*marks the default);cotal downremoves the entry.
The registry stores a path, never a secret; trust material stays in each project’s
.cotal/auth. If the mesh is down or won’t take your creds, spawn fails with one
sentence, never a raw NATS trace.
Meshes you did not start here
Section titled “Meshes you did not start here”A mesh running on another machine has no cotal up on this one, so register it by hand:
cotal meshes add # guided: asks for the broker, probes it, offers what it findscotal meshes add optiplex --server nats://100.90.12.34:4222 --root ~/meshes/optiplex \ --allow-unencrypted-overlay # see below: an overlay address needs thiscotal meshes rm optiplexOn a terminal, a bare cotal meshes add walks you through it: it probes the broker you name and
reports whether it is open or requires credentials, offers the spaces the folder already holds
credentials for, and shows the record before writing it. Scripts and agents keep the flag form -
without a terminal nothing prompts.
--root is the local folder holding that mesh’s .cotal/auth and .cotal/agents (its personas);
the mode is inferred from what that folder holds.
The instance identities of the manager and the user-auth service are not part of that folder.
Each root keeps its own in .cotal/space.<hex>/, so cotal supervise or cotal up --user-auth in
the root you copied the folder to starts an instance of its own. A root last run by an older Cotal
still holds them in .cotal/auth, as manager-instance.<hex>.json, manager-siblings.<hex>.json
and space.<hex>/.cotal/auth/auth-instance.<hex>.json. Delete those files from a copy of such a
folder before the first cotal supervise or cotal up there.
Know what you are copying. For an authenticated mesh that folder carries the space’s account
signing seed, which is the authority to mint any identity in the space. A machine holding it
is a certificate authority for the mesh rather than a client of it: anyone who reads it can
impersonate any agent, read every retained channel and DM, change ACLs, and keep issuing
themselves credentials. There is no per-machine revocation; undoing it means rotating the signing
key and re-minting every credential in the space. Copy it only to machines you would trust with
the whole mesh. cotal mint on its own does not substitute here: registering an auth mesh needs
signing material that composes, which a minted user credential is not. The
broker is probed before the record is written, so a bad address or a credential that mesh will not
accept fails at registration rather than at your first spawn (--force records it without verifying,
useful when the mesh is simply down right now).
Which addresses you may register
Section titled “Which addresses you may register”Registering a mesh is how this machine starts sending agent credentials to a broker it does not run. NATS announces itself in plaintext before anyone authenticates, so an attacker on the path can pose as the broker and read the credential out of the connect unless the connection requires TLS, which is recorded on the entry and enforced on every dial through it.
What the record will require decides what you may register:
- Without required TLS, the address is the gate: loopback (
127.0.0.0/8,::1), or your private overlay (100.64.0.0/10,fd7a:115c:a1e0::/48) with--allow-unencrypted-overlay. The tunnel provides the protection, and this command cannot check its state. Hostnames are refused because the lookup would choose which machine receives your credentials. - With required TLS, set
--tlsor use atls://URL. The recorded scheme enforces the TLS requirement. A hostname or public address is accepted because the certificate chain and hostname check identify the peer. A registration whose broker cannot complete the handshake fails unless you pass--force, which records the entry without verification.
Ordinary private ranges like 10.x and 192.168.x are refused in both modes. A café’s wifi
is private but does not belong to you, and no public CA issues certificates for those ranges. An
address spelling changes nothing: [::ffff:192.168.1.10], 3232235786, 0300.0250.01.012, and
192.168.257 all resolve to private addresses and receive the same refusal as the dotted form.
--force exists for a mesh that is down. It never permits an unsafe credential destination.
Registering a hosted user-auth mesh
Section titled “Registering a hosted user-auth mesh”A user-auth space’s IdP pins are established where the mesh runs and are never guessed. Register
one from supplied trust: --user-auth-file bundle.json (exported on the mesh’s machine), or
--from https://auth.example, which asks before it contacts the address at all, fetches the
discovery document at /.well-known/cotal-mesh under that address over HTTPS, shows you the pins,
and asks again before adopting them. A URL that already ends in /.well-known/cotal-mesh is
fetched as given. Redirects are refused because a 302 can walk a pinned fetch down to
plaintext or onto another host, and the pinned exchange must be an https:// URL too. The one
exception is an exchange on this machine, where nothing leaves the box: plain http:// is
accepted for a loopback literal (127.0.0.1, ::1, and any spelling of them), but not for
localhost, which a hosts entry or poisoned lookup could point elsewhere. Use the
literal. Registration checks that the pinned exchange
answers /health and /jwks as the pinned issuer. It also checks that the broker refuses a
bare connect; that refusal is the pass. The bundle’s sentinel credentials are written to a private (0600) file
under the entry’s root; the registry itself never carries the secret.
Without required TLS, an overlay address is refused unless you accept the dependency
explicitly, with --allow-unencrypted-overlay. The address is not the guarantee: it is protected
while the tunnel is up, and if the tunnel is down that range is ordinary carrier-grade NAT and
whoever answers the dial receives your credentials. Only you can know which it is, so the command
asks you to say so. Your acceptance is recorded on the mesh entry rather than printed and
forgotten, and the guided form asks the same question instead of taking the flag.
With required TLS (--tls, or a tls:// URL) that consent is no longer asked for, and the
flag is not needed: the handshake is what protects the connection, so the acceptance it stood in
for has been replaced by proof rather than promise. cotal meshes add <space> --server nats://100.64.0.1 --tls registers an overlay address with no prompt, no flag and no recorded
acceptance. This is the “the flag disappears once the broker can be served over TLS” case, and it
has now arrived.
This gate is on registration. cotal join --creds --server <url> deliberately takes an
explicit connection at face value and does not consult the registry, so it is not covered. Join
that way only to an address you would have registered.
The connection is still probed first, with the same second try at the longer budget the registry preflight uses, so a slow link reads as a connect that did not finish within that budget and a refused port reads as a broker that is not running.
Records added this way are removed only by something that names them. A failed liveness probe
does not delete any record: an unreachable broker, local or registered by hand, is shown as
offline in cotal meshes. A bare command does not count that offline record as running;
name it with --space to restart it. cotal down / cotal clean all still drop an up record for the
project they are tearing down, and they leave a hand-registered one alone even when --root
pointed at that project. A cotal up for that space refuses outright unless it is that same
endpoint: finding a broker already answering there is a refresh that starts nothing and leaves the
record’s provenance alone, while actually starting the broker for that space, server and root
makes this machine the one running it, so the record becomes an ordinary local one that
cotal down clears. The refusal names cotal supervise --space <s> --server <url> (plus cotal deliver) when the registered broker is on another host, and cotal meshes rm when it is local.
cotal meshes rm drops it and re-registering with --force replaces it. rm only forgets a
mesh. To stop one running here, use cotal down.
Watching
Section titled “Watching”cotal console is the terminal view (TUI on a real terminal, plain line stream when
piped); cotal web is the browser dashboard. Both are read-only observers; the
walkthrough is Watch a mesh.
History
Section titled “History”Retained history is operator-owned. cotal clean history --force purges a space’s
retained channel history; --dms also purges DMs (cotal history clear is an alias).
It is deliberately not an agent tool: agents cannot wipe the record
(identity & auth). For a stopped mesh, cotal clean store --force deletes the on-disk JetStream store outright, and cotal clean all --force
also resets the space identity (CLI reference).
Offline backup
Section titled “Offline backup”For a coherent durable cut, preserve the whole stack first, then create the artifact while it stays down:
cotal down --preserve-statecotal backup create ./space-backup # full by default# later: deliberately resume the unchanged sourcecotal up --detach# or, from another preserved cut, restore before the normal listener openscotal up --restore ./space-backup --detachA refused cut leaves the mesh running and unfenced: fix what the refusal names and run
cotal down --preserve-state again.
Use --store-dir on both preservation and backup for a custom JetStream store. A store cap set
with cotal up --max-file-store <bytes> travels with the preserved state, and the resume renders it
again. nats-server reads the cap once at start and refuses a config reload that changes it, so a new
cap always needs a restart. The cut records the chat stream’s frontier per retained seat, so a
resumed seat catches up from there instead of replaying its channels. registry is the
only partial selection (backup create ... --only registry; up --restore ... --restore-only registry). Backup never stops or restarts a mesh implicitly, never opens the original store, and
does not contain credentials or trust secrets. Backup/restore in every auth mode, open included,
uses isolated, operation-specific maintenance logins; normal agent credentials cannot enter that
listener. Full
restore requires the same space and exact current local trust continuity, recreates conservative
consumer checkpoints bound to their snapshot stream sequence state, and resumes retained agents under
their original principals. The trust commitment includes the cryptographically validated full
operator/system/data-account root chain as well as static/user authority state. A registry-only
restore completes canonical empty infrastructure but leaves retained agents stopped because their
DM/DLV/TASK/ACL state is outside that selection. Authenticated restore validates the complete space
trust bundle before staging or changing the preserved store. Interrupted ordinary resume retries the
same durable attempt after its prior listener is stopped. Restore re-entry can recover a surviving normal listener
only when its attempt nonce, NATS server name, process owner, endpoint, and target-store identity all
match the fsynced proof. A provably dead uncommitted owner is retired under lock and replaced with a
fresh attempt-bound listener; an occupied foreign listener or ambiguous owner is never adopted. The
manager commit validates while retained cleanup is still suppressed; the CLI durably records its
attempt-bound 64-hex token in manager-committed / resume-committed before finalizeResume can
release suppression. A retry from either committed state goes straight to exact-token finalization;
failure preserves the committed gate and retained cleanup suppression. Missing commit evidence,
interrupted finalization, a live recorded endpoint despite missing pidfiles, or ambiguous proof fails closed. See the CLI
backup and restore contract for artifact, checkpoint, fallback,
disaster-consent, and degraded-recovery details.
Personas from the CLI
Section titled “Personas from the CLI”cotal personas manages the local catalog offline: list (--running overlays live
markers), show <name>, edit <name> (re-validates on save), new <name>, rm <name> --force. The runtime write is cotal_persona; the runtime read is cotal_personas
(list / show), both over the wire with the manager’s ownership checks. Fields: agent files.
Gate recovery
Section titled “Gate recovery”A manager that dies mid-registration leaves its issuance gate frozen under that registration
op. The freeze is correct: it stops two incarnations serving at once. The successor now completes
that dead op on boot, using the same guard as cotal reconcile-gate: it
acts only when the freeze-holder is affirmatively gone under a complete CONNZ sweep (gone and
sweepComplete=true). If the dead op’s spec write committed, it finishes that same freeze
(promote and reopen at the committed registration revision). If the spec did not advance, it
abort-reopens the gate (generation+1, processEpoch unchanged) and continues the normal takeover.
Boot heal and the following re-registration use separate one-shot executor windows, so a large
predecessor family cannot spend the takeover’s credential lifetime. If that later registration
still crosses a connection lifetime, it retries the same frozen operation with fresh authority
and resumes verified-holder progress instead of freezing a new generation.
A live holder, an incomplete sweep, or an unreachable delivery daemon still
refuses. Silence is never evidence of death, and there is no TTL. If holder verification is
interrupted, the frozen operation resumes from its durable, operation-and-gate-revision-bound
progress after liveness is checked again. A later freeze cannot reuse that progress: the cursor
binds the exact op, gate revision, and holder set. Use cotal reconcile-gate when the boot path cannot run
(daemon down, a non-manager endpoint, or you want to lift the freeze without starting a manager). A spawn that hits the same frozen gate names that verb in the refusal
(blockedOp=registration, the holding opId, remedy=cotal reconcile-gate) instead of a
wait-timeout: the facts were always in the manager log; they now reach the spawn caller too.
Give reconciliation a quiet manager. Suspend systemd restart policies, watchdogs, health-check
restart loops, and any other automation that can start or kill cotal supervise while boot healing or
cotal reconcile-gate is running. Leave one recovery attempt in control until it finishes.
Restarting the manager during the walk interrupts the current authority window. Durable progress makes
that interruption resumable, but a quiet manager is still the fastest and safest incident procedure.
Last-resort JetStream store replacement
Section titled “Last-resort JetStream store replacement”Store replacement is not normal gate recovery, is never automatic, and is destructive to mesh history. Use it only after the retained store cannot be reconciled and after deciding that losing its durable contents is acceptable.
- Stop every actor touching the space: supervisor, watchdog, manager, delivery daemon, and broker. Confirm that no Cotal or NATS process still has the store open.
- Preserve the stopped store before changing anything. Move
.cotal/natsaside to a dated backup and archive bothnatsandauth. Do not delete the only copy. - Understand the loss: replacing the store removes JetStream message and control history and durable consumer state. Agent session files stored outside JetStream remain, but the mesh history they referenced does not.
- Start the broker against a new empty store, then start one manager. Wait until it reports serving successfully.
- Repopulate the mesh only after that manager is healthy. Re-enable supervisors, watchdogs, and other restart automation last.
Keep the preserved store until the incident is reviewed and any required forensic or manual recovery is complete. Restoring it later restores the old durable state, including the fault that led to this last resort, so do not swap it back into a live mesh casually.
When something looks absent
Section titled “When something looks absent”Permission denials are loud, never silent: an over-tight ACL rejects the endpoint call it
refuses instead of returning an empty or incomplete result that looks successful, and a denial no
call is waiting on, such as a refused subscription, shows up as a logged denial on the endpoint. Check
.cotal/manager.<key>.log, .cotal/delivery.<key>.log (one pair per space, keyed as
Config describes), and .cotal/nats.log; cotal status shows
what is actually running. Those files live under the project .cotal/, not ~/.cotal,
unless the mesh root is the home directory. cotal up --detach redirects delivery and manager
stdio onto those files, so an operator-created systemd unit around that launcher does not put
the child logs in that unit’s journal. journalctl -u <unit> can be empty while the crash
reason is already in the project log. Manager log lines start with the UTC time they were
written. The access rules are collected in
Channels & permissions.