Skip to content

Run a mesh

Guide (informative) · For: operators · Prereqs: Quickstart

Day-to-day operation of a local mesh: what cotal up actually runs, how spawning resolves personas, harnesses, and models, how to reach a mesh from any directory, and the operator-only maintenance verbs. Every command’s full flag set is in the CLI reference.

cotal up brings up the whole local stack and bare cotal down stops it. Managed agents stay running as unmanaged OS processes; pass --with-agents to take them with the stack. Seats of the built-in pty runtime run inside the manager process, so they stop with the manager either way. Ctrl-C on a foreground up stops the manager through the same stop as bare down and prints the same report; when that stop is refused, for example because the manager cannot prove it can spare, Ctrl-C leaves the stack running, and you end it with cotal down --with-agents. A current manager records what its stop does with its seats before bare down signals it. A pre-pin legacy manager instead receives a reduced-guarantee warning and is signalled according to the documented upgrade contract. Its running binary may still carry the older destructive SIGTERM handler, so the CLI does not claim its pre-signal agent inventory was spared; those agents may have been reaped.

  • Broker: a local nats-server (logs to .cotal/nats.log).
  • Delivery daemon: the durable backstop, auth mode only (what it does).
  • Manager: a detached supervisor answering the control plane, so cotal spawn --detach and the cotal_spawn tool work right after up.

Cotal creates the presence bucket in memory storage. Its records are liveness that every endpoint rewrites each heartbeat, so nothing is lost when a broker restart empties it, and nats-server’s file store write latch cannot reach it. A broker stop removes the memory stream itself, so every cotal up, including the resume after cotal down --preserve-state, creates it again before any daemon starts. JetStream fixes a stream’s storage class when it is created, so a presence bucket created file-backed by an older cotal stays file-backed until that stream is recreated.

A file-backed presence bucket can remain open and watchable while refusing every write. A bound endpoint reports this as presence-write-stuck after one full presence TTL of consecutive failures. The roster is last-known while that condition is active. Restarting the broker clears nats-server’s in-memory store latch and preserves the JetStream root. Current credentials split the required stream authority: the cotal up provisioner can create the presence stream but cannot delete it, while the teardown credential can delete it but cannot recreate it. Cotal therefore reports the condition but does not attempt an unsafe partial delete-and-recreate. Stop and restart the broker to recover. A broker below nats-server 2.14.5 carries the latch (nats-server fixed it in 2.14.5). When cotal up starts or finds such a broker and the space’s presence bucket is file-backed, it says so. A memory-backed bucket gets no warning. A broker below the SPEC §13.12 floor of 2.12 is refused at connect with the floor sentence.

Three modes:

  • Default (static auth). JWT-authed, on by default: sender authenticity and per-agent ACLs, enforced by the broker (how).
  • --user-auth --idp <url>. Per-user auth: people cotal login once, the operator grants their agents on the actor ledger, and every connect is authorized live against that grant. Starts the space’s auth service alongside the broker (how).
  • --open. An unauthenticated, live-only dev mesh (no auth, no delivery daemon). For quick local experiments.

The broker and local services bind loopback by default. --host 0.0.0.0 widens the broker bind independently of the auth mode, so “network-reachable” never silently means “unauthenticated”. With no explicit --server, cotal up auto-selects a free local port when the default address is already held by another project; an explicit --server fails loud on collision.

--host is a boot flag, not a live rebind. A fresh cotal up writes the generated .cotal/auth/server.conf (project-local, not ~/.cotal) with that bind and starts nats against it. If anything is already answering at the mesh URL, up refreshes the recorded mesh and leaves the running nats listener alone, so passing --host 0.0.0.0 on a live or orphaned broker does not change who can connect. To change the bind: cotal down, then cotal up --host <addr> against a stopped broker so the generated file is rewritten. Do not edit server.conf by hand; the next real boot overwrites it.

On a stopped shared broker, up renders every persisted space account and every enabled space’s auth-callout account into the resolver preload, regardless of which space starts the broker. A missing callout account for an enabled space stops the boot rather than starting with a reduced resolver. An already-running broker is refreshed without rewriting its config.

A broker-only host is a first-class up mode. cotal up --no-manager boots the broker and, in auth mode, the delivery daemon, and no local manager, so the broker host never has a manager to stop and never leaves a manager slot stale. A refresh under the flag of a mesh whose manager is live refuses rather than keeping or stopping it: cotal down manager first. Without the flag, auth-mode up still starts nats, the delivery daemon, and a local manager. A space may run more than one manager, addressed by instance id (control surface); putting no manager on the broker host is a topology choice, not a singleton invariant. A manager whose boot inventory has no available connector does not take unpinned spawn/launch on the class rail, so a sibling that can launch the harness can. describe still rides the class rail, so an unpinned spawn can bind-fence when that skip member answered describe; re-issue, or pin --on. Pin one instance with --on when a partial inventory still answers with a harness refusal. The supported split is:

Terminal window
# broker host (project root that owns the generated conf, pidfiles, and logs)
cotal up --detach --host 0.0.0.0 --space main --no-manager
# no local manager starts: the summary lists nats-server + delivery daemon, and there is no
# `.cotal/manager.<spaceKey>.log` to wait for on this host
# manager host (registered remote mesh, same space)
cotal meshes add --server nats://broker.example:4222 --root ~/meshes/main
cotal supervise --space main --server nats://broker.example:4222

Wait for ✓ manager up in .cotal/manager.<spaceKey>.log on the manager host before spawning agents. On a broker host started without --no-manager, cotal up --detach prints ✓ running in the background: with manager listed once the manager pidfile is live; stop that local manager only after the ✓ manager up line. A host started WITH --no-manager never runs one, so neither the wait nor the stop applies there. That detach stdout is not a safe teardown boundary: it is pidfile liveness, not ✓ manager up. ✓ manager up is supervise’s post-start line after await mgr.start(). cotal down manager after only the detach line can still default-terminate the child during registration after it has taken the governance slot. Stopping before that post-start log line can leave the endpoint governance slot held until the holder’s gate reopens past the stamp (the successor’s boot heal, or cotal reconcile-gate when that boot cannot run). See Gate recovery.

Standalone cotal deliver --creds is not a repair for that split. Production renewal needs the manager and the daemon to address one credential store. The manager renews its own service credential inside that credential’s own window and re-dials its service connection with the renewed credential; if the connection closes and cannot be restored within about forty seconds it releases its lease and exits so a restart can serve, while a broker that is briefly gone is waited out. Separate host filesystems still leave manager root A writing and the daemon reloading root B; that composition is refused while the daemon stays up. Before every remint the manager challenges the delivery daemon’s store identity, and the answer must come from the process holding the delivery lease: the reply names the answering endpoint and the manager reads the lease row itself under its own credential, so a non-holder answering on the queue-grouped admin rail is refused instead of counting as the daemon’s store. A rail that reports no responder is also settled from the lease row, so a live holder on record makes that outcome a refusal rather than an absent daemon. Keep delivery on the broker host under up, and share one store only when you are composing a hosted pair (embedding). On the --no-manager split above, the manager host’s manager stays off the daemon-credential renewal lease once its store check finds the daemon on another store. A filesystem store is named by its root and by a random id in .cotal/store.id, which the copied .cotal/auth does not carry, so this holds when both hosts use the same root path. cotal doctor auth --fix on the broker host then renews the daemon credentials once they pass their renewal point.

A remote manager cannot reach a loopback broker. After changing --host, confirm the generated host: in .cotal/auth/server.conf and that nats is listening on that address before registering the mesh on the manager host. Detached child logs stay under the project .cotal/ that up ran in (see When something looks absent); they are not ~/.cotal unless that directory is the mesh root.

A user-auth mesh can expose only its credential exchange through an operator-owned HTTPS reverse proxy while leaving the existing local exchange untouched:

Terminal window
cotal up --user-auth --idp https://idp.example/api/auth \
--exchange-public-port 7443 \
--exchange-public-url https://auth.example

The public listener itself still binds 127.0.0.1:7443; configure the proxy to terminate TLS and forward to it. It serves only /health, /jwks, /exchange, /manager-service-authority, and /.well-known/cotal-mesh with the documented methods. It needs no local file capability: the signed IdP JWT or managed-agent actor token is the proof, while the original loopback listener remains capability-gated. Add --exchange-trusted-proxy only when that listener is reachable exclusively through your trusted proxy; it keys failure throttling by the last X-Forwarded-For hop instead of the socket address. The well-known bundle includes IdP pins and a deny-all sentinel credential, so fetch it only from the configured HTTPS origin. To change these listener flags, stop and restart the mesh; a refresh of an already-running service does not replace its bind or proxy policy. See Identity & auth for the trust boundary.

A remote seat does not need to run cotal login when the mesh owner pre-mints a single-use enrollment for it. Mount the enrollment URL as a private file, place the seat persona on the remote machine, and launch the foreground seat:

Terminal window
COTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \
cotal spawn --config ./worker.md --space main

The URL is redeemed once with an unauthenticated GET. Redirects, off-machine plain HTTP, retries, and login fallback are refused. If the seat has no mesh record yet, the enrollment response’s stock user-bundle fields register it before the launch. The returned actor token then uses the same remote auth-service exchange as a login-provisioned agent. The enrollment URL and file path do not enter the preflight or harness environment. A failed or reused enrollment leaves no actor material on disk; ask the owner for a fresh enrollment. When the foreground seat exits, this machine’s credential files are removed and the mesh-side grant stays until the mesh operator revokes it; the launch line says so. The exact server contract is in Enrollment redeem.

cotal status prints the detailed setup, process, registry, and live mesh status. Its Machine section names the running CLI’s source checkout, installed package root, or npx package root beside the version. It has one row per installed connector, which reports whether the executables that connector declares in requires are on PATH. Status, setup and the manager’s preflight resolve them the same way: an entry written as a path is checked as given, and a directory never counts as the executable. A connector whose setup provider reports health adds its own rows above those. The Claude Code connector reports its plugin and its skills plugin, and a stale skills row names the installed and CLI versions it compared. cotal setup (after the first run) prints the compact card.

Before reporting ready, the manager resolves every installed connector’s declared harness binaries against its own environment. A missing binary does not stop unrelated manager work: boot continues, but prints a named connector <name> unavailable line and records that reason in the manager’s status response. Available connector rows record the absolute paths boot resolved. A spawned seat and a seat resumed after cotal down --preserve-state both launch from those paths, and both are refused with the recorded reason when their connector’s row is unavailable. cotal models takes the same rule and reports that reason in place of the catalog, so it agrees with a launch about a harness installed or removed after boot. The manager looks again only when it restarts. A connector registered after boot has no row, so spawn, resume and cotal models check its binaries on PATH when they run.

On an authenticated manager start, unfinished static lifecycle rows reconcile while the control endpoint is already serving. The manager status response reports the staticReconciliation state, the last sweep counts, and each failed alias with its durable phase and literal disposition. cotal status --components reports the state and per-alias failure details. A slot row carrying a DEL or PURGE marker stops the sweep before it plans any alias, and the manager log names the row. A failed exact terminal is retried in the same process after 1, 5, and 30 seconds. Each attempt re-reads the durable slot and re-enters the same deterministic terminal operation; the delays only schedule work and never release the lifecycle fence. The terminal’s cleanup removes the lifecycle’s credential file and its broker durables and read-ACL row as separate steps. A file that cannot be removed does not leave the broker footprint behind, and its failure keeps the alias held for the next attempt.

On shutdown, the manager fences new reconciliation work and waits for an exact terminal that already started. The current serial sweep stops before its next alias, and startup cannot publish the manager service after stop() completes.

The four-attempt budget is per manager process. An exhausted row stays held and reports retry-exhausted with the remedy to restart the manager. The next process derives a fresh budget from the still-authoritative durable row. A recovered row remains visible until the next static reconciliation sweep, then clears. This component reports reconciliation outcomes. It does not say whether footprint cleanup completed independently of the terminal result; that separate durable projection remains tracked by #1274.

cotal service install is the supported way to run the manager as a user service (CLI reference): a systemd user unit on Linux, a launchd agent on macOS, one per mesh, surviving logout and reboot. On Linux that needs user lingering: install refuses while it is off and prints the root command that enables it. It installs only the manager; the units below remain the process models for every other component, and they are still examples of process models for those: copy them only after you decide which processes the unit should own.

cotal up --detach is a launcher: it starts the broker, delivery daemon, and manager, reports what started, then exits. Do not wrap it in a systemd service with Type=oneshot and RemainAfterExit=yes and treat systemctl is-active as stack health. That unit becomes active (exited) when the launcher exits successfully and stays active even if every detached process dies. When up --detach can identify that exact unit shape, it prints a warning but keeps the requested startup behavior.

For a single-host stack, keep cotal up itself in the foreground so systemd tracks a long-running process and restarts the stack if that process fails:

[Service]
Type=simple
WorkingDirectory=/srv/cotal-mesh
ExecStart=/usr/bin/cotal up --space main --host 0.0.0.0
Restart=on-failure
RestartSec=5s

An active unit then proves the foreground launcher and broker are still running, but it still does not prove that every child component serves. Pair it with the component check below. Also remember that cotal up starts a local manager as well as the broker and delivery daemon; run cotal up --no-manager (add the flag to the unit’s ExecStart too) on a host intended to be broker-only, so the unit and the host agree.

Seats spawned by the built-in pty runtime run with oom_score_adj 500, so under memory pressure the kernel prefers a seat over the broker, manager and delivery daemon, which are left as they were started; the extension runtimes do not own the seat’s process and get no preference.

That Type=simple shape puts nats in the unit’s cgroup with the foreground up process. A Restart=always (or on-failure) of this unit therefore restarts nats as well, so remote managers drop for the time it takes the broker to come back. Wrapping cotal up --detach in Type=oneshot with RemainAfterExit=yes does not move nats out of that cgroup. Detached spawn starts a new process group, not a new systemd cgroup, and the default KillMode=control-group still signals every process left in the service cgroup on stop or restart, including the nats PID. Escaping that cgroup needs an explicit unit setting such as KillMode=process, or a separate nats unit; this CLI does not ship that escape. The Type=oneshot unit below is a cotal status --components liveness check, not a --detach launcher. Neither trade is universal from Type=simple alone; it follows from which processes the unit actually owns. cotal service install covers only the manager, so for the broker and its siblings pick the example that matches the ownership you want, and treat systemctl is-active as unit health, not mesh health.

A broker that crashes under that foreground up keeps its mesh record and exits non-zero, so the unit’s restart takes the repair path against the recorded store rather than starting a second one.

If the deployment deliberately uses cotal up --detach as a boot action, monitor observed state instead of the launcher’s exit:

[Unit]
Description=Check Cotal component liveness
[Service]
Type=oneshot
WorkingDirectory=/srv/cotal-mesh
ExecStart=/usr/bin/cotal status --components --space main

Run that check from a systemd timer or another monitor and alert on a nonzero exit. The command distinguishes absent, not-serving, and refused components and never treats a sibling’s health as proof. Its delivery-process check is local to the broker host, so run it there. On a split topology, also probe the broker URL from the manager host and monitor the manager’s own service there. A remote manager cannot observe the broker host’s delivery PID, and an active unit on either host says nothing about the other host.

Stop one part without tearing down the mesh by naming its registered component: cotal down manager, cotal down delivery, or cotal down web. Component names from installed extensions join the same surface; cotal down with no names retains whole-stack behavior and leaves managed agents running as unmanaged OS processes, except pty seats, which stop with the manager. cotal down --with-agents is the previous reap. If a pinned manager has no spare-capability record, stop its managed agents explicitly before running that whole-stack command. A current manager always publishes the record, so it is absent only for an older manager, which may not understand the reap request.

On a remote user-auth mesh, foreground cotal spawn remains the default participant path. A participant can run detached agents only after the host advertises and operates the remote manager authority service, and the participant’s actor-ledger row includes supervise. This is not implied by spawn or admin.

The participant’s loopback/operator exchange obtains one closed manager-service view for its ordinary derived owner, a fixed server-selected manager actor, and one opaque manager instance. The host, not the participant, issues the public-nkey JWT material via the replay-safe, lifecycle-bound prepare → activate → renew exchange, plus a one-shot target-pinned retirement request for a host-managed terminal. It never exports the space signer, a static provisioner credential, or generic storage authority. Remote registration publishes its service status at the registered revision and current process epoch, so manager-caller selection can find it.

Stock participant supervision asks its host to enroll a detached agent and to prepare its terminal retirement, over the same manager-authority transport. The stock auth service answers both when it runs with a public exchange face: it grants the agent under the participant’s owner at a lifecycle UID it picks, bounded by the participant actor’s own grant, provisions that UID’s durables, and on retirement releases them and revokes the grant before the manager’s terminal rail. It refuses a second enrollment of a name whose grant still stands until that agent’s retirement is prepared. A host platform that keeps these writers in its own storage intercepts both requests on its own route instead. Copying host secrets or actor-ledger files to a participant is not supported. Foreground spawning and operator-local hosted managers use their existing paths.

The remote manager that cotal supervise starts can host workflow runs through its host: the host admits each run and signs only the run’s own driver, mediator and operator credentials. A logged-in user’s cotal run start against it is admitted: the auth callout issues the user’s manager connection, and the host binds each run to the owner who registered the manager. The run spawns agents that user owns, enrolled by the host like any detached spawn, with the reach the user’s own row grants when the spawn runs. A spawn may be placed on that manager and on no other instance. The host’s own manager refuses user-auth runs by name. User-auth run start records the path.

The registry entry decides the broker URL supervise dials, so a mesh published over wss:// is dialed as a websocket. The manager-authority registration it runs first also takes its TLS requirement from that entry, so the prepare credential is not exchanged over a plaintext connection the record did not describe. cotal meshes add records both.

When the authority service, login, or renewal is unavailable, the remote manager degrades fail-closed: it refuses new agents, restarts, and credential replacement rather than pretending local authority exists. Existing agents remain live only while their independent credentials are valid. A hosted composition must revoke the managed grant and finish its resumable release before it requests terminal retirement. Deleting DM or delivery consumers is not retirement and must not reset a resumable lifecycle’s frontier or pending state. The alias remains held until the terminal barrier confirms. Restore service and renew successfully before asking it to recover an agent. See Identity & auth and the CLI reference.

Terminal window
cotal spawn # foreground: your default agent, in this terminal
cotal spawn reviewer --detach # supervised: the manager runs it in a PTY
cotal attach --name reviewer # watch/type into a detached agent (Ctrl-] detaches)
cotal ps # what the manager is running
cotal stop --name reviewer # stop one

How a spawn resolves:

  • Persona. A bare cotal spawn uses .cotal/agents/default.md; a positional name picks .cotal/agents/<name>.md; --config takes an explicit ref or path. Set COTAL_DEFAULT_PERSONA=<name-or-path> to change the fallback. Fields and format: agent files.
  • Harness. Resolution order is an explicit --agent or cotal_spawn agent argument, then the persona file’s agent: pin, then the invoking caller’s COTAL_DEFAULT_AGENT, then the manager’s COTAL_DEFAULT_AGENT, then the product default (Claude). Compared in Connectors; per-connector guides: Claude · OpenCode · Hermes · pi.
  • Model. --model overrides the persona file’s model: (Claude: opus / sonnet or a full id; OpenCode: provider/model). Connectors that expose a catalog report it via cotal models --agent opencode: model ids plus available variants; pick one with --model provider/model --variant high.
  • Tools. A spawned Claude Code agent gets the cotal tools plus the MCP servers the cotal config shares, which first-run cotal setup fills with your own; narrow them per spawn with --share-tools (config).
  • Launch options. --opt key=value (repeatable) passes a native harness flag straight through; a persona or manifest launchOptions: mapping does the same declaratively (a --opt wins per key). It is a raw passthrough, with no allow/deny list: Claude renders each as --key value (a bare --key for an empty value), OpenCode merges them into its agent config, and Hermes has no option surface so it fails loud. The trust boundary is the spawn capability itself, not the flag set, so granting spawn is host-launch authority (security). A key must be a plain flag name; malformed or prototype-polluting keys are refused.

Detach from an attached PTY with Ctrl-] (the agent keeps running); rebind it with COTAL_DETACH_KEY=ctrl-<char> when it clashes with a keybinding inside the agent’s TUI.

Runtimes. The manager spawns into a pty by default. It spawns the PTY in-process on every platform, so replacing the manager worker closes its seats and the pty runtime gives no hot update. Any manager stop, bare cotal down included, stops and deprovisions those seats. A stopping manager refuses new spawns and first waits for the ones it already accepted, so their seats stop too. On Linux it can still adopt and reap seats that an earlier manager launched under a detached per-seat custodian, so those seats drain under the new manager; it starts no new custodian. A custodian whose agent has exited exits a few seconds later on its own. cotal seats lists the custodians left on the machine, and cotal seats --drain retires the ones whose agent has exited while keeping every seat whose agent still runs (cli.md). When a pty agent exits on its own, in-process or under a custodian, the manager logs a seat reaped: line with the exit code and, for a signalled child, the signal number. The line ends with the last line the child printed that starts with a connector’s [cotal-<name>] or [cotal-<name>/<part>] prefix, cut to 240 characters, when it printed one. A custodian keeps the same record beside the seat’s custody record, so a later reap of that seat, including one by a successor manager, reports how the child ended. When the custodian cannot write that record, it says why in the seat’s custodian.log, and a later reap of a child that ended on its own reports the record as missing or unreadable. Optional runtimes are installed through the extension surface, for example cotal ext add @cotal-ai/orca, then selected with --runtime orca (similarly @cotal-ai/tmux, @cotal-ai/cmux, and @cotal-ai/herdr). They put teammates in native terminal surfaces rather than manager-owned PTYs. Those surfaces stream no exit, so the manager asks the runtime every five seconds whether a seat’s process has ended, and frees a seat that ended on its own once the runtime proves the exit. Its seat reaped: line says the exit detail is unavailable from that runtime. Runtime names are open-ended and resolved from the registry; a missing provider or app throws, never silently falls back (architecture).

cotal up records each running mesh in a machine-local registry (~/.cotal/meshes/space.<key>.json, named by a case-safe hex encoding of the space: broker URL, the project root holding its creds and personas, and its mode). So a bare cotal spawn <persona> from any directory joins the running mesh with the right credentials instead of mistaking the cwd for a space:

  • cotal use <name> sets the default from every directory, including inside another mesh’s project. --space <name> overrides it for one command.
  • When one broker has records for several spaces, cotal up --space <name> refreshes that named space.
  • A refresh rewrites only what that command decided: the server, root and mode, the user-auth endpoints, and an explicit --host or --max-sessions. Every other field, such as the TLS requirement, is kept as the record stands when the refresh writes it, so a change another command made during the refresh survives. If the record was removed during the refresh, up fails instead of writing it back.
  • With no live selected default, a project with its own .cotal/ resolves to that project’s mesh; otherwise one running mesh is used automatically and several are an error.
  • cotal meshes lists them (a * marks the default); cotal down removes the entry.

The registry stores a path, never a secret; trust material stays in each project’s .cotal/auth. If the mesh is down or won’t take your creds, spawn fails with one sentence, never a raw NATS trace.

A mesh running on another machine has no cotal up on this one, so register it by hand:

Terminal window
cotal meshes add # guided: asks for the broker, probes it, offers what it finds
cotal meshes add optiplex --server nats://100.90.12.34:4222 --root ~/meshes/optiplex \
--allow-unencrypted-overlay # see below: an overlay address needs this
cotal meshes rm optiplex

On a terminal, a bare cotal meshes add walks you through it: it probes the broker you name and reports whether it is open or requires credentials, offers the spaces the folder already holds credentials for, and shows the record before writing it. Scripts and agents keep the flag form - without a terminal nothing prompts.

--root is the local folder holding that mesh’s .cotal/auth and .cotal/agents (its personas); the mode is inferred from what that folder holds.

The instance identities of the manager and the user-auth service are not part of that folder. Each root keeps its own in .cotal/space.<hex>/, so cotal supervise or cotal up --user-auth in the root you copied the folder to starts an instance of its own. A root last run by an older Cotal still holds them in .cotal/auth, as manager-instance.<hex>.json, manager-siblings.<hex>.json and space.<hex>/.cotal/auth/auth-instance.<hex>.json. Delete those files from a copy of such a folder before the first cotal supervise or cotal up there.

Know what you are copying. For an authenticated mesh that folder carries the space’s account signing seed, which is the authority to mint any identity in the space. A machine holding it is a certificate authority for the mesh rather than a client of it: anyone who reads it can impersonate any agent, read every retained channel and DM, change ACLs, and keep issuing themselves credentials. There is no per-machine revocation; undoing it means rotating the signing key and re-minting every credential in the space. Copy it only to machines you would trust with the whole mesh. cotal mint on its own does not substitute here: registering an auth mesh needs signing material that composes, which a minted user credential is not. The broker is probed before the record is written, so a bad address or a credential that mesh will not accept fails at registration rather than at your first spawn (--force records it without verifying, useful when the mesh is simply down right now).

Registering a mesh is how this machine starts sending agent credentials to a broker it does not run. NATS announces itself in plaintext before anyone authenticates, so an attacker on the path can pose as the broker and read the credential out of the connect unless the connection requires TLS, which is recorded on the entry and enforced on every dial through it.

What the record will require decides what you may register:

  • Without required TLS, the address is the gate: loopback (127.0.0.0/8, ::1), or your private overlay (100.64.0.0/10, fd7a:115c:a1e0::/48) with --allow-unencrypted-overlay. The tunnel provides the protection, and this command cannot check its state. Hostnames are refused because the lookup would choose which machine receives your credentials.
  • With required TLS, set --tls or use a tls:// URL. The recorded scheme enforces the TLS requirement. A hostname or public address is accepted because the certificate chain and hostname check identify the peer. A registration whose broker cannot complete the handshake fails unless you pass --force, which records the entry without verification.

Ordinary private ranges like 10.x and 192.168.x are refused in both modes. A café’s wifi is private but does not belong to you, and no public CA issues certificates for those ranges. An address spelling changes nothing: [::ffff:192.168.1.10], 3232235786, 0300.0250.01.012, and 192.168.257 all resolve to private addresses and receive the same refusal as the dotted form. --force exists for a mesh that is down. It never permits an unsafe credential destination.

A user-auth space’s IdP pins are established where the mesh runs and are never guessed. Register one from supplied trust: --user-auth-file bundle.json (exported on the mesh’s machine), or --from https://auth.example, which asks before it contacts the address at all, fetches the discovery document at /.well-known/cotal-mesh under that address over HTTPS, shows you the pins, and asks again before adopting them. A URL that already ends in /.well-known/cotal-mesh is fetched as given. Redirects are refused because a 302 can walk a pinned fetch down to plaintext or onto another host, and the pinned exchange must be an https:// URL too. The one exception is an exchange on this machine, where nothing leaves the box: plain http:// is accepted for a loopback literal (127.0.0.1, ::1, and any spelling of them), but not for localhost, which a hosts entry or poisoned lookup could point elsewhere. Use the literal. Registration checks that the pinned exchange answers /health and /jwks as the pinned issuer. It also checks that the broker refuses a bare connect; that refusal is the pass. The bundle’s sentinel credentials are written to a private (0600) file under the entry’s root; the registry itself never carries the secret.

Without required TLS, an overlay address is refused unless you accept the dependency explicitly, with --allow-unencrypted-overlay. The address is not the guarantee: it is protected while the tunnel is up, and if the tunnel is down that range is ordinary carrier-grade NAT and whoever answers the dial receives your credentials. Only you can know which it is, so the command asks you to say so. Your acceptance is recorded on the mesh entry rather than printed and forgotten, and the guided form asks the same question instead of taking the flag.

With required TLS (--tls, or a tls:// URL) that consent is no longer asked for, and the flag is not needed: the handshake is what protects the connection, so the acceptance it stood in for has been replaced by proof rather than promise. cotal meshes add <space> --server nats://100.64.0.1 --tls registers an overlay address with no prompt, no flag and no recorded acceptance. This is the “the flag disappears once the broker can be served over TLS” case, and it has now arrived.

This gate is on registration. cotal join --creds --server <url> deliberately takes an explicit connection at face value and does not consult the registry, so it is not covered. Join that way only to an address you would have registered.

The connection is still probed first, with the same second try at the longer budget the registry preflight uses, so a slow link reads as a connect that did not finish within that budget and a refused port reads as a broker that is not running.

Records added this way are removed only by something that names them. A failed liveness probe does not delete any record: an unreachable broker, local or registered by hand, is shown as offline in cotal meshes. A bare command does not count that offline record as running; name it with --space to restart it. cotal down / cotal clean all still drop an up record for the project they are tearing down, and they leave a hand-registered one alone even when --root pointed at that project. A cotal up for that space refuses outright unless it is that same endpoint: finding a broker already answering there is a refresh that starts nothing and leaves the record’s provenance alone, while actually starting the broker for that space, server and root makes this machine the one running it, so the record becomes an ordinary local one that cotal down clears. The refusal names cotal supervise --space <s> --server <url> (plus cotal deliver) when the registered broker is on another host, and cotal meshes rm when it is local. cotal meshes rm drops it and re-registering with --force replaces it. rm only forgets a mesh. To stop one running here, use cotal down.

cotal console is the terminal view (TUI on a real terminal, plain line stream when piped); cotal web is the browser dashboard. Both are read-only observers; the walkthrough is Watch a mesh.

Retained history is operator-owned. cotal clean history --force purges a space’s retained channel history; --dms also purges DMs (cotal history clear is an alias). It is deliberately not an agent tool: agents cannot wipe the record (identity & auth). For a stopped mesh, cotal clean store --force deletes the on-disk JetStream store outright, and cotal clean all --force also resets the space identity (CLI reference).

For a coherent durable cut, preserve the whole stack first, then create the artifact while it stays down:

Terminal window
cotal down --preserve-state
cotal backup create ./space-backup # full by default
# later: deliberately resume the unchanged source
cotal up --detach
# or, from another preserved cut, restore before the normal listener opens
cotal up --restore ./space-backup --detach

A refused cut leaves the mesh running and unfenced: fix what the refusal names and run cotal down --preserve-state again.

Use --store-dir on both preservation and backup for a custom JetStream store. A store cap set with cotal up --max-file-store <bytes> travels with the preserved state, and the resume renders it again. nats-server reads the cap once at start and refuses a config reload that changes it, so a new cap always needs a restart. The cut records the chat stream’s frontier per retained seat, so a resumed seat catches up from there instead of replaying its channels. registry is the only partial selection (backup create ... --only registry; up --restore ... --restore-only registry). Backup never stops or restarts a mesh implicitly, never opens the original store, and does not contain credentials or trust secrets. Backup/restore in every auth mode, open included, uses isolated, operation-specific maintenance logins; normal agent credentials cannot enter that listener. Full restore requires the same space and exact current local trust continuity, recreates conservative consumer checkpoints bound to their snapshot stream sequence state, and resumes retained agents under their original principals. The trust commitment includes the cryptographically validated full operator/system/data-account root chain as well as static/user authority state. A registry-only restore completes canonical empty infrastructure but leaves retained agents stopped because their DM/DLV/TASK/ACL state is outside that selection. Authenticated restore validates the complete space trust bundle before staging or changing the preserved store. Interrupted ordinary resume retries the same durable attempt after its prior listener is stopped. Restore re-entry can recover a surviving normal listener only when its attempt nonce, NATS server name, process owner, endpoint, and target-store identity all match the fsynced proof. A provably dead uncommitted owner is retired under lock and replaced with a fresh attempt-bound listener; an occupied foreign listener or ambiguous owner is never adopted. The manager commit validates while retained cleanup is still suppressed; the CLI durably records its attempt-bound 64-hex token in manager-committed / resume-committed before finalizeResume can release suppression. A retry from either committed state goes straight to exact-token finalization; failure preserves the committed gate and retained cleanup suppression. Missing commit evidence, interrupted finalization, a live recorded endpoint despite missing pidfiles, or ambiguous proof fails closed. See the CLI backup and restore contract for artifact, checkpoint, fallback, disaster-consent, and degraded-recovery details.

cotal personas manages the local catalog offline: list (--running overlays live markers), show <name>, edit <name> (re-validates on save), new <name>, rm <name> --force. The runtime write is cotal_persona; the runtime read is cotal_personas (list / show), both over the wire with the manager’s ownership checks. Fields: agent files.

A manager that dies mid-registration leaves its issuance gate frozen under that registration op. The freeze is correct: it stops two incarnations serving at once. The successor now completes that dead op on boot, using the same guard as cotal reconcile-gate: it acts only when the freeze-holder is affirmatively gone under a complete CONNZ sweep (gone and sweepComplete=true). If the dead op’s spec write committed, it finishes that same freeze (promote and reopen at the committed registration revision). If the spec did not advance, it abort-reopens the gate (generation+1, processEpoch unchanged) and continues the normal takeover. Boot heal and the following re-registration use separate one-shot executor windows, so a large predecessor family cannot spend the takeover’s credential lifetime. If that later registration still crosses a connection lifetime, it retries the same frozen operation with fresh authority and resumes verified-holder progress instead of freezing a new generation. A live holder, an incomplete sweep, or an unreachable delivery daemon still refuses. Silence is never evidence of death, and there is no TTL. If holder verification is interrupted, the frozen operation resumes from its durable, operation-and-gate-revision-bound progress after liveness is checked again. A later freeze cannot reuse that progress: the cursor binds the exact op, gate revision, and holder set. Use cotal reconcile-gate when the boot path cannot run (daemon down, a non-manager endpoint, or you want to lift the freeze without starting a manager). A spawn that hits the same frozen gate names that verb in the refusal (blockedOp=registration, the holding opId, remedy=cotal reconcile-gate) instead of a wait-timeout: the facts were always in the manager log; they now reach the spawn caller too.

Give reconciliation a quiet manager. Suspend systemd restart policies, watchdogs, health-check restart loops, and any other automation that can start or kill cotal supervise while boot healing or cotal reconcile-gate is running. Leave one recovery attempt in control until it finishes. Restarting the manager during the walk interrupts the current authority window. Durable progress makes that interruption resumable, but a quiet manager is still the fastest and safest incident procedure.

Store replacement is not normal gate recovery, is never automatic, and is destructive to mesh history. Use it only after the retained store cannot be reconciled and after deciding that losing its durable contents is acceptable.

  1. Stop every actor touching the space: supervisor, watchdog, manager, delivery daemon, and broker. Confirm that no Cotal or NATS process still has the store open.
  2. Preserve the stopped store before changing anything. Move .cotal/nats aside to a dated backup and archive both nats and auth. Do not delete the only copy.
  3. Understand the loss: replacing the store removes JetStream message and control history and durable consumer state. Agent session files stored outside JetStream remain, but the mesh history they referenced does not.
  4. Start the broker against a new empty store, then start one manager. Wait until it reports serving successfully.
  5. Repopulate the mesh only after that manager is healthy. Re-enable supervisors, watchdogs, and other restart automation last.

Keep the preserved store until the incident is reviewed and any required forensic or manual recovery is complete. Restoring it later restores the old durable state, including the fault that led to this last resort, so do not swap it back into a live mesh casually.

Permission denials are loud, never silent: an over-tight ACL rejects the endpoint call it refuses instead of returning an empty or incomplete result that looks successful, and a denial no call is waiting on, such as a refused subscription, shows up as a logged denial on the endpoint. Check .cotal/manager.<key>.log, .cotal/delivery.<key>.log (one pair per space, keyed as Config describes), and .cotal/nats.log; cotal status shows what is actually running. Those files live under the project .cotal/, not ~/.cotal, unless the mesh root is the home directory. cotal up --detach redirects delivery and manager stdio onto those files, so an operator-created systemd unit around that launcher does not put the child logs in that unit’s journal. journalctl -u <unit> can be empty while the crash reason is already in the project log. Manager log lines start with the UTC time they were written. The access rules are collected in Channels & permissions.