`cotal` CLI reference
Reference: describes the TypeScript reference implementation (the
cotalCLI), not the wire contract. · For: operators · Wire contract: SPEC
cotal is the operator command line for the reference implementation: bring a mesh up, mint
identities, launch agents, watch what they do, and tear it all down. It is a thin client over the
wire contract: the normative subjects and schemas live in the SPEC; this page is
lookup material for the commands, not a walkthrough; if you are new, start with
Getting started.
Running it
Section titled “Running it”npm install -g cotal-ai # puts `cotal` on your PATH (needs Node 22+)cotal --help # every command, groupedcotal --version # cotal-ai version + each installed extension's (also `cotal -v`)cotal <command> --help # one command's flags and usagenpx cotal-ai <command> runs it without a global install; in a dev clone, pnpm cotal <command>
runs it through tsx with no build step. Bare cotal prints help. Every command generates its own
--help, usage, and shell completion from its declared flags.
An undeclared flag is a usage error, and so is a flag given more than once unless it is
repeatable, as --opt and down --session-store are. The command prints the error and its help,
exits 1, and does not run.
Command output, including error lines on stderr and the guided setup and meshes add prompts,
is colored only when stdout is a terminal, so piped or redirected output is plain text. A non-empty
NO_COLOR turns color off on a terminal too. FORCE_COLOR turns color on even when output is
piped, unless it is 0 or false, and it takes precedence over NO_COLOR.
Commands come from the surfaces the binary composes: the base mesh CLI, the manager
(supervise), and the delivery daemon (deliver), plus any operator-installed extensions.
cotal ext add <npm-package> installs any registry providers a package contributes: commands,
runtimes, and local process lifecycle descriptors. The web dashboard and optional manager
runtimes ship this way.
Commands
Section titled “Commands”| Area | Command | Purpose |
|---|---|---|
| Set up & lifecycle | setup |
Guided, configure-only setup (installs, seeds personas; launches nothing) |
| Set up & lifecycle | update |
Reconcile first-party extensions and check or opt into a coherent CLI upgrade |
| Set up & lifecycle | up |
Start a local mesh (nats-server + JetStream), or boot a whole manifest with -f |
| Set up & lifecycle | down |
Stop the whole stack, selected registered components, or a manifest deploy |
| Set up & lifecycle | backup |
Create an offline full-space or registry-only artifact from a preserved cut |
| Set up & lifecycle | clean |
Configurable cleanup: purge history (live), or wipe the local store / identity (stopped) |
| Set up & lifecycle | meshes |
List the running meshes on this machine |
| Set up & lifecycle | sync |
Refresh the signed-in account’s advertised spaces |
| Set up & lifecycle | use |
Set the default mesh a bare cotal spawn joins |
| Set up & lifecycle | status |
Read-only diagnostics for setup, processes, and the selected mesh |
| Agents & personas | spawn |
Launch an agent from a persona (foreground, or --detach via the manager) |
| Agents & personas | models |
List connector model catalogs and variants from the manager |
| Agents & personas | ps |
List managed agents and their mesh status |
| Agents & personas | stop |
Ask the manager to stop a managed agent |
| Agents & personas | attach |
Stream and drive a managed agent’s terminal (pty runtime) |
| Agents & personas | input |
Type one line into a managed agent’s terminal without attaching |
| Agents & personas | personas |
List, show, edit, create, or remove local personas |
| Agents & personas | supervise |
Run a manager daemon (the agent supervisor / control plane) |
| Agents & personas | service |
Run the manager as a user service (survives logout and reboot) |
| Agents & personas | runtimes |
List the agent runtimes the manager can spawn through and whether each is reachable |
| Agents & personas | seats |
List the pty seat custodians an earlier Linux manager left, and drain the ones whose agent has exited |
| Agents & personas | reconcile-gate |
Unfreeze an issuance gate left frozen by a crashed restart when the successor cannot boot-heal it (holder gone, complete CONNZ sweep) |
| Messaging & watching | endpoints |
List every endpoint in the live presence roster, including infrastructure |
| Messaging & watching | describe / invoke |
Resolve a v0.4 service’s command surface off the wire; invoke one command by name |
| Messaging & watching | send |
Send one message, then exit: DM a peer, post a channel, or ask a role |
| Messaging & watching | channels |
Inspect or set the channel registry |
| Messaging & watching | history |
Clear retained message history |
| Messaging & watching | console |
Live protocol view for a space (TUI, or --plain line stream) |
| Messaging & watching | web |
Browser dashboard (installed as the @cotal-ai/web extension) |
| Extensions & misc | linear |
Serve the official Linear MCP server as an endpoint agents can call (installed as the @cotal-ai/linear extension) |
| Auth & meshes | mint |
Mint a creds file for a space (static auth mode) |
| Auth & meshes | login |
Sign in to a per-user-auth mesh’s IdP (once per machine) |
| Auth & meshes | logout |
Revoke the IdP session and clear the cached login |
| Auth & meshes | actor |
Manage a user-auth space’s actor ledger (grant / revoke / list) |
| Auth & meshes | doctor |
Credential-health diagnosis and repair (doctor auth) |
| Auth & meshes | join |
Join a space as your own presence (interactive) |
| Manifest | topology |
Validate and view a mesh manifest’s access graph (read-only) |
| Extensions & misc | ext |
Install / remove operator CLI extensions |
| Extensions & misc | completion |
Print or install shell completion |
| Extensions & misc | feedback |
Send feedback to the Cotal developers |
| Extensions & misc | deliver |
Run the server-side Plane-3 delivery daemon |
| Workflow runs | run |
Operate durable workflow runs: start, resume, list, inspect, answer a checkpoint, check an edited program with migrate |
| Extensions & misc | feedback-intake |
Run a self-hosted feedback intake server |
The manifest modes of up, spawn, and down (-f <cotal.yaml>) plus topology are covered
together under Manifest deploys.
cotal setup [--full] [--demo] [--yes] [--skills]| Flag | Default | Meaning |
|---|---|---|
--full |
off | Redo the full guided flow (implies --demo) |
--demo |
off | Also seed the guided expert team (david, sven, me) |
--yes, -y |
off | Non-interactive accept-all (for agents / CI) |
--skills |
off | Reconcile Cotal skills only through installed connector providers, plus ~/.agents/skills. Refused with --full or --demo. |
Guided setup is configure-only: it checks prerequisites, invokes installed connectors’ declared setup providers, and
seeds persona files, and it launches nothing (no mesh, no web, no manager). First run gets the
narrated flow; later runs print a status card. By default it seeds one default persona; the
david/sven/me team is opt-in via --demo. cotal status points stale Claude skills and
out-of-date .agents skills at cotal setup --skills, not unscoped setup. See Getting started and, for
maintainers, setup internals.
When a mesh resolves, setup seeds that mesh’s recorded .cotal/agents catalog, the same catalog a
following cotal spawn reads. It prints the absolute destination. On a fresh machine with no mesh it
uses this folder and says why; when several meshes are available and none is selected, it refuses
rather than choosing a catalog.
update
Section titled “update”cotal update [--self] [--space <s>] [--server <url>] [--creds <path>]| Flag | Default | Meaning |
|---|---|---|
--self |
off | If a newer release exists, install that exact validated cotal-ai version globally and reconcile through the newly installed binary |
--space, --server, --creds |
resolved mesh | Select the running manager whose continuity state is reported |
Without --self, update keeps the installed first-party surfaces coherent with the running
binary: it force-reconciles the four built-in connectors, then reinstalls other @cotal-ai/*
operator extensions at the binary’s exact version. Each extension runs in an isolated child, so one
failure cannot poison later replays. It then checks npm; a newer binary is an informational notice
with cotal update --self as the next command, not an automatic install.
After disk reconciliation, update reads the selected running manager. A machine with no recorded
mesh has no running manager to observe, so that read is skipped and the command completes. The same
holds when every recorded mesh is down and none is selected. A remote user mesh, registered with
cotal meshes add --mode user, is named and skipped: its manager runs under another install, so
there is no custody on this machine to preserve, and a legacy verdict still comes only from a
manager this machine read. A recorded mesh that is down is still a
refusal when the command selects it, with --space or by running inside its project, and so is a
named space that is not running. With several meshes running and no --space, --server or
--creds, the install is machine-wide, so every running manager is reported in turn, each under its
space name, before anything is written; a legacy verdict on any of them makes the whole run not a
hot update. A selector flag still reports one manager. A manager without a
custody generation is reported as legacy: it cannot preserve its manager-owned PTYs, so the
command says that this is not a hot update and prints exact, fork, fresh, or drain-only
for every seat. This report sends no stop, preservation-commit, or replacement command.
It does not preserve a running PTY on a legacy manager. The built-in pty runtime spawns
in-process on every platform and reports legacy. On Linux it still adopts seats that an earlier
manager left under a detached custodian, but it starts no new custodian. An incompatible native
@lydell/node-pty or ConPTY ABI break remains an explicit per-seat maintenance cut.
With --self, the selected running manager is reported before any global install. When a newer
release exists, Cotal then installs the exact version it validated, resolves and verifies that
package in npm’s global root, then launches that binary with the same --space / --server /
--creds selection to reconcile connectors and first-party extensions to the new generation. An npx
or dev-clone invocation therefore installs and continues through a separate global copy; it never
claims the already-running process changed. If the binary is current, --self performs the normal
local reconcile without reinstalling it.
Third-party extensions are listed with their installed version and recorded spec but are not
auto-updated in v1. Floating third-party updates require @cotal-ai/* peer-range validation and are
a future follow-up. A failed connector/extension install, npm metadata check, or requested global
install is reported and makes the command exit nonzero. Independent extension attempts continue so
the output includes every failure; an unavailable npm registry does not undo a completed local
reconcile, but the command still exits nonzero because it could not establish that the install is
current.
cotal up [--detach] [--open] [--space <s>] [--server <url>] [--channels <path>] [--runtime <name>]cotal up --user-auth --idp <url> [--exchange-public-port <n> --exchange-public-url <https://…> [--exchange-trusted-proxy]]cotal up --tls-cert <cert.pem> --tls-key <key.pem> # serve broker TLS (both, or neither)cotal up --restore <dir> [--restore-only registry] [--accept-missing-source]cotal up -f <cotal.yaml> [--dry-run] [--runtime <name>]| Flag | Default | Meaning |
|---|---|---|
--server <url> |
auto (free local port) | Listen URL override |
--host <host> |
none | Bind host override for a fresh broker boot: an IP or hostname only, never a URL (that is --server) and never host:port (the port comes from --server or its default); a URL or port-bearing value is refused pointing at the right flag. With no --server, the broker URL is derived from it, so --host <addr> alone is enough to make a mesh reachable at that address; a --host/--server pair naming different addresses is refused. A wildcard bind (0.0.0.0, ::) keeps a dialable loopback URL. Recorded on the mesh and reused by every later manager launch, so a repair or resume keeps remote attach working. A live refresh (✓ mesh already running) does not rewrite .cotal/auth/server.conf or rebind nats; stop the broker, then re-run up --host |
--space <s> |
the folder’s name | Space name |
--store-dir <dir> |
none | JetStream store directory (recorded; a repair up reuses it) |
--max-file-store <bytes> |
nats-server’s dynamic cap | JetStream file storage cap in bytes (a positive integer, no unit suffix). Without it nats-server sizes the store at start as three quarters of the free space on its filesystem. The cap is fixed at broker start: a running broker cannot change it (cotal down first), down --preserve-state keeps it for the resume, and a resume with a different value is refused. Not accepted with -f |
--channels <path> |
.cotal/channels.json if present |
Channel-registry seed file (JSON). An explicit path that is missing is an error |
--restore <dir> |
none | Restore a completed offline backup before exposing the normal listener |
--restore-only registry |
artifact selection | Restore only the registry component |
--accept-missing-source |
off | Explicit disaster consent when the inode-bound preserved source is absent |
--accept-stale-checkpoint |
off | Explicit consent to resume a seat whose checkpoint was captured outside its recorded recency horizon |
--open |
off (auth) | Unauthenticated dev mesh: no JWT, no ACLs |
--user-auth |
off | Per-user auth: people cotal login; connects are authorized against the actor ledger |
--idp <url> |
none | With --user-auth: the IdP auth base URL to pin on first enable |
--exchange-public-port <n> |
none | With --user-auth: add the public exchange face on this loopback port, for an HTTPS reverse proxy to forward to |
--exchange-public-url <https://…> |
none | With --exchange-public-port: advertise the reverse proxy’s HTTPS URL in discovery |
--exchange-trusted-proxy |
off | With --exchange-public-port: attribute public failure buckets to the last X-Forwarded-For hop. Enable only when the listener is reachable solely through a trusted proxy; otherwise the socket address is used |
--detach |
off | Run in the background (stop with cotal down) |
--tls-cert <path> |
none | PEM certificate to serve TLS with. Must be given together with --tls-key. Before starting the broker, Cotal checks readability, private-key mode, key/certificate match, the validity window, and host coverage. nats-server accepts an expired certificate and leaves the failure to clients, so Cotal performs these checks first. The decision is recorded; a later bare cotal up keeps serving TLS |
--tls-key <path> |
none | PEM private key for --tls-cert. Refused if group- or other-readable (tighten to 600) |
--file <cotal.yaml>, -f |
none | Launch a whole mesh from a manifest |
--dry-run |
off | With -f: print the plan, mutate nothing |
--runtime <name> |
pty (or the manifest’s, with -f) |
Agent runtime for the mesh manager (pty built in; others are installed extensions, explicit-only). Resolved + probed before the broker starts; an uninstalled/unreachable runtime fails loud. With -f, overrides the manifest’s runtime |
--max-sessions <n> |
64 | Live-session ceiling for the mesh manager. Each console pane and each cotal attach is one session, so size for agents × panes, not agent count. Recorded on the mesh and reused by every later manager launch, so a repair or resume does not silently drop back to 64. A running manager cannot change it: cotal down first, then cotal up --max-sessions <n> |
--no-manager |
off | Broker-only boot: start the broker and, in auth mode, the delivery daemon, and no local manager. A refresh under the flag of a mesh whose manager is live refuses rather than keeping or stopping it (cotal down manager first). Cannot be combined with --runtime, --max-sessions, or an agent-declaring manifest |
--rotate-sys |
off | Rotate the space’s system account and re-mint its two $SYS creds. Needs a stopped mesh; refused with --open |
cotal up boots a local nats-server with JetStream and, in auth mode (the default), JWT auth and
per-agent ACLs; --detach records the mesh so cotal spawn from any directory can find it. With no
--server, it auto-selects a free port if the default address is taken; an explicit --server
stays fail-loud on collision. --detach also brings up the control plane (delivery daemon in auth
mode, then the manager). --no-manager is the broker-only mode: it boots
the broker (and the delivery daemon in auth mode) and starts no manager, so there is no manager
pidfile to leave stale. A refresh under the flag of a mesh whose manager is live refuses rather
than keeping or stopping it: cotal down manager first. For a split topology with a manager, wait for .cotal/manager.<spaceKey>.log to contain ✓ manager up, then cotal down manager on that
host and run supervise against the remote broker; see
Run a mesh. cotal up --detach prints ✓ running in the background: with
manager listed (pidfile liveness, not a teardown boundary); with --no-manager the line lists
only what actually started. Ctrl-C on a foreground up stops the manager through the same stop as
bare cotal down (see down), then the rest of the stack, and reports managed agents
under the same rule: when the manager stop is refused, Ctrl-C prints the refusal with the reap route
and leaves the stack running. The -f form is a
manifest deploy.
A repair up on a mesh whose broker died reopens the store its record names, and refuses a
different --store-dir rather than silently opening a second store.
The generated .cotal/auth/server.conf is written on a real broker boot and is not an
operator-owned config. --host changes that file only when nats is actually started. A unit
restart that leaves an answering listener in place is a refresh, not a rebind.
On an existing mesh, cotal up reconciles the presence and lease bucket TTLs. It writes a reserved
canary and waits for the bucket to expire it before reporting success. If the broker accepts the
stream update but the backing store does not persist or enforce it, up exits nonzero with a TTL
persistence error instead of trusting the value returned by stream info. A refresh that restores a
missing manager says so with its pid (✓ restored in the background: manager (pid N)); a refresh
that finds everything already running prints only the ✓ mesh "<space>" already running line. A
first boot starts its manager without the restore line.
--user-auth --idp <url> starts the space’s auth service alongside the broker: the NATS
auth callout plus its capability-gated local exchange, and optionally the closed public exchange
face configured by the three --exchange-* flags above. The service is torn down with cotal down,
and a re-run of cotal up heals a dead service on a running broker. up waits for the service to
finish binding: while the daemon it launched (or found running) stays alive, the wait extends past
the base 15s up to 60s; a daemon that exits is refused at once with “exited before becoming ready”,
and one alive past 60s is refused as “alive and still starting” (wedged), naming the pid record and
the service log. --user-auth and --open
contradict each other and are refused loudly; a running broker cannot change auth mode
without a cotal down first. See identity & auth.
--rotate-sys renews the two $SYS credentials (membership-observer, connection-evictor).
They carry a 30-day expiry and nothing re-signs them in place, because the system-account seed is
never persisted, so they are renewed by issuing a new system account under the same broker
operator and minting fresh creds against it. A plain re-up does not do this: it reuses the
existing trust record, and its $SYS creds along with it.
The rotation is safe to run on a real space, with one operational cost. The data account, the account
signing key, every agent credential minted from it, and the JetStream store are all untouched; what
dies is the retired system account, and with it any out-of-band copy of the old $SYS creds, on every
broker that loads the rotated config. The cost is that earlier full backups stop being restorable
(see below), so this is not a no-consequence operation. It needs the broker to restart on the rewritten
config, so it runs as part of a boot:
cotal downcotal up --rotate-sys --detach # agents reconnect; nothing is re-provisionedcotal doctor auth # both $SYS creds healthy again, 30 days outA rotation is a stopped, fresh boot, and anything that is not one refuses it, all for the same reason (the on-disk material and the broker it runs on must never end up on different generations):
- a live mesh, because the running broker would keep serving the retired account;
- an open mesh, whether that comes from
--openor frombroker.auth: falsein a manifest, which has no system account at all; --restore, because reinstating a trust root and superseding it in one command leaves no way to say which authority the mesh came up on;- an unfinished restore or resume attempt on this root, including one
cotal upwould recover on its own, because those paths can adopt a live listener and return without booting a broker; - a root that hosts more than one space, because the system account lives in the shared broker
record and a rotation would retire every tenant’s, while the root holds one
$SYScred pair pinned to one data account.
Two things to know before you run it:
- The retirement is config-load-bound. Old
$SYScreds are refused by any broker that loads the rotated config. A stalenats-serverstill running the previous config in memory would keep honouring them, so stop every broker for this root first.--rotate-sysrefuses if this root’s mesh is recorded as running, if anything unidentified is answering at the address it was given, or if the root’s pid file names a live (or unreadable) process. Those are Cotal’s own ownership records, not a scan of the process table: anats-serveryou started by hand against this root’sserver.confon some other port writes none of them and will not be seen. Do not run one. - It invalidates earlier full backups. A full artifact binds to the trust chain it was taken
against, and that commitment covers the operator JWT and the system account. Every full backup
taken before a rotation refuses to restore afterwards, so take a fresh
cotal backuponce the rotated mesh is up.cotal up --restorenames this case when the data account still matches.
The commit is not atomic (a trust-record write plus two credential writes), so an interrupted
rotation leaves the record ahead of the creds. That split is detected rather than silent: every
cotal up on an auth mesh, and every cotal doctor auth, compares each $SYS cred’s issuer against
the persisted record and names the retired account. up warns rather than refusing, because these
creds power the membership graph and live eviction, both of which degrade fail-soft; the mesh is not
worth taking down over them. Re-running the rotation heals it, at the cost of one generation.
While those creds are expired the mesh keeps delivering messages, but the
membership feed and live connection eviction stay down; cotal doctor auth
and the manager’s log both name the credential and this repair.
cotal downcotal down --with-agentscotal down --preserve-state [--store-dir <dir>] [--session-store <dir> …]cotal down manager [delivery auth web nats ...]cotal down web [--space <name>]cotal down -f <cotal.yaml> | --run <id> [--dry-run]| Flag | Default | Meaning |
|---|---|---|
--file <cotal.yaml>, -f |
none | Tear down this manifest’s deploy |
--run <id> |
none | Tear down one spawn -f run by id |
--space <name> |
current mesh | With components: the mesh whose target-addressed components (e.g. web) to stop |
--dry-run |
off | Print the manifest teardown or selected components, mutate nothing |
--with-agents |
off | Bare whole stack only: also stop and deprovision every managed agent |
--preserve-state |
off | Bare whole stack only: fence the manager, retain principals and durable state, stop and prove the stack down, then publish ready |
--store-dir <dir> |
.cotal/nats |
With --preserve-state: the actual store path (required for a custom store) |
--session-store <dir> |
none | With --preserve-state: a harness transcript store directory to capture with every continuation-capable retained seat. Repeatable. No default and never inferred from a connector name; a path that does not exist or is not a directory is refused before anything stops |
Bare cotal down stops the whole local stack in dependency order and leaves managed agents running
when their runtime lets them outlive the manager. Before signalling the manager it verifies the spare
capability of the exact recorded manager, which records what that manager’s stop does with its
seats, and it reports the agents left behind plus cotal down --with-agents as the explicit reap.
When the manager had no managed agents, it prints no report.
The built-in pty runtime keeps each PTY inside the manager process, so those seats cannot outlive
it: every manager stop stops and deprovisions them, and down reports them as stopped. Every manager
stop the CLI makes runs this one path: down, Ctrl-C on a foreground cotal up, the teardown after
that up’s broker exits, the leftover-manager stop before cotal up -f, and the delivery cutover.
Each holds the manager’s stop reservation, so a second stop while one is in flight is refused,
names the process holding it, and leaves that stop’s --with-agents policy in place. Each sends
SIGKILL to a manager still running 15s after SIGTERM. Ctrl-C stops the manager first; when that
stop is refused or the manager’s exit cannot be confirmed, Ctrl-C signals nothing else, prints the
refusal with the reap route, and leaves the stack running; end it with cotal down --with-agents.
--with-agents is a one-shot destructive policy bound to the exact verified manager process
and the exact live down stop reservation; a stale, malformed, crashed, or different stop attempt
cannot turn a later bare shutdown destructive. If a managed agent cannot be proven stopped within
the manager’s stop timeout, the manager logs which one, still closes its broker connections and
console listener, and exits with code 1. It does not release its pidfile, liveness lease or
service registration in that case, so no successor is handed authority while that agent may still
run; the lease lapses on its TTL. Positional component names stop
only those self-registered local processes; for example, cotal down manager leaves delivery and
the broker running, and cotal down web is available when the web extension is installed. A
component that starts target-resolved (the web dashboard) is stopped the same way: cotal down web
resolves the mesh the same way as cotal web (registry current mesh first, --space to name one), so
it works from any directory; the other components always stop under the folder you run it in. The
-f / --run forms tear down a manifest deploy without stopping the whole mesh
and cannot be combined with component names. Stopping nats alone is refused while an unselected
registered daemon is still live; include those components or use bare cotal down.
A pinned manager with no spare-capability record is not signalled by bare cotal down or cotal down manager. A current manager always publishes the record, so a missing one means an older
manager: one that predates capability reporting, or one whose pty runtime reported that it cannot
detach its agents. Stop each managed agent explicitly, then run cotal down --with-agents from the
mesh root to stop the whole stack. An older manager does not understand
the reap request, which is why the agents must already be stopped.
Bare cotal down inventories by pidfile. When this folder’s registered broker answers and no
nats.pid records it, the command does not say nothing is running. It names the space and the
broker address, says no pidfile records that process, says it will not stop a process it did not
start, and exits 1. Stop that broker with whatever started it (an init unit, a container, or the
hand-run process). cotal meshes rm <space> only drops the registration. The probe runs whether or
not other owned components were running: they stop and clear their artifacts first, then the broker
is named. A component stop and --dry-run stay pidfile-only and do not probe.
down reads each process record once. A component that exits and removes its own record while
down runs counts as having no record, so the stop goes on. Any other failed read is an error.
Teardown verifies pinned process identity before signalling. PIDs are recycled by every OS,
so a recorded pid alone is not a durable target identity. up and cotal web record each
process’s creation identity in a sibling <pidfile>.identity pin, which holds the pid and the
process start reported by the OS. Every stop path, including down for the broker, web and
extension components, and the manager, delivery and auth-service stops, applies the same rule. A pin
that names a different start means the pid was reused, so teardown refuses and preserves it. A torn
or unreadable pin also refuses.
The pidfile and its pin are published by renames, and the pidfile rename is the commit point. Just
before it, the pin holds two lines: the old process’s and the new one’s. A launcher that dies
mid-publish therefore leaves the old record or the new one, each checked against its own pin line,
never a pidfile without its pin. An old record with no pin is legacy, so its line holds - in place
of the token and it stays legacy until the commit. A CLI older than this change reads a two-line pin
as torn and refuses.
Publishes of one pidfile are serialized by a lock file beside it, <pidfile>.publish.lock, because
the launcher and the daemon it starts both publish the same record. The next publisher reclaims a
lock left by a crashed one. When no start token can be read for the new process, its pin line holds
- in place of the token, which reads as a legacy record, and the publish ends in the legacy shape:
a pidfile with no pin. Teardown, and a daemon removing its own record on exit, take the same lock and
remove the record only while the pidfile still names the pid they stopped, so a stop that races a
publish leaves the new record whole.
The web dashboard claims web.pid with an exclusive create, so a second dashboard for the same mesh
is refused, and writes its pin right after the claim. A stop that runs between the two reads a
legacy record. A refused claim calls the record stale, and points at cotal down web, only when the
file is empty or its pid is proven gone. Content that is not a pid, or a pid whose liveness the
kernel will not report, is refused as possibly fronting a running process, as cotal down web
refuses it.
The pidfile pid and the pin pid are two coordinates. Automatic cleanup follows proven death of
the pidfile target (ESRCH on that pid): a torn sibling pin does not wedge a dead pidfile pid.
A torn pairing where the pin names another pid, while the pidfile pid is still live or not proven
dead, still refuses. Inspect both pids with ps. Do not delete <pidfile>.identity to force a
stop; that weakens target-identity protection. Once the pidfile process is dead, rerunning
teardown clears the stale record automatically.
The first teardown after upgrading a running pre-pin stack has a narrower guarantee. A live record
with no identity pin is signalled after a loud warning that it predates identity pinning. Restarting
the component writes the pin, so later teardowns receive full match and mismatch protection. The
same warning applies on platforms where no stable start token is available. For a legacy manager,
bare cotal down also warns that agent sparing cannot be verified before it signals. Because the
CLI cannot establish which SIGTERM handler that already-running binary carries, it never presents
the pre-signal seat inventory as confirmed spared; a genuinely older destructive handler may still
reap those agents. --with-agents publishes a one-shot reduced-guarantee handoff bound to the
recorded manager pid and the live .stopping reservation’s inode, then signals unconditionally.
That handoff cannot be replayed by a later stop attempt. A pin that exists and does not match the
live process still refuses before signal.
--with-agents performs the old destructive logical teardown: managed processes stop and their
credentials, ACL rows, and delivery footprints are deprovisioned. --preserve-state is a different
maintenance transition: it stops retained processes while suppressing leave/deprovision cleanup, persists the manager’s
same-principal resume inventory, stops the entire stack without removing run/auth artifacts, and
publishes a stable inode-bound cut only after every recorded process is proven stopped and the exact
recorded NATS endpoint is unreachable. A missing or stale broker pidfile never counts as stopped. The
attempt is bound durably before the manager is fenced, the resume document and attempt-bound
cut-intent are fsynced before manager commit, and the manager’s commitment itself is journaled
(cut-committed) before any process stops. A retry after a crash at any of those boundaries reuses
the exact recorded attempt and finishes the remaining stop and endpoint proofs idempotently, without
needing the (by then intentionally dead) manager. A partial cut never publishes ready. It cannot
be combined with component names, manifest teardown, or --dry-run.
Seat checkpoints. After the stack is proven down, the cut writes one checkpoint per retained
seat under .cotal/maintenance/v1/checkpoints/<attempt>/<seat>/, and prints the path, the
continuity class and the generation for each. The path carries the preservation attempt because a
checkpoint is immutable once sealed: a shared directory would make the second cut in a root refuse
on the first cut’s leftovers, and clearing it would destroy an artifact a rollback still needs. The
capture happens only at that point because anything earlier races a harness that is still writing
its transcript and its working tree.
Each checkpoint directory is created 0700, refuses a destination that already exists, and holds:
repo.bundle, the seatcwd’s reachable history, anchored on the base commit the record names by full object id;repo.index.diffandrepo.worktree.diff, the staging state as two diffs, base to index and index to worktree. Two rather than one because a single combined diff restores a mixed tree with the right bytes and the wrong index: a source reportingMM READMEwould come back asM README;repo.untracked.tar, the untracked files in scope;- the harness session pointer, when the seat’s connector declares one, and the transcript store
files the operator named with
--session-store. Each records where the destination puts it back as an anchor (the workspace root, the account home, or the seat’scwd) plus a relative path, because the destination’s root and home are its own and the source host’s absolute spelling would either miss them or write outside them; checkpoint.json, written last, after every digest is computed over the bytes that landed.
The record carries the manager’s resume entry unchanged as its first field, then the space, the seat
name, the recovered lifecycleUid, the writer generation the cut was taken at, capturedAt, the
recency horizon, the applied profile revision, the seat’s git status --porcelain as the cut read
it, and the continuity class. Every captured file is
listed with its byte size and sha256, so an operator verifies the whole artifact with sha256sum
and git bundle verify. No secret values, no operator keys and no source-host launch material
enter it.
The generation is the highest one this root has claimed for the seat in .cotal/auth. A
generation entry there that is a symlink, a directory or anything else that is not a regular file
refuses the cut, as every other auth record does.
The continuity class is what the connector declares, capped by what the checkpoint carries. A
connector declaring session continuation classifies as exact, but reopening a session takes both
halves, the pointer that names it and the store that holds its transcript. A checkpoint missing
either one cannot reopen that session, so it is recorded as fresh when the connector declares a
fresh start and drain-only otherwise. A pointer with no store is capped the same way as a cut
carrying neither, because it names a session whose bytes the artifact does not contain. A class is a promise the destination is entitled
to act on, so it never describes bytes the artifact does not contain. The transcript store stays an
operator input: this repository does not know where a harness keeps its transcript, so exact
requires --session-store to name one.
The recorded status is read under the same selection rule as the untracked set, so it describes the state the captured bytes can reproduce. The destination re-reads it in the promoted tree and refuses a difference.
The untracked selection rule is recorded in the record and is
git ls-files --others --exclude-standard -z, excluding .cotal/. It honors .gitignore, so an
ignored file the seat needs does not travel and has to be moved separately. The .cotal/ exclusion
is a secrecy boundary rather than a size one: when a seat’s cwd is also the mesh root, the control
directory is untracked, and without the exclusion the broker trust material, the space account, the
manager instance identity’s private seed and the seat’s own credentials would land inside the
artifact. A checkpoint carries credential references only; the destination resolves that material
itself.
A seat whose launch options could not be resolved is refused rather than checkpointed, with the
manager’s own wording: imperative launch options have no non-secret durable source (<keys>). The
refusal arrives at prepare time, so the cut stops before any child does.
A delegated seat (SPEC §13.17) is refused at prepare time too, with
a delegated seat is not resumed by a later manager; stop it before preserving. A manager stop
after a refused cut retires that seat through its retirement path.
cotal clean <history|store|all> --forcecotal clean restore-attempt --attempt <id> --forcecotal clean restore-fallback --attempt <id> --force| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | history: target mesh |
--dms |
off | history: also clear DM history |
--store-dir <dir> |
.cotal/nats |
store/all: JetStream store directory |
--force |
none | Required: destructive, no prompting |
--attempt <id> |
none | restore-attempt: exact stale pre-commit attempt; restore-fallback: matching healthy committed restore |
One configurable cleanup verb; every target requires --force.
historypurges the retained message backlog on the running broker (channels, plus DMs with--dms). The same operation ashistory clear, which stays as an alias.storedeletes the stopped mesh’s JetStream store (.cotal/nats): streams, durable consumers, and messages. This is the reset for stale on-disk broker state, e.g. durables minted by an older, incompatible Cotal generation surviving adown/upcycle.allisstoreplus the space identity (.cotal/auth), the local creds and markers tied to it, any crash residue a normaldownwould have swept (stale pidfiles,run/), and the mesh’s registry entry; the nextcotal upmints a fresh identity.
history needs the mesh up; store and all refuse while any recorded mesh process is still
alive or any same-root recorded broker endpoint remains reachable (run cotal down first). They
also refuse outright on a root that holds accounts for several spaces: the store and the broker
trust record are shared by every space on the broker, so both targets would take out all of them
and no --space can narrow that. down, backup and up --restore refuse there for the same
reason. cotal status lists the tenants on such a root. Personas
(.cotal/agents) and logs are never touched. The mesh record now carries a custom
store location for up’s own repair, but clean still takes --store-dir itself; clean does
not read the record. Custom cleanup targets must contain either the Cotal store-generation marker or a
real jetstream/ store directory; filesystem roots, project roots, and Cotal auth/maintenance trees
are always refused.
store and all also refuse every maintenance journal state. After a healthy committed restore,
restore-fallback is the only supported way to remove the recorded unchanged old-store inode; it
never deletes the active target, requires both the exact attempt id and --force, and retires the
completed restore journal so a later down --preserve-state can start a new backup cycle.
Backups
Section titled “Backups”cotal down --preserve-state [--store-dir <dir>]cotal backup create <dir> [--only full|registry] [--store-dir <dir>]cotal up --restore <dir> [--restore-only registry] [--accept-missing-source]Backup is offline-only. It requires the stable ready record from down --preserve-state, an exact
store match, no live recorded process, and an unreachable exact endpoint from the recorded cut.
That endpoint is probed immediately before cloning, so a live broker with a missing or stale pidfile
is still refused. It claims the cut, reflink/copies the stopped source to a
private attempt clone, and opens only that clone on a random loopback bootstrap broker with an
independent parent/deadline watchdog. It validates the canonical stream and pull-consumer inventory,
writes native snapshots with consumers excluded, and stores conservative contiguous ACK-floor
checkpoints separately. The presence bucket is memory-backed, so it does not survive the cut and
the clone may lack it. Every other stream must be present. The original store is never opened by
the backup broker, and the stack is not restarted implicitly. Artifact destinations must not overlap
the preserved source or maintenance
attempt tree. Restore artifacts and targets likewise cannot nest inside or contain each other, the
preserved source, or the maintenance attempt tree.
Stopped client-managed KV ordered consumers are ephemeral read residue, not backup state. Backup ignores only the pinned client’s exact stopped shapes: ordinary last-value watchers and the whole-bucket scanner that uses all-history delivery to collapse concurrent tombstones. A bound consumer or any lookalike with a different filter, inbox, lifetime, or other config is still refused.
full is the default and indivisible: channel registry, CHAT/DM/TASK/INBOX/DLV, ACL, MEMBERS, and
validated durable checkpoints. registry is the sole partial artifact. Presence, derived membership
feed, leases, native ephemeral/history consumers, credentials, keys, tokens, owner secrets, and actor
ledger files are excluded. full means every transferable message and registry stream, not every
JetStream resource: endpoint submissions/facts/events/timers/workflow state, contract artifacts, and
the records/auth/session stores are nonportable control state. Restore recreates those streams empty
with their canonical configs before exposing the normal listener, so active endpoint runs,
lifecycles, and sessions do not cross a backup. Artifacts are exclusively created 0700;
snapshot/checkpoint files and
the manifest are 0600; manifest.json is written last with exact sizes and SHA-256 values. The
directory is trusted operator input: hashes detect corruption, not malicious rewriting.
Restore validates and stages the exact allowlisted artifact bytes before moving or creating a store.
It requires the same space and existing trust state. The whole pre-commit window holds a journaled
liveness claim (coordinator, watchdogs, brokers, absolute deadline): ordinary up and a repeated
up --restore refuse while the claim is live, and a stale attempt is recovered only after the
deadline has elapsed and every recorded owner is proven dead. A retried up --restore handles this
automatically; an operator can also recover it explicitly with cotal clean restore-attempt --attempt <id> --force. Nothing
ever rolls back a live attempt. A registry-only artifact restores as registry-only whether or not
--restore-only registry is passed; omitted infrastructure is always created and the exact
post-restore stream inventory is asserted before commit intent. Ordinary up from a preserved cut
resumes only the exact recorded source store and runtime; a contradicting --store-dir or
--runtime fails in preflight.
Admitting a seat checkpoint. An ordinary up from a preserved cut admits that cut’s seat
checkpoints before it journals the resume attempt and before any process starts, so a refusal costs
nothing. Three gates run in order, each naming what it saw.
- Integrity. Every file the record names must be present, a regular non-symlink file, the recorded byte size and the recorded sha256, re-stat’d after the read so a file that moved is a refusal. Failure here consults no other gate.
- Identity. The recorded space must match, the recorded
lifecycleUidmust not belong to a live incarnation, and the profile revision must match this host’s or be resumed under deliberately this host’s. A differing revision is refused with both digests and the remedy, and there is no override: the checkpoint carries the recorded digest and not the config bytes, so nothing could run the seat under the recorded revision, and the manager re-digests the same file and refuses drift on its own. This gate has no blanket override, which is the only reason the next one may have one. - Recency.
capturedAtis compared to this host’s clock against the horizon the record carries. Inside it, the seat resumes. Outside it,uprefuses and prints the capture instant, the clock reading and the horizon;--accept-stale-checkpointadmits it anyway and the exercised consent is printed with the actual age. An unreadablecapturedAtis refused with no override, because a freshness gate that fails open is not a gate.
Custody transfers only after all three pass. The destination claims the recorded generation plus one
by exclusive create, before it launches anything. A lost create means another destination is already
claiming that seat, and it refuses with seat-writer-generation-create-lost rather than adopting
the winner and becoming a second writer. The recorded lifecycleUid is reused and never minted, so
the resumed seat binds the same lifecycle-keyed durables.
Admission is reconciled against the inventory the resume is about to hand the manager, and that
reconciliation finishes before the restore moves a single tree. A checkpoint whose recorded
lifecycleUid is not the one the retained inventory carries describes a different incarnation of
that seat, and it refuses with both uids while every live working tree is still untouched and no
generation is claimed. A retained
seat with no admitted checkpoint refuses the resume by name: an absent checkpoint directory and an
absent record are indistinguishable from a seat that was never checkpointed, and a seat that starts
without passing the gates has claimed no generation. --accept-stale-checkpoint is recorded in the
resume journal with the seat, the capture instant, the admitted age and the horizon, so the consent
survives the terminal it was typed into.
The whole admission is all or nothing. Coverage is settled first, then every gate runs over every checkpoint, and only then is any generation claimed. A refusal at any point leaves every generation unclaimed, including a lost exclusive create during the claim itself: the claims that attempt made are removed before the refusal is raised, by the exact paths it wrote, so a generation another destination holds is never touched. A claim is a create that can never be made again, so a refusal that left one behind would consume the retry over the same checkpoint set.
Restoring a seat checkpoint. Once every gate has passed over every checkpoint, and before a
single generation is claimed, up puts each admitted seat’s captured bytes back. A refusal here
costs nothing for the same reason a gate failure does: no claim has been made and nothing has
started.
A restore never moves or replaces the destination’s own control directory. A checkpoint excludes
.cotal/ by design, so a seat whose cwd holds one, which is the layout an operator gets by
running up and spawn in a single directory, is refused before anything is staged: promoting a
tree that cannot contain .cotal/ over that cwd would carry this host’s live trust material and
maintenance state away with the superseded tree. The refusal names the control directory it found
and the remedy, which is to give the seat a working tree that is not a workspace root.
Each seat is staged beside its own cwd, in <cwd>.incoming:
- every recorded digest is verified again over the files as they are now;
- the bundle is cloned into
<cwd>.incoming, which is refused when that path already exists; - the recorded base commit is verified in the clone and checked out detached, so a bundle that does not contain it stops the resume instead of continuing against a different history;
- the index diff is applied with
--indexand the worktree diff without it, both--binary --allow-empty. That order is what puts staged content back in the index rather than only in the worktree, and--allow-emptyis why a seat with a clean tree is still restorable; - the untracked archive is extracted.
Every seat stages before any seat is promoted. Promotion moves an existing cwd aside to
<cwd>.superseded.<timestamp> and renames the staging directory into place, then puts the session
pointer and store files where the destination’s connector reads them, then re-reads
git status --porcelain in the promoted tree and compares it to the status the checkpoint recorded.
A restore that applied without error and produced a different index is a refusal, not a warning. The
two renames are the only steps that touch the path the seat will use, so a failure anywhere leaves
every seat’s live cwd as it was.
The rename itself claims the superseded name, and a taken name gets a numeric suffix. The timestamp has one-second resolution, so two promotions of the same seat within one second compute the same path; a rename onto a name that already holds a tree fails on every platform, and that failure is read as taken. Nothing creates the name ahead of the move, because Windows refuses to rename onto an existing directory at all. A superseded tree is the thing that rename exists to keep.
git and tar run as child processes with argument arrays, never a shell string.
A leftover <cwd>.incoming refuses the resume by name. A staging directory from a failed run is the
only record of what failed, so nothing removes one automatically: inspect it, remove it by hand, and
resume. A pre-existing cwd is renamed rather than deleted, so a wrong checkpoint costs a rename
instead of a tree. When a promotion fails, the renames that attempt made are undone and the staging
tree is left where it is, as the evidence for what did not verify.
A session pointer whose recorded sessionId is not the one the retained inventory reopens is
refused before anything is cloned. A session file already present at its destination is judged by
content: bytes equal to the recorded digest are already restored, and different bytes under the path
the connector is about to read are refused with both digests rather than clobbered.
up --restore <dir> reaches the same admission and the same restore, after the store is restored
and validated and before commit intent is journaled. A registry-only restore resumes no seat, so it
admits and restores nothing.
One limit is worth stating plainly. The writer generation is claimed by exclusive create inside one workspace root, so it fences two resumes on the same host and does not fence two independent destinations: copy a checkpoint to two roots and both claim the same successor. A real cross-host fence needs a coordinate neither root owns.
Authenticated restores validate the complete
space trust bundle before staging, including nkeys, seed matches, JWTs, signers, and space binding;
full restores commit to the validated operator, system-account, data-account, and active-signer root
chain in addition to the static/user authority fingerprint. Because the system account is part of that
commitment, a cotal up --rotate-sys makes every full artifact taken before it unrestorable
against this root: take a fresh full backup after each rotation. The composed commitment is revalidated
immediately before store mutation and never includes secret seeds. Restore never creates fresh auth.
Same-path restores atomically retain the old
source at the journaled fallback path; alternate targets retain it in place; a missing canonical
source needs explicit --accept-missing-source. Quarantine and target restores use current canonical
configs on isolated random-loopback brokers, never expose native snapshot consumers, and publish a
commit-intent immediately before the normal listener starts. Archive bytes never instantiate the real
target: after quarantine validation, every stream is re-snapshotted from the validated quarantine
state into attempt-owned sanitized files, and the target is restored solely from those. Before that boundary, failure rolls back
the attempt-owned target; after it, ambiguity preserves both stores and records forward-repair
recourse. The cooperative maintenance lock excludes Cotal commands, not arbitrary raw NATS processes.
Bootstrap brokers in every auth mode, including open, mount the store under a local account with
random operation-specific logins only, each carrying the exact per-phase subject permission matrix;
normal static credentials and user-auth sentinel/bearer connections are rejected, and no auth
service or callout starts. Open mode differs only in its account label, never in authority. Inventory, each stream snapshot,
restore initiation, exact upload id, validation, and each checkpoint recreation use separate exact
authorities. Every checkpoint carries the source stream’s message/first/last sequence state and must
match its snapshot record before mutation; core then derives and validates the only allowed start
policy. TASK is not a CLI exception: the same core checkpoint API recreates its canonical DeliverAll
WorkQueue durable because acknowledged tasks are absent from retention and NATS forbids a
start-sequence policy there. Registry-only restore creates every omitted canonical stream and transient
bucket on the isolated target before the normal listener is exposed. It deliberately does not resume
retained agents or recreate their DM/DLV/TASK/ACL state; their identity material stays retained and
stopped rather than being reprovisioned into a partial restore.
After listener readiness, the manager starts attempt-bound, validates retained credentials/tokens
without granting or reprovisioning, and resumes the exact persisted principals under cleanup
suppression. Registry-only restore uses the same flow with an empty agent set. On a user-auth mesh
these manager calls run as the logged-in operator’s cli actor, the caller the preserve cut used, so
that actor needs a current admin grant. commitResume is an
idempotent validation barrier only: success must be awaitingFinalize with an attempt-bound 64-hex
commit token and does not release suppression. Under the workspace lock, the CLI first fsyncs that
exact evidence as manager-committed (restore) or resume-committed (ordinary resume), then calls
token-bound finalizeResume; only an active response for the exact token releases suppression. The
CLI records the same token in finalization evidence before a restore becomes active, or before an
ordinary resume retires and consumes the marker. Re-entry from either committed state skips the prior
idempotent activation/commit phases, retries finalization with the durable token, and finishes the
workspace transition. Failure before finalization preserves the committed state and cleanup
suppression; it is not rewritten through a degraded transition. Re-entry between any two earlier
boundaries reuses the same attempt and may retry the idempotent phases without deleting retained state. A missing or
changed per-agent dependency is a named fail-closed result; the journal becomes degraded and remains
available for forward repair. A retry from resume-intent,
resume-active, or resume-degraded reuses the same attempt and inventory after the prior listener is
proven stopped. A retained agent the lost manager already launched can still be running, for example
in a tmux window, while the journal reads resume-intent. On a static mesh the replacement manager
closes that seat through the reference the lost manager recorded on the agent’s slot, waits for the
principal to leave presence, and launches it again. A live principal with no such record, or one that
stays live after the seat is closed, is refused. Every normal restore listener has an unguessable
attempt-bound NATS server name. The CLI fsyncs its exact name/nonce, canonical endpoint, process owner, and generation-bound target identity
immediately after spawn. Re-entry accepts a surviving listener only when its INFO server name, live PID
record, endpoint, and target identity all match that proof; degraded restore repair then moves through
the guarded workspace transition only after manager commit. If an uncommitted bound owner is provably
dead, recovery retires that exact proof under the maintenance lock and binds a fresh listener for the
same attempt, endpoint, and target with a new nonce and server name. A live foreign/mismatched listener
or ambiguous owner is preserved and refused, never adopted by reachability alone. A reconstructed
commit/degraded attempt without either the exact bound proof or a durable dead-listener replacement
record fails closed even when the recorded port is free. A later ordinary startup may pass an active
restore only when its details prove manager commit and its exact recorded listener is dead.
Mesh registry
Section titled “Mesh registry”cotal meshes [--json]cotal meshes add # guided, on a terminalcotal meshes add <space> --server <url> [--root <dir>] [--mode auth|open|user] [--tls] [--force]cotal meshes add <space> --mode user (--user-auth-file <bundle.json> | --from <https url>)cotal meshes rm <space> [<space> …] [--force]cotal sync [--idp <auth base URL>]cotal use <space>cotal status [--space <s>] [--server <url>] [--components]meshes lists the meshes this machine knows; a * marks the current default a bare
cotal spawn joins. Entries learned from a signed-in account are marked discovered. Their
registration trust is stored under the account’s private auth state, and the registry contains no
session token or sentinel credential bytes. Commands resolve the catalog slug; a different human
name is rendered only as a label.
meshes --json prints one JSON object per recorded mesh per line: space, server, mode,
root, default (the *), and origin (up, manual, or catalog for a discovered entry). A
local or hand-registered entry also carries offline. A discovered entry is never probed, so it has
no offline field. tlsRequired, events: "required" and a discovered entry’s catalogName appear
only when the record has them. An empty registry prints nothing and exits 0. The note about a default
that matches no record goes to stderr, so stdout carries only rows, on a first run too. The table is
presentation and is not a stable parsing target. meshes add and meshes rm refuse --json.
A registry record this build cannot use is refused by name, never rendered and never skipped. One
that does not parse, or is missing a field every consumer reads (server, mode, root, ts,
space), makes every registry command exit 1 with the file’s path and what is wrong with it.
Remove the file or restore the record; nothing repairs or invents a field for you.
An IdP may advertise a same-origin space catalog during login. Cotal reads the complete snapshot and
adds every valid registration without a separate meshes add. A snapshot younger than five seconds
is used without a request. After that, commands that resolve a mesh target conditionally refresh the
saved catalogs. An operation targeting a discovered space refreshes only that space’s account and
refuses if that account fails. Operations targeting local or manually registered meshes refresh every
account, print one warning for each failure, and continue. cotal status refreshes every account,
never refuses on a refresh failure, and lists each account as fresh, updated, not-modified,
no-catalog, or failed with its error. cotal sync bypasses freshness and reports added, changed,
removed, unchanged, and name collisions. --idp limits it to one signed-in account. It never connects
to a broker.
The registry is updated under the same lock that guards the catalog cache, so a command never lists a discovered space set that another command is still writing. The cache records a fetched snapshot as not yet applied before the first registry write and as applied after the last. If a command dies or is stopped in between, the next command applies that snapshot again before it can use it, with no request inside the freshness window.
The shared dispatcher applies this preparation to every command that declares both --space and
--server as mesh-target flags, including commands registered by other packages and commands that
declare their own equivalent flag objects. Daemon and startup commands that use those names only as
configuration explicitly opt out. Registry-local meshes add and meshes rm never refresh a catalog.
While the registry holds a record this build cannot use, the preparation neither refreshes nor
applies a catalog, so the command’s own checks run first. A snapshot left unapplied is applied by the
next preparation after the record is restored or removed. A command that resolves its target through
the registry still refuses the record by name.
Run on a terminal with the space or --server missing, meshes add is guided: it asks for the
one thing that cannot be derived (the broker URL), probes it, and tells you what answered - open or
requiring credentials. It then offers the spaces your --root already holds credentials for, states
the mode as a fact about that broker rather than asking, and shows the exact record before writing
anything. A broker that does not answer, or a space name already registered, becomes a choice rather
than an error. Anything you pass on the command line is taken as given and not asked again. Without
a terminal - a script, an agent, CI - nothing prompts and the flag form’s errors stand
(COTAL_NO_PROMPT=1 forces that too).
cotal up and cotal down maintain their own records. meshes add registers a mesh they cannot
speak for: one running on another machine, a shared broker, a hosted space. --root is the folder
whose .cotal/auth holds that mesh’s credentials and whose .cotal/agents holds its personas.
The default is the project you run it in. The registry stores that path, never a secret. --mode
defaults to auth when the root holds the space’s account record and to open otherwise. The
broker is probed before anything is recorded, so a wrong address, or credentials that mesh will
not accept, fails here instead of at the first spawn; --force records without verifying (and
replaces an existing record).
A hostname or public address is registrable only when the connection will require TLS. Pass
--tls, or use a tls:// URL. The scheme is recorded as enforced intent, so every later dial
through the record demands the handshake (and meshes add tls://… against a plaintext broker is
refused at registration). Without required TLS the fence admits loopback and private-overlay
literals only. RFC1918 addresses are refused in both modes because a cafe LAN is private but does not belong to you.
A user-auth mesh registers from supplied pinned trust, never guessed: --user-auth-file
takes the bundle exported where the mesh runs; --from asks before it dials the address at all,
then fetches the /.well-known/cotal-mesh discovery document under that address (HTTPS only; a URL
that already ends in that path is fetched as given), displays the pins, and asks again before
adopting them. Neither fetch follows redirects: a 302 can move a pinned fetch
onto plaintext or onto another host, so it is refused rather than followed, and the pinned
exchange must itself be an https:// URL, except for an exchange on this machine, where plain
http:// is accepted for a loopback literal (127.0.0.1, ::1, any spelling of them) but not
for localhost, which is a name rather than an address. Registration verifies that the exchange
answers /health and /jwks as the pinned issuer. It also verifies that the broker refuses a bare
connect; that auth-required refusal is the pass. The sentinel credentials land in a 0600 file under
the entry’s root; the registry records only the path.
meshes rm drops records. It never stops a mesh. For a mesh running on this machine cotal down
is the right verb, and rm says so unless you pass --force. A hand-added record is removed by
meshes rm, by an add --force replacement, or by a cotal up that actually starts the broker for that same space, server and root, which becomes that
mesh and so takes the record over (a cotal up for that space anywhere else refuses instead).
Nothing that merely infers a record is stale from a dead broker touches it: an
unreachable broker is listed offline and stays, whether cotal up or cotal meshes add
wrote the record; a foreground up whose broker exits unexpectedly keeps its record the same way.
A bare command does not treat that offline record as a running mesh;
name it with --space to restart it. cotal down / cotal clean all still drop an up record for the project
they tear down; a hand-added one they leave alone even when it shares a root, because nothing
on this machine could write it back.
A discovered entry belongs to the normalized IdP origin and proved subject that supplied it. Local teardown, cleanup, and liveness pruning do not remove it. A manual or locally started entry with the same name wins and remains untouched; that discovered name is reported as a collision. Logging out removes only the discovered entries owned by that account.
cotal meshes and cotal status print events: required for a registration carrying
policy: { events: "required" }. On that space, foreground spawn, detached spawn, manager starts,
and interactive join cannot opt out or join without an event plane. --no-events is refused with
the space named. A connector without an event plane is refused with both the space and connector
named. A session whose own grant omits events.<owner>.<actor> is refused before joining and the
message names a full-row actor grant repair. A running seat whose event plane stops for good on
that space stops too.
use <space> sets that default; the selection applies from every directory,
including inside another mesh’s project. status is a read-only report: machine prerequisites
(starting with the installed cotal-ai version), the installed extensions and their versions, this
folder’s .cotal/, the recorded meshes, and a live snapshot of the selected mesh (roster, channels,
membership feed). Stale Claude skills and out-of-date .agents skills recommend cotal setup --skills,
not unscoped cotal setup. status takes --space / --server to pick the mesh to inspect; it starts
nothing. The manager row asks the service endpoint once: a live process that does not answer is
not serving, and a probe that could not be made leaves the row running · service unchecked.
A process row whose PID record exists but cannot be read, or exists under both its current and a
pre-upgrade name, reads pidfile unreadable with the error, and the other rows still print. A live
manager whose delivery-aware marker cannot be read keeps its row and names the failure as
delivery-aware marker unreadable with the error. The Web process row
prints the address the selected mesh’s dashboard recorded in web.session once it was listening,
while the PID in its web.pid is alive. A web.pid or web.session that exists but cannot be read
is named on the row with the error. Otherwise it reads down, or not installed without the web
extension. Status and the cotal setup card count the web extension as installed only when a
package in the extensions manifest provides web and that package is on disk. A web package record
that cannot be read is named on the Web extension row and on the card’s web row, after the
dashboard’s address when it is listening and beside an unreadable web.pid or web.session. The
Web process row still prints, and reads down unless the dashboard is listening.
If a refresh fails, status may still show the kept catalog bytes for diagnosis. It labels them
stale with the last successful snapshot timestamp and the refresh error. It never calls that state
synchronized or online. If a selected discovered space vanishes from a successful snapshot, the
selection is cleared and the command reports that no default is selected.
For a user-auth mesh the selected-mesh section reports the login status works as: the signed-in
subject when this machine holds a cached session for the entry’s pinned IdP, or the exact cotal login --idp <url> line when it does not, with no network round trip either way. A locally
provisioned space also shows the actor grant row; a discovered or registered remote entry reports
the grant as not checkable on this machine, because the ledger runs where the space was
provisioned. --components on a user-mode target probes as that same signed-in login (ps’s
credential), never a static mint; when the login cannot supply a credential, the row says why
instead of printing the broker’s refusal of an unauthenticated probe.
Persona rows name the catalog they describe. If this folder and the selected mesh use different
catalogs, status names both and marks which one spawn launches from. A green default means the file
passes the same agent-file loader spawn uses; a present but invalid file is reported as invalid.
cotal status --components adds a fail-loud per-component health pass. It reads each
component’s own control surface, rather than treating a PID, a lease, or a successful probe of a
sibling as proof that the component serves. It prints one of serving, absent, not-serving, or
refused for each component and exits 0, 1, 2, or 3 respectively (the highest observed
state wins):
- manager: local PID record, its liveness-lease holder and PID, then the manager’s own typed
statusservice reachability from this host. Manager builds that do not report static reconciliation saystatic reconciliation not reported by this manager build; the line stays visible even when the manager is otherwiseserving. - delivery: local PID record, its ready lease (
readyis the daemon’s own bound-control signal), and the latestrenewal.<spaceKey>.jsonadoption verdict, the record of the space the command was asked about, keyed per space the way the pidfiles are. A re-signed credential and a broker-accepted adoption stay distinct facts. A root-onlyrenewal.jsonleft by an older build names no space and is never read as any space’s verdict (doctor authnames it as a leftover). - web: local PID record, then the
/api/metaresponse at the address the dashboard recorded inweb.sessiononce it was listening, which must name the same PID. The probe presents the readiness nonce recorded beside that address, the one credential the dashboard accepts on/api/meta. A live PID with no recorded address (the dashboard is still writing it, or an earlier build started it), or an unrecognizable process record, isrefused, not a green default-port guess. Aweb.sessionthat exists but cannot be read is alsorefused, and the row names the read error. - broker: the registered mesh URL dialed from this host with its recorded TLS requirement.
absent means Cotal has no live local component record (or has a stale record); not-serving
means the component record is live but its service/readiness surface did not answer or is not ready.
Those are intentionally separate exit cases. A failed or unreadable probe is refused, never an
absent component or a clean zero. A PID record that exists but cannot be read, or exists under both
its current and a pre-upgrade name, refuses only its own row. A record that its component removes
while the pass runs reads as absent.
cotal spawn [<persona>] [--detach] [--name <n>] [--agent <a>] [--model <m>] [--variant <v>] [--prompt <text>] [--cwd <dir>]cotal spawn -f <cotal.yaml> [--dry-run]For a foreground spawn onto a remote user-auth mesh, a launcher may supply a one-time enrollment instead of a cached human login. Prefer a private file:
COTAL_ENROLLMENT_FILE=/run/secrets/cotal-enrollment \ cotal spawn --config ./seat.md --space mainThe file contains only the enrollment URL, ending with at most one line terminator, and must be
mode 0600 on POSIX. An orchestrator that cannot mount a file may set COTAL_ENROLLMENT_URL
instead; that value is redeemed byte for byte, so a trailing newline in it is refused. Setting both
is refused. Enrollment input
requires --space and applies only to a foreground persona spawn. If the mesh is not registered yet,
the enrollment response must carry the stock user-bundle fields and the command needs
--config <persona-file> because there is no local remote-mesh persona catalog to read. The client
redeems the URL once, registers the returned mesh material, exchanges the returned actor token at the
pinned auth service, and removes both enrollment variables before starting any child process.
A cached login for the same IdP and an enrollment are conflicting proofs, so the command refuses rather than choosing one. An invalid enrollment never falls back to login provisioning. Unknown, expired, revoked, and already-used enrollments all produce one response: ask the owner for a fresh one. See Enrollment redeem for the HTTP contract.
A runtime that starts a managed seat outside the manager’s filesystem hands the child a managed
handoff instead: one 0600 file named by COTAL_MANAGED_HANDOFF_FILE, carrying the lifecycle the
manager already enrolled. The runtime builds the command with delegatedSeatCommand:
COTAL_MANAGED_HANDOFF_FILE=/run/seat/handoff.json \ cotal spawn --config ./seat.md --space main --name <actor> --agent claude \ --expect-owner <owner> --expect-lifecycle-uid <uid>The cotal entry reads the file, deletes it and drops the variable before it parses flags, prints
help or loads extensions, so every outcome leaves no file. The variable is read under any letter
case; spellings that name different files are refused after every one of them was deleted. The
spawn then refuses a malformed handoff, or one whose space, owner, actor or lifecycle UID differs
from --space, --expect-owner, --name and --expect-lifecycle-uid, before any broker
connection or exchange request. Every refusal on this path names the field and never a value from
the handoff. The registration’s server, exchange and enforcement checks, the local state this
machine keeps for the space (its mesh record, user-auth state and agent secret files), target
resolution, the policy refresh, the broker preflight and the agent auth preflight quote the space,
the server, the exchange URL, the actor or a path named for one of them in their own diagnostics and
in the filesystem errors under them. For a handoff each prints one fixed sentence that names the
field and the phase instead, whether its check fails or an error is thrown. When the agent auth
preflight’s rollback then fails to remove a secret or file, that sentence is followed by the names of
the cleanup steps that failed, without their errors. The event-plane policy
refusals name the handoff’s space field. An actor outside [A-Za-z0-9_] and a space that cannot
name local state, such as .., are refused as malformed before any plane. A handoff conflicts with
the enrollment variables, --detach, -f and --creds, and needs --config <persona-file>. From
there it runs the enrollment consumer above without redeeming anything. See
Delegated seats.
| Flag | Default | Meaning |
|---|---|---|
--space <s> |
resolved mesh | Target space |
--server <url> |
registry entry | Broker URL override |
--creds <path> |
none | Control-caller creds for an off-registry manager (--detach only) |
--name <n> |
persona’s name: |
Presence-name override (does not choose the persona) |
--config <persona-or-path> |
none | Persona catalog name or file path; wins over the positional |
--agent <a> |
persona’s agent:, else COTAL_DEFAULT_AGENT, else claude |
Connector type (claude, opencode, jcode, hermes, and so on) |
--role <r> |
persona’s role: |
Role override |
--model <m> |
persona’s model: |
Model override |
--variant <v> |
persona’s variant: |
Model variant override (connector-defined; e.g. OpenCode reasoning tiers) |
--cwd <dir> |
this cwd | Working directory to root the agent at. Refused before launch when the directory does not exist on the serving manager’s host. |
--prompt <text> |
none | Initial prompt auto-submitted at start |
--resume <id> |
none | Fork an existing session id into the mesh; only connectors that declare resume support accept it (see the matrix). The manager records the source session id, and ps --wide shows it. With --detach --on <instance>, a Claude session held on this host is carried to that instance first (Resume a session); carrying one needs --on |
--no-events |
event plane on where supported | Opt out of the session’s structured event plane (--events only restates the default) |
--share-tools <sel> |
none | Share named operator MCP servers with the agent |
--subscribe <a,b> |
persona’s | Channel read-set override |
--allow-subscribe <a,b> |
= subscribe | Read-ACL override |
--allow-publish <a,b> |
deny | Post-ACL override |
--detach, -d |
off | Launch via the manager into a detached PTY (reattach with cotal attach) |
--on <instance> |
class anycast | With --detach only: pin the launch to one manager instance id (the whole id, as ps prints it). Refused on a foreground spawn (no manager to pin), with -f (a manifest deploy launches through the manager class queue), and when empty |
--file <cotal.yaml>, -f |
none | Deploy a manifest onto the running mesh |
--dry-run |
off | With -f: print the plan, mutate nothing |
--allow-stale <a,b> |
none | With -f: waive named stale agents (apply-only) |
--runtime <name> |
manifest’s | With -f: override the manifest’s runtime |
--expect-owner <u_…> |
none | With COTAL_MANAGED_HANDOFF_FILE only, and required there: the owner the handoff must carry |
--expect-lifecycle-uid <uid> |
none | With COTAL_MANAGED_HANDOFF_FILE only, and required there: the lifecycle UID the handoff must carry |
Each session uses its connector’s event plane by default: a stream of structured events
describing what the agent did, rather than the prose it wrote, on a channel of its own. The channel is named after
the agent’s principal, events.<owner>.<actor>, never after its display name, because two live
agents are allowed to share a display name and would then share a stream. The launch grants publish
rights on that channel alone, foreground and detached alike. On an open mesh, which issues no
credentials, the launch still allocates the agent an id, so the channel names a stable actor.
--no-events is the explicit opt-out unless the selected registration says
policy: { events: "required" }. Required policy makes the
event arm and grant mandatory, so --no-events and connectors without an event plane are refused.
The launch decision and the grant are separate on purpose. Holding publish rights on a channel is
not a request to publish to it, so writing an event channel into an agent file’s allowPublish
does not override --no-events.
The persona (--config > positional > COTAL_DEFAULT_PERSONA > default) is loaded from the
target mesh’s .cotal/agents/ when it is a bare name. A reference that contains a path separator or
ends in .md is loaded from that file. A relative path resolves against the mesh root, except that
an enrollment or a managed handoff resolves --config against the working directory. A missing
persona is refused with the catalog directory or the file that was checked. The launch flags
override the file. On a user-auth mesh the
effective name is also the agent’s actor token, so it must match the token grammar (no -); the
spawn is refused with that explanation before any request is sent. Foreground runs the agent
attached to your terminal; --detach hands the launch to the running manager. Both modes get the
durable backstop on a mesh that runs the delivery daemon; --live-only skips it for a foreground
spawn (messages posted while it is disconnected are then not replayed). A foreground exit retires
the agent’s creds and broker footprint, like a manager despawn. On a user-auth mesh the two arms
differ: a spawn against a mesh this machine provisioned revokes the actor row on exit, while a
remote spawn (an enrollment or the advertised provisioning endpoint) removes only this machine’s
credential files; its grant stays until the mesh operator revokes it, and the launch line says
which arm you are on. A spawn through the advertised provisioning endpoint against a record that
pins no exchange URL is refused before the grant is requested, so no credential lands on this
machine. A --detach spawn is an
action: the manager accepts it and returns the allocated identity at once, then the launch
follows to a terminal outcome rather than blocking (see the control surface).
See Connect Claude Code and Agent files; -f is a
manifest deploy. (cotal start was merged into cotal spawn --detach.)
A --detach spawn onto a manager from another Cotal release is refused before any request is sent
when the manager’s contract does not declare a field this CLI sends. The refusal names the field,
calls it version skew, and gives this CLI’s version. A field you leave unset is not sent, so it
never causes that refusal.
A manager has 50 seat slots, and each seat counts once. A slot is held by a managed seat (a row in
that manager’s cotal ps, including a seat still joining), by a reserved launch the manager accepted
but has not started a process for, or by a cooling hold. A seat that ends within 10 seconds of
starting leaves its slot cooling until those 10 seconds pass, unless an operator stopped it. Such a
seat holds only that cooling slot, even while its launch is still reporting the failure. A spawn
refused at the limit states that split and whether waiting can free a slot:
at capacity (50 of 50 slots: 49 managed, 0 reserved, 1 cooling); waiting frees a cooling slot in 7s, or despawn oneA cooling slot frees at the stated time. A launch that has not settled frees its slot only if it fails, and a managed seat frees its slot only when it stops. The refusal counts a launch as pending only while it holds a slot, so a launch whose seat already ended is not counted. The roster counts presence, which also includes peers no manager owns, so its total is a different number.
Run from a managed seat’s own shell on a static or open mesh, cotal spawn --detach launches as
that seat when it targets the seat’s own space. Without --space it picks that target the way the
operator path does, so a recorded mesh that is not running is skipped. The CLI reads the seat’s
launch identity (COTAL_NAME, COTAL_ID, COTAL_LIFECYCLE_UID, COTAL_SPACE, and on a static
mesh the seat’s own credential), so the manager records the seat as the spawner, the same as for
the seat’s cotal_spawn tool. On a static mesh that credential also proves the seat’s space, so a
launch without COTAL_SPACE still runs as the seat, and a target space holding no credential for
the seat is refused. An open mesh acts as the seat only when COTAL_SPACE names its space. The
seat can then stop the child with cotal_despawn, and the manager stops the child when the seat
exits. On a static mesh a seat whose agent file lacks capabilities: [spawn] is refused, because
its credential holds no spawn subject.
--on <instance> keeps its pin: the seat’s own credential has no instance route, so on a static
mesh the CLI mints a one-shot manager-caller view for the seat, pinned to that instance and
carrying the spawn subject only when the seat’s credential holds it. On an open mesh the call keeps
the TLS requirement the mesh records. --creds, --server with an unregistered --space, and a
user-auth mesh keep the operator path.
models
Section titled “models”cotal models [--agent <connector>] [--refresh]| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Which manager to reach |
--agent <connector> |
all registered connectors | Connector whose catalog to list |
--refresh |
off | Ask the connector to refresh its provider cache |
Asks the running manager for each connector’s model catalog (model ids plus their variants)
for connectors that expose one. OpenCode and Codex query harness/provider surfaces; Jcode reads
providers that enable model_catalog = true in the operator Jcode config.toml. Jcode’s listed
effort tiers render as variants (declared, not provider-verified), and launch can still refuse one.
A connector without a catalog says so. A connector whose harness the manager did not find at boot
reports the reason boot recorded, as a spawn does, so restart the manager after installing it. Pick
a result with cotal spawn --model <id> --variant <v>, where <id> is the model id as the catalog
printed it. OpenCode and Codex ids are the full
provider/model; Jcode ids are bare (opus-5, not cliproxy/opus-5), because the provider is
selected by the operator’s Jcode config and a prefixed id is refused at launch with the bare form
named.
endpoints
Section titled “endpoints”cotal endpoints [--space <s>] [--server <url>] [--creds <path>]Lists the mesh presence roster: agents, the manager, and any other protocol endpoint, with each
endpoint’s role, kind, status, and current activity. Unlike ps, this is a read-only presence view;
it is not limited to child processes owned by the manager.
Endpoint control
Section titled “Endpoint control”cotal describe <endpoint> [--on <instance>] [--space <s>]cotal invoke <endpoint> <command> [--args '<json>'] [--space <s>]cotal invoke <endpoint> <command> --name <agent> [--admin] [--space <s>]The generic v0.4 service surface. describe resolves a registered endpoint’s command set off the
wire - the reserved describe command answers the registered contract digests, the schemas are
fetched from the space’s content-addressed contract store, recompiled, and verified against those
digests - and prints each command with its capability class and targeting shape. --on <instance>
pins describe to one manager instance’s rail (the whole id, as ps prints it under its
manager <id> headers), so an operator can read what that instance serves in a multi-manager space;
unpinned, the class queue answers and the attribution line names whichever instance did. invoke
calls one command by name: --args is a JSON object validated against the fetched input schema
before
publish; a targeted command takes --name <agent> (resolved to the agent’s current principal through
inspect) or --self. --admin uses the admin instrument credential, whose cross-agent reach rides
the operator-only any authorization mode. Neither command has compile-time knowledge of any
endpoint’s schemas - this is the same trust chain every built-in control command now uses. Needs an
auth mesh: the manager registers its service on both static and per-user meshes (a signed-in user
rides their bearer; each visible or invoked command still requires its existing grant, and cross-agent
reach needs the admin scope). An open mesh has no service registry.
Managed seats
Section titled “Managed seats”cotal ps [--on <instance>] [--wide | --json] [--slots] [--space <s>]cotal stop --name <n> [--on <instance>] [--space <s>]cotal attach --name <n> [--on <instance>] [--no-reconnect] [--space <s>]| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Which manager to reach |
--name <n> |
none | Managed agent to stop / attach (required) |
--on <instance> |
class anycast (ps: class scatter) |
Pin to one manager instance id (multi-manager space); takes the whole id as ps prints it, not a prefix. An empty value (--on "", an unset shell variable) is refused, never treated as absent. A roster principal id (local.…) is refused with a message naming the instance id ps prints |
--wide (ps) |
off | After each seat’s compact row, print extra operational facts the manager records: the provider the connector reported serving the model, cwd, pid, spawner, lifecycle uid, the owning manager’s instance id and host, and for a --resume seat the session it forked (forked from <id>, with the source title and transcript SHA-256 once a Hermes or Jcode seat has recorded its fork; a carried Claude session prints forked from <host>:<id> with its title, SHA-256 and carried <time>, the time its bytes reached the manager). Model and requested variant stay in the identity row rather than printing twice. A fact the manager did not record (for example a runtime with no real process, or a connector that reported no provider) prints nothing, never a placeholder |
--json (ps) |
off | Machine-readable: one JSON object per seat per line, copied unchanged from the manager row. Instance headers and errors go to stderr, so stdout contains only rows. Mutually exclusive with --wide |
--slots (ps) |
off | List the durable static slot rows this manager owns instead of live seats, through the slots command. Mutually exclusive with --wide. A row that is not in the live roster still prints, with live=false; a retired row never prints |
--no-reconnect (attach) |
off | End the attach when its session ends, instead of re-establishing it. For scripts that want one run and one exit code |
A raw --creds file is refused by ps, stop, attach and the other control commands, because
that route mints no endpoint-caller triple; the project folder, or --space against the registry
entry, is the route that does.
The human ps row is presentation text and is not a stable parsing target. Scripts use --json,
which is the machine-readable row contract.
--slots --wide is refused: --slots lists durable static slot rows, --wide prints live seat facts, and the two answer different questions. Across a multi-manager scatter, --slots prints each manager’s rows under its own instance header, the same way the plain ps scatter does.
These are operator clients over the running manager’s control plane. The default row includes the
connector, model pin, optional requested variant, and runtime as operational descriptors for the
managed row. They do not make a shared display name a unique protocol identity; use --json when
unambiguous owner+actor attribution is required. An omitted variant means no override was requested;
Cotal does not invent an effective provider default it cannot observe. ps also prints two state
facts per managed agent, because they answer different questions: the process fact from the manager’s
own runtime handle (running with its uptime, or exited with how long it ran), and the mesh fact
from the roster (idle / working / waiting / mesh offline, or not in roster when the seat has
no presence row at all: a seat that has not joined yet, or one that never did). When the seat’s
connector relays a harness-reported condition, the mesh fact carries its code and how long it has
held, so a seat whose turn died on a provider rate limit reads waiting (rate_limit for 40m) rather
than a bare waiting, and --json carries the whole condition object. When the connector reports
the seat’s last work event (presence activeAt), the mesh fact ends with its age, such as
· active 3s ago, and --json carries activeAt. A seat whose turn stopped advancing keeps
heartbeating, so its presence row stays fresh and this age is what shows the stall. A seat can be
running and mesh offline at once: the process is alive and its presence has lapsed. That row says
how long, as in mesh offline for <age> with an age such as 3.5h, counted from the seat’s last
presence heartbeat, which --json carries as offlineSince (epoch ms). The age is read only from
the seat’s own presence record, matched on its principal and lifecycle uid, so a same-named peer or
an older lifecycle never dates it. The manager log names each managed seat that is offline on the
mesh while its slot is held
(seat offline on the mesh: <name> - last heartbeat <time>; process <state>), including one its
watch first sees offline after a reconnect, and each one that comes back
(seat back on the mesh: <name>), so a watchdog that only checks process liveness has a line to
act on. The manager does not reap or re-key such a seat. The mesh fact is only a verdict while the
manager’s own presence watch is fresh: when that watch has been silent past the liveness window, or
has not replayed the bucket yet, every row prints mesh unknown with the reason instead (--json
carries it as meshView: stale | unpopulated), because offline and not in roster would then
describe the manager’s watch rather than the seat. The manager rebinds a watch that goes
silent under a live connection on its own, so mesh unknown normally clears within a liveness window.
On a user-auth mesh ps also renders each managed agent’s last credential-refresh outcome, fail-closed.
Mode split (chosen up front, never try-scatter-then-degrade):
- Static / open mesh. Bare
psis a class scatter: it freezes the live manager class from the records registry, merges every registered instance’s agents grouped and attributed per instance, and a non-answering instance is shown asregistered, no answer within the deadline(never silently omitted). A refused list or a missing answer makes the census incomplete: rows from other instances remain visible, butpsprints an incomplete-census warning on stderr and exits non-zero, including with--json. Those rows are not a complete seat count. A contract mismatch prints one plain comparison of the requested and served input/output digest pairs and advises aligning manager versions. The no-answer label means only that the instance is registered and did not answer. It does not say the host is down, because a dead host never deregisters itself and a live one can be slow; if it is gone, deregister it.--on <instance>pins the read to one exact instance id instead. A wrong pin fails loud rather than falling through: a well-formed id that no live manager carries is reported asmanager instance <id> did not answer(nothing else is asked), and a credential without that instance’s rail is reported as refused by the broker, not as an unresponsive manager. A manager that answers with a refusal is shown with its own cause; “no manager reachable” is said only when nothing answered at all. If the scatter’s own registry read fails (the freeze or the reconcile),pssays the manager registry could not be read rather than pronouncing on the managers, which may all be up.
The verdict is scoped to the endpoint rail the request rode. An issued caller rides the
versioned ep.v1 rail, a separate subject space from the legacy ep rail, and an endpoint serves
both (SPEC 13.15). A manager older than the versioned rail serves ep alone, so it can be running,
registered and answering while an issued caller’s request reaches nobody. Silence on ep.v1 is
reported as no manager answered on the <rail> rail with ep.v1 as the rail, and names both causes
it is consistent with: no manager running, or one older than the rail. The CLI cannot tell them
apart, because the service registry records no package version, so check whether a manager is
running and, if it is, its version. The same scoping applies to cotal run’s hosted verbs, which
drop the --local suggestion there, since --local drives the run from the calling process and
names the caller as its answerer.
stop and attach route by seat locality. A seat can only be stopped or attached by the
manager actually running it, and the class queue does not know which one that is. So on a
static/open mesh both verbs first ask every registered instance which one hosts the named seat, then
address that instance directly. This happens by default; you do not need --on.
--on <instance> remains the override, for when you already know where the seat lives or the
lookup itself is degraded. On a user-auth mesh, the exchange selects one authorized manager
for a short-lived manager-caller view. --on requests a specific instance; without it, selection
must be unique. Discovery and the command use that instance route. The caller gains no registry
read or scatter permission. An absent, ambiguous or unauthorized selection refuses before sending
the command.
A seat is reported as not found only when every reachable instance answered for itself. An
instance that stayed silent past the deadline, or that refused the read rather than answering, said
nothing about which seats it hosts, so the seat may be running on it. That case reports that the
location could not be established, names the instances that did not answer, and states outright
that it is not a report that the seat is gone. Read it as unknown and retry with
--on <instance>; a retry loop that treats it as “already gone” stops looking for a seat that is
still running. A single manager cannot tell “hosted elsewhere” from “does not exist”: it answers
not-found for both, which is why the search asks all of them and why an incomplete search
concludes nothing.
- User-auth mesh.
cotal psreports what one authorized manager knows about your agents (an instance-addressed read against its in-memory roster, owner-filtered). It does not report other manager instances or establish whether they are reachable. Completeness across a multi-manager user-auth space is not claimed. A manager that does not answer fails the command outright (exit non-zero), rather than printing an empty list that could be read as “no agents”. Your ledger row needs theadminscope to reachpsat all;spawnalone is refused by the broker (the ep tier boundary).
attach streams and drives an agent’s terminal on the pty runtime; detach with the escape key
(Ctrl-] by default; see COTAL_DETACH_KEY). The key is recognised as the legacy
control byte and as the kitty keyboard protocol and xterm modifyOtherKeys encodings of the same
press, so a terminal with either protocol enabled detaches too. It does so over a one-use, holder-bound
mesh session (SPEC §13.6): the manager replies with a signed session grant (never a
127.0.0.1 URL), the CLI redeems it once over the broker, and the browser console (cotal console)
drives the same session. stop and attach need a running manager to talk to. On a static mesh
they are cross-agent admin operations. On a user-auth mesh, your own agents (any agent under your
owner) need only the spawn scope; another owner’s agent needs admin on your ledger row
(identity & auth). Launch detached agents with spawn --detach.
attach reconnects when the link dies. A session lives on a network link, and a laptop that
sleeps, a VPN that drops or a wifi handover kills it. When that happens attach prints
[cotal: connection lost, reconnecting] on stderr and starts asking the manager for a new session:
a fresh grant, a fresh per-session credential, a fresh connection, so every attempt re-runs the same
authorization the first attach did. On success it prints [cotal: reconnected], the manager repaints
the seat’s current screen the way it does for any attach, and you carry on in the same terminal.
Retries wait 1s, 2s, 5s, 10s, then 30s, for as long as the seat exists. The detach key is read the
whole time the loop runs, the waits and the attempts alike, so a reconnect never traps you: press it
while a session is being established and the attach ends there, and a session that lands behind the
press is handed back to the manager rather than left holding a slot. Everything else you type while
there is no session is dropped rather than queued, so keystrokes aimed at a terminal that turned out
to be frozen, Ctrl-C included, are not delivered to the agent by a reconnect you did not know had
happened. That starts before the first session, not at the first reconnect: at a terminal, attach
reads and drops what you type while it is still resolving the mesh, so a key struck at a prompt that
has not come up yet does not reach the agent when it does.
The terminal is in raw mode for the whole reconnect, including when the link died before the first
session finished opening, so the detach key works there too instead of echoing as ^].
A pipe carries script input. For example, printf 'ls\n' | cotal attach --name web is
buffered until the session opens. Buffering continues across reconnects, so
tail -f log | cotal attach --name web does not lose the part of its feed written while the link was
down. Only a terminal gets the reader; --no-reconnect keeps the old behaviour on both.
It stops on its own when reconnecting cannot help, and says why: a manager that refuses the attach
exits non-zero with the manager’s own message, and a reconnect that finds the seat no longer there
(despawned, or its agent exited while the link was down) exits cleanly with seat <name> is gone.
A local connect refusal that retrying cannot fix, such as a static-auth mesh whose seed is now
missing, also exits non-zero with the refusal’s own sentence. A broker that is still unreachable
keeps the loop trying in silence.
A refusal that could still pass, such as a manager at its session ceiling, is relayed in the
manager’s own words while the loop keeps trying, once per refusal rather than once per attempt.
Pressing the detach key, or the agent’s process exiting while you are attached, ends the attach as
it always did. --no-reconnect turns all of this off and restores the single-session behaviour,
which is what a script wants.
Each reconnect also hands the abandoned session back to the manager, over the first link that can
carry the message, so an attach that flaps does not eat the manager’s session slots one outage at a
time. If that message never gets a link, the attach says so when it ends. The live-session ceiling
defaults to 64 concurrent sessions (--max-sessions); the browser console opens one session per
pane, so a dashboard over a large mesh should size for agents × panes. Hitting the ceiling refuses
before a credential is minted and names --max-sessions.
Which mesh attach resolves also decides how it redeems the grant. On a registered open mesh
there is no local seed. The CLI connects bare, the same way other control commands already do, and
the session rail is the caller rail that a real open-mode connection already reaches. Telling the
operator to re-register the root is false: the registered root is already the contract. On a
static-auth mesh the grant is still redeemed by minting a short-lived
session-scoped credential from the seed at the root the mesh resolved to, never from a .cotal
found by walking up from whichever directory you happen to be standing in. The difference is not
hypothetical: ~/.cotal exists on every install because the mesh registry lives there, so a command
run anywhere under your home directory but outside a project used to mint from your home
directory’s trust and present it to a broker that trusts a different chain, which surfaced as a
bare authorization failure that named nothing. A directory that does hold another chain for the
same space is now reported on the way past, and not obeyed:
! this directory resolves to /Users/you, whose .cotal/auth holds a DIFFERENT trust chain for space "team". attach used /Users/you/projects/app, the root this mesh resolved to. The other one is not being used, and is worth a look.When a static-auth mesh holds no seed at the resolved root, attach refuses and names what it
resolved, the broker and the root, instead of describing a directory it did not use and instead of
taking the open-mode path. An authenticated registry entry with a missing seed is still
authenticated. On a USER-AUTH mesh attach reads no seed. It sends your login and the session grant to
the auth service, which issues a session-caller bearer only if your owner and actor hold that
session. The connection it opens expires with the session grant.
Terminal bytes stream over the mesh; the manager’s own HTTP/WS face serves the console. That endpoint binds
loopback by default, so nothing is exposed by accident; cotal up --host <addr> passes its bind
address down, which is what lets you reach the browser console (cotal console) for an agent whose manager runs on another machine.
attach does not use that face: it redeems a signed mesh session grant over the broker instead (see above), so it reaches a
remote manager regardless of the bind address. A
bare cotal supervise and an embedded manager stay machine-local. Set it directly with
supervise --console-host <host>.
That address is recorded on the mesh and carried forward, because it is a decision rather than
something later commands can work out for themselves (a broker dial address is not a manager bind
address). Every later manager launch for the same mesh reuses it, including a same-root cotal up repair,
an adopted preserved or restored listener, and a spawn -f manifest deploy. A manager replacement
does not quietly move a reachable attach face back to loopback. Passing --host again overrides it,
so you can widen or narrow exposure whenever you like; a mesh that never asked stays loopback-only
and records nothing.
Because that face mints terminal read and write authority for every managed agent’s browser session, it is credentialed in two
tiers. A mesh caller receives a ticket bound to the single agent the manager just authorized,
single-use and short-lived, so one authorized attach can never be re-pointed at someone else’s
agent. The console token is the operator’s own, reaches every agent, and is printed only to the
manager’s output. The roster, the live feed, and the PTY stream all answer 401 without one; the
static console shell is served openly, since it describes no agent.
cotal input --name <n> --text <text> [--no-enter] [--on <instance>] [--space <s>]| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Which manager to reach |
--name <n> |
Managed agent to type into (required) | |
--text <text> |
The text to type, taken verbatim (required) | |
--no-enter |
off | Type the text and stop there, without pressing Enter |
--on <instance> |
class anycast | Pin to one manager instance id using the same rules as attach |
Types one line into a running agent’s terminal, as if you had typed it there, and returns. This is
the half of attach that a program wants: attach is a live stream that holds a
session open and expects a terminal on your side, so a script, a cron job or a web UI cannot use it
to send a single line. input is one authorized call.
What it is for is harness commands. A line beginning with / is not chat and not a message: it
is something the agent’s own harness handles, and the only way in is the keyboard.
cotal input --name reviewer --text "/compact" # ask the harness to compact its contextcotal input --name reviewer --text "/model opus" # switch its modelcotal input --name reviewer --text "hold on that PR" # ordinary typing works tooQuoting. --text takes a value, so a payload starting with / survives as written. A payload
starting with a dash needs the = form, because the shell-style --text --foo is ambiguous and is
refused rather than guessed:
cotal input --name reviewer --text=--verbose # dash-leading text: use --text=<value>Enter is pressed by default, since a command typed but never submitted has not been delivered.
--no-enter types the text and leaves it sitting at the prompt, which is how you stage a line and
send it later.
Nothing comes back but a delivery receipt (✓ sent 9 bytes to reviewer, counting the trailing
carriage return). Whatever the agent does next shows up where its output already goes: the mesh, its
transcript, or an attach.
This one is operator-only, and more narrowly than stop or attach. Those two are granted to
anything holding spawn, so an agent can stop and attach to seats under its own owner. input is
not: it is granted only to operator credentials, which on a user-auth mesh means your ledger row
needs the admin scope, the same scope ps already needs there. The reason is
that a write into a terminal is control of whatever is running in it, and on a user-auth mesh the
own-owner rule covers every seat under you, not only the ones you launched: a spawn-scoped agent
could otherwise type into a sibling it never started. Seat locality is still resolved for you.
Only the pty runtime can be typed into. The external terminal runtimes (tmux, cmux, orca,
herdr) attach to a process they do not own, so they have no input stream for it and the command
refuses by name rather than dropping the keystroke.
personas
Section titled “personas”cotal personas list [-v] [--running]cotal personas show <name>cotal personas edit <name>cotal personas new <name> (--prompt <t> | --from <f>) [--role <r>] [--model <m>]cotal personas rm <name> --force| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Which mesh’s persona catalog |
--role <r> |
none | new: the persona’s role |
--model <m> |
none | new: the persona’s model |
--prompt <t> |
none | new: the persona’s prompt text |
--from <f> |
none | new: seed the prompt from a file |
--verbose, -v |
off | list: include role / model / description |
--running |
off | list: mark personas live on the mesh |
--force |
none | rm: required, delete without prompting |
Personas are the local agent files under the resolved mesh root’s .cotal/agents/, the same catalog
cotal spawn launches from. --space and --server therefore move every list, read, write, delete
and completion operation to the selected mesh. An unresolved target refuses rather than falling back
to the current directory. See Agent files for the file format.
supervise
Section titled “supervise”cotal supervise [--runtime <name>] [--space <s>] [--server <url>] [--spawn <names>]| Flag | Default | Meaning |
|---|---|---|
--space <s> |
this folder’s auth space | Space to supervise |
--server <url> |
hosting mesh, or matching registered mesh | Broker URL. A registered mesh supplies it when omitted; a different explicit value is refused before anything is dialed. |
--runtime <name> |
pty |
Agent runtime (pty built in; extension runtimes are explicit-only) |
--console-port <n> |
none | Protocol-console port |
--console-host <host> |
loopback | Bind host for the console endpoint. Loopback keeps it machine-local; cotal up passes the address it bound the broker to, which is what lets the browser console reach this manager from another machine. cotal attach does not use this face: it redeems a mesh session grant over the broker |
--max-sessions <n> |
64 | Live-session ceiling. Each console pane and each cotal attach is one session, so size for agents × panes, not agent count. A capacity refusal names this flag. cotal up --max-sessions records the same number on the mesh so a later supervise started by repair or spawn -f keeps it |
--roster <file> |
none | Declarative roster to boot at startup. See Roster files |
--launch <spec> |
none | Resolved manifest launch spec (from up -f / spawn -f) |
--spawn <names> |
none | Comma-separated personas to pre-spawn at startup |
The manager is the agent supervisor and control plane: it answers spawn --detach, stop, ps,
attach, and the cotal_* manager tools. cotal up --detach starts one for you; run supervise
directly to recover a dead manager or drive a custom runtime. Default runtime is pty; install an
optional provider first (cotal ext add @cotal-ai/orca, @cotal-ai/tmux, @cotal-ai/cmux, or @cotal-ai/herdr) and
select it explicitly. A missing provider or app fails loudly; there is no fallback. See Deploy.
Boot inventory decides whether this process takes unpinned spawn/launch on the class rail:
if every declared connector is unavailable, those commands stay on this instance rail only
(status reports classSpawn: false). describe still answers on the class rail, so an
unpinned spawn can bind-fence against a skip member; re-issue, or pin --on. A partial
inventory keeps the class rail and names --on on a harness refusal, because sibling
inventories are not readable from the serve credential. See control surface.
On a normal SIGINT/SIGTERM, the manager stops every seat and requires the selected runtime to
prove the seat is gone before it releases the manager lease or service registration. A stop that
cannot prove exit fails loud and keeps manager authority instead of reporting a clean shutdown while
an orphan still holds broker rails. When the runtime refuses the stop itself, the shutdown fails at
once and its error names the refusal as stop failed: <message>. After an abrupt manager death, the same logical successor
terminalizes only its own durable static slots, verify-evicts the predecessor’s broker principal,
records that result in the lifecycle’s caller-readable audit detail, reaps the predecessor’s seat
process through the runtime’s custody reference recorded on the slot (the pty runtime verifies the
process start identity in its seat record, so a reused pid is never signalled), and only then
retires the lifecycle and frees the alias. A runtime that custodies its seats reserves that
reference before it launches one, and the manager records it on the slot’s first durable row, so a
manager that dies part-way through a spawn also leaves a seat its successor can address. A
same-lifecycle restart or a resume records the new seat’s reference on the slot the same way, and
when the slot does not take it the restart or resume fails and stops any seat it started, so the
slot never names a seat that has already exited while its replacement runs. A resumed seat keeps
its retained credentials, so the resume frees it only once its exit is proved; a seat whose stop
cannot be proved stays managed, and the resume’s error says so, naming a stop the runtime refused
as stop failed: <message>. A spawn
that launched its seat and then failed is rolled back by the manager that launched it, and that
rollback reaps the seat through the same reserved reference before the lifecycle retires. Missing or unverified broker evidence keeps the slot
terminalizing, and so does a runtime that cannot reap by reference.
A meshes add --mode user entry is a participant registration, not hosting authority. A
participant may run supervise only when the host advertises the remote manager authority service
and the signed-in actor has the dedicated supervise ledger scope. The CLI obtains the closed,
loopback-only manager-service view; spawn and admin do not substitute for that scope. The
host issues the manager’s public-nkey JWT material through its lifecycle-bound prepare → activate
→ renew protocol, never by handing the participant a signer or static provisioner credential.
The host also performs instance-scoped eviction and guarded gate reconciliation. A remote manager
refreshes its short-lived registration executor before clean deregistration, so a long-running
process removes its service row on SIGINT or SIGTERM. After an unclean stop, the same instance
verify-evicts its superseded family and advances the process epoch. If an abandoned frozen gate
holds the manager governance slot, a different supervise-scoped manager asks the host to reconcile
that holder after a complete gone verdict, then retries its registration once.
A remote supervise never uses local signing trust: with host-issued authority in hand, the manager mints from that authority alone and consults local records only to refuse a conflict, namely the supervised space’s own trust records under the cwd root. A root that hosts another static space beside the sign-in is a normal configuration and is never read as this space’s trust.
The broker URL in the registry entry decides the transport. A remote broker is often published
over a wss:// edge rather than a raw nats:// port, and supervise dials whichever scheme the
record holds, starting with the manager-authority registration it runs before the manager exists.
The record also decides whether that registration requires TLS, so a participant never downgrades
the credential exchange to a plaintext connection the registry did not describe.
Without that advertised host service or scope, supervise refuses before it starts a manager.
Run cotal spawn without --detach to launch a foreground agent, or ask the space host to enable
the authority service and grant supervise for detached agents. If a running remote manager loses
renewal, it reports degraded state and refuses unsafe new starts and restarts; live agents are not
silently replaced. Do not run cotal down or cotal up on a participant machine to repair this
condition.
service
Section titled “service”cotal service install [--mesh <name>] [--linger]cotal service status [--mesh <name>] [--json]cotal service uninstall [--mesh <name>]| Flag | Default | Meaning |
|---|---|---|
--mesh <name> |
this folder’s mesh | The mesh whose manager the service runs; one unit per mesh |
--linger |
off | install: when lingering is off, ask logind to enable it so the user manager starts at boot and the service survives logout. Never enabled silently |
--json |
off | status: machine-readable output |
Runs the manager as a user service so it survives logout and reboot. On Linux this installs a
systemd user unit (~/.config/systemd/user/cotal-manager@<key>.service, where <key> is the
case-safe mesh key); on macOS a launchd agent plist under ~/Library/LaunchAgents/. Any other
platform, or an absent systemd/launchd user session, fails with a message naming what is missing.
install resolves the mesh from the registry and binds the unit to that entry’s root and broker
address, so it can be run from any directory. The mesh must be registered (cotal up or
cotal meshes add) before installing; an unregistered name refuses before anything is written.
The unit’s ExecStart is the bare supervise command. The mesh facts travel in the unit’s
environment (COTAL_SPACE, COTAL_SERVER pinned to the registered broker URL, whatever port it
listens on) rather than the command line, because command lines are readable by every user on a
multi-user host. On Linux that environment is a 0600 EnvironmentFile; on macOS it is the
plist’s EnvironmentVariables. The same environment gives the service a private COTAL_HOME and
XDG_CONFIG_HOME under the unit directory, so the service manager never touches the login
user’s ~/.cotal. First-run connector seeding runs synchronously inside service install,
against that private config root; the unit itself starts with COTAL_SKIP_CONNECTOR_SEED=1
so a manager is never interrupted mid-seed by a restart. An install whose pre-seed cannot
complete (network unreachable, registry error) refuses instead of deferring.
The same environment pins PATH to the PATH of the shell that ran install.
Without it the unit inherits the service manager’s own short
PATH, which usually lacks ~/.local/bin and Homebrew, so the manager’s boot inventory would
report a harness unavailable that your shell resolves. Install from a shell that resolves every
harness the service should launch, and reinstall after moving one. A relative entry, including
an empty one, is resolved against the directory you ran install from, because the unit starts in
the mesh root where the same spelling names another directory. An entry with a .. segment is
pinned as the directory your shell reaches through it, with symlinks followed, and refuses when it
reaches none. A PATH set to the empty string is one empty entry, so it pins that directory. An
unset PATH refuses.
Every value the unit derives from a path (WorkingDirectory, the EnvironmentFile path, the
ExecStart tokens) is escaped for systemd specifiers (% becomes %%), so a mesh root that
contains % starts over its real path instead of a path systemd rewrote by expanding it. The
provenance comment records the root unescaped.
On Linux a user unit starts at boot and survives logout only while the user lingers. Without
lingering, systemd starts no user manager at boot, so an enabled unit stays inert until the next
login and stops at the last logout. install checks lingering before it writes anything, and when
lingering is off it fails with the root command that turns it on (sudo loginctl enable-linger <user>). With --linger it first asks logind to enable lingering for the current user, and fails
with the same command when logind refuses (unprivileged users over SSH get Access denied).
service status prints that command while lingering is off. A Linger query that does not answer
yes or no (logind unreachable, no loginctl) is never read as off: install refuses with
the query’s own error and enables nothing, and service status shows lingering as unknown with
that error (--json gives "linger": { "error": ... }).
service install also refuses while a manager is already running for the mesh (cotal down manager first). The restart policy is Restart=always with RestartSec=20s, chosen for
manager units in production: a manager exits for reasons that are not failures (broker
restarts, host suspend), where on-failure with a short interval thrashes.
The unit also sets a start limit (StartLimitIntervalSec=30min, StartLimitBurst=20). A manager
that keeps failing to start stops after 20 attempts, about seven minutes at 20 seconds apart, and
the unit is left failed instead of restarting forever. One such failure is deliberate. After an
unclean stop, a manager that cannot verify eviction of its predecessor’s credentials exits 1 and
leaves the issuance gate frozen, because starting without that proof could let two incarnations
serve at once (SPEC 13.1). It first waits up to 60 seconds for the delivery daemon to answer, so a
daemon that is still starting does not fail the start. The log names the cause. When the delivery
daemon is down, it says the daemon is not reachable on the ctl.delivery-admin rail. When the
daemon answers and refuses, for example because the space is missing a $SYS cred, it prints the
daemon’s own reason and repair step. Fix that cause, then run systemctl --user reset-failed <unit> and systemctl --user start <unit>. The macOS agent has no start limit: launchd’s
ThrottleInterval only spaces restarts.
service status reports the unit state from systemd/launchd, the manager’s own health read from
its pidfile at the unit’s recorded root, and the machine facts a hosting side asks for:
architecture, OS (the platform, never the hostname), whether /dev/kvm is present and
accessible, CPU count, and total memory. --json returns the same fields as one object. The
manager row names the recorded pid, and the command it runs when another program has reused that
pid. --json also gives the command of a live recorded pid whenever it can be read.
service uninstall stops and disables the unit and removes it plus the private state directory.
It works from any directory: the unit’s own records name the mesh and root it serves, and an
explicit --mesh <name> selects it. It refuses any unit that was not written by service install (the files carry a provenance comment), whose recorded mesh is missing, or that was
installed for a different mesh, so operator-written units are never destroyed; service status
applies the same rule and never reports a mesh a unit does not record. service status also
refuses a unit that records no absolute root, because the manager’s health is read at that root
and is never guessed from the current directory. Remove such a unit with service uninstall
and install it again.
This command installs only the manager. The per-space auth service and the delivery daemon are
not installed by it: on a shared broker an operator runs three units per space with After=
edges (auth service, then manager, then delivery) and stops them in reverse. A broker-side cotal up unit is a separate unit documented in Run a mesh.
reconcile-gate
Section titled “reconcile-gate”cotal reconcile-gate [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]| Flag | Default | Meaning |
|---|---|---|
--space <s> |
this folder’s auth space | Space the frozen gate lives in |
--server <url> |
the local mesh | Broker URL |
--endpoint <e> |
manager |
Endpoint whose gate is frozen |
--instance <id> |
this folder’s persisted manager instance | Instance id |
When you need this. A manager restart killed after deregistration begins but before the new
incarnation finishes leaves the endpoint’s issuance gate frozen, held by a
process that no longer exists. The freeze is what stops two incarnations serving at once, which is
correct. The successor manager now completes that dead registration itself on boot, including on
the remote user-auth path. A foreign remote manager blocked by this gate also asks the host to repair
it before one registration retry. Both use the same guard this command uses: they act only when the freeze-holder is affirmatively gone under a complete
CONNZ sweep (gone and sweepComplete=true). If that registration’s spec write already committed,
it finishes the same freeze at the committed registration revision. If the spec did not advance, it
abort-reopens the gate at generation+1 with processEpoch unchanged and continues the normal takeover.
Live, unknown, unestablishable, and
wrong-op-kind still refuse; there is no TTL.
Use this command when the automatic path cannot run: the delivery daemon is down, the repair targets a non-manager endpoint, or you want to lift the freeze without starting a manager. It checks that the holder really is gone, prints what it found, and then finishes the dead operation the same way as the interrupted restart would have: revoke the old credentials, evict their holders with verification, and reopen the gate.
The command revokes the old credentials 16 at a time. It then verifies the holders’ eviction in
shared sweeps of up to 256 holders on the delivery daemon. Each sweep scans the broker a fixed number
of times and kicks live connections 16 at a time, so holders that are already gone add almost
nothing and live ones add one broker round trip per 16 connections. The daemon must serve the
evictPrincipals verb; an older daemon refuses it and the gate stays frozen.
Each sweep durably records the holders it verified before the next sweep starts. If a holder is not verified gone, the command leaves the gate frozen with those records kept. An interrupted sweep records nothing, and the sweeps before it stay recorded. A retry still repeats the freeze-holder liveness check, then skips only progress bound to the same registration operation, frozen-gate revision, and holder set. The output reports holders completed before this attempt, completed now, and still remaining. A new freeze or changed holder set starts from zero. Cursor cleanup happens only after reopen; a retained cursor is harmless because its old gate revision cannot authorize a later freeze.
It refuses far more often than it acts, on purpose, and always says which check stopped it:
| Refusal | What it means | What to do |
|---|---|---|
holder-alive |
The freeze-holder still has a live connection: a manager is running | Stop that process first. Reconciling would evict a live manager’s credentials |
holder-unknown |
The connection sweep could not prove the holder absent | Not safe to proceed: an unprovable holder is treated as a live one. Re-run once the broker answers completely |
liveness-unestablishable |
The delivery daemon gave no verdict: it was unreachable, timed out, or refused | Act on the delivery lease line in the refusal (below). Silence is never read as death |
not-frozen / no-gate |
The gate is open, or there is no gate at that coordinate | Nothing to repair: check --endpoint / --instance |
wrong-op-kind |
Frozen under a takeover or retirement, not a registration | Out of scope for this command; it will not reinterpret another operation’s intent |
eviction-unverified |
The holder looked gone but eviction could not be verified | The gate is left frozen, unchanged. Investigate the broker before retrying |
raced |
A newer manager moved the gate mid-repair | Re-run cotal doctor and look again |
When the daemon gives no verdict, the refusal also reads the delivery lease (lease.0) and names
what is blocking the rail:
| Lease reading | What to do |
|---|---|
| absent | No daemon is running. Start it (cotal up runs it) and re-run |
| unreadable | The daemon cannot be named, so do not assume none is running. Fix the lease read, then re-run |
| held, not ready | That holder claimed the shard and has not bound its rails. Wait for it, or stop it so its lease lapses |
| held, ready, no answer | The query may have gone to another daemon still subscribed to the rail, such as a stopped one whose lease lapsed. Re-run before stopping anything. If no run gets an answer, stop any other delivery daemon for the space, then stop or restart the holder |
| changed hands | The holder took the shard after the query was sent, so it was never asked. Re-run before stopping anything |
The command reads the lease before it sends the query and again after the query fails. It names a holder as the blocker only when the same run of the same daemon held the lease both times, and two rows from a daemon too old to record its run never count as the same run. Even then a ready holder may not have been asked: the rail is queue-grouped, so any daemon still subscribed to it can take the query. A row whose times are not valid dates reads as unreadable.
A daemon that answered and refused keeps its own reason, followed by the same lease line. The lease line names the holder, whether it is ready, the space account that holds the lease bucket, when that holder acquired the shard, and when the row was last written. A ready holder rewrites the row on every renewal and keeps its acquisition time, which only a successful acquisition sets. A row written by a daemon that predates the acquisition time reports it as unknown. The lease reads never change the outcome: the gate stays frozen and the command exits 2. A manager’s boot self-heal uses the same check and reports the same line.
There is no --force, and no path that discards gate state: the only way this reopens a gate is by
proving the holder is gone and then completing the operation properly.
What reopening the gate does for the endpoint’s governance slot. A registration takes the endpoint-wide governance slot before it publishes its spec, and holds it until its gate reopens. An instance that died between those two points leaves the slot held with no registration behind it. This command does not write that slot and never has; the registration path is its only writer. What the reopen does is advance the holder’s gate past the generation the slot is stamped with, which is what marks the slot abandoned. The next registration for that endpoint then reclaims it as part of its ordinary start. So the repair here is still one command followed by starting the manager, and the slot needs no separate step.
deregister-instance
Section titled “deregister-instance”cotal deregister-instance [--space <s>] [--server <url>] [--endpoint <e>] [--instance <id>]| Flag | Default | Meaning |
|---|---|---|
--space <s> |
this folder’s auth space | Space the instance is registered in |
--server <url> |
the local mesh | Broker URL |
--endpoint <e> |
manager |
Endpoint the instance serves |
--instance <id> |
this folder’s persisted manager instance | Instance id, the whole id as cotal ps prints it |
When you need this. The service registry records registration, not liveness, and nothing in
the model expires a row. A manager that stops cleanly removes its own registration. One whose host
died without writing anything cannot, so its record goes on claiming a live instance forever: every
class scatter in that space freezes the dead slot in, and cotal ps, stop and attach each pay
their whole deadline waiting for a machine that is never coming back. A laptop that was reimaged, a
container that was deleted, a box that will not be back on the network: those registrations have no
other exit.
This command is that exit. It asks the instance first, and it removes a record only when the broker affirms the instance’s own rail is empty: nothing subscribed there. Then it deletes the registration’s two records keys, each pinned to the revision it read, and prints what it removed.
Silence alone never passes. An unanswered describe is what a dead host, a wedged process and a slow one all look like, and a hung process still holds its subscriptions, so the broker sees interest on its rail. That instance is refused and the observation is printed. A dead process holds no connection and therefore no subscription, so a real corpse is still removed.
Every refusal names the failed check:
| Refusal | What it means | What to do |
|---|---|---|
instance-answered |
The instance answered a pinned describe. It is alive | Nothing to repair. If it is wedged rather than gone, stop the process first; its own clean stop removes the record |
instance-not-affirmed-gone |
It did not answer, and the broker did not report its rail empty, which is what a held subscription looks like: slow or hung, not affirmed gone | Nothing was removed. Stop the process; its record goes on its own clean stop, or re-run this once it is down |
liveness-unestablishable |
The probe itself failed, so nothing was learned | Fix the probe’s path (credential, broker) and re-run. A probe that could not run is never read as death |
not-registered |
No registration at that coordinate | Check --instance and --endpoint. This takes the whole id, never a prefix |
registration-in-flight |
The instance holds the endpoint governance slot at the live issuance-gate generation, so a registration is still completing | Nothing was removed. Wait for that registration to finish, then re-run |
superseded |
The record moved between the read and the delete | Something is writing to it. Nothing was removed; re-observe before retrying |
There is no --force and no sweep: silence is not death, and a rule that removed rows on silence
would eventually remove a live instance that was merely slow. An operator names one instance, the
broker’s verdict on its rail is what authorizes the removal, and the guard’s job is to show them
they named a dead one. Removal is not a one way door either. The same instance re-registers over
the tombstone on its next start, under the same identity.
runtimes
Section titled “runtimes”cotal runtimesLists every agent runtime the manager can spawn through: the built-in pty, the official providers
(orca, tmux, cmux, herdr), and any custom provider installed via cotal ext add. Each installed
provider is probed so you can see what is actually reachable on this machine before selecting it:
pty built inorca installed · reachable @cotal-ai/orcatmux available · cotal ext add @cotal-ai/tmuxcmux available · cotal ext add @cotal-ai/cmuxherdr available · cotal ext add @cotal-ai/herdrinstalled · reachable / unreachable is the provider’s own available() probe; available means
it is a known runtime you can add with the shown command. Selecting an unknown or uninstalled runtime
via up/spawn --runtime <name> fails loud and, for a known one, points at the exact cotal ext add
package. There is no silent fallback to pty.
cotal seats [--drain]| Flag | Default | Meaning |
|---|---|---|
--drain |
off | Retire every seat whose agent has exited. A seat whose agent still runs is kept |
The pty runtime used to start a detached custodian process for every Linux seat. It now spawns
in-process, but custodians that an earlier manager started keep running, and one whose agent has
exited stays resident while a manager still holds its connection. This command lists the custody
records under COTAL_SEAT_ROOT (default ~/.cotal/seats), one line per seat:
| State | Meaning |
|---|---|
live-child |
The agent process still runs. The seat is never signalled, and a manager can still adopt it |
childless |
The agent has exited, or the record comes from an earlier boot. --drain retires the seat |
drained |
--drain proved the custodian and the agent gone and removed the record |
refused |
The record cannot be read, carries no start or boot identity, this host publishes no boot identity, or the reap could not prove the processes gone. The record stays on disk |
A drain signals only a custodian whose recorded start identity still matches the live process, so
a reused pid is never touched. No process outlives a reboot, so a record from an earlier boot is
reported childless and --drain removes it without signalling anything. A record with no start or
boot identity is refused with or without --drain, and is never reported as running or exited.
On a host that publishes no boot identity (/proc/sys/kernel/random/boot_id) every record is
refused the same way, because no record can be tied to this boot.
That refusal and an unreadable record signal nothing. A refusal from the reap itself can come after the drain already
sent SIGKILL to the custodian. Its detail names the pid or process group the reap could not prove
gone, so check those processes before you retry. The command exits non-zero when any record is
refused. It is Linux-only and throws on other platforms.
cotal send dm <agent> "<text>" [--space <s>] [--server <url>] [--creds <path>]cotal send msg <channel> "<text>"cotal send ask <role> "<text>"A send dm prints one line naming three facts: → <name> stored seq <N>; recipient <status> at send; delivery not confirmed <text>. stored seq N is the JetStream sequence the broker
assigned to the publish; recipient <status> at send is the roster status (idle, working,
or offline) resolved right before the publish, which can change the instant after; the send
never prints delivered, because the sender’s credential cannot read the recipient’s durable
to confirm it. Inspect what the broker actually holds for a recipient with
cotal deliver pending.
| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Which mesh, and (off-registry) which credential |
One-shot messaging: connect, send a single direct message (dm), channel post (msg), or role
ask/anycast (ask), then exit. For a running conversation, agents use the mesh tools instead
(MCP tools).
cotal send works from an operator shell or from a seat. Its display name is <login>@<host> of
the shell that ran it, so the recipient can tell one operator’s send from another’s; it is taken
from the operating system, never from COTAL_NAME. The wire principal comes from the resolved
operator credential or user bearer, not from COTAL_NAME, COTAL_ID, COTAL_OWNER, or
COTAL_ACTOR. On an open mesh the transient endpoint self-mints its principal.
The transient endpoint never joins the roster and binds no inbox. A recipient can still answer a
send dm or send ask with cotal_dm, by the sender’s name or by the id on the message it
holds: the reply is stored under the sender’s id in the space’s DM history, which an operator’s DM
view such as the dashboard’s Direct messages lens shows. The cotal send that asked has already
exited, so the reply never reaches that shell.
channels
Section titled “channels”cotal channels listcotal channels set <name> [--replay | --no-replay] [--window <n>] [--desc <s>] [--instructions <s>]cotal channels default --replay | --no-replay| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Target mesh |
--replay / --no-replay |
none | set/default: replay history to new joiners, or not |
--window <n> |
none | set: replay window size |
--desc <s> |
none | set: one-line channel description |
--instructions <s> |
none | set: instructions shown to joiners |
Inspects and edits the channel registry: replay policy, description, and joiner instructions. ACL
semantics (who may read or post) are set at mint / provision time, not here; see
Channels and permissions. On a user-auth mesh, list rides your
own login as is; set and default edit the registry over a short-lived
channel-writer view, which needs ledger scope admin (Identity & auth).
On a remote user-auth mesh that view is served by the public exchange; space-history purger
and the read-only admin view are not.
history
Section titled “history”cotal history clear --force [--dms] [--space <s>]| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Target mesh |
--dms |
off | Also clear DM history |
--force |
none | Required: clear without prompting |
Purges retained channel history; --dms extends it to direct-message history. An alias of
clean history. On a user-auth mesh the purge rides a short-lived purger view over
your login, which needs ledger scope admin (Identity & auth).
console
Section titled “console”cotal console [--plain] [--space <s>]| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Space to watch |
--plain |
off | Line stream instead of the TUI |
A live protocol view for a space: a lazygit-style TUI, or a plain line stream on --plain. On a
user-auth mesh it rides the read-only admin view over your login, which needs ledger scope
admin. Inside the TUI, operator control (D kill, :spawn, :status, :purge) rides the
same per-action instrument path as cotal stop and cotal ps, never the observer; a raw
--creds file cannot drive it. a (or :attach <agent>) runs
cotal attach in place and returns to the console on detach. See
Watch a mesh.
cotal web [--detach] [--host <host>] [--port <n>] [--no-open] [--space <s>]| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Space to serve |
--host <host> |
127.0.0.1 |
Concrete HTTP bind and browser host; wildcard addresses are refused |
--port <n> |
7799 |
HTTP port, a decimal number from 1 to 65535 |
--detach |
off | Run in the background; stop with cotal down web or bare cotal down |
--no-open |
off | Don’t open the browser |
The browser observability dashboard: presence, channels, and a live feed. It is not part of
cotal up: it ships inside cotal-ai as the @cotal-ai/web extension, seeded automatically on first
run (like the built-in connectors) so it always matches your CLI version. It self-registers cotal web
into this surface and serves
http://cotal.localhost:7799 by default (loopback; *.localhost resolves in Chrome/Firefox/Edge; for Safari
or a system resolver such as WSL2’s, the launch link is also printed at http://127.0.0.1:7799).
On a user-auth mesh the dashboard rides the read-only admin view
over your login, and a channel purge asks for its own channel-purger view per click; both need
ledger scope admin. The public exchange serves channel-purger for a remote owner; it still
refuses the startup admin view, so a remote cotal web is not a complete channel-management
surface. Detached mode re-execs the current Cotal installation, writes diagnostics to
the mesh root’s .cotal/web.log, and reports success only after the HTTP server answers. It requires
a recorded mesh root, but can be launched from any directory once cotal up has recorded the mesh.
See Watch a mesh.
linear
Section titled “linear”cotal ext add @cotal-ai/linearcotal linear account add <name> --mode <write|readonly> (--token-stdin | --token-file <path>)cotal linear account login <name> --mode <write|readonly>cotal linear account show <name>cotal linear inventory <account> [--json]cotal linear call <account> <tool> [--args '<json>'] [--inventory <digest>] [--timeout <ms>]cotal linear resource <account> <uri> [--inventory <digest>] [--timeout <ms>]cotal linear prompt <account> <name> [--args '<json>'] [--inventory <digest>] [--timeout <ms>]cotal linear serve <account> --endpoint <reverse-dns-name> [--space <s>] [--server <url>]cotal linear caller <name> --endpoint <reverse-dns-name> --out <path> [--channels <a,b>] [--expires-in <s>]Serves one Linear account as a registered Cotal endpoint that agents call with cotal_describe and
cotal_invoke, and gives the operator a direct client for the same server. It only talks to https://mcp.linear.app/mcp
(mode write) or https://mcp.linear.app/mcp/readonly (mode readonly); there is no URL or header
option. Use a separate account per Linear workspace, and prefer a readonly account with a
restricted read key wherever writes are not needed.
An account’s credential is sent as Authorization: Bearer and is never taken from argv. account add
takes a Linear API key: --token-stdin stores it under the cotal home as a 0600 file, and
--token-file points at an existing file that must not be readable by other users. The file is read
at each use, so rotating it needs no restart. account login runs Linear’s OAuth flow instead: it
registers a client, prints the authorize URL, and waits on a loopback redirect. It asks for scope
read in readonly mode and read write in write mode, keeps the tokens in a 0600 file, and
refreshes them before a request is sent. OAuth requests only go to https://mcp.linear.app.
Redirects are refused so the credential is never forwarded.
inventory reads the server’s capabilities and every page of its tools, plus resources, resource
templates and prompts when the server advertises them. Names and schemas are printed as the server
sent them. The digest covers the whole inventory; pass it to --inventory so a call is refused
before dispatch if the inventory changed. Capabilities this command does not represent, such as
resource subscriptions or logging, are listed under unsupported.
call, resource and prompt print the server’s reply as JSON. A tool result with isError: true
is a normal result (exit 0). A refusal before dispatch exits 2 with outcome not-executed; that
includes a deadline that expires while queued, discovering or connecting. A protocol error, a
timeout, Ctrl-C, a response over the size cap, an expired session, an HTTP 401, 403 or 429 answer,
or a transport failure after the request was sent exits 3 and is never retried. Its outcome is
unknown, for read-only tools too: read-only means a repeat is safe, not that the first call did
not run. Errors carry an HTTP status or an error class, never the response body.
Serving the endpoint
Section titled “Serving the endpoint”serve registers the account as an endpoint named by --endpoint and serves it until Ctrl-C. The
name must be a reverse-DNS name in a namespace you own, such as com.example.linear; it is
authorized for this mesh’s local owner only because you configured it. serve reads the inventory
first, so an account that cannot be reached is never registered. It then walks the same path every
registered endpoint does: it publishes the contract to the contract store, opens the issuance gate,
registers a fresh instance, writes its ready status, and mints a scoped serve credential through
the gate. Every step runs on short-lived executor credentials scoped to this one endpoint instance.
The space signer stays in the serve process for renewal at 75% of the credential’s life and is
never written out, printed, or handed to the serving connection. Ctrl-C stops serving, closes the
Linear session and removes the instance’s service record, each within 10 seconds. A start that
fails after registration removes the record before it exits and says whether that removal
completed; a registration that fails part-way rolls its own record back and reports when it could not.
The endpoint serves five fixed commands: inventory, call-tool, read-resource, get-prompt and
complete. Each needs the one linear.mcp capability. That capability is permission to call the
endpoint; which tools a call can reach is decided by the Linear account and its mode. inventory
replies in pages that fit a broker message: the first page carries the server, its capabilities,
its instructions and per-section counts, and every page carries entries verbatim with an opaque
cursor bound to the inventory digest. An entry is never cut: discovery refuses an inventory whose
first-page head (server, instructions, capabilities) or any entry prints over 48 KiB the way an
agent tool prints it (indented JSON, which grows with nesting), naming the entry, so the first page
and every entry it accepts fit one reply and the tool an agent reads it with. A cursor or inventoryDigest from an older inventory is refused before anything is sent to
Linear. call-tool, read-resource, get-prompt and complete require the digest the caller read.
Callers
Section titled “Callers”caller provisions an isolated caller on a static-auth mesh: an agent credential with its own
mailboxes and request rows for only the five Linear commands on that endpoint, plus the channels
in --channels. It carries no spawn, run, admin or provisioner capability. The credential is
written 0600 to --out and never printed. Start a hand-driven connector session with it, using the
printed COTAL_CREDS and COTAL_LIFECYCLE_UID; that session then sees the endpoint through
cotal_describe and calls it through cotal_invoke, and the broker refuses any command its
credential does not name. A Claude Code or jcode session also serves a local control endpoint and
refuses to start without one: set COTAL_CONTROL_SOCKET to a socket path you choose and
COTAL_CONTROL_TOKEN to a random secret. OpenCode, Codex and pi sessions start without them. Managed cotal spawn agents on a static-auth mesh still cannot carry
endpoint capabilities, so a Linear caller is a hand-launched seat.
A per-user-auth mesh is refused by both serve and caller: there, endpoint and caller credentials
come from the remote auth service, which this package does not support yet. Nothing falls back to a
local signer. An open mesh is refused too, because it has no credentials to scope.
Not represented yet, and listed under unsupported when the server advertises them: resource
subscriptions, logging, progress notifications and experimental capabilities. The server cannot
call back for sampling, elicitation or roots. Every call acts as the account’s one Linear identity.
deliver
Section titled “deliver”cotal deliver [--space <s>] [--server <url>] [--tls] [--creds <file>] [--root <dir>] [--shard <n>] [--shards <n>] [--dev-mint]cotal deliver pending <name> [--limit <n>] [--durable <name>] [--json]With no positional, cotal deliver runs the delivery daemon (see
the delivery daemon). deliver pending <name> never starts the daemon: it
is an operator-only read over one recipient’s DM durable, for the moment after a send when the
question is “what does the broker actually hold for them.” It resolves <name> against a short
presence watch (an offline card still counts, since the recipient may be dead, that is what
the verb exists to inspect); when neither a card nor the durable can be found, it prints
✗ not-found: no agent "<name>" and no DM durable for it in space <s> and exits non-zero, never
pending 0. On a match it prints the durable name and one fact per line: pending,
ack-pending, delivered, ack-floor, created, frontier, and the stream’s max_age /
max_msgs_per_subject / discard limits (--json prints the same facts as one object), followed
by a bounded, unacked read of up to --limit (default 20) recent candidate message ids under the
heading recent candidate ids (from the ack floor; not proof of a hole), a list of what is
there, not proof that nothing was lost.
The verb needs the admin credential profile: it runs through the same static-mesh route as
cotal mint --profile admin, and refuses a user-mode mesh, naming the retired static credential,
because there is no user-mode inspection authority yet. Pass --creds <file> for an off-registry
admin credential. A same-name respawn never inherits a predecessor’s held DMs (the durable is
lifecycle-keyed); an old lifecycle’s durable is reachable only by the name a live read printed
(the <durable> line on the first line of this verb’s output). Pass that name with --durable <name> to read it directly once the lifecycle’s card is gone from the roster. This skips the
presence watch on <name> entirely, so <name> is required but only echoed in error text.
cotal mint <name> [--profile <agent|observer|admin>] [--out <path>] [--signer]cotal mint <name> --provision [--role <role>] [--space <s>] [--server <url>]cotal mint <name> --expires-in <seconds> | --expires-at <unix-seconds>cotal mint <name> --identity <creds> [--expires-in <seconds>]| Flag | Default | Meaning |
|---|---|---|
--profile <agent|observer|admin> |
agent |
Credential profile |
--out <path> |
.cotal/auth/creds/space.<key>/<name>.creds |
Output path - the default sits under the resolved space’s segment (<key> is that space’s hex encoding, as in Project files) |
--signer |
off | Emit a stripped account-signing file instead |
--force |
off | With --signer: overwrite an existing file |
--allow-subscribe <a,b> |
the agent file’s, else subscribe | Read-ACL override, agent profile only: observer and admin carry a fixed read set, and mint refuses this flag there rather than narrowing nothing |
--allow-publish <a,b> |
the agent file’s, else deny | Post-ACL override, agent profile only |
--role <role> |
the agent file’s | Agent profile: the anycast task queue the identity pulls (svc_<role>) |
--provision |
off | Agent profile: also pre-create the identity’s bind-only DM/deliver durables (and its role’s task queue) on the live mesh, so the credential can consume |
--expires-in <seconds> |
unbounded | Bound the credential’s lifetime: the JWT exp is iat + <seconds>. A positive integer; refused together with --expires-at |
--expires-at <unix-seconds> |
unbounded | Bound the credential to an absolute exp (unix seconds). Refused together with --expires-in |
--identity <creds> |
a fresh identity | Re-mint for the nkey carried by this creds file, keeping the principal and every durable keyed to it. The file is read by the same loader the endpoint uses; a file with no seed is refused by name |
--space <s>, --server <url> |
the resolved mesh | Which root supplies the agent file, static trust and default credential storage; with --provision, also which live mesh receives the durables |
Mints a NATS creds file for a space in static auth mode, scoped to a profile and (optionally)
explicit read/post ACLs. --signer emits an account-signing file for delegating minting to another
host. A per-user-auth space refuses mint: agents there join under a logged-in user
(login + actor grant), never via a handed-out creds file. See
Identity and auth.
For an agent profile, the resolved mesh root supplies the persona ACL, the signing material and the default credential destination as one authority. If the current folder also holds trust for a different space or account, mint refuses before writing and names both roots. It never combines a persona from one root with credentials signed or stored under another.
A plain mint is creds only: the identity can publish within its post ACL at once, but on an authed
mesh its DM inbox and task queue are provisioner-pre-created and bind-only, so a consuming
connect fails until they exist. --provision performs that pre-create in the same command (a
provisioner cred is minted from the space’s trust material, used, and dropped), so a long-running
client you start yourself can receive DMs and role anycasts like a spawned seat. The command prints
the identity’s principal (its wire id) and lifecycle uid; a consuming client passes that uid as its
lifecycleUid. Agent profile only; an open mesh needs none of this (peers self-create there). The
same resolved authority is used for both the credential and --provision, so the broker
footprint cannot be created under a different root’s trust material.
The CLI-mintable profiles carry no default TTL: without a lifetime flag the credential is
unbounded, and a standing-renewal consumer refuses it. --expires-in <seconds> (or
--expires-at) is the door the renewal seam’s own error names. --identity <creds> re-mints for
the nkey the file already carries, so the new credential presents the SAME principal and every
durable keyed to it survives; combine it with a lifetime flag to rotate an expiring credential
without churning the identity.
cotal login --idp <auth base URL> [--client-id <id>]cotal logout --idp <auth base URL>Signs you in to a per-user-auth mesh’s IdP (device code flow) and caches the session; run it
once per machine. The IdP URL must use https://. Plain http:// is accepted only on a loopback IP
literal such as 127.0.0.1 or ::1, for an IdP on this machine. localhost is refused because it
is a name: a hosts entry would choose the IdP. It prints your IdP subject, the id the operator
grants against. When the trusted
/token response advertises a same-origin space catalog, login validates and records that account’s
spaces immediately. After a
login, every command on that mesh works under your identity: each connect takes a fresh IdP
proof, exchanges it locally for a short-lived bearer, and is authorized against the actor
ledger at connect time. logout revokes the IdP session, clears its cache, and removes only that
account’s discovered registry entries. See
identity & auth.
# an upsert of the WHOLE row: name all three ACL flags, or pass --full for the wide defaults belowcotal actor grant <actor> --sub <IdP subject> --scope a,b --allow-subscribe a,b --allow-publish a,b [--role <r>] [--label <l>]cotal actor grant <actor> --sub <IdP subject> --full [--scope a,b] [--allow-subscribe a,b] [--allow-publish a,b] [--role <r>] [--label <l>]cotal actor revoke <actor> (--sub <IdP subject> | --owner <u_…>)cotal actor list| Flag | Default | Meaning |
|---|---|---|
--space <s> |
the folder’s | Space whose ledger to manage |
--sub <subject> |
none | The IdP subject (shown by cotal login) the actor belongs to |
--owner <u_…> |
none | The derived owner token (alternative to --sub) |
--full |
off | Fill each ACL flag left off with its wide default; without it, grant refuses unless all three are named |
--scope <a,b> |
spawn,role:default with --full |
Capability scope ('' = none; spawn = may run agents; role:<r> = may delegate role r; admin = cross-agent control; supervise = eligible for the closed remote manager-service view when the host enables it) |
--allow-subscribe <a,b> |
> (all channels) with --full |
Channel read ACL; the user’s envelope, their agents can never read beyond it |
--allow-publish <a,b> |
> (all channels) with --full |
Channel post ACL; also the envelope for their agents’ posting |
--role <r> |
none | Role (scopes the task-queue consumer) |
--label <l> |
none | Display label for actor list (never the IdP subject) |
The actor ledger is the single authorization source of a user-auth space: no row, no access.
grant --full is the full envelope (all channels; scope spawn,role:default, so it may spawn and may delegate the default role). A
grant that leaves off --scope, --allow-subscribe or --allow-publish without --full is
refused and writes nothing. A re-grant replaces the whole row, not the one field you name, so to add a capability spell
every field out: the new scope plus the row’s current read set, post set, role and label
(cotal actor list shows what a row holds). Under --full, a field left off does not stay as it
was: it reverts to the wide default in the table above. A re-grant retires the current interactive lifecycle through the running auth
service before it rotates the row, so copied bearers cannot cross an authorization update. If that
retirement cannot be confirmed, the row is left unchanged and the command fails with the recovery
action. revoke uses the same retirement before deleting the row, which lets a later grant create a
real successor instead of colliding with a live predecessor. supervise is separate from spawn and admin: it only makes a signed-in
person eligible for the host-provided closed remote manager-service view; it does not grant
management of another owner or a general host profile. revoke denies the next exchange and
the next connect with no restart, and evicts the principal’s live connections. Managed-agent rows
(written by the spawn path) live in a disjoint row space this command never touches. See
identity & auth.
doctor
Section titled “doctor”cotal doctor auth [--fix]Credential-health diagnosis and repair for this folder’s mesh: renders every managed
credential as healthy / near-expiry / expired and ends in healthy or the exact next
command; --fix applies the repairs it can. The one surface every stale-credential error
points at. --fix takes the mesh’s renewal lease when the broker answers and refuses while
a manager or another doctor holds it; with no broker it repairs offline and says so.
cotal join --space <s> --name <n> [--role <r>] [--channel <c>]cotal join --link <url> | --token <t>| Flag | Default | Meaning |
|---|---|---|
--space <s> / --server <url> / --creds <path> |
resolved mesh | Which mesh, and which credential |
--name <n> |
none | Your presence name |
--role <r> |
none | Your role |
--channel <c> |
none | Channel to join |
--kind <k> |
agent |
Endpoint kind |
--link <url> |
none | Join link (cotal://…) |
--token <t> |
none | Join token |
--lifecycle-uid <uid> |
none | Required with --creds: the lifecycle UID minted alongside the credential (COTAL_LIFECYCLE_UID works too). A credential’s durable grants name exact lifecycle-keyed resources, so join refuses to invent one |
--tls |
off | Connect over TLS |
An interactive presence: join a space under your own name and role, without launching an agent
harness. A --link or --token supplies the where and the auth in one value. See
Spaces and Identity and auth.
Manifest deploys
Section titled “Manifest deploys”A cotal.yaml manifest declares a whole mesh (channels, personas, roles, and ACLs) in one file.
Three commands consume it, plus a read-only validator:
cotal up -f cotal.yaml # boot a fresh mesh from the manifestcotal spawn -f cotal.yaml # deploy the manifest additively onto a running meshcotal down -f cotal.yaml # tear that deploy down (or --run <id> for one run)cotal topology view -f cotal.yaml # validate + view the access graph, change nothingup -f and spawn -f differ in target: up -f brings up a new broker and applies the manifest;
spawn -f requires an already-reachable mesh and applies additively (ownership-scoped). On a
user-auth mesh, spawn -f deploys over your own login (the deployer view, gated on ledger scope
spawn): the manifest’s agents land under your owner, a manifest claiming another owner is
refused, and seeding new channels additionally needs scope admin. Both take
--dry-run to print the plan without mutating anything. topology validates the manifest and
renders its channel / role / ACL graph. See Define a team and the
manifest reference.
cotal ext # same as `list`cotal ext add <npm-package>cotal ext remove <name>cotal ext list [--json]cotal ext root # print just the install prefix (scriptable)cotal ext seed [--repair|--reset|--force]Operator-installed extensions: add installs an npm package into a cotal-owned prefix and records
every registry provider it contributes. Commands appear in help, completion, and dispatch; runtime
providers are lazy-loaded by commands such as supervise; local process providers participate in
status and selective down. remove and list manage them. The @cotal-ai/web dashboard is the
canonical command/process example. Installed packages and their location are described in
config. When a package needs an export its linked @cotal-ai/* peer does not have,
add rolls back and names which install is behind, as a later load of an installed one does.
Bare cotal ext lists the inventory, headed by the install prefix. That prefix is a cotal-owned npm
root kept separate from npm’s own global tree. These packages never show up in npm list -g,
cotal ext (or the Extensions section of cotal status) is the canonical inventory. cotal ext root
prints only the path, for scripts. The versions shown are the manifest pin recorded at add time.
ext list --json (or bare ext --json) prints one JSON object per installed extension per line:
pkg, version, spec (what ext add was given), seeded (true for an entry the built-in seed
installed) and provides (the kind:name refs the table shows). No header or footer reaches stdout,
and an empty prefix prints nothing and exits 0. The table is presentation and is not a stable parsing
target. The other ext subcommands refuse --json.
Removing an extension that owns a running local process is refused with the mesh root and its
cotal down <component> command; stop it first so uninstalling the package never strands a process
whose lifecycle provider is gone.
Built-in connectors are seeded extensions
Section titled “Built-in connectors are seeded extensions”The first-party agent connectors (claude, opencode, codex, hermes, jcode, pi) are not compiled into
the binary. They are seeded on first run through the same ext add path a third party uses, and
appear in cotal ext list like any other extension. So you can remove one you do not want
(cotal ext remove @cotal-ai/connector-hermes), and a deliberately-removed connector STAYS removed
across upgrades. cotal ext add <your-package> adds a third-party connector the same way. The web
dashboard (@cotal-ai/web, providing command:web) is the seventh built-in seeded on the same path.
cotal ext seed is the maintenance entry for that seeding (it runs automatically on the first real
command of each boot, so you rarely call it). Each seeded connector’s ✓ added line goes to stderr,
so the command that triggered the seed keeps stdout to itself:
| Flag | Meaning |
|---|---|
| (none) | Reconcile: seed any never-seeded built-in, refresh a seeded one whose version the binary bumped, leave a removed one removed. A no-op once current. |
--repair |
Recover after an interrupted seed or a lost authority (rebuilds the interrupted connector; restores the removed-vs-never-seeded record from its durable backup). |
--reset |
Discard the record and re-seed all seven built-ins (the six connectors plus the web dashboard). Resurrects any you removed. Rebuilds cleanly over corrupt seed state. |
--force |
Re-seed the built-ins even when the version stamp is current or a downgrade. |
When a newer cotal advances the operator-global seed store to its generation, it prints one
migration line naming the old and new generations, the exact CLI entry that wrote the store, the
commit timestamp, and seed/stamp.json. That writer and timestamp are kept in the stamp, so a later
older CLI refusal can say which executable wrote the generation it will not overwrite and when.
Legacy generation-only stamps remain readable; their refusal simply has no writer provenance to add.
An older cotal refuses a seed store written by a newer version. When it can verify a sufficient
cotal executable on PATH or at the installer’s ~/.local/bin/cotal location, the refusal names
that absolute path so a reduced service PATH does not select the older binary again. Otherwise it
keeps the generic newer-version instruction. --force rebuilds the store for the running older
version without discarding the ever-seeded authority. --reset still exists for corrupt state and
resurrects deliberately-removed connectors.
A source-checkout CLI (pnpm cotal, tsx bin/cotal.ts, node bin/cotal.ts, or a suite child of
those) refuses to write or garbage-collect that store. The refusal names the path, the generation
it declined, and COTAL_SKIP_CONNECTOR_SEED=1 as the way to run other commands from a checkout,
because pointing $XDG_CONFIG_HOME at a scratch dir alone does not lift it. With the skip set, a fresh
config gets no built-in connectors (cotal ext list shows none), so add one from the checkout with
cotal ext add <checkout>/extensions/connector-claude-code against that scratch $XDG_CONFIG_HOME. COTAL_HOME does not relocate this
store. An entry that cannot be proven as a released install is refused the same way. Isolated
release tests that must seed from a checkout-shaped bin/ set COTAL_ALLOW_CHECKOUT_SEED=1 after
pointing $XDG_CONFIG_HOME at a scratch dir; that override is documented here, not on the refusal
line. An opt-in write still records the checkout path in seed/stamp.json as writtenBy.
The default connector for a bare cotal spawn (no --agent) is the persona’s agent: pin if it
has one, else claude; set COTAL_DEFAULT_AGENT (e.g. opencode) to change the fallback. It is
a default, so a persona that pins its harness still wins over it. An --agent naming a removed
connector fails loud with the exact
cotal ext add to restore it. Set COTAL_SKIP_CONNECTOR_SEED=1 to turn off the automatic first-run
seed/refresh entirely (for a controlled or offline setup that manages connectors by hand); cotal ext seed still runs on request. cotal agent-bearer never takes the seed at all: it is exec’d by
spawned seats on every bearer refresh, so it neither reconciles nor is refused by the store’s
generation (see Plumbing).
completion
Section titled “completion”cotal completion <bash|zsh|fish|powershell> # print a stub to eval / sourcecotal completion install [shell] # install it persistentlyPrints or installs shell completion. Completion candidates come from each command’s declared flags and, where useful, live mesh state (spaces, personas, managed agents) resolved offline.
feedback
Section titled “feedback”cotal feedback "<summary>" [--type <t>] [--email <e>] [--details <text>]| Flag | Default | Meaning |
|---|---|---|
--type <t> |
none | bug | idea | friction | praise | other |
--details <text> |
none | Longer free-form details |
--severity <s> |
none | low | medium | high |
--area <a> |
none | The part of Cotal this concerns |
--email <e> |
git email | Contact email (required on the keyless public path) |
--name <n> |
none | Your name (optional) |
--url <url> |
keyed / public intake | Intake URL override |
--key <k> |
COTAL_FEEDBACK_KEY |
Feedback key |
Sends feedback to the Cotal developers. With a key (--key / COTAL_FEEDBACK_KEY) it routes to the
keyed beta intake; without one it goes to the public cotal.ai intake and requires a contact email
(--email / COTAL_FEEDBACK_EMAIL, else your git email). Run a self-hosted intake with
feedback-intake.
Operate durable workflow runs (cotal-lang programs) from the terminal.
cotal run start --file <program> [--timeout <dur>] [--local]cotal run resume <runId> [--local --file <program>]cotal run ps [--endpoint <ep>] [--json]cotal run journal <runId> [--endpoint <ep>] [--json]cotal run answer <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]cotal run amend <runId> <stepKey> [--value <json>] [--artifact <ref>] [--endpoint <ep>] [--local --by <who>]cotal run migrate <runId> --local --file <program> [--endpoint <ep>]start hands the program to the mesh’s manager, which validates it, mints the run id (the record
never takes a caller-supplied one), drives it in its own process, and answers with the id once the
run is recorded; a program that does not validate is refused with every problem listed. resume
asks the manager to take an existing run back and continue it from its step journal; the source is
the recorded program, so no --file is taken. Neither takes --endpoint: the manager records
its runs under its own endpoint, and naming another is refused. ps lists the run records and
journal renders one run’s durable records; both only inspect. An open pause prints its question.
A pause settled with an accepted answer prints its value as JSON plus the recorded answerer,
artifact when present, time, and answer id, then one amended line per later amendment, in the
order the store committed them, so the last is the current position.
Expired pauses and ordinary steps print no answer line.
--json on ps or journal prints each row the manager answers with (or --local reads) as one
JSON object per line. A ps row carries runId, endpoint, state, holder, epoch,
journalHigh, forkedFrom, startedAt and programHash (the values the program’s run()
reports; programHash is absent for a run with no recorded program), and revoked or
revocationUnreadable when the marker says so. A journal row is an activation or a step. A
step row carries its step key, the effect kind and its name, state, outcome, the recorded
status and errorCode once settled, and startedAt and endedAt in epoch milliseconds. An open
pause adds its asks, its deadlineAt, and for a checkpoint the onExpiry it was armed with; a
settled pause adds its answer and its amendments, as the text view prints them. A field the
journal does not record is absent: a checkpoint opened before onExpiry was recorded carries none.
The run header and errors go to stderr, so stdout carries only rows; an unreadable revocation marker
prints its reason there and still exits 1. The text view is presentation and is not a stable
parsing target. --json on any other verb is refused.
answer resolves an open
checkpoint through the manager, presenting as the holder that armed it; the manager records the
answerer from your credential, so no --by is taken there. A settled step refuses a second
answer. amend records a changed position on a settled checkpoint or ask: it files a new
answer beside the accepted one, naming it, and the journal lists it under the step. The pause stays
settled and the run keeps the answer it acted on. A step that is still open or settled without an
answer refuses an amend. A spawned seat may amend only an answer recorded under its own name. migrate runs the migrate check of an
edited program against a run’s journal, from this terminal under a read credential (--local
only; the manager serves no run-migrate command): it prints whether the migration is admissible,
every orphaned step with its verdict and code, and exits 0 on admissible and non-zero on not. It
writes nothing: the commit that would file the migration is not reachable yet, and the report
says so. --timeout sets the default
checkpoint timeout for a drive (default 1h). --local drives in this process instead, over one
connection per invocation under the run’s own credential minted from the project folder’s trust
material, and is the path on a bare broker with no manager or for a run with no recorded program
(cotal run resume <runId> --local --file <program>); answer --local and amend --local take
--by <who>. On a
user-auth mesh the host’s own manager refuses the family by name, and --local has no credential
there. A participant’s manager started with cotal supervise hosts a logged-in user’s runs through
its issuing host: the auth callout issues the user’s manager connection, and every run verb rides
the versioned rail under that issuance.
User-auth run start
records the path. The guide is workflows.
Server daemons
Section titled “Server daemons”Two long-lived infra roles ship with the CLI. They are not part of everyday operation; the delivery
daemon comes up automatically with cotal up --detach in auth mode.
cotal deliver --space <s> [--server <url>] [--creds <file>] [--root <dir>]cotal auth-service --space <s> --server <url> [--port <n>] [--exchange-public-port <n>] [--exchange-public-url <https://…>] [--exchange-trusted-proxy]cotal feedback-intake --keys <keys.json> [--port <n>] [--creds <file>]auth-service runs a user-auth space’s identity plane: the NATS auth callout, the
capability-gated local exchange and JWKS, and, when --exchange-public-port is set, the closed public
exchange/discovery face forwarded by an HTTPS reverse proxy. --exchange-public-url is the proxy URL
advertised to clients; --exchange-trusted-proxy opts into last-hop X-Forwarded-For attribution.
cotal up --user-auth starts and supervises the service for you, so you run it directly only to
recover one by hand.
deliver runs the server-side Plane-3 delivery daemon: the durable backstop and membership/ACL
authority. It is auth-mode-only and single-instance (--shard/--shards accept only N=1);
--dev-mint mints a scoped cred from the local signer for standalone dev. --creds can start a
daemon that already looks healthy, but production renewal is not that file alone: the manager and
the daemon must address one credential store. On a stock split host with two project roots, a
direct deliver is not an independent repair; keep the daemon under cotal up on the broker
host, or inject the same store into both processes (embedding).
Typed by hand on the workstation, deliver dials the broker recorded for --space in the mesh
registry (a mismatching --server is refused before any dial, and a record for a different
workspace root is refused outright); with no record for the space it falls back to the local mesh.
The daemon serves the workspace root that --root <dir> names, which must hold .cotal/, or else
the nearest .cotal/ above its working directory. With neither, it refuses at start and names the
directory it searched from, before it reads a credential or dials a broker.
See the delivery daemon. feedback-intake runs a self-hosted feedback server
(requires --keys and a scoped --creds), announcing submissions into a space channel; flags
include --host/--port, --store, --space/--channel, --max-bytes, and --rate-limit.
--port takes a decimal port from 1 to 65535 and is checked before the intake dials the broker.
Plumbing
Section titled “Plumbing”cotal __complete <words…> is the internal entry the shell-completion stubs call to emit candidates
for the current command line; you never run it directly. cotal agent-bearer is machine-facing
plumbing on user-auth meshes: spawned agents exec it to print a fresh short-lived bearer from their
spawn-time secret; you never run it directly either. Its local arm uses --dir to discover the
capability-gated loopback service. A remotely enrolled, already-granted agent instead receives
--exchange-url <https://base> in its launch argv: that arm sends {owner, actor, actorToken} to the
pinned public exchange with no local capability, follows no redirects, and refuses every non-HTTPS
URL because the actor token is the credential in the request body. Because a seat execs it on every
bearer refresh, it skips the connector-seed boot gate entirely: it reads one 0600 token file,
exchanges it and prints the bearer without consulting or writing the operator-global seed store, so
a newer store generation cannot refuse a live seat’s refresh. --manager-call asks for the
instance-bound manager-caller view; --manager-instance <id> selects an explicit live candidate.
That mode still prints only the raw token and does not update --health-file. A spawn runs it once as the agent auth
preflight. When it fails there without printing a sentence of its own, the refusal names the cause: the 30 second
timeout, the signal that killed it, or its exit code. (cotal start is a removed tombstone: it
errors and points you to cotal spawn --detach.)