Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Control Plane & Persistent Daemon

This chapter covers the machinery that turned Intendant’s controller from a single agent loop into a multi-session orchestration host: the single-writer control plane, the long-lived session supervisor, the task dispatcher, the file-watcher that powers rewind/redo, the headless daemon that an idle --web launch becomes, Agenda scheduling, and cost accounting.

For the system-wide picture and the EventBus that ties these together, start with Architecture.

Why a Single-Writer Control Plane

Intendant has three frontends (web dashboard, MCP server, control socket) and they can all be live at once against the same daemon. If each frontend wrote shared mutable state directly — “the dashboard sets autonomy to High, MCP sets it to Low” — the truth would depend on event ordering and which handler happened to run. Worse, some state (the active external-agent backend, Codex sandbox/model config) also has to persist to intendant.toml, and you do not want three frontends racing to rewrite the same file.

The fix is a hard rule, stated at the top of control_plane.rs:

Frontends remain display-only — they render state changes but never write to shared state from ControlMsg handlers.

Frontends emit intents (ControlMsg, defined in event.rs) onto the EventBus. Exactly one subscriber — the control plane — interprets the state-mutating ones and applies them. Everyone else (including the frontend that sent the intent) learns the result by observing the broadcast event the control plane emits afterward (AutonomyChanged, ExternalAgentChanged, CodexConfigChanged, …).

Stated honestly, the invariant covers the intent path: no frontend mutates shared state from a ControlMsg handler. It is not a literal single-writer over every piece of shared state — three documented paths write from their own tasks: approval side effects (apply_user_approval applies approve-all escalation and the first display-control grant to the shared autonomy state, identically from every approval surface), the MCP autonomy/display tools, and platform display activation. Those are deliberate carve-outs with one owner each; new shared-state writers belong in the control plane.

  Web ─┐  emit ControlMsg          ┌──────────────┐  write    ┌──────────────┐
  MCP ─┼───────────────▶ EventBus ─┤ Control Plane │──────────▶│ shared state │
 Sock ─┘                  (bus)    │ (sole writer) │           │  + intendant │
        ◀───── observe ────────────┤               │◀──────────│    .toml     │
         AutonomyChanged etc.      └──────────────┘            └──────────────┘

What the control plane owns

ControlPlaneState (in control_plane.rs) holds the shared, mutable runtime state:

FieldTypeNotes
autonomySharedAutonomyGlobal autonomy level + the user-display grant flag
external_agentArc<RwLock<Option<AgentBackend>>>Active backend: Codex / Claude Code / Kimi / Pi, or None (internal)
codex_configSharedCodexConfigRuntime Codex config, including ordinary/managed commands, sandbox, approval policy, model/reasoning/service tier, web/network/write access, and managed-context/archive modes
claude_configSharedClaudeConfigRuntime Claude Code config (model, permission mode, allowed tools)
kimi_configSharedKimiConfigRuntime Kimi config (command, model, thinking, permission mode, plan mode, swarm mode)
project_rootOption<PathBuf>When set, state changes also persist to intendant.toml

Pi is deliberately not another shared config field. Its global defaults live in [agent.pi]; new sessions seed from that TOML, while reattached sessions merge the current TOML base with their persisted launch overlay (the overlay wins). Live model/thinking changes use Pi RPC and persist in that overlay; command/tool changes require restart. This keeps a replaceable engine’s configuration below the platform boundary instead of mirroring it throughout the control plane.

It is spawned once during each controller startup path (startup/daemon.rs or startup/headless.rs). spawn takes the state and opens the bus’s lossless intent subscription internally:

#![allow(unused)]
fn main() {
let _control_plane_handle = control_plane::spawn(
    control_plane::ControlPlaneState {
        autonomy, external_agent, codex_config, claude_config, kimi_config,
        bus, project_root
    },
);
}

The “applies on the NEXT task” rule

A subtlety worth internalizing: external-agent launch settings latch at process spawn. Codex locks its sandbox / approval / model / tool configuration at thread/start; Claude Code likewise latches model, permission mode, and allowed tools when the CLI process is launched. Kimi latches its command at process start, but can apply model, thinking, permission, plan, and swarm changes live through its session profile. Pi latches command and active tools but applies model and thinking live through RPC. So when a frontend flips, say, the Codex sandbox mode or the Claude Code permission mode, the control plane updates the shared config and persists it, but an already-running external-agent thread keeps its old values for the rest of its life. The change takes effect on the next task — the daemon loop re-reads the shared config at the start of each task and, for changes that cannot be applied mid-session, tears down the persistent agent so the next launch picks them up. Each ControlMsg::SetCodex*, ControlMsg::SetClaude*, and ControlMsg::SetKimi* variant documents this in its doc comment. Pi’s global Settings save goes straight to TOML rather than adding a fourth family of daemon-wide backend-specific setters.

Thread actions apply immediately rather than next-task when the selected backend supports them: /new, /compact, /fast, /fork, /undo, /review, goals, and the backend-neutral subsets exposed by Claude Code, Kimi, and Pi. The wire variant retains its historical name CodexThreadAction, but the advertised operation catalog is universal. The control plane does not own the persistent agent, so it merely re-broadcasts these as CodexThreadActionRequested for the daemon-side watcher that does own it.

Session Supervisor

session_supervisor/ is the long-lived owner of every session launched from the control plane at runtime. Where the control plane owns settings, the supervisor owns sessions — their lifecycle, their per-session resources, and the graph of how they relate. mod.rs holds the supervisor/state types and shared helpers; the behavior is sliced into dispatch.rs (ControlMsg intake), exec.rs (the off-intake per-session executor), launch.rs (new/resume flows and the shared session spawner), sub_agents.rs (delegation), fork.rs (anchor forks), routing.rs (follow-up/steer/edit/stop), agent_config.rs (per-session agent config and identity), and registry.rs (lifecycle observation and registration).

The intake is two-tier. The supervisor drains its lossless intent lane strictly in order, but dispatch_control_msg runs only the fast, ordering-critical work inline: validate, mint/reserve session identity, and dedup repeat requests. Slow launch bodies — session create (including a worktree checkout), resume, restart, fork, dashboard delegation — execute on a per-session ordered executor (exec.rs) under a small global concurrency bound, so one session’s multi-second checkout no longer head-of-line-blocks another session’s approvals, steers, or interrupts. Commands for one session still run in arrival order: while a session’s launch body (or anything deferred behind it) is pending, later commands for that session defer onto its queue, and the identity minted at intake keeps the session addressable — and its peer-delegation id deduplicated — before the slow body completes.

It subscribes to the bus and handles the session-lifecycle ControlMsgs:

ControlMsgBehavior
CreateSessionExplicitly create a new managed session and submit its first task (the forward-compatible primitive for parallel sessions). A task of exactly /fast is special-cased into a new idle Codex session with the fast service tier enabled.
StartTask { session_id: None }Start a new managed session (legacy clients). A task of exactly /fast follows the same idle fast Codex bootstrap path as CreateSession.
StartTask { session_id: Some(id) }Route the text as a follow-up turn into the named session
ResumeSessionRe-attach an existing session by source (intendant/codex/claude-code/kimi/pi) + id, optionally with a prompt
FollowUpRoute text to a session’s follow-up channel (target id, or the active session)
EditUserMessageRewind a session to a user turn and submit replacement text (rollback-capable backends only)
Interrupt / SteerMid-turn control of a running session. If a steer body is a supported Codex slash command, the supervisor converts it to a Codex thread action instead of injecting it as model text.
Approve/Deny/Skip/ApproveAllResolve a pending approval against the right session’s ApprovalRegistry
RenameSessionRename via the cross-backend naming abstraction (see Session Logging)

Per-session state

Internally the supervisor keeps a SupervisorState:

SupervisorState {
    sessions:         HashMap<String, ManagedSession>,   // canonical id → session
    session_aliases:  HashMap<String, String>,           // alias id → canonical id
    related_sessions: HashMap<String, RelatedSession>,   // child id → {parent, relationship}
    active_session_id: Option<String>,                   // the "current" session for un-targeted commands
    next_session_instance: u64,                          // rejects stale task completions
    restart_dedupe / external_attach_dedupe: …,
    known_external_sessions / advertised_thread_actions: …,
    delegation_receipts: …,                             // at-least-once peer-task dedup
    unmanaged_user_halts: …,                            // stop-vs-auto-attach race guard
}

Each ManagedSession carries its session_id, source (intendant or the external backend’s short name), optional display name, phase, project_root, session_dir, a follow_up_tx channel, its own ApprovalRegistry, an instance id and finish receiver for lifecycle races, delegation depth, and (for native sessions) the child registry shared with wait_sub_agents.

When start_new_session runs, it:

  1. resolves a fresh session log directory (SessionLog::resolve_path(None)<state-root>/logs/<uuid>/, honoring INTENDANT_HOME) and opens the SessionLog;
  2. resolves the project root (per-session override or the daemon’s default) and loads the Project. On a projectless daemon (see below) there is no default: a CreateSession without an explicit project_root fails with SessionEnded { error_kind: "no_project" } — a structured class the dashboard turns into the project picker — instead of silently rooting the session at the daemon’s launch directory;
  3. writes session_meta.json and activates the shared active-session handle;
  4. resolves the backend (one-shot agent override → configured default → internal) and applies the runtime external-agent config;
  5. resolves any dashboard attachments (frames/uploads) into agent content;
  6. spawns the agent session loop and emits SessionStarted.

The session graph

The supervisor tracks relationships between sessions, not just a flat list. The live supervisor’s alias graph accepts side and subagent (apply_session_relationship folds those through apply_related_session). That map lets a follow-up addressed to such a child id resolve to its canonical parent session, and removing a parent prunes those children. Forks are independent sessions rather than supervisor aliases: their fork edge is emitted and persisted for the log/catalog/lineage views, alongside specialized edges such as anchor-fork, rewind-restore/rewind-backout, and fission-branch/fission-detached/fission-imported. This lets the dashboard render related work without redirecting a fork’s messages to its parent. Identity and relationship updates arrive over the bus as SessionIdentity and SessionRelationship events and are also persisted to session logs.

active_session_id is the fallback target: an un-targeted FollowUp or Interrupt resolves to it, which is how single-session frontends keep working unchanged while multi-session clients address sessions explicitly by id.

Task Dispatch

task_dispatch.rs decides which channel a task goes to (it was pulled out of the TUI frontend back when that existed, as part of making frontends display-only). The Dispatcher owns up to three senders — presence_tx, task_tx, follow_up_tx — and routes a StartTask/FollowUp like this:

  1. If the task is not direct and presence_tx exists → send the text to the presence layer, which decides whether to forward it as a real task (via its own submit_task tool) or answer in-line.
  2. Else if task_tx exists → wrap in a TaskEnvelope and send (preserving attachments / frame refs / display target). direct (and legacy orchestrate == Some(false)) forces this path.
  3. Else if follow_up_tx exists → send a follow-up message (metadata dropped; non-presence mode has no CU-first routing anyway).
  4. Else → warn and drop.

A task carrying metadata (attachments, reference frames, a display target) is always forced onto task_tx even when non-direct, because presence’s text channel cannot carry that metadata.

The dispatcher and the session supervisor coexist: the dispatcher serves the legacy single-session loop and routes into channels it already owns, while CreateSession and targeted multi-session commands are left to the supervisor. A targeted command for a session the dispatcher does not own is simply ignored by it, so the supervisor picks it up.

File Watcher, Snapshots, and Rewind

file_watcher.rs is a live filesystem watcher rooted at the project directory (a projectless daemon starts no watcher — there is nothing to watch). It works for all agent types — internal, Codex, Claude Code, Kimi, Pi — because it watches the filesystem directly rather than diffing git, so an external CLI’s edits show up the same as Intendant’s own. It does two jobs:

  1. Live change events. It emits AppEvent::FileChanged { Created / Modified / Deleted } so the dashboard’s activity view can show per-file diffs as the agent works.
  2. Per-round content-addressed history for rewind / redo / branching.

On every AppEvent::RoundComplete that belongs to its root, the watcher records a HistoryRound capturing the full restorable state — supported, non-ignored text files no larger than 1 MB — as a path → sha256 map, plus the subset of paths that changed. A broader display/count mirror can include inspected files that are not restorable. Rounds are routed by the event’s project_root: a round emitted by a session working a different root (a worktree sub-agent, an external session supervised elsewhere) is skipped, while a round with no resolvable root fails open and records as before. Content blobs are stored once in a content-addressed objects/ directory, so identical content across rounds costs no extra disk. Each round also records turn_count and native_message_count, which conversation rollback uses to truncate the native conversation correctly.

The snapshot store lives inside the session log dir:

<state-root>/logs/<uuid>/file_snapshots/
├── baseline/            # initial text-file snapshot at session start
├── baseline_manifest.json # baseline metadata/fingerprints
├── objects/             # content-addressed blobs (sha256-named)
├── rounds/round_<id>/
│   └── manifest.json    # full round maps or a maps_from_round backreference
├── history.json         # slim format-2 index, heads/rounds/branches
└── store.lock           # advisory snapshot-store lock

The public API on FileWatcher:

  • rollback(target_round_id) — restores every tracked path to that round’s recorded state, moves current_head_id back without truncating history (so redo stays available), and emits AppEvent::RolledBack { from_id, to_id, files_reverted }. A new action after a rollback branches off the abandoned path and stores it in abandoned_branches for later pruning.
  • redo() — moves current_head_id forward along the linear path, restoring file state, emitting AppEvent::Redone { to_id }.
  • prune_abandoned() — drops abandoned branches and garbage-collects orphaned blobs, emitting AppEvent::HistoryPruned { branches_removed, bytes_freed }.

A soft byte cap bounds total snapshot size; once exceeded, pruning kicks in. The dashboard exposes this as the rewind/redo UI; on session resume or controller restart, history.json is reloaded so history survives the restart.

Headless Daemon Mode (idle --web)

A bare --web launch with no task and no --task-file is special. should_start_idle_web_daemon returns true when --web is set, it is not an MCP-stdio run, no --task-file was supplied, and no inline task was supplied:

#![allow(unused)]
fn main() {
fn should_start_idle_web_daemon(use_web: bool, flags: &CliFlags) -> bool {
    use_web
        && !flags.mcp
        && flags.task_file.is_none()
        && flags.task.as_ref().map(|t| t.trim().is_empty()).unwrap_or(true)
}
}

When that holds, the controller runs the daemon arm (startup/daemon.rs), which constructs and runs the session supervisor:

run_daemon_loop(DaemonConfig { bus, project_root, autonomy,
                               shared_external_agent, shared_codex_config,
                               frame_registry, web_port, flags_direct,
                               shared_session })
    └─▶ SessionSupervisor::new(..).run()   // owns every launch

The daemon then sits idle waiting for CreateSession / StartTask / ResumeSession / follow-up intents arriving over the dashboard WebSocket (or the control socket). This is the persistent-daemon mode: one long-lived controller, many sessions over its lifetime, driven entirely from frontends. Passing a task on the command line instead runs it as the foreground session under the same gateway (the headless arm), which falls through to this daemon loop when the session ends.

Projectless daemons. The daemon path’s project_root is an Option<PathBuf>. When the daemon starts in a directory with no project marker — no .git (directory or worktree file) and no intendant.toml; the single definition lives in project::root_has_project_marker, which the boot-scan gate (file_watcher::root_is_snapshot_worthy) also uses — it runs projectless instead of adopting cwd as an implicit project. Concretely: no file-watcher baseline (nothing is watched at all), no cwd-derived sandbox write scope (scratch + session logs + Unix toolchain caches + explicit absolute grants only; never the daemon state root wholesale), no project-root .env layer at startup, GET /api/project-root returns project_root: null (the dashboard’s New Session pane requires a project before submitting), and there is no default session project — each CreateSession/resume carries its own project_root, and a create without one gets the structured no_project failure. This is the normal shape for installed-app and service deployments, whose launch cwd ($HOME, /Applications, …) is an accident; CLI invocations keep cwd-as-project, which is correct for intendant "task" inside a repo.

Ctrl+C in this mode is handled by the global signal handler installed in main (it marks the session interrupted and exits 130); the daemon loop deliberately does not also listen for it, to avoid racing two handlers.

Controller stdout/stderr tee. daemon_log_tee::install tees the controller process’s stderr and stdout into its bootstrap session’s <state-root>/logs/<uuid>/daemon.log (with per-line timestamps) while still mirroring to the original terminal. This is Unix-only; on Windows install is a no-op. It is what lets the dashboard’s “Download session report” bundle carry controller-side eprintln!/panic/tracing output alongside session.jsonl.

Cost Accounting

app_state_pricing.rs provides server-side per-model USD cost estimation, mirroring the pricing table in presence-web/app_state.rs so the native daemon and the browser agree. Each entry gives per-token prices for input, cache-write, cached, and output tokens; estimate_session_cost(...) combines those with a session’s token usage, and estimate_live_usage_cost(...) covers live-audio usage. The dashboard surfaces these as the running session cost.

Agenda Reminders and One-Shot Scheduled Sessions

One AgendaHandle lives under the daemon state root (<state-root>/agenda/), not under a project. Its append-only JSONL operation log is the durable ledger for parked notes, tasks, questions, due reminders, and scheduled-session effects. Frontends and agents submit commands; the daemon serializes, validates, appends, folds, and broadcasts the resulting state.

Due agenda items feed an owner-controlled reminder policy. The scheduler honors enablement, quiet hours, per-item urgency ceilings, and a staleness window; fresh reminders travel the existing UserNotification ladder, while older missed reminders collapse into a digest. Reminder delivery does not execute the item’s body as instructions.

Scheduled sessions are deliberately narrower than cron:

  • An effect manifest binds a goal and one fire_at_ms instant (plus execution shape intent). Any agenda writer may propose or revise it, but only a dashboard or local-process owner surface may approve/revoke the exact manifest digest; editing it invalidates the approval.
  • At the due instant the scheduler fsyncs a prepared occurrence, then dispatches a normal supervised session through StartTask. The spawned session still has its ordinary agent-session principal, sandbox, autonomy, and approval gates — schedule approval is not permission to bypass them.
  • Occurrence identity and the delegation-receipt ledger make launch idempotent. A session missed while the daemon was down is marked missed; one prepared without launch confirmation or one that crossed a daemon restart resolves unknown. These states are terminal and never auto-retry; the owner must revise/re-approve to schedule another occurrence.

There is still no recurring cadence or cron vocabulary. The separate one-shot ScheduleControllerRestart (event.rs / mcp/) carries a goal and handoff across a controller restart; it is a continuity mechanism, not an Agenda session occurrence.

Graceful Daemon Handover (drain + takeover)

Co-homed daemons coordinate through the active-scheduler lease (scheduler-lease/holder.lock, an advisory file lock; see the agenda chapter for the firing rules). A running holder can hand its role to a successor without kill-and-relaunch:

  • intendant --takeover boots a successor that asks the current holder to drain (POST /api/daemon/takeover — owner-grade, loopback admission-token trust class; intendant ctl takeover is the standalone verb). Drain is one-way and idempotent.
  • The draining daemon stops all standing automations (its scheduler performs the ordered entry between passes, so a firing pass never straddles the release), flips lease.json to draining, frees the flock, and keeps serving in-flight work only: follow-ups, steers, interrupts, approvals, and ordinary agenda writes all serve; session-creating intents (create, untargeted start_task, the resume family) refuse with a structured daemon_draining error carrying the successor’s port, and the agenda immediacy verbs (start_now, request_occurrence) refuse the same way. The classification lives in one place (ControlMsg::creates_session) and every surface consults it.
  • When its last work-holding session finishes (sessions parked after done do not hold; parked conversations resume on the successor), the drainer records an exited presence state and the process exits. If the successor dies while the drainer still drains, ONE loud notification says standing automations are paused until someone relaunches or takes over — the drainer never reclaims the lease.

intendant ctl status shows the whole story under scheduler_lease (role, drain state, and every co-homed boot with probed liveness).

The update-available surface

The daemon watches its own binary image on disk: a boot-time identity stamp (length + mtime, plus dev/inode on Unix) beside the compiled-in build provenance, a 60 s stat poll, and — when the image changes — one bounded, environment-scrubbed <binary> --version probe that reads the NEW build’s provenance. The daemon never execs a successor daemon; takeover stays an explicit gesture. What the watch produces:

  • An update block on the handover status payload (GET /api/daemon/handover, its api_daemon_handover tunnel twin, and ctl status under scheduler_lease): the running and on-disk builds side by side (git_sha, built_at, versions), or an honest probe_error when the changed image’s provenance is unreadable.
  • One info-urgency notification per distinct on-disk commit sha (in-memory dedup; a restarted daemon states the fact once more).
  • The dashboard’s update chip (bottom corner, suppressed while a drain banner is up). Inside the packaged macOS app the chip offers Update now: the app’s BackendSupervisor spawns the successor from the on-disk binary on a fresh port, waits for readiness, re-points the webview, and only then asks the predecessor to drain — the old child is never killed; it finishes its in-flight sessions and exits on its own. On a CLI-launched daemon the chip is honest about its reach: it can offer Hand off to :PORT when a live co-homed daemon is already running (draining this daemon toward it), and when none is running it says it cannot launch one itself.
  • On macOS, a non-Developer-ID on-disk build carries the keychain/TCC honesty line: item ACLs and TCC grants key on the signing identity, so the new build’s first custody or capture access may re-prompt.

The SPA also learns the daemon’s boot_id from every config lane and the handover poll; a tab that hears a new daemon process answering on its origin offers “The daemon updated — Reload” (auto-reloading only when hidden and composer-safe), so a stale tab can no longer misrepresent a replaced daemon.

Producing the update: the self-update lane

Everything above assumes a newer binary already ON disk. The self-update lane produces one, on the owner’s explicit click, from the Daemon update panel in Access → Daemons (beside the daemons list’s provenance rows). It classifies the install once at boot from the watched binary path:

  • Source install — the binary is a checkout’s target/release output, or the app bundle carries the source-checkout stamp scripts/bundle-macos.sh writes (a stamp whose recorded path no longer exists, e.g. a release app built on CI, falls down to the consumer lane). The panel runs a bounded behind-origin/main compare (a timeout-bounded fetch, rev-list --count capped at 500, and a dirtiness probe) at boot, every 12 h, and on Check now. The click then runs git pull --ff-only plus cargo build --release (plain binary) or scripts/bundle-macos.sh (app shape — builds, signs with the stable local identity, installs to /Applications) as supervised child processes. The build rides the machine’s rustc governor (the child env never sets RUSTC_WRAPPER, so the box-wide cargo-config wrapper stays engaged) and is headroom-gated: under memory pressure (macOS kern.memorystatus_vm_pressure_level > 1, Linux MemAvailable under a 3 GiB floor) the job refuses to start instead of joining an OOM spiral. A dirty checkout or a non-fast-forwardable branch refuses honestly — the lane never stashes, resets, or merges over local work.
  • Consumer install — an installed release app with no reachable checkout. The automatic cadence only runs when Connect is configured (the tripwire’s posture — an unprompted check reaches the rendezvous and GitHub); Check now always may. The check verifies the latest logged release through the transparency-log ritual (hosted_verify: inclusion proof, signed tree head, append-only pin, the compiled-in PGP identity and signature-coverage checks) and compares versions. The click then downloads the platform’s app zip and its detached .asc, verifies sha256 against the log’s committed digests, verifies the signature with gpg --verify in a throwaway GNUPGHOME that trusts only the compiled-in release signing key (gpg absent = the lane refuses; it never installs unverified bytes), probes the new app’s provenance, and swaps it in beside the running one (the old bundle stays as Intendant.app.previous). Fail closed everywhere: any verify failure deletes the staging bytes and reports the reason.

Both lanes end the same way: a newer binary sits at the watched path, the update watch above announces it, and the existing chip/one-click swap lane performs the actual handover — produce and swap stay two honest phases, and the daemon never execs a build, a fetch ritual, or a successor into its own process. Progress and failure render live on the panel (phase, a bounded child-process log tail, and the outcome) served inside the update_lane block of GET /api/daemon/handover; the two actions (POST /api/daemon/update-lane/{check,produce}) are owner-grade and deliberately HTTP-only — remote tunnel-primary surfaces watch progress but cannot click a build onto the box. Checking is automatic and read-only; producing only ever happens on the click — there is no auto-update.

Where to Go Next

  • Architecture — the EventBus, the execution shapes, and the corrected frontend-parity model.
  • Session Logging — the on-disk layout these components read and write, and the cross-backend session naming the supervisor uses.
  • Web Dashboard — the primary consumer of these events (activity diffs, rewind UI, session graph, cost).
  • Multi-Agent Orchestration — orchestration sessions and supervised sub-agents the supervisor launches.