Orchestrator: one session talks to the owner, fresh workers do the tasks
Why: cost is dominated by cache reads, which grow with context size. A fresh worker per task costs about as much less as a cheaper model does; researching in one context and implementing in another is NOT cheaper (rediscovery, handbacks). So: the orchestrator slices and rules, and one fresh worker per task researches and implements.
Sessions
- Orchestrator (interactive, opus,
claude --autocompact 120k): owner talk about runs, Awaiting rulings, slicing, spawning workers, post-checks, unattended batches. Readswfoutput and worker reports, never code. Holds no lane: neverwf next --as(usewf lanes,wf list --runner --lane <lane>). - Design session (optional, big window): brainstorming, specs, hard rulings. Ends each topic with a committed spec/task; the orchestrator picks it up from TASKS.md.
- Workers: the
wf-workersubagent (shared/agents/wf-worker.md; install:ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md), one task each.
Starting a session
Skills (install once: ln -s /projects/public/workflow/shared/skills/wf-orchestrate ~/.claude/skills/ and the same for
wf-design and wf-pilot): after /clear, type /wf-orchestrate [go | batch N], /wf-pilot [--batch 4] [--lanes a,b] [--for 8h] or /wf-design [topic | id] (not
"continue": that is the normal worker cold start). Agents and skills installed while a session runs appear
only in sessions started afterwards: restart it.
Runner-ready tasks (slicing rule)
Model:line set to the cheapest lane that fits (sonnet: exact Done + a pattern to copy; haiku: mechanical).- Done checkable headless: a file, test result or artifact, plus
wf addfor follow-ups. Never "report to owner" / "owner confirms". After:satisfied, noSessions: owner.Sessions: solo→ run it alone.
One worker
Two commands do steps 1-3 and 5-7 (one call each): wf orch pick <lane> [--id ID] [--recovery WHY] (stop file,
pick, claim, free worktree, recorded in .wf/orch/<id>.json, prints the spawn line and the prompt below) and
wf orch post <id> <lane> --result "<report result line>" [--commit SHA] [--agent A] [--duration S] [--no-pick]
(post-check, merge if needed, leftover commit, orch + cost log lines, then alert: … for a parked task, and the next pick or stop lane <lane>: <why> (none, stop file, push-failed, wip, no report)).
The steps they automate:
- Pick:
wf list --runner --lane <lane>→ first id (lanes:wf lanes). Claim:wf status <id> progress "worker". No commit of its own:wf mergeor any later bookkeeping commit takes TASKS.md (and the claim) along; harmless. - Worktree:
<main>/.worktrees/<lane>(split project: under the code repo,<code_root top>/.worktrees/<lane>); if a live session or another worker uses it,<lane>-2,-3, … (live interactive sessions:~/.claude/sessions/*.jsonfieldcwd). Branch<lane>/<id>. - Spawn: Agent tool,
subagent_type: wf-worker,model: <task Model, none = opus>, background, no worktree isolation (it blockswf mergeon the main checkout). Prompt:
Task: <id> Lane: <lane> Model: <model>
Main tree: <main> Worktree: <path> Branch: <lane>/<id>
Final message: the 4 report lines only.
(the last line is needed: the agent rule alone did not stop prose before the report.)
Slice job (row starts slice: / wf show effort > slice_above): same spawn; the worker only slices.
(no wf ctx output in the prompt: the worker fetches it; the prompt stays in your context forever.)
4. At most one worker per lane at a time; lanes in parallel. wf writes and wf merge take the project
lock (.wf/lock), so parallel workers are safe.
5. On completion read only the 4-line report (followups = ids the worker added). Post-check:
- done: wf log -n 3 shows the id; git -C <main> branch --list <lane>/<id> empty; wf check 0 errors;
git -C <path> status --porcelain empty; git -C <path> merge-base --is-ancestor HEAD <master> exit 0
(worktree HEAD in master: no impl commit left on a detached HEAD; else cd <path> && wf merge).
- Everything else: commit leftover bookkeeping
git -C <main> commit -m "<id> <outcome> (orchestrator)" -- TASKS.md <archive> if they changed.
6. Outcome:
| Report / state | Action |
|---|---|
| done + post-check ok | next task in the lane |
| model raised by the worker | continue; next pick spawns it on the new model |
| sliced | next task in the lane |
| done+gate-red <culprit> <fix-id> (async gate red, earlier task's break) | post-check as done; next pick in the lane = the P0 fix (no Done/Model → add them first); never stop the lane |
| awaiting / needs-owner / handback / post-check red | wf orch post parks only that task (needs-owner: After: its Needs-human h-task, pickable again after the owner's wf done; awaiting: blocked on its a-id, else Sessions: owner + note; handback clears the claim), prints alert: …; tell the owner, the lane keeps picking (never wait on the owner while tasks are pickable) |
| wip (after a wrap-up) | stop the lane; reslice or rerun later |
| crashed / killed (no report) | one fresh worker, prompt plus Recovery: <why> (it inspects git status, git log master..HEAD, git diff -- TASKS.md — often just the claim — and wf show <id>), then stop |
7. Log one line per task to out/wf-orch.log (git-ignored): time, lane, model, id, outcome, commit, duration.
Cost: wf usage --agent <agent id> --log <id> <outcome> --effort <task estimate: <1h|1h|5h|10h|100h> --lane <lane> appends one key=value line to
out/wf-cost.log (tokens + API-price $ from the subagent transcript). --duration <s> adds dur=; wf usage --report (med_dur) = per lane/model and
effort: n, done, median/total $, $ per done task, median turns, over every project's log.
The completion notice's token count is NOT the cost (one worker: notice 67k, transcript 5.0M cache reads);
real usage is in the subagent transcript ~/.claude/projects/<project>/<session>/subagents/agent-*.jsonl
(wf usage per session/agent; one request is several transcript entries: counted once;
subagent transcripts lack final output counts → out estimated from content, shown ~, log est=N).
Cloud lane
Virtual lane cloud (projects with cloud = true; design: cloud-lane spec §4.6): tasks run as
claude --cloud sessions billed to the cloud budget, not as local workers.
- wf orch pick cloud [--id ID] = stop file → ledger (wf cloud ledger) → first fitting task across all lanes
(runner-ready, not in progress, wf cloud fit rules; opus only, Cloud: yes first, then prio, then larger effort) → wf cloud send
(claim in progress: cloud:<sid>, record .wf/orch/<id>.json lane cloud). Nothing to spawn: the session runs
remotely. Otherwise stop lane cloud: ledger (…) (balance < reserve, or send refused) | none fit (…) |
max parallel (…). A stop of the cloud lane never stops the local lanes.
- Local lanes skip tasks claimed cloud: (every in progress task is skipped by wf orch pick <lane>).
- On each wake-up, at most every 10 min: wf cloud pull --all (teleport ~30 s, no tokens). Per ended task it prints
<id>: <state> (<sid>, $… ) (state done | awaiting | handback | lost; done also report: commit <sha>) →
wf orch post <id> cloud --result <state> [--commit <sha>] = the worker-report path (done: archive line, branch
<task lane>/<id> gone, wf check; else leftover commit of the note/claim pull left), log line, next cloud
pick. running (exit 4) → nothing; pull again next wake. awaiting / handback / lost → post parks only that
task (awaiting: blocked on its a-id; handback / lost: Cloud: no, so the cloud lane never re-sends it; it is back
in its local lane with the note Recovery: cloud attempt …), prints alert: …; lane cloud keeps picking, tell the
owner, next cloud pick.
- Fill: one wf orch pick cloud per free slot (up to max_parallel), each in its own call.
- Batch (wf batch on a cloud = true project) pulls itself: the job runs wf_res.py batch-sidecar around
claude -p; every 10 min while it lives (until out/wf-batch.stop) it runs wf cloud pull --all and per ended
task wf orch post <id> cloud --result <state> [--commit <sha>] --no-pick + a cloud <id> <state> <sha> (sidecar pull)
line in the summary, so pulls never starve while the batch waits on foreground workers. The batch orchestrator
only picks cloud, never pulls. Two pulls of one id at once: the second prints <id>: pulled by another wf cloud
pull, skipped (rc 0; flock on .wf/cloud/<id>.json).
Tail: the orchestrator exited with .wf/cloud/*.json records left (sent at batch end) → the sidecar keeps
pulling every 10 min (no picks, no sends) until no record is left, the stop file appears or 24 h passed; the job's
ledger reservation shrinks to 0.2 GB / 1 cpu for it (ledger only, the unit's MemoryMax stays); the summary ends
with cloud tail ended: <why>; <n> pulls; left: <ids|none> (sidecar). A next batch's sidecar may overlap it (record
flock).
Unattended batch (nightly, long runs)
The orchestrator does not babysit long runs; a headless batch orchestrator (separate claude -p process,
fresh context per batch) does. The owner starts it (auto mode denies an agent launching another unattended agent): ! + the command, or
once an allow rule Bash(python3 /projects/public/workflow/wf.py batch:*) in ~/.claude/settings.json.
wf batch N [--lanes sonnet,haiku] [--model opus] [--for 6h] [--mem 4G] # --dry-run prints the command; default mem/for: wf res hist
wf batch --status # newest job + out/wf-batch-*.md
wf batch K --prep [--lanes fast] # prep batch (below)
= wf res run --mem 4G --for 6h --title wf-batch -- claude -p --model opus --permission-mode auto
--permission-prompts none "<templates/batch-prompt.md>" with CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS = --for.
The prompt: follow 'One worker' for up to N tasks, never implement; each round one wf-worker per lane in ONE
message, foreground, post-check, one line per task to out/wf-batch-<stamp>.md; stop a lane on its first stop
outcome; never ask (follow-ups → wf add -p 2, decisions → wf add -s awaiting); end with the ids added.
Why foreground: claude -p ends when the orchestrator's turn ends and kills background workers after a
wait ceiling (default 10 min; the env above raises it as a backstop). Workers spawned in one message run in
parallel and the turn waits for all. Background subagents cannot spawn subagents, so the batch
orchestrator is never a subagent. The interactive orchestrator checks wf batch --status.
Headless batch = short fire-and-forget runs (a few tasks, no steering). Long runs: pilot below.
Prep batch (--prep): makes the backlog runner-ready, never implements. Targets = up to K Pending tasks without
Done, not Sessions: owner, no status, no open slices (pickable first, then priority); none → prints prep: …
nothing started, starts nothing. Prompt templates/prep-prompt.md: sonnet subagents (≤ 3 tasks each) read the
task text, wf set <id> --done "…" (+ --model if missing); unclear → wf add -s awaiting + wf status <id>
blocked <a-id>. Summary lines <id> → Done: … | <id> → awaiting <a-id>; then it commits the task file.
Pilot runs one when idle with not runner-ready tasks (once per idle spell).
Pilot (nested)
/wf-pilot [--batch 4] [--lanes a,b] [--for 8h], typed into an idle project session (phone via Remote Control),
keeps that session tiny: subagents cannot spawn subagents, so it pilots headless batch orchestrators instead of
workers. Loop: wf batch K --left <time to deadline> (a wf res job, returns at once) → background Bash wf res wait <rid> → its exit
re-invokes the pilot → wf batch --status → one line to its context and out/wf-orch.log → next batch. A batch =
up to K tasks, rounds of one worker per lane in parallel (2 lanes, K=4 → two rounds). Nothing pickable →
prep batch if tasks are not runner-ready (once per idle spell), then background wf lanes --wait 1800, until --for. Deadline fit from history: task p90 = p90 of done/handback durations in out/wf-orch.log (local lanes, > 0; < 3 runs → 30 min); --left launches K' = min(K, floor(left / p90)) tasks with --for = left (K' = 0 → fit: 0 …; nothing started, the pilot stops); the batch prompt runs wf batch --time-left <deadline> before each round and spawns nothing more on : stop (left < p90). Steering: the owner talks to the pilot only; stop →
wf batch --stop (stop file, checked before every spawn: latency ≤ the running tasks), skip → wf move <id> deferred. Batch rc≠0 or no summary →
retry once, then stop. PushNotification on new awaiting items, on a stop and at run end; each alert also an ALERT <text> line in out/wf-orch.log and in the summary (push may be off). Phone push needs agentPushNotifEnabled (settings) / app notifications on. Needs the allow rule
Bash(python3 /projects/public/workflow/wf.py batch:*).
Context
Orchestrator large (after design talk) → finish the topic, commit, ask the owner to /clear. State lives in
TASKS.md, the log and the batch summaries.
Safe to clear only with no background worker running. Tested (2026-10-05): /clear leaves a background
subagent alive and its hand-back + completion notice reach the new context, but SIGKILLs its running Bash
command (exit 137): a worker mid-test or mid-commit loses that step and reports a false failure or half state.
Split projects
code_root = another git repo: worktrees and branches live in the code repo, wf start makes them (with .wf-home
pointing at the private project). wf orch post checks the code repo. A done worker whose push (push key) failed
leaves .wf/push-failed: the lane stops (push-failed); fix the push, delete the marker, continue.