1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
|
# Orchestrator: one session talks to the owner, fresh workers do the tasks
Why: cost is dominated by cache reads, which grow with context size. A fresh worker per task costs about
as much less as a cheaper model does; researching in one context and implementing in another is NOT cheaper
(rediscovery, handbacks). So: the orchestrator slices and rules, and one fresh worker per task researches
and implements.
## Sessions
- **Orchestrator** (interactive, opus, `claude --autocompact 120k`): owner talk about runs, Awaiting rulings,
slicing, spawning workers, post-checks, unattended batches. Reads `wf` output and worker reports, never
code. Holds no lane: never `wf next --as` (use `wf lanes`, `wf list --runner --lane <lane>`).
- **Design session** (optional, big window): brainstorming, specs, hard rulings. Ends each topic with a
committed spec/task; the orchestrator picks it up from TASKS.md.
- **Workers**: the `wf-worker` subagent (`shared/agents/wf-worker.md`; install:
`ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md`), one task each.
## Starting a session
Skills (install once: `ln -s /projects/public/workflow/shared/skills/wf-orchestrate ~/.claude/skills/` and the same for
`wf-design` and `wf-pilot`): after `/clear`, type `/wf-orchestrate [go | batch N]`, `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]` or `/wf-design [topic | id]` (not
"continue": that is the normal worker cold start). Agents and skills installed while a session runs appear
only in sessions started afterwards: restart it.
## Runner-ready tasks (slicing rule)
- `Model:` line set to the cheapest lane that fits (sonnet: exact Done + a pattern to copy; haiku: mechanical).
- Done checkable headless: a file, test result or artifact, plus `wf add` for follow-ups. Never "report to
owner" / "owner confirms".
- `After:` satisfied, no `Sessions: owner`. `Sessions: solo` → run it alone.
## One worker
Two commands do steps 1-3 and 5-7 (one call each): `wf orch pick <lane> [--id ID] [--recovery WHY]` (stop file,
pick, claim, free worktree, recorded in `.wf/orch/<id>.json`, prints the spawn line and the prompt below) and
`wf orch post <id> <lane> --result "<report result line>" [--commit SHA] [--agent A] [--duration S] [--no-pick]`
(post-check, merge if needed, leftover commit, orch + cost log lines, then `alert: …` for a parked task, and the next pick or `stop lane <lane>: <why>` (none, stop file, push-failed, wip, no report)).
The steps they automate:
1. Pick: `wf list --runner --lane <lane>` → first id (lanes: `wf lanes`). Claim: `wf status <id> progress "worker"`. No commit of
its own: `wf merge` or any later bookkeeping commit takes TASKS.md (and the claim) along; harmless.
2. Worktree: `<main>/.worktrees/<lane>` (split project: under the code repo, `<code_root top>/.worktrees/<lane>`); if a live session or another worker uses it, `<lane>-2`, `-3`, …
(live interactive sessions: `~/.claude/sessions/*.json` field `cwd`). Branch `<lane>/<id>`.
3. Spawn: Agent tool, `subagent_type: wf-worker`, `model: <task Model, none = opus>`, background, no worktree isolation
(it blocks `wf merge` on the main checkout). Prompt:
```
Task: <id> Lane: <lane> Model: <model>
Main tree: <main> Worktree: <path> Branch: <lane>/<id>
Final message: the 4 report lines only.
```
(the last line is needed: the agent rule alone did not stop prose before the report.)
Slice job (row starts `slice:` / `wf show` effort > slice_above): same spawn; the worker only slices.
(no `wf ctx` output in the prompt: the worker fetches it; the prompt stays in your context forever.)
4. At most one worker per lane at a time; lanes in parallel. `wf` writes and `wf merge` take the project
lock (`.wf/lock`), so parallel workers are safe.
5. On completion read only the 4-line report (`followups` = ids the worker added). Post-check:
- `done`: `wf log -n 3` shows the id; `git -C <main> branch --list <lane>/<id>` empty; `wf check` 0 errors;
`git -C <path> status --porcelain` empty; `git -C <path> merge-base --is-ancestor HEAD <master>` exit 0
(worktree HEAD in master: no impl commit left on a detached HEAD; else `cd <path> && wf merge`).
- Everything else: commit leftover bookkeeping
`git -C <main> commit -m "<id> <outcome> (orchestrator)" -- TASKS.md <archive>` if they changed.
6. Outcome:
| Report / state | Action |
|---|---|
| done + post-check ok | next task in the lane |
| model raised by the worker | continue; next pick spawns it on the new model |
| sliced | next task in the lane |
| done+gate-red <culprit> <fix-id> (async gate red, earlier task's break) | post-check as done; next pick in the lane = the P0 fix (no Done/Model → add them first); never stop the lane |
| awaiting / needs-owner / handback / post-check red | `wf orch post` parks only that task (needs-owner: `After:` its Needs-human h-task, pickable again after the owner's `wf done`; awaiting: blocked on its a-id, else `Sessions: owner` + note; handback clears the claim), prints `alert: …`; tell the owner, the lane keeps picking (never wait on the owner while tasks are pickable) |
| wip (after a wrap-up) | stop the lane; reslice or rerun later |
| crashed / killed (no report) | one fresh worker, prompt plus `Recovery: <why>` (it inspects git status, git log master..HEAD, `git diff -- TASKS.md` — often just the claim — and `wf show <id>`), then stop |
7. Log one line per task to `out/wf-orch.log` (git-ignored): time, lane, model, id, outcome, commit, duration.
Cost: `wf usage --agent <agent id> --log <id> <outcome> --effort <task estimate: <1h|1h|5h|10h|100h> --lane <lane>` appends one key=value line to
`out/wf-cost.log` (tokens + API-price $ from the subagent transcript). `--duration <s>` adds dur=; `wf usage --report` (med_dur) = per lane/model and
effort: n, done, median/total $, $ per done task, median turns, over every project's log.
The completion notice's token count is NOT the cost (one worker: notice 67k, transcript 5.0M cache reads);
real usage is in the subagent transcript `~/.claude/projects/<project>/<session>/subagents/agent-*.jsonl`
(`wf usage` per session/agent; one request is several transcript entries: counted once;
subagent transcripts lack final output counts → out estimated from content, shown `~`, log `est=N`).
## Cloud lane
Virtual lane `cloud` (projects with `cloud = true`; design: cloud-lane spec §4.6): tasks run as
`claude --cloud` sessions billed to the cloud budget, not as local workers.
- `wf orch pick cloud [--id ID]` = stop file → ledger (`wf cloud ledger`) → first fitting task across all lanes
(runner-ready, not in progress, `wf cloud` fit rules; opus only, `Cloud: yes` first, then prio, then larger effort) → `wf cloud send`
(claim `in progress: cloud:<sid>`, record `.wf/orch/<id>.json` lane cloud). Nothing to spawn: the session runs
remotely. Otherwise `stop lane cloud: ledger (…)` (balance < reserve, or send refused) | `none fit (…)` |
`max parallel (…)`. A stop of the cloud lane never stops the local lanes.
- Local lanes skip tasks claimed `cloud:` (every `in progress` task is skipped by `wf orch pick <lane>`).
- On each wake-up, at most every 10 min: `wf cloud pull --all` (teleport ~30 s, no tokens). Per ended task it prints
`<id>: <state> (<sid>, $… )` (state done | awaiting | handback | lost; done also `report: commit <sha>`) →
`wf orch post <id> cloud --result <state> [--commit <sha>]` = the worker-report path (done: archive line, branch
`<task lane>/<id>` gone, `wf check`; else leftover commit of the note/claim pull left), log line, next cloud
pick. `running` (exit 4) → nothing; pull again next wake. awaiting / handback / lost → `stop lane cloud`, tell the
owner; the task is back in its local lane (handback note `Recovery: cloud attempt …`).
- Fill: one `wf orch pick cloud` per free slot (up to `max_parallel`), each in its own call.
- Batch (`wf batch` on a `cloud = true` project) pulls itself: the job runs `wf_res.py batch-sidecar` around
`claude -p`; every 10 min while it lives (until `out/wf-batch.stop`) it runs `wf cloud pull --all` and per ended
task `wf orch post <id> cloud --result <state> [--commit <sha>] --no-pick` + a `cloud <id> <state> <sha> (sidecar pull)`
line in the summary, so pulls never starve while the batch waits on foreground workers. The batch orchestrator
only picks cloud, never pulls. Two pulls of one id at once: the second prints `<id>: pulled by another wf cloud
pull, skipped` (rc 0; flock on `.wf/cloud/<id>.json`).
Tail: the orchestrator exited with `.wf/cloud/*.json` records left (sent at batch end) → the sidecar keeps
pulling every 10 min (no picks, no sends) until no record is left, the stop file appears or 24 h passed; the job's
ledger reservation shrinks to 0.2 GB / 1 cpu for it (ledger only, the unit's MemoryMax stays); the summary ends
with `cloud tail ended: <why>; <n> pulls; left: <ids|none> (sidecar)`. A next batch's sidecar may overlap it (record
flock).
## Unattended batch (nightly, long runs)
The orchestrator does not babysit long runs; a headless batch orchestrator (separate `claude -p` process,
fresh context per batch) does. The owner starts it (auto mode denies an agent launching another unattended agent): `!` + the command, or
once an allow rule `Bash(python3 /projects/public/workflow/wf.py batch:*)` in `~/.claude/settings.json`.
```
wf batch N [--lanes sonnet,haiku] [--model opus] [--for 6h] [--mem 4G] # --dry-run prints the command; default mem/for: wf res hist
wf batch --status # newest job + out/wf-batch-*.md
wf batch K --prep [--lanes fast] # prep batch (below)
```
= `wf res run --mem 4G --for 6h --title wf-batch -- claude -p --model opus --permission-mode auto
--permission-prompts none "<templates/batch-prompt.md>"` with `CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS` = `--for`.
The prompt: follow 'One worker' for up to N tasks, never implement; each round one wf-worker per lane in ONE
message, foreground, post-check, one line per task to `out/wf-batch-<stamp>.md`; stop a lane on its first stop
outcome; never ask (follow-ups → `wf add -p 2`, decisions → `wf add -s awaiting`); end with the ids added.
Why foreground: `claude -p` ends when the orchestrator's turn ends and kills background workers after a
wait ceiling (default 10 min; the env above raises it as a backstop). Workers spawned in one message run in
parallel and the turn waits for all. Background subagents cannot spawn subagents, so the batch
orchestrator is never a subagent. The interactive orchestrator checks `wf batch --status`.
Headless batch = short fire-and-forget runs (a few tasks, no steering). Long runs: pilot below.
Prep batch (`--prep`): makes the backlog runner-ready, never implements. Targets = up to K Pending tasks without
Done, not `Sessions: owner`, no status, no open slices (pickable first, then priority); none → prints `prep: …
nothing started`, starts nothing. Prompt `templates/prep-prompt.md`: sonnet subagents (≤ 3 tasks each) read the
task text, `wf set <id> --done "…"` (+ `--model` if missing); unclear → `wf add -s awaiting` + `wf status <id>
blocked <a-id>`. Summary lines `<id> → Done: …` | `<id> → awaiting <a-id>`; then it commits the task file.
Pilot runs one when idle with `not runner-ready` tasks (once per idle spell).
## Pilot (nested)
`/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]`, typed into an idle project session (phone via Remote Control),
keeps that session tiny: subagents cannot spawn subagents, so it pilots headless batch orchestrators instead of
workers. Loop: `wf batch K --left <time to deadline>` (a wf res job, returns at once) → background Bash `wf res wait <rid>` → its exit
re-invokes the pilot → `wf batch --status` → one line to its context and `out/wf-orch.log` → next batch. A batch =
up to K tasks, rounds of one worker per lane in parallel (2 lanes, K=4 → two rounds). Nothing pickable →
prep batch if tasks are not runner-ready (once per idle spell), then background `wf lanes --wait 1800`, until `--for`. Deadline fit from history: task p90 = p90 of done/handback durations in `out/wf-orch.log` (local lanes, > 0; < 3 runs → 30 min); `--left` launches K' = min(K, floor(left / p90)) tasks with `--for` = left (K' = 0 → `fit: 0 …; nothing started`, the pilot stops); the batch prompt runs `wf batch --time-left <deadline>` before each round and spawns nothing more on `: stop` (left < p90). Steering: the owner talks to the pilot only; stop →
`wf batch --stop` (stop file, checked before every spawn: latency ≤ the running tasks), skip → `wf move <id> deferred`. Batch rc≠0 or no summary →
retry once, then stop. PushNotification on new awaiting items, on a stop and at run end; each alert also an `ALERT <text>` line in `out/wf-orch.log` and in the summary (push may be off). Phone push needs `agentPushNotifEnabled` (settings) / app notifications on. Needs the allow rule
`Bash(python3 /projects/public/workflow/wf.py batch:*)`.
## Context
Orchestrator large (after design talk) → finish the topic, commit, ask the owner to `/clear`. State lives in
TASKS.md, the log and the batch summaries.
Safe to clear only with no background worker running. Tested (2026-10-05): `/clear` leaves a background
subagent alive and its hand-back + completion notice reach the new context, but SIGKILLs its running Bash
command (exit 137): a worker mid-test or mid-commit loses that step and reports a false failure or half state.
## Split projects
`code_root` = another git repo: worktrees and branches live in the code repo, `wf start` makes them (with `.wf-home`
pointing at the private project). `wf orch post` checks the code repo. A done worker whose push (`push` key) failed
leaves `.wf/push-failed`: the lane stops (`push-failed`); fix the push, delete the marker, continue.
|