aboutsummaryrefslogtreecommitdiffziptar.gz
path: root/docs/orchestrator.md
diff options
context:
space:
mode:
authorgodosa <godosa@godosa.eu>2026-10-07 07:27:17 +0200
committergodosa <godosa@godosa.eu>2026-10-07 07:27:17 +0200
commit81d4e80fd5aabe4e80f58e960affa795cf7d34ec (patch)
treee98eeac2af6af63aa4287bba1f6d4a3af26b5727 /docs/orchestrator.md
downloadworkflow-81d4e80fd5aabe4e80f58e960affa795cf7d34ec.tar.gz
workflow-81d4e80fd5aabe4e80f58e960affa795cf7d34ec.zip
workflow: initial public history
Diffstat (limited to 'docs/orchestrator.md')
-rw-r--r--docs/orchestrator.md152
1 files changed, 152 insertions, 0 deletions
diff --git a/docs/orchestrator.md b/docs/orchestrator.md
new file mode 100644
index 0000000..845c527
--- /dev/null
+++ b/docs/orchestrator.md
@@ -0,0 +1,152 @@
+# Orchestrator: one session talks to the owner, fresh workers do the tasks
+
+Why: cost is dominated by cache reads, which grow with context size. A fresh worker per task costs about
+as much less as a cheaper model does; researching in one context and implementing in another is NOT cheaper
+(rediscovery, handbacks). So: the orchestrator slices and rules, and one fresh worker per task researches
+and implements.
+
+## Sessions
+- **Orchestrator** (interactive, opus, `claude --autocompact 120k`): owner talk about runs, Awaiting rulings,
+ slicing, spawning workers, post-checks, unattended batches. Reads `wf` output and worker reports, never
+ code. Holds no lane: never `wf next --as` (use `wf lanes`, `wf list --runner --lane <lane>`).
+- **Design session** (optional, big window): brainstorming, specs, hard rulings. Ends each topic with a
+ committed spec/task; the orchestrator picks it up from TASKS.md.
+- **Workers**: the `wf-worker` subagent (`shared/agents/wf-worker.md`; install:
+ `ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md`), one task each.
+
+## Starting a session
+Skills (install once: `ln -s /projects/public/workflow/shared/skills/wf-orchestrate ~/.claude/skills/` and the same for
+`wf-design` and `wf-pilot`): after `/clear`, type `/wf-orchestrate [go | batch N]`, `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]` or `/wf-design [topic | id]` (not
+"continue": that is the normal worker cold start). Agents and skills installed while a session runs appear
+only in sessions started afterwards: restart it.
+
+## Runner-ready tasks (slicing rule)
+- `Model:` line set to the cheapest lane that fits (sonnet: exact Done + a pattern to copy; haiku: mechanical).
+- Done checkable headless: a file, test result or artifact, plus `wf add` for follow-ups. Never "report to
+ owner" / "owner confirms".
+- `After:` satisfied, no `Sessions: owner`. `Sessions: solo` → run it alone.
+
+## One worker
+Two commands do steps 1-3 and 5-7 (one call each): `wf orch pick <lane> [--id ID] [--recovery WHY]` (stop file,
+pick, claim, free worktree, recorded in `.wf/orch/<id>.json`, prints the spawn line and the prompt below) and
+`wf orch post <id> <lane> --result "<report result line>" [--commit SHA] [--agent A] [--duration S] [--no-pick]`
+(post-check, merge if needed, leftover commit, orch + cost log lines, then the next pick or `stop lane <lane>: <why>`).
+The steps they automate:
+
+1. Pick: `wf list --runner --lane <lane>` → first id (lanes: `wf lanes`). Claim: `wf status <id> progress "worker"`. No commit of
+ its own: `wf merge` or any later bookkeeping commit takes TASKS.md (and the claim) along; harmless.
+2. Worktree: `<main>/.worktrees/<lane>`; if a live session or another worker uses it, `<lane>-2`, `-3`, …
+ (live interactive sessions: `~/.claude/sessions/*.json` field `cwd`). Branch `<lane>/<id>`.
+3. Spawn: Agent tool, `subagent_type: wf-worker`, `model: <task Model, none = opus>`, background, no worktree isolation
+ (it blocks `wf merge` on the main checkout). Prompt:
+
+ ```
+ Task: <id> Lane: <lane> Model: <model>
+ Main tree: <main> Worktree: <path> Branch: <lane>/<id>
+ Final message: the 4 report lines only.
+ ```
+ (the last line is needed: the agent rule alone did not stop prose before the report.)
+ Slice job (row starts `slice:` / `wf show` effort > slice_above): same spawn; the worker only slices.
+ (no `wf ctx` output in the prompt: the worker fetches it; the prompt stays in your context forever.)
+4. At most one worker per lane at a time; lanes in parallel. `wf` writes and `wf merge` take the project
+ lock (`.wf/lock`), so parallel workers are safe.
+5. On completion read only the 4-line report (`followups` = ids the worker added). Post-check:
+ - `done`: `wf log -n 3` shows the id; `git -C <main> branch --list <lane>/<id>` empty; `wf check` 0 errors;
+ `git -C <path> status --porcelain` empty; `git -C <path> merge-base --is-ancestor HEAD <master>` exit 0
+ (worktree HEAD in master: no impl commit left on a detached HEAD; else `cd <path> && wf merge`).
+ - Everything else: commit leftover bookkeeping
+ `git -C <main> commit -m "<id> <outcome> (orchestrator)" -- TASKS.md <archive>` if they changed.
+6. Outcome:
+ | Report / state | Action |
+ |---|---|
+ | done + post-check ok | next task in the lane |
+ | model raised by the worker | continue; next pick spawns it on the new model |
+ | sliced | next task in the lane |
+ | done+gate-red <culprit> <fix-id> (async gate red, earlier task's break) | post-check as done; next pick in the lane = the P0 fix (no Done/Model → add them first); never stop the lane |
+ | awaiting / handback / post-check red / no report | stop the lane, tell the owner |
+ | wip (after a wrap-up) | stop the lane; reslice or rerun later |
+ | crashed / killed (no report) | one fresh worker, prompt plus `Recovery: <why>` (it inspects git status, git log master..HEAD, `git diff -- TASKS.md` — often just the claim — and `wf show <id>`), then stop |
+7. Log one line per task to `out/wf-orch.log` (git-ignored): time, lane, model, id, outcome, commit, duration.
+ Cost: `wf usage --agent <agent id> --log <id> <outcome> --effort <task estimate: <1h|1h|5h|10h|100h> --lane <lane>` appends one key=value line to
+ `out/wf-cost.log` (tokens + API-price $ from the subagent transcript). `--duration <s>` adds dur=; `wf usage --report` (med_dur) = per lane/model and
+ effort: n, done, median/total $, $ per done task, median turns, over every project's log.
+ The completion notice's token count is NOT the cost (one worker: notice 67k, transcript 5.0M cache reads);
+ real usage is in the subagent transcript `~/.claude/projects/<project>/<session>/subagents/agent-*.jsonl`
+ (`wf usage` per session/agent; one request is several transcript entries: counted once;
+ subagent transcripts lack final output counts → out estimated from content, shown `~`, log `est=N`).
+
+## Cloud lane
+Virtual lane `cloud` (projects with `cloud = true`; design: cloud-lane spec §4.6): tasks run as
+`claude --cloud` sessions billed to the cloud budget, not as local workers.
+- `wf orch pick cloud [--id ID]` = stop file → ledger (`wf cloud ledger`) → first fitting task across all lanes
+ (runner-ready, not in progress, `wf cloud` fit rules; opus only, `Cloud: yes` first, then prio, then larger effort) → `wf cloud send`
+ (claim `in progress: cloud:<sid>`, record `.wf/orch/<id>.json` lane cloud). Nothing to spawn: the session runs
+ remotely. Otherwise `stop lane cloud: ledger (…)` (balance < reserve, or send refused) | `none fit (…)` |
+ `max parallel (…)`. A stop of the cloud lane never stops the local lanes.
+- Local lanes skip tasks claimed `cloud:` (every `in progress` task is skipped by `wf orch pick <lane>`).
+- On each wake-up, at most every 10 min: `wf cloud pull --all` (teleport ~30 s, no tokens). Per ended task it prints
+ `<id>: <state> (<sid>, $… )` (state done | awaiting | handback | lost; done also `report: commit <sha>`) →
+ `wf orch post <id> cloud --result <state> [--commit <sha>]` = the worker-report path (done: archive line, branch
+ `<task lane>/<id>` gone, `wf check`; else leftover commit of the note/claim pull left), log line, next cloud
+ pick. `running` (exit 4) → nothing; pull again next wake. awaiting / handback / lost → `stop lane cloud`, tell the
+ owner; the task is back in its local lane (handback note `Recovery: cloud attempt …`).
+- Fill: one `wf orch pick cloud` per free slot (up to `max_parallel`), each in its own call.
+- Batch (`wf batch` on a `cloud = true` project) pulls itself: the job runs `wf_res.py batch-sidecar` around
+ `claude -p`; every 10 min while it lives (until `out/wf-batch.stop`) it runs `wf cloud pull --all` and per ended
+ task `wf orch post <id> cloud --result <state> [--commit <sha>] --no-pick` + a `cloud <id> <state> <sha> (sidecar pull)`
+ line in the summary, so pulls never starve while the batch waits on foreground workers. The batch orchestrator
+ only picks cloud, never pulls. Two pulls of one id at once: the second prints `<id>: pulled by another wf cloud
+ pull, skipped` (rc 0; flock on `.wf/cloud/<id>.json`).
+ Tail: the orchestrator exited with `.wf/cloud/*.json` records left (sent at batch end) → the sidecar keeps
+ pulling every 10 min (no picks, no sends) until no record is left, the stop file appears or 24 h passed; the job's
+ ledger reservation shrinks to 0.2 GB / 1 cpu for it (ledger only, the unit's MemoryMax stays); the summary ends
+ with `cloud tail ended: <why>; <n> pulls; left: <ids|none> (sidecar)`. A next batch's sidecar may overlap it (record
+ flock).
+
+## Unattended batch (nightly, long runs)
+The orchestrator does not babysit long runs; a headless batch orchestrator (separate `claude -p` process,
+fresh context per batch) does. The owner starts it (auto mode denies an agent launching another unattended agent): `!` + the command, or
+once an allow rule `Bash(python3 /projects/public/workflow/wf.py batch:*)` in `~/.claude/settings.json`.
+
+```
+wf batch N [--lanes sonnet,haiku] [--model opus] [--for 6h] [--mem 4G] # --dry-run prints the command; default mem/for: wf res hist
+wf batch --status # newest job + out/wf-batch-*.md
+wf batch K --prep [--lanes fast] # prep batch (below)
+```
+= `wf res run --mem 4G --for 6h --title wf-batch -- claude -p --model opus --permission-mode auto
+--permission-prompts none "<templates/batch-prompt.md>"` with `CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS` = `--for`.
+The prompt: follow 'One worker' for up to N tasks, never implement; each round one wf-worker per lane in ONE
+message, foreground, post-check, one line per task to `out/wf-batch-<stamp>.md`; stop a lane on its first stop
+outcome; never ask (follow-ups → `wf add -p 2`, decisions → `wf add -s awaiting`); end with the ids added.
+
+Why foreground: `claude -p` ends when the orchestrator's turn ends and kills background workers after a
+wait ceiling (default 10 min; the env above raises it as a backstop). Workers spawned in one message run in
+parallel and the turn waits for all. Background subagents cannot spawn subagents, so the batch
+orchestrator is never a subagent. The interactive orchestrator checks `wf batch --status`.
+
+Headless batch = short fire-and-forget runs (a few tasks, no steering). Long runs: pilot below.
+
+Prep batch (`--prep`): makes the backlog runner-ready, never implements. Targets = up to K Pending tasks without
+Done, not `Sessions: owner`, no status, no open slices (pickable first, then priority); none → prints `prep: …
+nothing started`, starts nothing. Prompt `templates/prep-prompt.md`: sonnet subagents (≤ 3 tasks each) read the
+task text, `wf set <id> --done "…"` (+ `--model` if missing); unclear → `wf add -s awaiting` + `wf status <id>
+blocked <a-id>`. Summary lines `<id> → Done: …` | `<id> → awaiting <a-id>`; then it commits the task file.
+Pilot runs one when idle with `not runner-ready` tasks (once per idle spell).
+
+## Pilot (nested)
+`/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]`, typed into an idle project session (phone via Remote Control),
+keeps that session tiny: subagents cannot spawn subagents, so it pilots headless batch orchestrators instead of
+workers. Loop: `wf batch K --left <time to deadline>` (a wf res job, returns at once) → background Bash `wf res wait <rid>` → its exit
+re-invokes the pilot → `wf batch --status` → one line to its context and `out/wf-orch.log` → next batch. A batch =
+up to K tasks, rounds of one worker per lane in parallel (2 lanes, K=4 → two rounds). Nothing pickable →
+prep batch if tasks are not runner-ready (once per idle spell), then background `wf lanes --wait 1800`, until `--for`. Deadline fit from history: task p90 = p90 of done/handback durations in `out/wf-orch.log` (local lanes, > 0; < 3 runs → 30 min); `--left` launches K' = min(K, floor(left / p90)) tasks with `--for` = left (K' = 0 → `fit: 0 …; nothing started`, the pilot stops); the batch prompt runs `wf batch --time-left <deadline>` before each round and spawns nothing more on `: stop` (left < p90). Steering: the owner talks to the pilot only; stop →
+`wf batch --stop` (stop file, checked before every spawn: latency ≤ the running tasks), skip → `wf move <id> deferred`. Batch rc≠0 or no summary →
+retry once, then stop. PushNotification on new awaiting items, on a stop and at run end; each alert also an `ALERT <text>` line in `out/wf-orch.log` and in the summary (push may be off). Phone push needs `agentPushNotifEnabled` (settings) / app notifications on. Needs the allow rule
+`Bash(python3 /projects/public/workflow/wf.py batch:*)`.
+
+## Context
+Orchestrator large (after design talk) → finish the topic, commit, ask the owner to `/clear`. State lives in
+TASKS.md, the log and the batch summaries.
+Safe to clear only with no background worker running. Tested (2026-10-05): `/clear` leaves a background
+subagent alive and its hand-back + completion notice reach the new context, but SIGKILLs its running Bash
+command (exit 137): a worker mid-test or mid-commit loses that step and reports a false failure or half state.