diff options
| author | godosa <godosa@godosa.eu> | 2026-10-07 07:27:17 +0200 |
|---|---|---|
| committer | godosa <godosa@godosa.eu> | 2026-10-07 07:27:17 +0200 |
| commit | 81d4e80fd5aabe4e80f58e960affa795cf7d34ec (patch) | |
| tree | e98eeac2af6af63aa4287bba1f6d4a3af26b5727 /docs/orchestrator.md | |
| download | workflow-81d4e80fd5aabe4e80f58e960affa795cf7d34ec.tar.gz workflow-81d4e80fd5aabe4e80f58e960affa795cf7d34ec.zip | |
workflow: initial public history
Diffstat (limited to 'docs/orchestrator.md')
| -rw-r--r-- | docs/orchestrator.md | 152 |
1 files changed, 152 insertions, 0 deletions
diff --git a/docs/orchestrator.md b/docs/orchestrator.md new file mode 100644 index 0000000..845c527 --- /dev/null +++ b/docs/orchestrator.md @@ -0,0 +1,152 @@ +# Orchestrator: one session talks to the owner, fresh workers do the tasks + +Why: cost is dominated by cache reads, which grow with context size. A fresh worker per task costs about +as much less as a cheaper model does; researching in one context and implementing in another is NOT cheaper +(rediscovery, handbacks). So: the orchestrator slices and rules, and one fresh worker per task researches +and implements. + +## Sessions +- **Orchestrator** (interactive, opus, `claude --autocompact 120k`): owner talk about runs, Awaiting rulings, + slicing, spawning workers, post-checks, unattended batches. Reads `wf` output and worker reports, never + code. Holds no lane: never `wf next --as` (use `wf lanes`, `wf list --runner --lane <lane>`). +- **Design session** (optional, big window): brainstorming, specs, hard rulings. Ends each topic with a + committed spec/task; the orchestrator picks it up from TASKS.md. +- **Workers**: the `wf-worker` subagent (`shared/agents/wf-worker.md`; install: + `ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md`), one task each. + +## Starting a session +Skills (install once: `ln -s /projects/public/workflow/shared/skills/wf-orchestrate ~/.claude/skills/` and the same for +`wf-design` and `wf-pilot`): after `/clear`, type `/wf-orchestrate [go | batch N]`, `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]` or `/wf-design [topic | id]` (not +"continue": that is the normal worker cold start). Agents and skills installed while a session runs appear +only in sessions started afterwards: restart it. + +## Runner-ready tasks (slicing rule) +- `Model:` line set to the cheapest lane that fits (sonnet: exact Done + a pattern to copy; haiku: mechanical). +- Done checkable headless: a file, test result or artifact, plus `wf add` for follow-ups. Never "report to + owner" / "owner confirms". +- `After:` satisfied, no `Sessions: owner`. `Sessions: solo` → run it alone. + +## One worker +Two commands do steps 1-3 and 5-7 (one call each): `wf orch pick <lane> [--id ID] [--recovery WHY]` (stop file, +pick, claim, free worktree, recorded in `.wf/orch/<id>.json`, prints the spawn line and the prompt below) and +`wf orch post <id> <lane> --result "<report result line>" [--commit SHA] [--agent A] [--duration S] [--no-pick]` +(post-check, merge if needed, leftover commit, orch + cost log lines, then the next pick or `stop lane <lane>: <why>`). +The steps they automate: + +1. Pick: `wf list --runner --lane <lane>` → first id (lanes: `wf lanes`). Claim: `wf status <id> progress "worker"`. No commit of + its own: `wf merge` or any later bookkeeping commit takes TASKS.md (and the claim) along; harmless. +2. Worktree: `<main>/.worktrees/<lane>`; if a live session or another worker uses it, `<lane>-2`, `-3`, … + (live interactive sessions: `~/.claude/sessions/*.json` field `cwd`). Branch `<lane>/<id>`. +3. Spawn: Agent tool, `subagent_type: wf-worker`, `model: <task Model, none = opus>`, background, no worktree isolation + (it blocks `wf merge` on the main checkout). Prompt: + + ``` + Task: <id> Lane: <lane> Model: <model> + Main tree: <main> Worktree: <path> Branch: <lane>/<id> + Final message: the 4 report lines only. + ``` + (the last line is needed: the agent rule alone did not stop prose before the report.) + Slice job (row starts `slice:` / `wf show` effort > slice_above): same spawn; the worker only slices. + (no `wf ctx` output in the prompt: the worker fetches it; the prompt stays in your context forever.) +4. At most one worker per lane at a time; lanes in parallel. `wf` writes and `wf merge` take the project + lock (`.wf/lock`), so parallel workers are safe. +5. On completion read only the 4-line report (`followups` = ids the worker added). Post-check: + - `done`: `wf log -n 3` shows the id; `git -C <main> branch --list <lane>/<id>` empty; `wf check` 0 errors; + `git -C <path> status --porcelain` empty; `git -C <path> merge-base --is-ancestor HEAD <master>` exit 0 + (worktree HEAD in master: no impl commit left on a detached HEAD; else `cd <path> && wf merge`). + - Everything else: commit leftover bookkeeping + `git -C <main> commit -m "<id> <outcome> (orchestrator)" -- TASKS.md <archive>` if they changed. +6. Outcome: + | Report / state | Action | + |---|---| + | done + post-check ok | next task in the lane | + | model raised by the worker | continue; next pick spawns it on the new model | + | sliced | next task in the lane | + | done+gate-red <culprit> <fix-id> (async gate red, earlier task's break) | post-check as done; next pick in the lane = the P0 fix (no Done/Model → add them first); never stop the lane | + | awaiting / handback / post-check red / no report | stop the lane, tell the owner | + | wip (after a wrap-up) | stop the lane; reslice or rerun later | + | crashed / killed (no report) | one fresh worker, prompt plus `Recovery: <why>` (it inspects git status, git log master..HEAD, `git diff -- TASKS.md` — often just the claim — and `wf show <id>`), then stop | +7. Log one line per task to `out/wf-orch.log` (git-ignored): time, lane, model, id, outcome, commit, duration. + Cost: `wf usage --agent <agent id> --log <id> <outcome> --effort <task estimate: <1h|1h|5h|10h|100h> --lane <lane>` appends one key=value line to + `out/wf-cost.log` (tokens + API-price $ from the subagent transcript). `--duration <s>` adds dur=; `wf usage --report` (med_dur) = per lane/model and + effort: n, done, median/total $, $ per done task, median turns, over every project's log. + The completion notice's token count is NOT the cost (one worker: notice 67k, transcript 5.0M cache reads); + real usage is in the subagent transcript `~/.claude/projects/<project>/<session>/subagents/agent-*.jsonl` + (`wf usage` per session/agent; one request is several transcript entries: counted once; + subagent transcripts lack final output counts → out estimated from content, shown `~`, log `est=N`). + +## Cloud lane +Virtual lane `cloud` (projects with `cloud = true`; design: cloud-lane spec §4.6): tasks run as +`claude --cloud` sessions billed to the cloud budget, not as local workers. +- `wf orch pick cloud [--id ID]` = stop file → ledger (`wf cloud ledger`) → first fitting task across all lanes + (runner-ready, not in progress, `wf cloud` fit rules; opus only, `Cloud: yes` first, then prio, then larger effort) → `wf cloud send` + (claim `in progress: cloud:<sid>`, record `.wf/orch/<id>.json` lane cloud). Nothing to spawn: the session runs + remotely. Otherwise `stop lane cloud: ledger (…)` (balance < reserve, or send refused) | `none fit (…)` | + `max parallel (…)`. A stop of the cloud lane never stops the local lanes. +- Local lanes skip tasks claimed `cloud:` (every `in progress` task is skipped by `wf orch pick <lane>`). +- On each wake-up, at most every 10 min: `wf cloud pull --all` (teleport ~30 s, no tokens). Per ended task it prints + `<id>: <state> (<sid>, $… )` (state done | awaiting | handback | lost; done also `report: commit <sha>`) → + `wf orch post <id> cloud --result <state> [--commit <sha>]` = the worker-report path (done: archive line, branch + `<task lane>/<id>` gone, `wf check`; else leftover commit of the note/claim pull left), log line, next cloud + pick. `running` (exit 4) → nothing; pull again next wake. awaiting / handback / lost → `stop lane cloud`, tell the + owner; the task is back in its local lane (handback note `Recovery: cloud attempt …`). +- Fill: one `wf orch pick cloud` per free slot (up to `max_parallel`), each in its own call. +- Batch (`wf batch` on a `cloud = true` project) pulls itself: the job runs `wf_res.py batch-sidecar` around + `claude -p`; every 10 min while it lives (until `out/wf-batch.stop`) it runs `wf cloud pull --all` and per ended + task `wf orch post <id> cloud --result <state> [--commit <sha>] --no-pick` + a `cloud <id> <state> <sha> (sidecar pull)` + line in the summary, so pulls never starve while the batch waits on foreground workers. The batch orchestrator + only picks cloud, never pulls. Two pulls of one id at once: the second prints `<id>: pulled by another wf cloud + pull, skipped` (rc 0; flock on `.wf/cloud/<id>.json`). + Tail: the orchestrator exited with `.wf/cloud/*.json` records left (sent at batch end) → the sidecar keeps + pulling every 10 min (no picks, no sends) until no record is left, the stop file appears or 24 h passed; the job's + ledger reservation shrinks to 0.2 GB / 1 cpu for it (ledger only, the unit's MemoryMax stays); the summary ends + with `cloud tail ended: <why>; <n> pulls; left: <ids|none> (sidecar)`. A next batch's sidecar may overlap it (record + flock). + +## Unattended batch (nightly, long runs) +The orchestrator does not babysit long runs; a headless batch orchestrator (separate `claude -p` process, +fresh context per batch) does. The owner starts it (auto mode denies an agent launching another unattended agent): `!` + the command, or +once an allow rule `Bash(python3 /projects/public/workflow/wf.py batch:*)` in `~/.claude/settings.json`. + +``` +wf batch N [--lanes sonnet,haiku] [--model opus] [--for 6h] [--mem 4G] # --dry-run prints the command; default mem/for: wf res hist +wf batch --status # newest job + out/wf-batch-*.md +wf batch K --prep [--lanes fast] # prep batch (below) +``` += `wf res run --mem 4G --for 6h --title wf-batch -- claude -p --model opus --permission-mode auto +--permission-prompts none "<templates/batch-prompt.md>"` with `CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS` = `--for`. +The prompt: follow 'One worker' for up to N tasks, never implement; each round one wf-worker per lane in ONE +message, foreground, post-check, one line per task to `out/wf-batch-<stamp>.md`; stop a lane on its first stop +outcome; never ask (follow-ups → `wf add -p 2`, decisions → `wf add -s awaiting`); end with the ids added. + +Why foreground: `claude -p` ends when the orchestrator's turn ends and kills background workers after a +wait ceiling (default 10 min; the env above raises it as a backstop). Workers spawned in one message run in +parallel and the turn waits for all. Background subagents cannot spawn subagents, so the batch +orchestrator is never a subagent. The interactive orchestrator checks `wf batch --status`. + +Headless batch = short fire-and-forget runs (a few tasks, no steering). Long runs: pilot below. + +Prep batch (`--prep`): makes the backlog runner-ready, never implements. Targets = up to K Pending tasks without +Done, not `Sessions: owner`, no status, no open slices (pickable first, then priority); none → prints `prep: … +nothing started`, starts nothing. Prompt `templates/prep-prompt.md`: sonnet subagents (≤ 3 tasks each) read the +task text, `wf set <id> --done "…"` (+ `--model` if missing); unclear → `wf add -s awaiting` + `wf status <id> +blocked <a-id>`. Summary lines `<id> → Done: …` | `<id> → awaiting <a-id>`; then it commits the task file. +Pilot runs one when idle with `not runner-ready` tasks (once per idle spell). + +## Pilot (nested) +`/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]`, typed into an idle project session (phone via Remote Control), +keeps that session tiny: subagents cannot spawn subagents, so it pilots headless batch orchestrators instead of +workers. Loop: `wf batch K --left <time to deadline>` (a wf res job, returns at once) → background Bash `wf res wait <rid>` → its exit +re-invokes the pilot → `wf batch --status` → one line to its context and `out/wf-orch.log` → next batch. A batch = +up to K tasks, rounds of one worker per lane in parallel (2 lanes, K=4 → two rounds). Nothing pickable → +prep batch if tasks are not runner-ready (once per idle spell), then background `wf lanes --wait 1800`, until `--for`. Deadline fit from history: task p90 = p90 of done/handback durations in `out/wf-orch.log` (local lanes, > 0; < 3 runs → 30 min); `--left` launches K' = min(K, floor(left / p90)) tasks with `--for` = left (K' = 0 → `fit: 0 …; nothing started`, the pilot stops); the batch prompt runs `wf batch --time-left <deadline>` before each round and spawns nothing more on `: stop` (left < p90). Steering: the owner talks to the pilot only; stop → +`wf batch --stop` (stop file, checked before every spawn: latency ≤ the running tasks), skip → `wf move <id> deferred`. Batch rc≠0 or no summary → +retry once, then stop. PushNotification on new awaiting items, on a stop and at run end; each alert also an `ALERT <text>` line in `out/wf-orch.log` and in the summary (push may be off). Phone push needs `agentPushNotifEnabled` (settings) / app notifications on. Needs the allow rule +`Bash(python3 /projects/public/workflow/wf.py batch:*)`. + +## Context +Orchestrator large (after design talk) → finish the topic, commit, ask the owner to `/clear`. State lives in +TASKS.md, the log and the batch summaries. +Safe to clear only with no background worker running. Tested (2026-10-05): `/clear` leaves a background +subagent alive and its hand-back + completion notice reach the new context, but SIGKILLs its running Bash +command (exit 137): a worker mid-test or mid-commit loses that step and reports a false failure or half state. |
