From 81d4e80fd5aabe4e80f58e960affa795cf7d34ec Mon Sep 17 00:00:00 2001 From: godosa Date: Wed, 7 Oct 2026 07:27:17 +0200 Subject: workflow: initial public history --- docs/orchestrator.md | 152 +++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 152 insertions(+) create mode 100644 docs/orchestrator.md (limited to 'docs/orchestrator.md') diff --git a/docs/orchestrator.md b/docs/orchestrator.md new file mode 100644 index 0000000..845c527 --- /dev/null +++ b/docs/orchestrator.md @@ -0,0 +1,152 @@ +# Orchestrator: one session talks to the owner, fresh workers do the tasks + +Why: cost is dominated by cache reads, which grow with context size. A fresh worker per task costs about +as much less as a cheaper model does; researching in one context and implementing in another is NOT cheaper +(rediscovery, handbacks). So: the orchestrator slices and rules, and one fresh worker per task researches +and implements. + +## Sessions +- **Orchestrator** (interactive, opus, `claude --autocompact 120k`): owner talk about runs, Awaiting rulings, + slicing, spawning workers, post-checks, unattended batches. Reads `wf` output and worker reports, never + code. Holds no lane: never `wf next --as` (use `wf lanes`, `wf list --runner --lane `). +- **Design session** (optional, big window): brainstorming, specs, hard rulings. Ends each topic with a + committed spec/task; the orchestrator picks it up from TASKS.md. +- **Workers**: the `wf-worker` subagent (`shared/agents/wf-worker.md`; install: + `ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md`), one task each. + +## Starting a session +Skills (install once: `ln -s /projects/public/workflow/shared/skills/wf-orchestrate ~/.claude/skills/` and the same for +`wf-design` and `wf-pilot`): after `/clear`, type `/wf-orchestrate [go | batch N]`, `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]` or `/wf-design [topic | id]` (not +"continue": that is the normal worker cold start). Agents and skills installed while a session runs appear +only in sessions started afterwards: restart it. + +## Runner-ready tasks (slicing rule) +- `Model:` line set to the cheapest lane that fits (sonnet: exact Done + a pattern to copy; haiku: mechanical). +- Done checkable headless: a file, test result or artifact, plus `wf add` for follow-ups. Never "report to + owner" / "owner confirms". +- `After:` satisfied, no `Sessions: owner`. `Sessions: solo` → run it alone. + +## One worker +Two commands do steps 1-3 and 5-7 (one call each): `wf orch pick [--id ID] [--recovery WHY]` (stop file, +pick, claim, free worktree, recorded in `.wf/orch/.json`, prints the spawn line and the prompt below) and +`wf orch post --result "" [--commit SHA] [--agent A] [--duration S] [--no-pick]` +(post-check, merge if needed, leftover commit, orch + cost log lines, then the next pick or `stop lane : `). +The steps they automate: + +1. Pick: `wf list --runner --lane ` → first id (lanes: `wf lanes`). Claim: `wf status progress "worker"`. No commit of + its own: `wf merge` or any later bookkeeping commit takes TASKS.md (and the claim) along; harmless. +2. Worktree: `
/.worktrees/`; if a live session or another worker uses it, `-2`, `-3`, … + (live interactive sessions: `~/.claude/sessions/*.json` field `cwd`). Branch `/`. +3. Spawn: Agent tool, `subagent_type: wf-worker`, `model: `, background, no worktree isolation + (it blocks `wf merge` on the main checkout). Prompt: + + ``` + Task: Lane: Model: + Main tree:
Worktree: Branch: / + Final message: the 4 report lines only. + ``` + (the last line is needed: the agent rule alone did not stop prose before the report.) + Slice job (row starts `slice:` / `wf show` effort > slice_above): same spawn; the worker only slices. + (no `wf ctx` output in the prompt: the worker fetches it; the prompt stays in your context forever.) +4. At most one worker per lane at a time; lanes in parallel. `wf` writes and `wf merge` take the project + lock (`.wf/lock`), so parallel workers are safe. +5. On completion read only the 4-line report (`followups` = ids the worker added). Post-check: + - `done`: `wf log -n 3` shows the id; `git -C
branch --list /` empty; `wf check` 0 errors; + `git -C status --porcelain` empty; `git -C merge-base --is-ancestor HEAD ` exit 0 + (worktree HEAD in master: no impl commit left on a detached HEAD; else `cd && wf merge`). + - Everything else: commit leftover bookkeeping + `git -C
commit -m " (orchestrator)" -- TASKS.md ` if they changed. +6. Outcome: + | Report / state | Action | + |---|---| + | done + post-check ok | next task in the lane | + | model raised by the worker | continue; next pick spawns it on the new model | + | sliced | next task in the lane | + | done+gate-red (async gate red, earlier task's break) | post-check as done; next pick in the lane = the P0 fix (no Done/Model → add them first); never stop the lane | + | awaiting / handback / post-check red / no report | stop the lane, tell the owner | + | wip (after a wrap-up) | stop the lane; reslice or rerun later | + | crashed / killed (no report) | one fresh worker, prompt plus `Recovery: ` (it inspects git status, git log master..HEAD, `git diff -- TASKS.md` — often just the claim — and `wf show `), then stop | +7. Log one line per task to `out/wf-orch.log` (git-ignored): time, lane, model, id, outcome, commit, duration. + Cost: `wf usage --agent --log --effort --lane ` appends one key=value line to + `out/wf-cost.log` (tokens + API-price $ from the subagent transcript). `--duration ` adds dur=; `wf usage --report` (med_dur) = per lane/model and + effort: n, done, median/total $, $ per done task, median turns, over every project's log. + The completion notice's token count is NOT the cost (one worker: notice 67k, transcript 5.0M cache reads); + real usage is in the subagent transcript `~/.claude/projects///subagents/agent-*.jsonl` + (`wf usage` per session/agent; one request is several transcript entries: counted once; + subagent transcripts lack final output counts → out estimated from content, shown `~`, log `est=N`). + +## Cloud lane +Virtual lane `cloud` (projects with `cloud = true`; design: cloud-lane spec §4.6): tasks run as +`claude --cloud` sessions billed to the cloud budget, not as local workers. +- `wf orch pick cloud [--id ID]` = stop file → ledger (`wf cloud ledger`) → first fitting task across all lanes + (runner-ready, not in progress, `wf cloud` fit rules; opus only, `Cloud: yes` first, then prio, then larger effort) → `wf cloud send` + (claim `in progress: cloud:`, record `.wf/orch/.json` lane cloud). Nothing to spawn: the session runs + remotely. Otherwise `stop lane cloud: ledger (…)` (balance < reserve, or send refused) | `none fit (…)` | + `max parallel (…)`. A stop of the cloud lane never stops the local lanes. +- Local lanes skip tasks claimed `cloud:` (every `in progress` task is skipped by `wf orch pick `). +- On each wake-up, at most every 10 min: `wf cloud pull --all` (teleport ~30 s, no tokens). Per ended task it prints + `: (, $… )` (state done | awaiting | handback | lost; done also `report: commit `) → + `wf orch post cloud --result [--commit ]` = the worker-report path (done: archive line, branch + `/` gone, `wf check`; else leftover commit of the note/claim pull left), log line, next cloud + pick. `running` (exit 4) → nothing; pull again next wake. awaiting / handback / lost → `stop lane cloud`, tell the + owner; the task is back in its local lane (handback note `Recovery: cloud attempt …`). +- Fill: one `wf orch pick cloud` per free slot (up to `max_parallel`), each in its own call. +- Batch (`wf batch` on a `cloud = true` project) pulls itself: the job runs `wf_res.py batch-sidecar` around + `claude -p`; every 10 min while it lives (until `out/wf-batch.stop`) it runs `wf cloud pull --all` and per ended + task `wf orch post cloud --result [--commit ] --no-pick` + a `cloud (sidecar pull)` + line in the summary, so pulls never starve while the batch waits on foreground workers. The batch orchestrator + only picks cloud, never pulls. Two pulls of one id at once: the second prints `: pulled by another wf cloud + pull, skipped` (rc 0; flock on `.wf/cloud/.json`). + Tail: the orchestrator exited with `.wf/cloud/*.json` records left (sent at batch end) → the sidecar keeps + pulling every 10 min (no picks, no sends) until no record is left, the stop file appears or 24 h passed; the job's + ledger reservation shrinks to 0.2 GB / 1 cpu for it (ledger only, the unit's MemoryMax stays); the summary ends + with `cloud tail ended: ; pulls; left: (sidecar)`. A next batch's sidecar may overlap it (record + flock). + +## Unattended batch (nightly, long runs) +The orchestrator does not babysit long runs; a headless batch orchestrator (separate `claude -p` process, +fresh context per batch) does. The owner starts it (auto mode denies an agent launching another unattended agent): `!` + the command, or +once an allow rule `Bash(python3 /projects/public/workflow/wf.py batch:*)` in `~/.claude/settings.json`. + +``` +wf batch N [--lanes sonnet,haiku] [--model opus] [--for 6h] [--mem 4G] # --dry-run prints the command; default mem/for: wf res hist +wf batch --status # newest job + out/wf-batch-*.md +wf batch K --prep [--lanes fast] # prep batch (below) +``` += `wf res run --mem 4G --for 6h --title wf-batch -- claude -p --model opus --permission-mode auto +--permission-prompts none ""` with `CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS` = `--for`. +The prompt: follow 'One worker' for up to N tasks, never implement; each round one wf-worker per lane in ONE +message, foreground, post-check, one line per task to `out/wf-batch-.md`; stop a lane on its first stop +outcome; never ask (follow-ups → `wf add -p 2`, decisions → `wf add -s awaiting`); end with the ids added. + +Why foreground: `claude -p` ends when the orchestrator's turn ends and kills background workers after a +wait ceiling (default 10 min; the env above raises it as a backstop). Workers spawned in one message run in +parallel and the turn waits for all. Background subagents cannot spawn subagents, so the batch +orchestrator is never a subagent. The interactive orchestrator checks `wf batch --status`. + +Headless batch = short fire-and-forget runs (a few tasks, no steering). Long runs: pilot below. + +Prep batch (`--prep`): makes the backlog runner-ready, never implements. Targets = up to K Pending tasks without +Done, not `Sessions: owner`, no status, no open slices (pickable first, then priority); none → prints `prep: … +nothing started`, starts nothing. Prompt `templates/prep-prompt.md`: sonnet subagents (≤ 3 tasks each) read the +task text, `wf set --done "…"` (+ `--model` if missing); unclear → `wf add -s awaiting` + `wf status +blocked `. Summary lines ` → Done: …` | ` → awaiting `; then it commits the task file. +Pilot runs one when idle with `not runner-ready` tasks (once per idle spell). + +## Pilot (nested) +`/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]`, typed into an idle project session (phone via Remote Control), +keeps that session tiny: subagents cannot spawn subagents, so it pilots headless batch orchestrators instead of +workers. Loop: `wf batch K --left