diff options
| author | godosa <godosa@godosa.eu> | 2026-10-07 07:27:17 +0200 |
|---|---|---|
| committer | godosa <godosa@godosa.eu> | 2026-10-07 07:27:17 +0200 |
| commit | 81d4e80fd5aabe4e80f58e960affa795cf7d34ec (patch) | |
| tree | e98eeac2af6af63aa4287bba1f6d4a3af26b5727 | |
| download | workflow-81d4e80fd5aabe4e80f58e960affa795cf7d34ec.tar.gz workflow-81d4e80fd5aabe4e80f58e960affa795cf7d34ec.zip | |
workflow: initial public history
73 files changed, 18930 insertions, 0 deletions
diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..7ac043d --- /dev/null +++ b/.gitattributes @@ -0,0 +1 @@ +CHANGES.md merge=union diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..30e97e5 --- /dev/null +++ b/.gitignore @@ -0,0 +1,5 @@ +__pycache__/ +inbox.md +.worktrees/ +.superpowers/ +publish-denylist.local diff --git a/CHANGES.md b/CHANGES.md new file mode 100644 index 0000000..5f02c1e --- /dev/null +++ b/CHANGES.md @@ -0,0 +1,184 @@ +# Changes (newest first) +- 2026-10-07 Tests use generic fixtures (/h/u, 10.0.0.x) instead of real home-dir and LAN paths (publish-scan clean). Projects: nothing. + +- 2026-10-06 Tool path is now `/projects/public/workflow` (was `/projects/workflow`): docs, shared CLAUDE.md, skills, wf-worker agent, messages. `wf projects` / `wf usage --report` scan the grandparent when the parent holds no project. Projects: point `~/.claude` symlinks/settings, `wf-res.service` and any script calling `wf.py` at the new path. +- 2026-10-06 New `scripts/publish_snapshot.py [DEST]`: scrubbed fresh-history snapshot of the master tree into a separate local repo (default ~/src/wf-public; drops inbox.md, .worktrees/, __pycache__/, out/); refuses on any word of the git-ignored `publish-denylist.local` in a path or file; one commit per run (new CHANGES lines as message); never pushes or adds a remote (docs/manual.md). Projects: nothing. +- 2026-10-06 docs/manual.md: owner-facing manual (setup, daily use per session type, writing good tasks, best practices, troubleshooting); linked from the README top (replaces 'no installation guide'). Projects: nothing. +- 2026-10-06 `wf orch post`: `--duration 0` (orchestrator lacked duration_ms) falls back to the time since the pick like an omitted one (was logged `0m00s`, starving the task p90 of `wf batch --left/--time-left`); a post-check-red post keeps the orch record, so the re-post after the fix logs the real duration. Projects: nothing. +- 2026-10-06 `wf res` entry project = main tree name (git common dir's parent), not the lane worktree dir (was 'slow'/'fast-2': history kinds split per worktree); `wf res hist`, history hints, `wf batch --status` use it too. Older history records keep their worktree name. Projects: nothing. +- 2026-10-06 `wf batch K --left DURATION`: K' = min(K, floor(left / task p90)) (p90 of done/handback durations in `out/wf-orch.log`, local lanes, > 0; < 3 runs → 30 min), prints `fit: K' of K tasks (…)`, K' = 0 → nothing started (rc 0); `--for` defaults to left. New `wf batch --time-left DEADLINE` → `time left X, task p90 Y (n runs): spawn|stop`; batch prompt runs it before each round (deadline = start + ceiling) and stops with 'stopped: deadline (…)'. wf-pilot launches with `--left <time left>` (was `--for <time left, ≥ 1h>`: a pilot of a few hours closed ~1h early). Projects: nothing. +- 2026-10-06 `wf res` reservations from history: finished runs (rc=0, peak) go to `~/.local/state/wf/resources-history.jsonl` when pruned from the ledger; new `wf res hist [--project P]` = per project and title kind (title minus trailing sha/hex words and `(…)`) n, request, peak p50/p95, estimate, duration p50/p90, suggestion (≥ 3 runs: p95 peak ×1.15 ≥ 0.2 GB, p90 duration ×1.5 ≥ 5 min); `run`/`note` asking > 2× it print one `hint: history says ~X GB / Y min (n runs)` line on stderr (no auto-resize); `wf batch` default `--mem`/`--for` = the project's `wf-batch` suggestion when it exists (ceiling stays 6h; not `--prep`). Projects: nothing. +- 2026-10-06 `wf batch` cloud tail: when the orchestrator exits with `.wf/cloud` records left, the batch sidecar keeps pulling every 10 min (`wf orch post … --no-pick`, no picks/sends) until none remain, `out/wf-batch.stop` or 24 h (`batch-sidecar --cap`); the job's ledger entry shrinks to 0.2 GB / 1 cpu (via `WF_RES_ID`; unit MemoryMax unchanged); summary line `cloud tail ended: <why>; <n> pulls; left: …`. Cause: tasks sent at batch end stayed unpulled. Projects: nothing. +- 2026-10-06 `wf batch` on a `cloud = true` project pulls cloud tasks itself: the job runs `wf_res.py batch-sidecar` around `claude -p`, every 10 min while it lives (until the stop file) `wf cloud pull --all` → per ended task `wf orch post <id> cloud --result <state> [--commit <sha>] --no-pick` + a `(sidecar pull)` summary line; batch prompt: the orchestrator only picks cloud, never pulls (pulls starved while it blocked on foreground workers). `wf cloud pull`: flock on the task record, a second concurrent pull of one id prints `<id>: pulled by another wf cloud pull, skipped` (rc 0) instead of `branch … exists (local WIP?)`. Projects: nothing. +- 2026-10-06 Cloud routing per task: `wf add --cloud yes|no` (like `wf set --cloud`); `Cloud: yes` still needs Model opus (sonnet/haiku never fit, was: yes overrode Model) and skips only the regex list; `wf orch pick cloud` order = `Cloud: yes` first, then prio, then larger effort (was: opus first). shared/CLAUDE.md "Lanes and models": Cloud bullet (classify at add); wf-worker follow-ups get `--cloud` in cloud projects. Projects: with `cloud = true`, mark runner-ready opus tasks `Cloud: yes|no`. +- 2026-10-06 `wf finish` no longer half-finishes: a PATH outside the repo (e.g. a tool/other-project file), missing, unchanged (lane worktree) or the worktree's own TASKS/archive copy is refused before done; an id already archived resumes (commit + merge) instead of `unknown id`; a git add/commit failure after done says rerun resumes. `wf done`/`wf finish` in the main tree refuse a task `in progress: <branch>` checked out in a linked worktree (run it there, or `wf status <id> clear` first). wf-worker.md: paths = worktree files only, never hand-edit/commit the books. Cause: workers passed other-repo paths → done ran, commit failed, retry `unknown id` → books hand-edited on the branch / left uncommitted in the main tree. Projects: nothing. +- 2026-10-06 `wf res run --lock KEY`: one queued/running job per KEY and main tree (lane worktrees share it); titles starting with `gate` lock `gate` unasked, so overlapping gates in one shared gate checkout are refused (exit 3 `busy: lock 'gate' held by r-N …`, also with --force) or wait with --queue (queued lock-blocked entries are skipped by the FIFO, not blocking others). Ledger field `lock`. The 'ambiguous argument %s' in a-game-project-builder gate logs r-671/r-672 = nested `\"%h %s\"` quoting in a hand-written bash -c copy of gate-bg.sh (no wf/systemd expansion; test pins verbatim argv). Projects: nothing (gate scripts titled `gate …` get the lock). +- 2026-10-06 `wf cloud pull` archives each session it ends (done/awaiting/handback/lost) in the claude.ai app via the CLI's internal `POST /v1/code/sessions/<sid>/archive` (login token read in-process, never printed); failure → one `archive <sid> failed: …; archive by hand (wf cloud archive --ended)` line, pull's exit code unchanged. New `wf cloud archive <sid>|--ended`. Ledger rows get `archived: true`. Projects: nothing. +- 2026-10-06 `wf cloud pull`: no `● WF-RESULT` but the last assistant message ends in a bare `WF-PATCH-END` (session dropped the key words) → handback `bad result: no WF-RESULT key` (usage charged, patch kept when a `sha256=… bytes=…` header decodes) instead of running until 24h lost. Projects: nothing. +- 2026-10-06 Cloud prompt: final-message lines written literally (`WF-RESULT <…>`, …; the `[WF-RESULT R]` form made a live session drop every key word → pull saw no result), the usage script prints the `WF-USAGE` line and the patch command prints the `WF-PATCH-BEGIN sha256=… bytes=…` / base64 / `WF-PATCH-END` block to copy; commit only needed paths (no __pycache__/bin/obj). `wf cloud send --dry-run` snapshots into `out/cloud/<id>.dry-run`: it no longer deletes the live snapshot of a sent task. Projects: nothing. +- 2026-10-06 `wf orch pick cloud [--id ID]`: virtual lane cloud = stop file, ledger check (`stop lane cloud: ledger|max parallel (…)`), first task across all lanes that is runner-ready, not in progress and fits the cloud rules (opus first; none → `stop lane cloud: none fit (…)`) → `wf cloud send`, orch record lane cloud (task_lane, branch); `wf orch post <id> cloud --result <state>` after `wf cloud pull --all` (branch check uses the recorded `<task lane>/<id>`). Local lanes already skip `in progress: cloud:*`. docs/orchestrator.md "Cloud lane", wf-orchestrate skill. Projects: opt in with `cloud = true`. +- 2026-10-06 `wf cloud pull <id> | --all [--export FILE] [--no-push]`: teleport + /export, parse the final `● WF-RESULT` message (wrapped header/base64 rejoined, sha256 + bytes checked, gunzip); done → patch paths checked (outside the project / TASKS / archive / .wf / .worktrees / out / cloud_include refused), `git am --3way` on `<lane>/<id>` from the recorded master in a free lane worktree, worktree_setup, quick_gate, wf done -m "<WF-REPORT> (cloud <sid>)" + wf merge, `report: commit <sha>`; awaiting → wf add -s awaiting + blocked; handback / red gate / bad patch → note `Recovery: cloud attempt <sid> — <why>`, status clear, patch kept in out/cloud/<id>.patch; no result → running (exit 4), > 24h → lost. Every end charges the ledger and deletes snapshot + record; branch exists / worktree setup failed → exit 1, session kept. `wf.lane_worktree` shared with orch pick. Projects: nothing. +- 2026-10-06 `wf cloud send <id> [--dry-run] [--project DIR]`: snapshot of the project's master (git archive minus TASKS/archive/.wf/.worktrees/out, plus `cloud_include`) in `out/cloud/<id>` as a one-commit repo, 90 MB packed cap, prompt `templates/cloud-prompt.md` (task + area notes + `cloud_note`, WF-RESULT/REPORT/USAGE/PATCH final-message rules, usage script), ledger reserve (refused: exit 3), `claude --cloud`, claim `in progress: cloud:<sid>`, record `.wf/cloud/<id>.json`. `cloud.final_message(export)` = the marker parser (only an assistant `● WF-RESULT` message). New workflow.toml keys `cloud`, `cloud_include`, `cloud_note`. Projects: nothing (opt in with `cloud = true`). +- 2026-10-06 Cloud pty driver (library, no command yet): `wf_cloud.send(folder, prompt) -> sid` runs `claude --cloud` on a pty (pre-accepts the trust dialog in ~/.claude.json, Down+Enter fallback), `wf_cloud.export(folder, sid, out) -> path` = `--teleport` → `/export` (Enter retries) → `/exit`, teleport retried once; `cloud.strip_ansi/parse_sid/trust` pure. Env `WF_CLAUDE`, `WF_CLAUDE_JSON`. Projects: nothing. +- 2026-10-05 `wf wip <id> -m NOTE [--commit MSG] [PATHS]` (lane worktree): wrap-up in one call = commit paths, note, status clear; wf-worker prompt uses it. Projects: nothing. +- 2026-10-05 Lane worktree config: wf run inside a linked worktree whose copy of the project has a workflow.toml reads `worktree_setup`, `quick_gate`, `areas`, `area_stale_commits`, `area_ignore` from that copy, and the areas file (default CLAUDE.md) resolves there: a branch can test its setup/gate/area edits before merge; `wf areas --mark` writes the worktree copy (commit it on the branch). Tasks/archive/ledgers/other keys stay the main tree's. New `config.load_at` / `find_local`. Projects: nothing. +- 2026-10-05 `wf orch pick <lane> [--id ID] [--recovery WHY]` (stop file, runner pick skipping in-progress, claim, free worktree `.worktrees/<lane>`/-2… (no live session cwd, no other picked worker, clean detached or own branch), prints the wf-worker prompt; record `.wf/orch/<id>.json`) and `wf orch post <id> <lane> --result LINE [--commit] [--agent] [--duration] [--no-pick]` (post-check: archive line, branch gone, worktree clean + in master else wf merge, wf check; leftover commit otherwise; out/wf-orch.log + wf-cost.log lines; model-raised detection; next pick or `stop lane`). wf-orchestrate 'One worker', docs/orchestrator.md, batch prompt use them. Projects: nothing. +- 2026-10-05 `wf res` ledger fix: entries vanished (ledger moved to .bad, restarted empty at r-1) when a process running older code (a `wf res wait` started before the `by`-field release) read entries with a key it did not know. Unknown entry keys are now ignored; a corrupt ledger restarts with its salvaged id counter; starting a job removes stale logs/<id>.rc/.peak. Projects: nothing. +- 2026-10-05 wf finish prints `report: commit <sha> [tool <sha>]` (+ --tool-commit); wf-worker.md Report says copy it. Projects: nothing. +- 2026-10-05 `wf ctx <id>` ends with the notes (code map, test recipe…) of every area the task names (area name, code-map anchor or Paths entry, whole word). workflow.toml `code_root`: areas' anchors / Paths / Checked refer to another repo (wf areas, --mark, check, stale t-map tasks; no uncovered nudge). Projects: nothing. +- 2026-10-05 /wf-orchestrate-loop retired (duplicate of /wf-pilot): the skill only points to /wf-pilot for one release, then is removed; docs/orchestrator.md section dropped. `.gitattributes`: CHANGES.md merge=union (parallel releases no longer conflict on the top line). Projects: nothing. +- 2026-10-05 `wf start <id> --worktree PATH --branch B [--recovery]` (main tree): worker setup in one call = worktree add/switch (dirty refused; --recovery keeps a dirty WIP on B and prints it with master..HEAD), worktree_setup, status progress B, wf ctx, a ready `cd wt && (verify) && wf finish …` line. wf-worker.md Start section → that one call. Projects: nothing. +- 2026-10-05 shared/CLAUDE.md trimmed 7.1 → 4.9 KB (same rules, less wording; TASKS item format left to wf check; new line: batch independent tool calls). Projects: nothing. +- 2026-10-05 `wf res` entries name their owner: ledger field `by` (name: `--by` on run/note, else WF_SESSION_NAME, else `<lane> session` from .wf/sessions by CLAUDE_PID; task: WF_TASK, else `<lane>/<id>` branch; batch: WF_RES_ID, now set in every job unit; address: uds:<messaging socket>). `wf res status` lines end `[by … , message uds:…]`, throttle warning ends `; started by …`. Projects: nothing. +- 2026-10-05 `wf res run/note --force`: fit against MemAvailable − user reserve − headroom only (ledger claims + queue ignored; entry still recorded); busy line ends with '; X GB really free beyond the reserve: --force starts it past the ledger (…)' when --force would fit; refused → 'busy even with --force: …' exit 3. `wf batch --prep` / `wf prep` default --mem 1G (was 4-8G). Projects: nothing. +- 2026-10-05 Batch stop at once: batch prompt, wf-orchestrate-loop, wf-orchestrate 'One worker' check `out/wf-batch.stop` before EVERY spawn (not per round; running workers finish, latency ≤ the running tasks); batch prompt gets the absolute stop path; `wf lanes --wait` exits 2 ('stop requested') as soon as the stop file exists; pilot/loop treat exit 2 as stop. Projects: nothing. +- 2026-10-05 Uncovered-area nudge: when a project's areas use `- Paths:`, `wf done` adds `t-map-<folder>` (≤3, once per id, also not when archived) for folders (≤2 levels) the task's diff (branch since merge-base with master/main + uncommitted + untracked; bookkeeping files skipped) touches outside every area's Paths (an area without Paths: folders of files holding its Code map anchors); `wf areas` prints `uncovered: <folder> (n files)` lines. New workflow.toml `area_ignore` (default tests/test/docs/doc; dot folders and root files always skipped). Projects: nothing (optional area_ignore). +- 2026-10-05 `wf finish <id> -m ENTRY [--commit MSG PATH…]` = quick_gate → done → commit explicit paths → wf merge in one call (main tree: TASKS/archive join the commit); refuses before done on gate red / stray uncommitted files. `wf ctx <id>` prints a `Verify (run before wf finish):` line. wf-worker.md Done uses finish. `wf usage --explore`: `git worktree list` / `git branch --list` no longer count as bookkeeping. Projects: nothing. +- 2026-10-05 Prep batch: `wf batch K --prep [--lanes …]` = headless run (templates/prep-prompt.md) where sonnet subagents write Done (+Model) for up to K pending tasks without Done (not owner-bound, no status/open slices; pickable first); unclear → awaiting + blocked; summary lines `id → Done: …` | `id → awaiting a-…`; none → 'nothing started'. New `wf set <id> --done TEXT` ("" removes). wf-pilot / wf-orchestrate-loop run a prep batch when idle with not-runner-ready tasks (once per idle spell). Projects: nothing. +- 2026-10-05 wf-pilot / wf-orchestrate-loop: every PushNotification also logged as 'ALERT ...' line in out/wf-orch.log + summary; docs/orchestrator.md names the push setting (agentPushNotifEnabled). Projects: nothing. +- 2026-10-05 `wf batch` writes the `out/wf-batch-*.md` "# wf-batch <project> <time> (N, lanes)" header itself before launch (also --dry-run); prompt only appends. Projects: nothing. +- 2026-10-05 `wf merge` on a detached-HEAD lane worktree with commits not in master now rebases + ff-merges them (message 'bookkeeping', output 'merged detached HEAD'), never only the bookkeeping; dirty worktree refused in both modes; `wf done` there warns 'detached HEAD has N commits not in master'. Orchestrator post-check (wf-orchestrate, docs/orchestrator.md): worktree HEAD ancestor of master. Projects: nothing. (`wf check` unaffected.) +- 2026-10-05 `wf usage --log ... --duration <s>` writes `dur=` (agent wall time); `wf usage --report` adds med_dur column; wf-orchestrate / loop pass duration_ms/1000. Projects: nothing. +- 2026-10-05 Stopped lane (handback/awaiting) resumes only after the task/fix that stopped it is done, or the owner says so: wf-orchestrate-loop, wf-pilot, batch prompt text. Projects: nothing. +- 2026-10-05 `wf add --done "<text>"` writes the Done line (follow-ups born runner-ready); `wf add -p 0` without Done warns on stderr naming --done; wf-worker adds follow-ups / gate-fix tasks with --done. Projects: nothing. +- 2026-10-05 `wf merge` run again in the (detached) worktree after a merge commits TASKS.md/archive changes made since (late notes, follow-ups; message 'bookkeeping'); nothing dirty → still 'nothing to merge'. Workers: run `wf merge` once more after post-merge `wf add`/`note`. Projects: nothing. +- 2026-10-05 Gate-red outcome: new optional workflow.toml `quick_gate = [...]` + `wf gate` (runs it in the current tree, red = exit 1); wf-worker runs `wf gate` before `wf done`. Async slow gate red not caused by the worker's diff → worker adds a runner-ready P0 fix task and reports `done+gate-red <culprit> <fix-id>`; orchestrate / loop / batch prompt treat it as done and pick the fix next (lane not stopped); `wf usage --report` counts it as done. Projects: optional — set `quick_gate` to a fast regression check (e.g. contract tests). +- 2026-10-05 `wf lanes` / `lanes --wait` count runner-pickable (= `list --runner --lane`: needs a usable Done); not-ready ones shown as "· N not runner-ready (no Done)". `wf next` lanes block unchanged. Projects: nothing. +- 2026-10-05 New skill `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]`: nested overnight pilot running `wf batch K` in a loop (wait via `wf res wait`, idle-poll `wf lanes --wait`, stop via `wf batch --stop`, PushNotification on stop/awaiting/end); docs/orchestrator.md 'Pilot (nested)'. Install: `ln -s /projects/workflow/shared/skills/wf-pilot ~/.claude/skills/`. Projects: nothing. +- 2026-10-05 `wf batch --stop`: creates out/wf-batch.stop; batch checks it before each round, stops ("stopped: stop file"), removes it; `--status` shows it. Projects: nothing. (`wf check` unaffected: no change to check code.) +- 2026-10-05 New skill `/wf-orchestrate-loop [N] [--lanes a,b] [--for 10h]` (replaces the `loop N` argument of wf-orchestrate, same release day): piloted loop, typed into any idle project session (e.g. phone Remote Control); docs/orchestrator.md 'Piloted loop session'. Install: `ln -s /projects/workflow/shared/skills/wf-orchestrate-loop ~/.claude/skills/`. Projects: nothing. +- 2026-10-05 `/wf-orchestrate loop N [--lanes a,b] [--for 10h]`: piloted overnight loop mode (background workers, never asks, idle-poll via `wf lanes --wait`, summary in out/wf-batch-*.md); docs/orchestrator.md 'Piloted batch session'. Projects: nothing (skill symlink picks it up at next session start). +- 2026-10-05 `wf lanes --wait SECS`: block until any lane has a pickable task (exit 0) or timeout (exit 1); prints the lanes line either way. Projects: nothing. +- 2026-10-05 `wf body -h` names stdin, kept lines (Model/Sessions/After/Ref), heredoc example; `wf body ID text` errors with stdin hint. Projects: nothing. +- 2026-10-05 `wf list --runner` no longer lists sliced-out parents (effort > slice_above, slices all archived); row set + matches `wf lanes` pickable. Projects: nothing. +- 2026-10-05 Model-lane leftovers removed: `wf batch --lanes <model>` warning (a model name is now just an + unknown lane), old `{haiku,sonnet,opus}.json` registry cleanup. docs/model-lanes.md folded into docs/lanes.md + (now a current reference: Model line §0, registry/claims §4). Projects: nothing. +- 2026-10-05 Speed lanes: lanes by size instead of model (`[lanes.*]` in workflow.toml; default fast = <1h + unblock-first + slice jobs, slow = 1h priority-first, falls back to fast), `slice_above` (default 1h: bigger = + slice job, split never implement), sessions take Model ≤ their own (`wf next --as M [--lane L]`), `wf lanes`, + orchestrator/worker/batch by lane, `wf areas` (code-map anchors, staleness, `--mark`) + `wf done` auto-adds + `t-map-<area>` refresh tasks. Old: `--lanes <model>` warns; old model registry files dropped. Spec: docs/lanes.md. + Projects: rollout — area sections get `Paths:` + first `wf areas --mark`, remove merged model worktrees. +- 2026-10-05 `wf setup` (in a lane worktree) runs new workflow.toml key `worktree_setup = [cmds]` (cwd = worktree, + env `WF_MAIN` = main tree; first non-zero exit stops, exit 1). wf-worker runs it after worktree create/reuse; + `wf next` Multi-session hint appends `&& wf setup` when set. Projects: optional — list commands that provide + git-ignored inputs (e.g. link `out/…` from `$WF_MAIN`); must be idempotent. +- 2026-10-05 `wf init` also seeds CLAUDE.md (templates/CLAUDE.md: Layout · Loop additions · Gotchas · Areas with code map / test recipe / tools / worktree setup); kept if it exists. Projects: nothing. +- 2026-10-05 `wf usage --explore`: per subagent of a session, % billed input exploring / before the first edit, median calls before it, result tokens per tool (grep/Read/sed-cat/…), files read by >= 2 agents. Projects: nothing (replaces per-project session-cost scripts). +- 2026-10-05 re-learning rules: shared CLAUDE.md grep discipline + code-map/test-recipe notes, worker/design/orchestrate lines; `wf done` always prints a checklist line "re-learned anything…". Projects: keep a code map + test recipe per area in their notes. +- 2026-10-05 `wf usage --effort` accepts only task estimates (<1h|1h|5h|10h|100h), not reasoning effort. + Projects: nothing; fix old cost-log lines with effort=medium by hand. +- 2026-10-05 `wf usage`: subagent output tokens were undercounted (transcripts keep only the stream-start + placeholder); now estimated from content (fit ±5% per session), shown `~`, cost-log lines get `est=N`. + Projects: nothing; older `out/wf-cost.log` lines undercount out/usd. +- 2026-10-05 `wf check` warns about tool worktrees (wf repo `.worktrees/*`) whose branch is merged into master (leftover after release; `WF_TOOL_ROOT` overrides the repo). Projects: nothing. +- 2026-10-05 `wf lanes --unregister` drops this session's lane record (orchestrator holds none); wf-orchestrate cold start runs it. Projects: nothing. +- 2026-10-05 `wf list --runner [--model M]` = ready + Done line, no Sessions: owner, Done not about owner/report to/confirm, + no open slices. `wf add` of a task without Done prints a stderr hint. Projects: nothing. Orchestrator uses --runner. +- 2026-10-05 `wf list --stale` lists in-progress tasks with no live claim (dead holder or none); `wf status --clear-stale` + clears them. Projects: nothing. Orchestrator cold start may use it. +- 2026-10-05 `wf next` / `wf done`: last line `context ~Nk tokens (> 100k): ask the owner to /clear, then continue` + when this session's last prompt is over `ctx_hint` (workflow.toml, tokens, default 100000, 0 = off). + Projects: nothing (optional key). +- 2026-10-05 `wf batch N [--lanes …] [--model] [--for] [--mem] [--dry-run]` starts the unattended batch + orchestrator (claude -p under `wf res run`, bg-wait ceiling = --for, prompt in templates/batch-prompt.md); + `wf batch --status`. Owner: allow rule `Bash(python3 /projects/workflow/wf.py batch:*)` once. Projects: nothing. +- 2026-10-05 docs: `/clear` safe only with no background worker (it SIGKILLs the worker's running Bash; + notice still arrives). orchestrator.md Context + wf-orchestrate skill. Projects: nothing. +- 2026-10-05 `wf usage`: tokens + API-price $ per agent of a session (`--session`, `--agent`, `--since`), + deduped per request (transcripts repeat usage per content block: earlier 153 turns / 8.9M cr = 86 / 5.0M); + `--log TASK OUTCOME --effort E` appends to `out/wf-cost.log`; `--report` = lane/effort table over all + projects. Orchestrator skill step 6 logs the cost line. Projects: nothing to do. + +- 2026-10-04 shared CLAUDE.md compacted (orchestrator → `/wf-orchestrate`, reflowed); no rule changes. + +- 2026-10-04 `wf add --parent` no longer chains After to a deferred slice; rules: never delete what you did + not create, owner-done parts / approved "(brainstorm)" titles updated same turn, post-merge gates run from + the main tree, batch orchestrator wf-adds follow-ups instead of asking. + +- 2026-10-04 Pilot 4 fixes: wf-worker 4-line report with `followups` (loose ends → `wf add`), background + gates run after `wf merge` on the merged sha, `Recovery:` prompt mode, wf spelled out; spawn prompt ends + with "Final message: the 4 report lines only."; status-note parentheses error has a hint. + +- 2026-10-04 Skills `/wf-orchestrate [go | batch N]` and `/wf-design [topic | id]` (`shared/skills/`, symlink + into `~/.claude/skills/`): cold start of an orchestrator or design/rulings session after /clear. + +- 2026-10-04 Orchestrator + workers (docs/orchestrator.md): project lock `.wf/lock` around every wf write + (parallel sessions/workers safe); `wf merge` in a lane worktree = rebase, ff-merge, TASKS/archive commit, + detach, push home, under the lock (`wf done` there now points to it); `wf-worker` subagent definition + `shared/agents/wf-worker.md` (install: symlink into `~/.claude/agents/`). Shared rule: an orchestrator session + holds no lane and spawns workers. Projects: `out/` git-ignored for the orchestrator log. + +- 2026-10-04 Sessions field: task line `Sessions: parallel|solo|owner [— why]` (none = parallel; `- Sessions:` + note form works too; `wf add/set --sessions`, `wf list` mark `[solo]`/`[owner]`, `wf check` validates). + `wf next` skips solo while another session is live; a solo task in progress → other sessions pick nothing; + its `wf done` notifies them. Owner tasks only with `wf next --owner`. Header flag `, interactive` = owner + (`wf check` warns: `wf set <id> --sessions owner`); `--interactive` = old spelling, one more release. + Projects: nothing required; replace `interactive` flags when `wf check` warns. + +- 2026-10-04 Multi-session (docs/multi-session.md): `wf status … progress` claims the task for the session + (`.wf/claims/`); `wf next` skips tasks held by another live session and warns when another live session holds + your lane. More than one live session → `wf next` prints a `Multi-session` block: work in `.worktrees/<lane>` + on a branch `<lane>/<task>`. `wf` run in a linked git worktree uses the main tree's project (TASKS.md, archive, + docs, `.wf`). `wf done` there prints the merge-back steps; ff-merging one's own task branch is allowed. + Shared rule: commit explicit paths, never `git add -A` / `commit -a`. Projects: `.worktrees/` git-ignored. +- 2026-10-04 Model lanes (docs/model-lanes.md): task line `Model: haiku|sonnet|opus` (none = opus; + `wf add/set --model`, `wf list` column + `--model`, `wf check` validates). `wf next --as <model>` picks only + that lane (no `--as` = haiku, with a hint), registers the session in `.wf/sessions/` (self-ignored) and shows + other lanes' sessions and what waits on them; new `wf lanes`; `wf done` names the lanes to notify. Pick order + for everyone: inherited priority (a task counts at the best P of all that wait on it), then work other lanes + wait on, then more waiting. Shared rule: cold start `wf next --as <your model>`, own lane only. `wf res run` + keeps the caller's environment (variables the user manager lacks; secret-named ones are not stored). + Projects: flag tasks with `wf set <id> --model …` (plain `- Model: sonnet ok` notes → check warning). +- 2026-10-03 shared rule: agents may push to remote `home` (the owner's private server) after a Done commit or + a group of commits; every other remote still only when asked. Projects: nothing (repos without `home` unchanged). +- 2026-10-02 `wf res`: each Claude session in agents.slice is capped at `session_mem_gb` (default 6 GB, + MemoryHigh: throttled, not killed) by the timer tick; `status` warns when a session sits at its cap. A runaway + foreground test can no longer take all RAM. Projects: big work via `wf res run` (as before). +- 2026-10-02 `wf res` automatic clean also sweeps `/tmp` and `/var/tmp` litter: own empty dirs idle 2 h, .NET debug + pipes of dead processes (thousands accumulated). Files with content are untouched. Projects: tests should + still delete their temp dirs. +- 2026-10-02 `wf res run --queue` says how to cancel (`cancel: wf res release r-N`); `release -h` too. Projects: nothing to do. +- 2026-10-02 `wf res` clean: never deletes the scratch dir of a running Claude session, however idle (was: after + 2 h, taking its scratchpad and background-task output). Projects: nothing to do. +- 2026-10-01 Shared rule: `/tmp` is RAM; temp dirs made by code/tests must be deleted after use. Machine tip: + `TMPDIR=/var/tmp` (disk) in the Claude Code settings `env` moves agents' temp files off RAM. Projects: nothing to do. +- 2026-10-01 `wf res hook`: Claude Code PreToolUse hook; while game mode is on, agent Bash commands run without + a display (no windows over the game). Shared rule line added. Machine: add the hook to + `~/.claude/settings.json` (docs/resource-ledger.md §4.6). Projects: nothing to do. +- 2026-10-01 tests: the `wf res` lock race test fails instead of hanging when a child crashes. Projects: nothing to do. +- 2026-10-01 `wf res status`: `unreserved agent memory` says how much of it is reclaimable file cache. + Projects: nothing to do. +- 2026-10-01 `wf res status` warns about jobs started with bare `systemd-run` (transient services outside + agents.slice); shared rule: also not from project scripts. Projects: a script that calls `systemd-run` + itself → start it through `wf res run`. +- 2026-10-01 `wf res`: `run`, `note`, `status` and `-h` say what `--for` means: a job is never killed when it + runs over (`ETA overdue (still running, not killed)`); a note is freed. Projects: nothing to do. +- 2026-10-01 `wf res status`/`wait`: warn when a running job is stuck at its own memory limit (≥ 85 % of + `--mem` and ≥ 20 % memory stall) — `r-N throttled at its memory limit …: likely too small`. Projects: nothing to do. +- 2026-10-01 `wf res`: a job killed with its wrapper (OOM, signal, timeout) now says why in `status`/`wait`, + e.g. `done rc=? peak 3.9 GB … (killed: oom-kill by systemd-oomd, limit 4.0 GB; raise --mem)`, read from + the journal; was `rc=? peak ?`. Projects: nothing to do. +- 2026-10-01 `wf projects` skips a wf checkout (its `templates/workflow.toml` is no project). Projects: nothing to do. +- 2026-10-01 Repo split: the tool repo holds only the tool (fresh history); the workflow's own tasks + and history moved to a separate private wf project. Reference docs: `docs/design.md`, + `docs/resource-ledger.md`. `wf report` / inbox unchanged. Projects: nothing to do. +- 2026-10-01 `wf res`: background Claude sessions (process named by version, e.g. `2.1.283`) are now + recognised: `adopt`/timer move them into agents.slice, and their `note`s are owned by the session, not + its shell. Projects: nothing to do. +- 2026-10-01 `wf res`: shared memory/CPU ledger for agent jobs — `run` (exit 3 = busy, with ETA; `--queue`), + `status`, `wait`, `release`, `note`, `game on|off` (gaming reserve, auto-off 4h), `clean` (stale + /tmp/claude-* scratch auto-deleted after 2h), `adopt` (running Claude sessions → agents.slice, no restart), + `timer on|off`, `shell-init`. Rule in shared CLAUDE.md: big/long jobs go through `wf res run`. + Projects: nothing to change; move ad-hoc systemd-run jobs to `wf res run`. +- 2026-10-01 `wf check` errors on bare ids in `After:` / `Slices:` (were silently ignored, so + `next` could pick a task before its dependency). Project: write them as `[[id]]`. +- 2026-09-30 Inbox fixes: `wf rename OLD NEW` (id + all links); `add` takes a leading `<id>: ` + as the id; `--parent --id` slices chain `After:` via `Slices:` and count for `done`/`next`; + `next` skips parents with open slices; `set --title "Title. Goal."` / question replaces the text. + Projects: nothing to do. +- 2026-09-29 First release: `wf` (next list show ctx search log projects check add done prio + move status set note body tick report init migrate), shared CLAUDE.md, TASKS format 1. + Projects: `wf init` or `wf migrate --write`, pointer line in CLAUDE.md. diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..d588d7f --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,30 @@ +# CLAUDE.md — workflow tool repo (sessions changing wf) + +This repo = the tool: `shared/CLAUDE.md` (symlinked as `/projects/CLAUDE.md`, loaded by every agent) ++ `wf.py` / `wf_res.py` / `wflib/`. Reference: `docs/design.md`, `docs/resource-ledger.md`. +No personal data here (project names, tasks, paths under home): the workflow's own tasks live in a +separate private wf project; publishable as is. + +## Everything here is live +Workers run the main tree's `master` the moment it changes. So: +- `wf.py`, `wf_res.py`, `wflib/`, `shared/CLAUDE.md`, `templates/` → git worktree + `.worktrees/<topic>` on a branch, never the main tree. +- Release = fast-forward merge to `master`, only after: + 1. `python3 -m unittest discover -s tests` green in the worktree; + 2. `wf check` run with the new code in every project of `wf projects` gives the same + result as the old code (or only the intended differences); + 3. a line on top of `CHANGES.md`: date, what changed, what a project must do (if anything). +- The user sees each release. Push only to remote `home` (private), never elsewhere. + +## Compatibility +- Commands and options are only added. Rename/remove → old spelling works one more release + and prints a one-line notice. +- TASKS.md format change → raise `FORMAT` (`wflib/config.py`) + `wf migrate` step; older + projects keep read commands, write commands refuse until migrated. +- `shared/CLAUDE.md` reaches a worker at its next session start: a rule change must not make + files written under the old rule invalid. Keep it short; project needs go to the project. + +## Code rules +- Python ≥ 3.11 stdlib only. `wflib/` pure (text in, text out); files and printing in `wf.py` / `wf_res.py`. +- TDD; oracles = hand-written literals; see each new test fail first. +- Expected failures: one `wf: …` line on stderr, exit 1; usage errors exit 2; no tracebacks. diff --git a/CREDITS.md b/CREDITS.md new file mode 100644 index 0000000..cde46d7 --- /dev/null +++ b/CREDITS.md @@ -0,0 +1,8 @@ +# Credits + +| Kind | Who | Note | +|---|---|---| +| Author | godosa | | +| Libraries | none (Python standard library only) | | +| Origin | the author's earlier per-project task scripts | unified here | +| Tooling | AI assistance (Claude, Anthropic) | engineering tool; statement in README | @@ -0,0 +1,21 @@ +MIT License + +Copyright (c) 2026 godosa + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. diff --git a/README.md b/README.md new file mode 100644 index 0000000..421b856 --- /dev/null +++ b/README.md @@ -0,0 +1,156 @@ +# wf — a shared workflow for many coding agents on one machine + +`wf` is the task workflow I use to run several long-lived Claude Code sessions in parallel, each +working in its own project under one folder (`/projects`). How to set it up and use it: +[docs/manual.md](docs/manual.md). It has three parts: + +1. **One rules file for every agent.** `shared/CLAUDE.md` sets out the loop each agent follows: + cold start, test-first implementation, one commit per finished task, when to ask and when to + decide, and how to report problems with the workflow itself. Claude Code loads `CLAUDE.md` from + the working directory and every parent, so the file is symlinked once as `/projects/CLAUDE.md` + and every project gets it. A project's own `CLAUDE.md` holds only project specifics and wins on + conflict. +2. **A task tool, `wf`.** Each project keeps its open work in a `TASKS.md` with stable ids. Agents + never read or rewrite that file whole; they use one short command per operation (`wf next`, + `wf add`, `wf done`, `wf ctx`, …), and `wf check` makes sure no reference rots. +3. **A memory/CPU ledger, `wf res`.** Agents reserve memory and CPUs before a big or long job and + get a clear "busy, retry after HH:MM" instead of out-of-memory kills. The owner keeps a + guaranteed reserve, and a larger one in gaming mode. Built on systemd user units and cgroups. + +> **Status: a personal setup, published as a reference.** It is shaped around one person's +> machine and habits: Linux, systemd, Claude Code, one `/projects` folder. It is provided as is. +> Setup and daily use: [docs/manual.md](docs/manual.md). There is no support and no plan for +> Windows or macOS. Take whatever ideas or code are useful to you. + +## Why + +With one agent, a task list in a markdown file works fine. With five or more agents running for +days, things go wrong: + +- Each project grows its own task scripts and workflow prose, and they drift apart. +- Agents spend context reading and rewriting long task files, renumber items, and break + references ("see task 14"). +- A workflow fix has to be copied into every project. +- Two agents start a 10 GB build at the same time and one of them is killed, or the desktop + freezes. + +`wf` answers these with one rules file, a task format with stable ids, a CLI that makes every +routine edit one command, and a ledger that rations memory and CPU. + +## The loop (shared/CLAUDE.md) + +Short version. The file itself is the authority. + +- **Cold start** ("continue"): `wf next` shows the questions waiting for the owner, items that + need a human, plans in flight, and the next pickable task with its context. The agent works on + it without asking. +- **Task kinds**: an implementation slice (at most about 1 h) is test red → implement → verify + green. Design work goes through brainstorming → spec → owner approval → `wf add` slices. Bugs: + evidence → failing test → fix. +- **Autonomy**: decide, record the decision in the spec or task, report it. Ask only at real + forks. Blocked on the owner → `wf add -s awaiting "<question>"` and move on to another task. +- **Done**: `wf done <id> -m "<entry>"` archives the task and prints the project's checklist, + `wf check` must report 0 errors, then one commit. Never merge or push unasked. +- **Resources**: a job over 2 GB RAM, over 4 cores or over 10 minutes goes through + `wf res run`. Exit 3 means busy: the agent notes it and works on something else. +- **Workflow problems**: `wf report "<what>"` appends to a shared inbox. The agent works around + the problem and keeps going. The owner's workflow session triages the inbox later. + +## TASKS.md + +```markdown +# Tasks — myproject + +## Awaiting your decision + +- **a-save-format**: JSON or binary saves? Blocks [[t-save-load]]. + +## Pending + +- **t-cave-seams** [P1] (<1h): Terrain: cave entrance seams. Close the slit beside the lintel. + - Tests (red first): seam test on the cave fixture. + - After: [[t-terrain-cave-render]] + Ref: DESIGN.md#terrain + +- **t-save-load** [P2] (1h) (blocked: [[a-save-format]]): Save and load. Round-trip the world. + +## Needs human + +- **t-playtest-feel** [P1] (<1h): Play-test feel and tuning. + - [ ] jump height + +## Deferred + +- **t-android-port** [P3] (10h): Android port. +``` + +- Four sections. Items in Needs human are never picked by an agent. +- Ids (`t-…` for tasks, `a-…` for questions) are stable and never reused. Links are written as + `[[id]]`. +- Priority P0–P3, effort from a fixed list (`<1h`, `1h`, `5h`, `10h`, `100h`). Anything over + 1 h is split into slices with `--parent`. +- `After:` lists dependencies. `Ref:` points into the project's docs, down to a heading anchor. + `wf ctx` and `wf next` print the referenced sections, so an agent loads only what it needs. +- Finished tasks become one dated line in `tasks/archive.md`. The full history stays in git. + +## Commands + +Run inside a project (a folder with `workflow.toml`). Every write command accepts `--dry-run` +and refuses a change that would add a `wf check` error. + +| Read | | +|---|---| +| `next [--brief]` | what to work on now, with context | +| `list` · `show` · `ctx` | items, filtered · raw items · an item or doc anchor with everything it links to | +| `search <words>` · `log` | ranked search over tasks, archive and docs · newest archive lines | +| `projects` | every project under the root: counts, next task, check errors | +| `check` | validate format, ids, links, refs and anchors; exit 1 on any error | + +| Write | | +|---|---| +| `add` · `done` | new item (id from the title or `--id`; `--parent` for slices) · archive finished tasks | +| `prio` · `move` · `status` · `set` | reprioritise · move or reorder · in progress / blocked · header fields | +| `rename` · `note` · `body` · `tick` | change an id and its links · append a line · replace the body · check a box | +| `report` · `init` · `migrate` | file a workflow problem · start a project · convert a numbered task list | + +| `wf res` | | +|---|---| +| `run --mem 10G [--cpus N] --for 40m --title T [--queue\|--force] -- CMD` | reserve and start a job as a systemd user unit; exit 3 = busy (`--force`: fit against really free memory, ignoring ledger claims; the busy line names it when it would fit) | +| `status` · `wait` · `release` · `note` | reservations and budget · wait for a job · drop an entry · reserve for foreground work | +| `game on [--for 4h]` / `off` | raise the owner's reserve; agents slow down, never get killed; with `hook` installed, agent commands get no display | +| `hook` | Claude Code `PreToolUse` hook for `~/.claude/settings.json` (see docs/resource-ledger.md §4.6) | +| `clean` · `adopt` · `timer on/off` · `tick` | stale tmpfs scratch · move running sessions into the agents slice · 1-minute housekeeping | + +`wf -h` and `wf <command> -h` give the details. + +## Design notes + +- **Pure core.** `wflib/` is text in, text out, with no file access. `wf.py` and `wf_res.py` do + the I/O. Python ≥ 3.11 standard library only. +- **Live tool, safe releases.** Every agent runs the main tree of this repo the moment it + changes. Changes are made in a git worktree and fast-forwarded only when the tests pass and + `wf check` gives the same result in every project as before. Commands are only ever added; a + task-format change comes with `wf migrate`. +- **Concurrent writers.** Task files are written whole via temp file + rename, and a write is + refused if the file changed since it was read. `wf report` appends with one `O_APPEND` write. + The resource ledger is a JSON file under a lock. +- **Ledger, not scheduler.** `wf res` keeps no daemon. Every call reads `/proc/meminfo` and the + ledger, admits or refuses, and starts the job as a transient systemd unit with `MemoryMax` + inside `agents.slice`. A one-minute timer starts queued jobs, ends gaming mode and cleans up. + +Full reference: [docs/design.md](docs/design.md) (format, config, every command, release rules) +and [docs/resource-ledger.md](docs/resource-ledger.md). + +## Tests + + python3 -m unittest discover -s tests + +## Licence and credits + +MIT, see [LICENSE](LICENSE). Credits: [CREDITS.md](CREDITS.md). + +## From the author + +These projects are things that have been tumbling about in my head for a long time, and I now feel like trying to do something with. The bulk of my focus here is on some of the games that I love, but every time I boot them up, I just fiddle about in the menu and set up mods and fixes for a couple hours, with my drive to play them fizzling out. This is my hope to solve that, and while I am a professional software engineer this would not have happened were it not for the rise of (relatively) cheap AI that could do the bulk of the work with me. If you don't approve, that's fine. I did these things for me, sharing them is something I do in the hopes to help others in similar situations, that just want to play the games they love. Thank you to all that made these playable in the first place, with some luck this finds you, and can bring some joy. + +Written with AI assistance (Claude, by Anthropic), used as an engineering tool. diff --git a/docs/cloud-lane.md b/docs/cloud-lane.md new file mode 100644 index 0000000..d48c1ed --- /dev/null +++ b/docs/cloud-lane.md @@ -0,0 +1,195 @@ +# wf cloud — cloud worker lane on promo credits — design + +Status: owner decisions 2026-10-06 (design session). +Tool code goes to `/projects/public/workflow` (public tool repo); this spec was originally developed locally, +then a copy `docs/cloud-lane.md` in the tool repo becomes the live one (same as resource-ledger). + +## 1. Goal + +1. Use a promo cloud credit for worker tasks, so the weekly subscription limit goes to design/orchestration/local-only work. +2. No GitHub, no hosting: code goes up as a `claude --cloud` bundle upload from a local folder, results come + back through the session transcript. Push stays local-only. +3. A local ledger keeps cloud spend under a cap; the orchestrator routes fitting work to the cloud only while + the ledger is positive; otherwise everything runs locally as today. +4. Cloud work never touches TASKS.md / archive: bookkeeping stays local (`wf` is not in the VM). + +## 2. Owner decisions (2026-10-06) + +| # | Decision | +|---|----------| +| C1 | No GitHub (no mirror, no GitHub App). Bundle upload only. | +| C2 | Cloud ledger with budget set by owner (balance minus margin against overrun). | +| C3 | The orchestrator may put fitting work in the cloud on its own while the ledger balance is positive; at ≤ 0 it runs locally, no question asked. | +| C4 | Opus in the cloud is fine (most work is opus). | +| C5 | Purpose-built snapshot repos (code + the resource files a task needs) are the unit sent up; no hosted git. | + +<a id="facts"></a> +## 3. Facts (measured 2026-10-06, CLI 2.1.290) + +| # | Fact | +|---|------| +| F1 | `claude --cloud "<prompt>"` in a folder with no git remote uploads a bundle (history + tracked files; untracked and `.env`-like files left out; ≤ 100 MB, else branch-only, else squashed snapshot). It prints `Created cloud session: …`, `View: https://claude.ai/code/session_…`, `Resume with: claude --teleport session_…` and **exits** — scriptable. | +| F2 | `--cloud` needs a TTY (`-p` refused; piped stdout refused). A pty (tmux, python `pty`) works. | +| F3 | A never-seen folder shows the trust dialog first (blocks; "Yes" = Down, Enter). Pre-accepting it (`~/.claude.json` `projects[<path>].hasTrustDialogAccepted = true`, atomic rewrite) works — verified 2026-10-06, no dialog. | +| F4 | Cloud sessions run **Opus 5.5 at effort medium** by default (commit trailer, status line). | +| F5 | The VM can't reach local networks behind a firewall → it can't push to local git. Custom allowlist = claude.ai environment settings (owner action). | +| F6 | Large projects with native build tools may require environment setup; e.g., .NET SDK not in the base image → download from package repos (minutes, cached with environment setup script). | +| F7 | `claude --teleport <sid>` for a bundle session: validates, fetches logs, then **fails silently at "Checking out branch"** (branch only in the VM) — sometimes exits, sometimes still resumes the conversation. It never applies the branch. Teleporting a *running* session does not interrupt it (local copy shows "Interrupted"; cloud kept working). | +| F8 | After a teleport resumes, `/export <file>` writes the rendered conversation (80-col wrapped; tool output truncated `… +N lines`; **assistant text complete**) with **no model call**. The local `.jsonl` transcript only appears after a local message (= a full-context model call, avoid). | +| F9 | So the result channel is the session's **final assistant text**: a gzip+base64 patch between markers with a sha256, whitespace-stripped on read. A patch inside tool output is truncated and useless. | +| F10 | `claude -p "<msg>" --cloud <sid>` queues a follow-up into a running/idle session and exits (no TTY needed). | +| F12 | The export contains the prompt too: marker parsing must take only assistant lines (`● WF-RESULT …` = bullet prefix) after the prompt, never a bare substring match. | +| F13 | Environment configuration (owner, 2026-10-06): custom network allowlist for required services; environment variables for telemetry and timeouts; setup script for large dependencies. The exact setup depends on project needs and should be configured in claude.ai environment settings. | +| F11 | Cloud sessions share the subscription rate limits only after the promo credit is used up; until then they draw the credit (promo terms). No CLI shows the balance: owner reads it on claude.ai. | + +Baseline local cost (`wf usage --report`, API-price $, 2026-10-06): +opus worker ≈ $1.10–1.32 per done task (median $0.58–0.96, n=64), sonnet ≈ $0.14–0.30, orchestrator +≈ $0.25 per dispatched task. Transcript output counts are estimates. + +## 4. Design + +<a id="snapshot"></a> +### 4.1 Unit of work: snapshot repo + +`wf cloud send <id>` builds `<project>/out/cloud/<id>/` (git-ignored, deleted by `pull` after apply): +- `git archive <master HEAD>` of the project → plain files; plus every path in workflow.toml + `cloud_include` (git-ignored inputs, e.g. large data files), force-added; one commit "base <sha>". +- Packed size > 90 MB → refuse (`wf: snapshot <n> MB > 90 MB; trim cloud_include`). +- Never in the snapshot: `TASKS.md`/archive/`.wf`/`.worktrees`/`out/` except `cloud_include`. +- Pre-accept the trust dialog for that path (F3), then run `claude --cloud "<prompt>"` under a python pty with + a timeout (5 min); parse `session_…` from the output. No session id → exit 1 with the last 5 output lines. + +<a id="prompt"></a> +### 4.2 Prompt (template `templates/cloud-prompt.md`, filled by `send`) + +Fixed text (≈ 25 lines) + `wf show <id>` body + the project CLAUDE.md test recipe of the areas the task names +(`wf ctx` areas section). Rules in the template: +- `wf` is absent: ignore wf/TASKS/lane/gate rules in CLAUDE.md; never edit TASKS.md/archive. +- Never ask; blocked → final text `WF-AWAITING: <one-line question>` and stop. +- Test red → implement → green; commit on `main` (snapshot repo) with one-line messages. +- Network blocked for a needed source → say so in the report line, continue with what the repo has. +- **Final message text** (not a tool call), exactly: + ``` + WF-RESULT done|awaiting|handback + WF-REPORT <≤2-line archive entry> + WF-USAGE in=<n> cw=<n> cr=<n> out=<n> model=<id> (summed from its own transcript, script given in the template) + WF-PATCH-BEGIN sha256=<hex of the gzip> bytes=<n> + <base64 of `git format-patch --binary <base>..HEAD --stdout | gzip -9`> + WF-PATCH-END + ``` +- Built (cloud send task): the template writes each marker line in brackets (`[WF-RESULT R]`) so no prompt line + starts with a marker; `cloud.final_message(export)` = the F12 parser: lines of the last message starting + `● WF-RESULT` (assistant bullet; the prompt renders as `❯ `/` `, never `● `), bullet/indent stripped, + ends at the next non-indented line. `pull` parses fields from that list only. Snapshot record + `.wf/cloud/<id>.json`: id, sid, project, lane, master (sha for `git am`), base (snapshot commit = patch base), + folder, sent, model, bytes. Ledger row reserved before `claude --cloud` (sid `pending:…`), removed on failure. + +<a id="pull"></a> +### 4.3 Return: `wf cloud pull <id> | --all` + +- Teleport under a pty in the snapshot folder, `/export <tmp>`, `/exit` (F7, F8); no model call. +- Parse the export: last line starting `● WF-RESULT` (F12); none → `running` (exit 4, prints age since send). + None, but the last assistant message (no non-empty prompt after it) ends in a bare `WF-PATCH-END` line + (key words dropped) → handback `bad result: no WF-RESULT key`, patch kept when a `sha256=… bytes=…` header decodes. +- `done`: join base64 between markers (strip whitespace and the 2-space indent), check sha256 and bytes, + gunzip → patch. Apply in the task's lane worktree on branch `<lane>/<id>` from the recorded base: + `git am --3way`; files outside the project's tracked tree or in `cloud_include` → refuse. + Then `quick_gate`; green → `wf done <id> -m "<WF-REPORT>"` + commit + `wf merge` (the `wf finish` path, run by + `pull` itself, so a pull costs the orchestrator one Bash call). Red / am conflict → handback (below). +- `awaiting` → `wf add -s awaiting "<q>"` + `wf status <id> blocked <a-id>`. +- `handback` / red gate / bad patch → `wf note <id> "cloud <sid>: <why>"`, `wf status <id> clear`, task goes back + to the local queue with `Recovery: cloud attempt <sid> — <why>` in its note (a local worker may take the patch + from `out/cloud/<id>.patch`, kept on handback). +- Every pull that ends the session charges the ledger (4.4) and deletes the snapshot folder. +- Archive (owner ruling 2026-10-06, built as part of cloud infrastructure): each pull that ends a session then archives it in the + app (`POST https://api.anthropic.com/v1/code/sessions/<sid>/archive`, the CLI's internal archiveRemoteSession, + CLI 2.1.291; 200/409 ok). Failure → one `archive <sid> failed: …` line, pull exit unchanged; ledger row + `archived: true`; `wf cloud archive <sid>|--ended` retries. +- Built (cloud pull task): sections of the final message open at a key line, other lines continue it (export + wraps); the patch header + base64 are joined with all whitespace removed (`sha256=HEXbytes=N` then base64: + gzip base64 starts `H4sI`, never a digit). Refused paths: outside the tree / project folder, TASKS, archive, + `.wf`, `.worktrees`, `out`, `cloud_include`. One handback note `Recovery: cloud attempt <sid> — <why>[; patch …]`. + Lost = no result (or no export) and > 24 h since send. Branch `<lane>/<id>` already there / worktree_setup red → + exit 1, session kept (pull again). `--export FILE` parses a given export (tests, manual). Exit 0 ended, 4 running. + +<a id="ledger"></a> +### 4.4 Ledger: `~/.local/state/wf/cloud.json` (same state dir and lock style as `wf res`) + +``` +{"budget": 240.0, "spent": 0.0, "reserve_per_task": 4.0, "max_parallel": 3, + "entries": [{"id", "project", "sid", "model", "sent", "state": "running|done|awaiting|handback|lost", + "usd": null|float, "usd_source": "self|est|owner"}]} +``` +- balance = budget − spent − reserve_per_task × running. `send` refuses (exit 3, one line) when + balance < reserve_per_task or running ≥ max_parallel. +- Charge on end: `WF-USAGE` × `usage.PRICES` (reuse `wflib/usage.py`), plus a 15 % overhead factor for + what the session can't see (title generation, setup); missing WF-USAGE → charge reserve_per_task, + `usd_source=est`. +- `wf cloud ledger` prints budget/spent/balance/running; `--set-balance <usd from claude.ai>` sets + spent = budget − balance (owner reconcile, `usd_source=owner` row); `--budget N` changes the cap. +- A running entry older than 24 h without `WF-RESULT` → `lost` (charged reserve), task back to local. + +<a id="fit"></a> +### 4.5 What fits the cloud (routing) + +Project opt-in: workflow.toml `cloud = true` (+ `cloud_include = [...]`, optional `cloud_note = "..."` appended +to the prompt) only enables the lane; routing is per task (owner ruling 2026-10-06, amended: opus only). +Whoever adds a task (design session, orchestrator, worker follow-ups) classifies it: `Cloud: yes|no` +(`wf add --cloud` / `wf set --cloud`); yes = Done checkable from the repo alone (code, synthetic data, +`cloud_include` data), no GUI/display, live game/server/capture, LAN, Wine or local installs, `wf`/`wf res` steps, +owner; opus only, prefer larger (1h); unsure → no. Task fits when all hold: +- project opted in; task runner-ready (Done + Model), not `Sessions: owner|solo`, no `Cloud: no` line; +- effort ≤ `slice_above` (no slice jobs); Model opus (sonnet/haiku never, also with `Cloud: yes`: overhead ≈ their cost, §5); +- `Cloud: yes` → fits (regex skipped); no Cloud line → body must not name `wf ` commands as steps + (workflow-upgrade / bookkeeping tasks need `wf`), GUI, LAN hosts, `wf res`, or a live server — checked by a small + regex list in `wflib/cloud.py`. +`wf orch pick cloud` order (`cloud.pick_key`): `Cloud: yes` first, then prio, then effort desc (1h before <1h). +A worker that finds a task unfit for the cloud adds `Cloud: no` (`wf set <id> --cloud no`). + +<a id="orch"></a> +### 4.6 Orchestrator + +`/wf-orchestrate` and `wf batch` gain a virtual lane `cloud`: +- `wf orch pick cloud` = first fitting task across the project's lanes, ledger allows → `wf cloud send`, claim + `status progress cloud:<sid>`. Prints `stop lane cloud: <ledger|none fit|max parallel>` otherwise. +- The orchestrator calls `wf cloud pull --all` on each wake-up (no faster than every 10 min; a teleport + takes ~30 s and no tokens). Results feed `wf orch post` like a worker report. +- Local lanes skip tasks claimed `cloud:`; the cloud lane takes opus tasks only (the expensive ones locally), `Cloud: yes` first (§4.5). +- Ledger ≤ reserve → the cloud lane stops; the local lanes run as today (C3). + +## 5. Economics (measured, see §6) + +Value of local work = its API-price $ (what the weekly limit is spent on). Per opus task: + +| Item | $ | Source | +|---|---|---| +| Local opus worker avoided | 1.20 | Historical API-price per task (range 1.10–1.32, n=64) | +| Local overhead of a cloud task | 0.10 | send + pull = 2 orchestrator calls at ~150k cached ctx (~0.05 each); poll wakes amortised ~0.03; gate = CPU only | +| Failure (handback → local redo) | 0.24 | assumed p_fail 20 % × 1.20 | +| **Local $ saved** | **0.86** | | +| Cloud credit used | 1.40 | typical replay scenario with setup overhead | + +Net per opus task = 0.86 − v × 1.40: + +| Credit value v | Net per task | $240 credit → tasks / net | +|---|---|---| +| 0 % (use-or-lose, time-limited) | **+0.86** | ~150 tasks / **+$129 local** | +| 50 % | +0.16 | ~150 / +$24 | +| 100 % (as paid usage) | −0.54 | loss: only worth it when the local limit is hit | + +The analysis shows that using cloud credits for opus-only tasks is economical only when the credit value is low (time-limited or not yet consumed). +Sonnet tasks: local $0.14–0.30 vs overhead + failure ≈ $0.15 → ≈ 0 net even at v = 0 → **cloud lane takes opus +only** (sonnet only via `Cloud: yes`). Build cost: 6 slices worth of engineering. + +## 6. Validation + +Cloud worker send/pull was validated with: +- Manual pty driver tests against the real CLI on throwaway repos (no project content). +- Round-trip end-to-end tests from send → export → pull with real snapshot repos. +- Synthetic test fixtures with the same message shape as real cloud sessions (for unit tests in the tool repo). + +Key findings: +- Marker parsing must handle export word wrapping (spaces, indents). +- ANSI codes in the TUI output must be stripped before marker matching. +- Trust dialog pre-acceptance prevents interactive blocks. +- Export `/exit` may take multiple Enter presses to ensure the file is written. diff --git a/docs/design.md b/docs/design.md new file mode 100644 index 0000000..5f691d7 --- /dev/null +++ b/docs/design.md @@ -0,0 +1,329 @@ +# wf — design reference + +Shared task workflow for many coding agents (Claude Code sessions) working in parallel on the +projects under one folder (`/projects`). This file is the reference for the TASKS.md format, +`workflow.toml`, every `wf` command and the release rules. The memory/CPU ledger `wf res` has its +own document: [resource-ledger.md](resource-ledger.md). Overview: [README](../README.md). + +## 1. Goal + +One workflow definition and one tool, used by every agent working in a subfolder of +`/projects`, instead of task scripts and workflow text copied (and drifting) per project. +Project `CLAUDE.md` files keep project specifics only. + +Success: +- An agent started in any project answers "continue" with the same steps. +- A workflow fix is made once, in one repo, under revision control. +- Task references cannot rot silently (`wf check` fails on them). +- Less always-loaded instruction text per project. +- A worker that hits a workflow problem reports it in one command and keeps working; fixes + reach all workers without breaking one that is mid-task. +- Routine task upkeep (list, find, add, reprioritise, flag, finish) is one short command + each; an agent never reads or rewrites TASKS.md whole to do it. + +Non-goals: +- No runner: the task tool never builds, tests, commits or pushes. + +## 2. Layout + +``` +/projects/ + CLAUDE.md -> workflow/shared/CLAUDE.md symlink + workflow/ this repo + shared/CLAUDE.md shared loop (section 6), loaded by every project + CLAUDE.md rules for sessions changing this repo + wf.py CLI entry + wf_res.py `wf res` (files, /proc, systemd); logic in wflib/res.py + wflib/tasks.py pure text: parse, insert, remove, archive + wflib/refs.py pure text: ids, links, anchors, doc sections + wflib/config.py find project root, load workflow.toml + wflib/check.py everything `wf check` reports + wflib/ledgers.py in-flight plan ledgers for `wf next` + wflib/search.py pure text: ranked search over tasks, archive, docs + wflib/migrate.py one-time converter from a numbered task list + wflib/res.py pure logic of the resource ledger + tests/test_*.py unittest, stdlib + inbox.md reports from workers, append-only, not in git (section 7) + templates/ workflow.toml, TASKS.md, archive.md for `wf init` + CHANGES.md release notes, newest first + <group>/<project>/ any depth ≤ 3; a project = a folder with workflow.toml +``` + +Loading: Claude Code reads `CLAUDE.md` from the working directory and every parent, so +sessions in a project or in its `.worktrees/*` get the shared file without a pointer. If a +symlink is not loaded, a real one-line file `/projects/CLAUDE.md` containing +`@workflow/shared/CLAUDE.md` does the same. The shared file sits in `shared/` so that a session +inside `/projects/public/workflow` does not load it twice. + +Each project `CLAUDE.md` starts with: +`Workflow: /projects/CLAUDE.md (shared loop, wf). Below = project specifics; this file wins on conflict.` + +The folder scanned by `wf projects` is the parent of this repo, or its grandparent when the parent holds no project (`WF_ROOT` overrides); the inbox +is `inbox.md` here (`WF_INBOX` overrides). The tasks of the workflow itself can live in a +separate (private) wf project; nothing in this repo needs them. + +Requirements: Linux, Python ≥ 3.11 (stdlib only). `wf res` also needs systemd user units with +cgroup v2. + +## 3. TASKS.md format + +```markdown +# Tasks — <project> + +<free prose: commands, pointers> + +## Awaiting your decision + +- **a-smoke-screens**: needs you present, screen unlocked. ... + +## Pending + +- **t-cave-seams** [P3] (<1h): Terrain: cave entrance seams. Close the slit beside the lintel. + - Steps: ... + - Tests (red first): ... + - Done: ... + - After: [[t-terrain-cave-render]] + Ref: DESIGN.md#terrain + +## Needs human + +- **t-playtest-feel** [P1] (<1h): Play-test feel + tuning. ... + - [ ] ... + +## Deferred + +- **t-android-port** [P3] (10h): ... +``` + +Rules: +- Item = a line `- **<id>**` at column 0 plus every following line that is blank or indented. + A flush-left non-item line ends the item list of the section (prose after items is kept). +- Id grammar: `t-[a-z0-9-]+` for tasks (Pending, Needs human, Deferred), `a-[a-z0-9-]+` for + Awaiting items. A task keeps its id when it moves between sections. + (Changed from the chat draft, which used `h-` for Needs human: a prefix tied to a section + would force an id change on a move.) +- Ids are unique across TASKS.md and the archive, and never reused. +- Task header: `- **<id>** [P0-P3] (<effort>) [status]: Title. Goal.` (old `(<effort>, interactive)` + still reads as `Sessions: owner`; `check` warns.) + Effort is one of `<1h`, `1h`, `5h`, `10h`, `100h`. Awaiting items have no priority or effort. +- Status, optional, fixed words: `(in progress: <branch or note>)`, `(blocked: [[a-id]])`. +- `After: [[id]], [[id]]` lists dependencies. `Ref:` lists `path` or `path#anchor`, comma + separated, relative to the project root; text in parentheses after a ref is a free note. +- `Model: haiku|sonnet|opus` (none = opus): the model for the task (docs/lanes.md §0). + `Sessions: parallel|solo|owner [— why]` (none = parallel; `- Sessions:` works too): + solo = only while no other session is live in the project (lib/submodule bump, repo-wide refactor, + shared runner), and while it is in progress no other session picks anything; owner = needs the owner + present (real display, playtest, choices): `next` picks it only with `--owner`. +- Slices of a task over 1h: own items `t-<parent>-1`, `t-<parent>-2` (or any id given with + `--id`), each `After:` the previous one. The parent item stays as the summary and lists its + slices (`Slices:` line); it is done when its last slice is done, and `next` does not pick it + while a slice is open. +- Pending holds open work only, in pick order. +- Sections other than the four named ones are prose: the tool keeps them and checks only + the `[[id]]` links in them. + +Archive file: newest first, one line per finished task: +`- 2026-09-29 **t-cave-seams** Terrain: cave entrance seams — <entry, ≤2 lines>`. +Older lines without an id stay as they are. + +## 4. workflow.toml + +In the project root. Its presence marks the project root for `wf`. + +```toml +format = 1 # required: TASKS format version +tasks = "TASKS.md" # required +archive = "tasks/archive.md" # required +docs = ["DESIGN.md", "docs/"] # files/dirs whose links and anchors `check` validates +verify = ["dotnet build X.slnx", "dotnet test X.slnx"] # printed by `done`, never run +done = ["Feature changed -> docs/CATALOG.md"] # printed by `done` +ledgers = ".superpowers/sdd" # optional: in-flight plan ledgers +ctx_hint = 100000 # optional: tokens; 0 = off (see below) +worktree_setup = ["mkdir -p out && ln -sfn \"$WF_MAIN/out/data\" out/data"] # optional: `wf setup` (below) +quick_gate = ["dotnet test X.slnx --filter ContractTests"] # optional: `wf gate` (below) +cloud = true # optional: project opts in to the cloud lane (default false) +cloud_include = ["out/data/"] # optional: git-ignored paths force-added to the cloud snapshot +cloud_note = "No GUI here." # optional: appended to the cloud prompt + +[anchors] # optional: index/spec anchor scheme +index = "DESIGN.md" # every task Ref anchor must be a heading here +index_section = "Subsystems" +specs = "docs/superpowers/specs" # explicit <a id> here needs an index entry +``` + +Only `format`, `tasks` and `archive` are required. Unknown keys are an error (catches typos). + +`worktree_setup`: shell commands `wf setup` runs inside a lane worktree (cwd = the worktree's project +folder, env `WF_MAIN` = the main tree's project folder), in order, stopping at the first non-zero exit +(`wf: …`, exit 1). They provide git-ignored inputs a fresh worktree lacks (e.g. link `out/…` from the +main tree) and must be idempotent: wf-worker runs it after every worktree create/reuse. + +`cloud` fit (cloud-lane spec §4.5): with `cloud = true` a task fits when runner-ready, not `Sessions: owner|solo`, +no `Cloud: no` line, effort <= `slice_above`, Model opus, and its title/body does not match the unfit regex list +in `wflib/cloud.py` (`wf` / `wf res` steps, GUI, LAN hosts, live server). Routing is per task: whoever adds it sets +`Cloud: yes|no` (`wf add/set --cloud`); `Cloud: yes` skips the regex list only (still opus only), `Cloud: no` +always excludes, no line = the rules above. `wf orch pick cloud` order (`cloud.pick_key`): `Cloud: yes` first, +then prio, then larger effort. `wf check` accepts the `Cloud:` line (yes|no, one per task). + +`quick_gate`: fast regression commands (minutes, not the full slow gate) `wf gate` runs in the current +tree's project folder (worktree or main; env `WF_MAIN`), stopping at the first red (`wf: …`, exit 1); +none configured → exit 0. wf-worker runs it before `wf done`, so a break is caught by the task that made +it, not blamed on a later task by the async slow gate. The project picks the commands. + +## 5. wf CLI + +Run as `python3 /projects/public/workflow/wf.py <cmd>` from anywhere inside a project. The +project root is the nearest ancestor of the working directory that holds `workflow.toml`; +`--project DIR` overrides. Output is plain text, exit 0 on success. + +Principle: every routine operation is one command with short output, so agents spend no +context on reading or rewriting task files. Hand edits are for body prose only. + +Read (never write): + +| Command | Does | +|---|---| +| `next [--brief]` | Awaiting items, Needs-human titles, in-flight ledgers, then the next task with `ctx` of its refs. `--brief`: the task only. Skips blocked items, open `After:`, parents with open slices, tasks held by another live session, `Sessions: owner` (unless `--owner`), `Sessions: solo` while another session is live. A solo task in progress by another live session → picks nothing, exit 1. Exit 1 if nothing is pickable. | +| `list [filters]` | One line per item: `id P2 1h status Title` (title cut to fit 100 columns). Default: Pending. Filters: `-s pending\|human\|awaiting\|deferred\|all`, `-p 0-3` (that priority or higher), `--ready` (pickable now), `--runner` (ready + Done line, not owner-bound, no open slices), `--blocked`, `--progress`, `--ref <path[#anchor]>`, `-n N`. Last line: counts per section. | +| `show <id>...` | The raw item(s), nothing resolved. | +| `ctx <id \| path#anchor>` | The item, its refs resolved, its dependencies and the items that depend on or link to it, then the notes of every area (`## Areas`) the item names (area name, code-map anchor or Paths entry as a whole word). For an anchor: that doc section plus the tasks that reference it. | +| `search <words>... [filters]` | Ranked hits over tasks, archive and `docs`: one line each, `id` or `path:line heading`, plus the matching line. Filters: `--tasks`, `--archive`, `--docs`, `-n N` (default 15). | +| `log [-n N] [words]` | Newest archive lines (default 10), optionally only those matching the words. | +| `projects` | Every project under the folder that holds the workflow repo (`/projects`) with a `workflow.toml` (depth ≤ 3): counts per section, next task id + title, number of `check` errors. | +| `check` | Validates; prints every problem; exit 1 on any error. | + +Write (each changes only the items named, and prints the changed header lines): + +| Command | Does | +|---|---| +| `add "<Title. Goal.>" -p N -e EFFORT [opts]` | New item. Id made from the title unless `--id` given or the text starts with `<id>: `; printed. Opts: `-s SECTION`, `--after ID,…`, `--ref REF,…`, `--parent ID` (slice: id `t-<parent>-N` or `--id`, `After:` the previous slice, parent's slice list updated), `--model M`, `--sessions solo|owner`, `-b` (body lines from stdin; never read unasked, a harness may hold stdin open). `add -` reads a whole item block from stdin. | +| `done <id>... [-m "<entry>"]` | Removes the task(s), puts the dated line(s) on top of the archive, prints the `verify` and `done` lists once. `-m` is required for a single task and applies to it; several ids without `-m` use each task's goal sentence. | +| `prio <id> <0-3>` | Sets the priority and moves the item to its place by the insert rule. | +| `move <id> <section>` or `move <id> --before\|--after <id>` | Moves between sections, or reorders inside one (refused if it breaks priority or `After:` order, unless `--force`). | +| `status <id> progress "<note>" \| blocked <a-id> \| clear` | Sets or clears the status words. | +| `set <id> [--title T] [--effort E] [--after IDS] [--ref REFS] [--model M] [--sessions S] [--done TEXT]` | Changes header fields and the `After:` / `Ref:` / `Model:` / `Sessions:` / `Done:` lines (`""` removes). `--interactive` = old spelling of `--sessions owner`. `--after ""` removes the line. `--title` keeps the goal, unless it has its own (`Title. Goal.`) or ends in `?`/`!`: then it replaces the whole text. | +| `rename <id> <new>` | Changes an open item's id (same kind, unused, not archived) and every `[[id]]` link in TASKS.md. | +| `note <id> "<line>"` | Appends one `- <line>` to the item body. | +| `body <id>` | Replaces the item body with stdin (header, `After:`, `Ref:` kept). | +| `tick <id> <n \| text>` | Checks the n-th (or the matching) `- [ ]` box of a Needs-human item. | +| `report "<what happened>" [--kind bug\|idea\|friction] [--cmd "<command>"]` | Appends one entry to the workflow `inbox.md` (section 7). Touches nothing in the project. | +| `batch N [--lanes L,…] [--prep] [--model M] [--for 6h] [--mem 4G] [--dry-run]` · `batch --status` | Starts an unattended batch orchestrator (`claude -p`, auto mode, no prompts, prompt `templates/batch-prompt.md`) as a `wf res run` job with `CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS` = `--for`; it writes `out/wf-batch-<stamp>.md`. `--status` = this project's newest batch job + newest summary file. `--prep` = prep batch (`templates/prep-prompt.md`): sonnet workers write Done (+Model) for up to N tasks without Done, unclear → awaiting; none → nothing started. See `docs/orchestrator.md`. | +| `init` | In a folder without `workflow.toml`: writes it, TASKS.md and the archive from the templates. | +| `migrate [--write]` | One-time converter from the numbered format. Without `--write`: prints the diff, the id map and notes, changes nothing. `--write` also raises `format` in `workflow.toml`. | + +Every write command accepts `--dry-run` (prints a unified diff, writes nothing) and ends by +comparing the full check before and after: a change that adds a problem is refused and +nothing is written; problems that were there before do not block other edits. + +Details: + +- **Pick rule (`next`)**: first Pending item, in file order, that has no `(blocked: …)` and + whose `After:` ids are all in the archive. Items it skipped are listed in one line each + with the reason. +- **Ref resolution (`ctx`, `next`)**: `path#anchor` → the section under the heading whose + slug or explicit `<a id>` equals the anchor, up to the next heading of the same or higher + level, cut at 80 lines with a `… (N more lines: <path>:<line>)` note. `path` alone → the + file's first heading and its `**Goal:**` line if it has one. With `[anchors]` set, an + index anchor also prints the spec section that carries the same `<a id>`. +- **Context hint (`next`, `done`)**: with `CLAUDE_CODE_SESSION_ID` set, the last request's prompt size + (input + cache write + cache read) from the last 1 MB of `<CLAUDE_CONFIG_DIR|~/.claude>/projects/*/<id>.jsonl`; + over `ctx_hint` → last line `context ~Nk tokens (> Mk): ask the owner to /clear, then continue`. No id, + no transcript → nothing. Subagents share the id, so the line says it is about the main session. +- **Ledgers (`next`)**: for each `<ledgers>/*/progress.md`: plan path from its first line, + done tasks, where to resume, count of `Ruling:` lines. +- **Insert rule (`add`, `prio`, `move <section>`)**: position in the section = after the + last item of the same or higher priority, and after every item named in its `After:`. + Body is re-indented to two spaces. A block given to `add -` must carry an id and, for + tasks, a priority; a leading `N.` from the old format is refused with a message. +- **Id from title (`add`)**: `t-` (`a-` for Awaiting) + slug of the title (the text up to + the first `. `), at most 40 characters, cut at a word; if taken (TASKS.md or archive), `-2`, `-3`, … is appended. +- **Search ranking**: words are matched case-insensitively as whole words or word prefixes. + Score per hit = words matched × weight of where (id or title 5, doc heading 4, header + goal 3, `Ref:`/body/doc text 1); all words matched ranks before some. Ties: open tasks, + then docs, then archive. No index file; everything is read per call. +- **Exact ids only** in write commands; a unique id prefix is accepted by read commands. + An unknown id prints the three nearest ids. +- **`done`** refuses when the id is unknown or is a parent with open slices. On an `a-` id + it removes the item without an archive line (`-m` is optional there) and clears the + `(blocked: [[a-id]])` status of the tasks that waited on it, listing them. +- **`check`** errors: + - duplicate id; id reused from the archive; bad id grammar; id prefix not allowed in its section + - task without priority or with an effort outside the vocabulary + - `[[id]]` that is neither in TASKS.md nor in the archive + - `(blocked: [[x]])` where `x` is not an open `a-` item + - `After:` cycle; Pending item placed before an open item it is `After:` + - `Ref:` path missing; anchor missing in that file + - old-format leftovers: numbered items + - item after a flush-left prose line inside a task section; malformed item header + - with `[anchors]`: task anchor without index heading; two index headings with the same + slug; index link to a missing file or anchor; explicit spec `<a id>` without index entry + - config: missing required key, unknown key, configured path that does not exist + In `docs` only `[[t-…]]` / `[[a-…]]` links are validated (other `[[x]]` may be the + project's own wiki links). + Warnings (exit stays 0): `#N` number refs in TASKS.md; task in progress with no branch/note; Awaiting item that no task + references and that is older than 30 days by `git blame`. +- **Writes** are whole-file: read, change in memory, write to a temp file in the same + directory, rename. If the file changed on disk between read and write (mtime + size), the + command stops without writing. `done` writes the archive first, then TASKS.md; if the + second write fails it reports that the archive line must be removed. +- **Errors** go to stderr as one line each, naming file and id; no tracebacks for expected + failures. + +## 6. Shared CLAUDE.md + +Text: `shared/CLAUDE.md` (the file is the authority). Sections: Style (concise everywhere; +compaction summaries, specs, plans stay full and exact), Cold start, Loop, Done, TASKS.md, +Context economy, Memory / CPU, Workflow problems. + +## 7. Improving the workflow without disturbing workers + +Problem: `wf.py` and the shared CLAUDE.md are live for every worker the moment they change. +A worker that edits them mid-task, or a half-finished change, would hit all the others. + +**Report (any worker, any time)** +- `wf report` appends to `inbox.md` with a single `O_APPEND` write, so parallel workers + never clash and nobody waits. Entry: date, project, kind, the text, the command if given, + wf version (git short hash). + `- 2026-09-29 proj-a friction: prio change needs two commands (cmd: wf prio …) @84decd1` +- The worker then works around the problem and continues its own task. It does not edit + anything under `/projects/public/workflow` and does not wait for a fix. +- `wf next` in any project prints one line when the inbox has entries, for the user's eyes: + `workflow inbox: 3 reports (triage: workflow session)`. + +**Triage and fix (a workflow session, started by the owner)** +- The workflow's own tasks live in a wf project of their own (e.g. a private + `/projects/workflow`), whose CLAUDE.md holds the rules of this section. Cold start there: + each inbox entry becomes a task (`wf add`, duplicates merged) or is rejected with a line in the + archive; the entry is then deleted from `inbox.md`. `inbox.md` is git-ignored; tasks and + archive carry the record. Bookkeeping there changes nothing a worker runs. +- Changes to `wf.py`, `wflib/`, `shared/CLAUDE.md` or templates happen in a git worktree + `/projects/public/workflow/.worktrees/<topic>` on a branch, never in the main tree. Workers keep + running the main tree's `master` meanwhile. +- Release = fast-forward merge to `master` in the main tree after: tests green, and + `wf check` run with the new code against every project listed by `wf projects` gives the + same result as before the change (or differences that are the intended ones). +- The user starts workflow sessions and sees the release; no worker does it on its own. + +**Compatibility rules for a release** +- Commands and options are only added. Renaming or removing one keeps the old spelling + working for one release, printing a one-line notice with the new spelling. +- A change to the TASKS.md format raises `format` and ships with a `wf migrate` step. + `wf` run in a project with an older `format` refuses write commands with + `TASKS format 1, wf needs 2: run wf migrate --write (idle project, one commit)`; read + commands keep working on the old format. +- Shared CLAUDE.md text changes reach a worker at its next session start; a rule change + must therefore not make files written under the old rule invalid. +- Every release adds a line to `CHANGES.md` (newest first): date, what changed, what a + project has to do, if anything. + +## 8. Testing + +- `wflib` functions are pure (text in, text out); unit tests use literal before/after texts + written by hand, never output of the code under test. +- CLI tests run `wf.py` with `--project <tempdir>` on fixture projects: pick rule, skip + reasons, insert position with priority and `After:`, `done` refusals, every `check` error + once, write-conflict stop, `list` filters, `prio`/`move` placement and refusals, `status`, + `set`, `note`, `body`, `tick`, search ranking order on a fixed corpus, `report` from two + processes at once (both entries whole), `format` gate, migrate on a fixture with the known oddities (stray `N.`, + `N/M` titles, `After:` by title, prose after items). +- Each new test is seen failing first, and again with the tested line broken by hand. diff --git a/docs/lanes.md b/docs/lanes.md new file mode 100644 index 0000000..2fc6833 --- /dev/null +++ b/docs/lanes.md @@ -0,0 +1,115 @@ +# Lanes, Model line, slice jobs, area maps + +Reference (released 2026-10-05). Lanes split work by task **size** and **pick policy**: a fast lane clears +small, unblocking work and slices big tasks; a slow lane takes 1h slices; more lanes by config. The model is +a per-task attribute (`Model:` line): interactive sessions take tasks with Model ≤ their own, orchestrated +workers run on the task's Model. Area maps (code map + test recipe per area) are checked and refreshed by +auto-added tasks. + +## 0. The Model line +- Task body line ` Model: haiku|sonnet|opus` (2-space indent like `Ref:`; `- Model:` also parsed), in the + body tail before `After:` / `Ref:`. No line → `opus`. Order haiku < sonnet < opus. +- Set by the task's writer at `wf add --model M`; `wf set ID --model M`; `--model ""` removes it (→ opus). + - `haiku`: mechanical edit, exact instructions, no judgement. + - `sonnet`: exact Done criteria + an existing pattern to copy + a real-data pass/fail test. + - `opus`: design forks, debugging, reverse engineering / oracles, cross-source analysis, anything unclear. +- Wrong level found mid-task: `wf set ID --model <higher>` + `wf note ID "<why>"` + `wf status ID clear`. +- `wf check`: unknown value or two Model lines → error; extra words after the value → warning; missing → fine. + + +## 1. Lane config +`workflow.toml`, optional; absent → built-in default equal to: +```toml +slice_above = "1h" # effort above this → slice job (§3); one of EFFORTS + +[lanes.fast] +efforts = ["<1h"] # efforts this lane implements +order = "unblock" # pick policy (§2) +slices = true # this lane also takes slice jobs (exactly one lane may set it) + +[lanes.slow] +efforts = ["1h"] +order = "priority" +fallback = "fast" # nothing pickable here → pick from that lane +``` +- Lane name: `[a-z][a-z0-9-]*`; used in worktree `.worktrees/<lane>`, branch `<lane>/<id>`, registry file. +- `efforts`: subset of `EFFORTS`, ≤ `slice_above`; lanes' efforts are disjoint; every effort ≤ `slice_above` + belongs to some lane. Tasks without effort → the lane holding `<1h`. +- `order`: `unblock` | `priority`. `fallback`: another lane name, optional, no cycles. +- `slices = true` on exactly one lane; default: the lane holding `<1h`. +- Validation errors in `config.load` (unknown keys are errors). Defining `[lanes.*]` replaces the + whole default set (no merge). +- Extra lanes: add tables (e.g. `[lanes.long] efforts = ["5h"]` with `slice_above = "5h"`). Several workers in + one lane = existing `<lane>-2`, `-3` worktrees. + +## 2. Pick order +Pickable set (Pending, not blocked, no open slices, all `After:` archived, no live foreign claim), filtered +to the lane's efforts (plus slice jobs on the `slices` lane) and to Model ≤ session model (interactive only). +Effective priority = best (lowest P) of the task's own and that of every open task waiting on it, +transitively ("waits on" = `After: [[x]]`, or a parent waiting on its open slices). Keys: +- `unblock`: (1) waited-on count > 0 first, (2) more open tasks waiting on it (transitively) first, + (3) effective priority, (4) file order. +- `priority`: (1) effective priority, (2) waited on by another lane first, (3) waited-on count, + (4) file order (section order, then position). + +## 3. Slice jobs +- A Pending task with effort > `slice_above` and no open slices = slice job. It is not pickable as an + implementation task in any lane; it is pickable in the `slices` lane, ranked by that lane's order. +- Shown with `slice:` prefix in `wf next`, `wf list --runner`, `wf lanes` counts (`2 pickable (1 slice)`). +- `wf next` prints for it: `Slice job: wf add "<title>" -e <1h|1h> --parent <id> --model M` per slice, each + with Steps/Done/Ref; no code; `wf note <id> "sliced into …"`; `wf status <id> clear`. Not `wf done` (parent + stays open until its slices are done). +- Model of a slice job = the task's Model (default opus; slicing is judgement). +- A task with open slices is never a slice job (already sliced). All slices done → parent pickable + only if its effort ≤ `slice_above`; otherwise `wf done` of the last slice prints `parent <id>: all slices done + → wf done <id> or add slices` (it is not offered as a slice job again while it has done slices: rule = no + open slices AND no archived slices). + +## 4. Sessions, registry, model +- `wf next [--lane L] [--as M]`: `--as` = the session's model (haiku|sonnet|opus); none → haiku, with a + hint line. `--lane` = lane name; no `--lane` → all lanes merged, `priority` order, slice jobs included. +- Registry `.wf/sessions/<lane>.json` (`{"lane", "model", "socket", "pid", "session", "at"}`, from env + `CLAUDE_CODE_MESSAGING_SOCKET`, `CLAUDE_PID`, `CLAUDE_CODE_SESSION_ID`; skipped without them); no `--lane` + → `all.json` (`all` is reserved). `.wf/.gitignore` = `*`. Newest wins per lane unless another live session + holds it: then kept, and `wf next` first prints `another live <lane> session holds this lane: uds:…`. + Alive = pid and socket file exist. `wf lanes --unregister` drops this session's records. +- Claims: `wf status ID progress …` writes `.wf/claims/ID.json` (`{"id", "lane", "model", "socket", "pid", + "session", "at"}`); `status clear|blocked` and `wf done` remove it. `wf next` skips a task whose claim is + alive, not mine and still `in progress`. +- `wf lanes [--lane L]`: one line per configured lane: + `fast: 3 pickable (1 slice) · 1 waiting · session uds:… (alive)` / `… · no session`. +- Waiting block: `t-y waits on t-x (slow lane) → message uds:…: "t-x blocks my t-y, please take it"` / + `… → no slow session: tell the owner` — lane of a task = lane by effort (slice lane for slice jobs). +- `wf done`: tasks of other lanes that became pickable → `notify <lane> uds:…: now pickable t-y` / + `<lane> work now pickable: t-y (no session: tell the owner)`. +- Messages go with the agents' SendMessage tool, `to` = `uds:<socket>`. +- `wf list`: model and lane columns; filters `--model M`, `--lane L`; `--runner --lane L` = ids in pick order + for that lane (no session model filter; the orchestrator spawns per task model). + +## 5. Orchestrator, batch, worker +- `docs/orchestrator.md`, `shared/skills/wf-orchestrate`: lanes from `wf lanes`; per lane + `wf list --runner --lane <lane>` → first id; spawn `wf-worker` with `model: <task's Model>`; worktree + `.worktrees/<lane>`, branch `<lane>/<id>`. At most one worker per lane worktree; lanes in parallel. +- `wf batch N [--lanes fast,slow]` (default: all configured lanes). +- `shared/agents/wf-worker.md`: prompt line `Lane: <lane> Model: <model>`; slice job → slice only, no code. +- `wf usage --log … --lane L`: cost line per worker; `--report` rows by lane × model × effort. + +## 6. Area maps +- Notes file: config `areas = "<path>"`, default project `CLAUDE.md`. Section `## Areas`, one `### <area>` each + with lines (all optional): + - `- Code map: …` — anchors = backticked tokens (`` `parse_item` ``). + - `- Test recipe: …` + - `- Paths: <globs, space-separated>` — default: whole repo. + - `- Checked: <short sha>` — last refresh; set by `wf areas --mark`. +- `wf areas [AREA]`: per area: missing anchors (`git grep -F -q <token> -- <paths>` fails), commits touching + Paths since Checked (`git rev-list --count <sha>..HEAD -- <paths>`), stale yes/no. Exit 0. +- `wf areas --mark AREA`: write `Checked: <HEAD short sha>` into the section (add the line if missing). +- Stale = any missing anchor, or Checked set and commits ≥ `area_stale_commits` (config, default 20). + No Checked → only the anchor test. +- `wf check`: missing anchor → warning `area <a>: anchor <tok> not found`. Unknown Checked sha → warning. +- `wf done`: after bookkeeping, for each stale area with no open task whose id is `t-map-<slug>` → add + `t-map-<slug>` P2 `<1h` Model: sonnet, title `Refresh <area> area map`, Steps: `wf areas <area>`; fix missing + anchors and test recipe from changed files (`git log --stat <Checked>..HEAD -- <paths>`); `wf areas --mark + <area>`. Done: `wf areas <area>` shows not stale. Printed `added t-map-<slug> (area map stale)`. + Id taken by an archived task → `t-map-<slug>-2`, … (ids never reused). +- No `## Areas` section → feature silent. Template `templates/CLAUDE.md` area block has `Paths:`/`Checked:`. diff --git a/docs/manual.md b/docs/manual.md new file mode 100644 index 0000000..2b4594b --- /dev/null +++ b/docs/manual.md @@ -0,0 +1,239 @@ +# wf — owner's manual + +How to set up and run `wf` day to day. The README tells you what it is and why it exists. The +reference docs give the details: [design.md](design.md) (format, config, every command), +[orchestrator.md](orchestrator.md), [lanes.md](lanes.md), [multi-session.md](multi-session.md), +[resource-ledger.md](resource-ledger.md). Every command below also has `wf <command> -h`. + +This is still a personal setup. The manual describes how it runs on the author's machine. It is +not a supported product. + +## 1. Setup + +### Requirements +- Linux with systemd. `wf res` needs systemd user units and cgroup v2. Everything else only needs a shell. +- Python ≥ 3.11 (standard library only, nothing to `pip install`). +- git, and [Claude Code](https://claude.com/claude-code) (`claude` on `PATH`). +- Optional: the `superpowers` Claude Code plugin. The shared rules refer to its brainstorming, + debugging and plan-writing skills. Without it, agents just follow the plain rules. + +### One root folder +All projects live under one folder. The shipped rules, skills and agent file use the absolute +path `/projects` (for example `python3 /projects/public/workflow/wf.py`). Keep that path: create it once +and own it, or point it at a folder in your home directory. + +```sh +sudo mkdir /projects && sudo chown "$USER": /projects # or: sudo ln -s ~/projects /projects +git clone <this repo> /projects/public/workflow +ln -s workflow/shared/CLAUDE.md /projects/CLAUDE.md # every agent under /projects loads it +echo "alias wf='python3 /projects/public/workflow/wf.py'" >> ~/.bashrc +``` + +Claude Code reads `CLAUDE.md` from the working directory and from every parent folder, so every +session in `/projects/<anything>` gets the shared rules. If the symlink is not picked up, use a +one-line `/projects/CLAUDE.md` with the content `@workflow/shared/CLAUDE.md` instead. +`wf projects` scans the parent folder of the tool repo, or its grandparent when the parent holds no project (tool under e.g. `public/`; override: `WF_ROOT`). Workflow reports +go to `/projects/public/workflow/inbox.md` (override: `WF_INBOX`). + +### Agent and skills +The orchestrator spawns the worker subagent. The skills are the session modes in section 2. +Install them as symlinks, so that `git pull` updates them: + +```sh +mkdir -p ~/.claude/agents ~/.claude/skills +ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md +for s in wf-orchestrate wf-pilot wf-design; do ln -s /projects/public/workflow/shared/skills/$s ~/.claude/skills/; done +``` + +Agents and skills appear only in sessions started after the install, so restart any running session. + +### A project +```sh +cd /projects/games/mygame # any depth ≤ 3 under /projects; a git repo is recommended +wf init # writes workflow.toml, TASKS.md, tasks/archive.md, CLAUDE.md (keeps existing files) +wf check # 0 errors +``` +If the project already has a numbered task list, `wf migrate` shows the conversion (dry run) and +`wf migrate --write` applies it. + +Edit the generated `CLAUDE.md`: layout, verify command, gotchas and one `### <area>` block per +code area (code map with grep anchors, test recipe, paths). Agents read these area notes before +they read code. + +`workflow.toml` (the template lists every key with a comment; unknown keys are errors): + +| Key | What it does | +|---|---| +| `tasks`, `archive` | file names (defaults `TASKS.md`, `tasks/archive.md`) | +| `docs` | files/folders whose `[[id]]` links `wf check` validates | +| `verify` | commands `wf done` reminds of (it never runs them) and `wf ctx` prints as Verify | +| `done` | extra checklist lines `wf done` prints | +| `quick_gate` | fast regression check; `wf gate` and `wf finish` run it and refuse when it is red | +| `worktree_setup` | idempotent commands `wf setup` runs in a fresh lane worktree (link git-ignored data, venv…; env `WF_MAIN` = main tree) | +| `slice_above`, `[lanes.*]` | task sizes per lane and when a task must be sliced (default lanes: fast `<1h`, slow `1h`); see lanes.md | +| `areas`, `code_root`, `area_stale_commits`, `area_ignore` | where the area notes live and when they count as stale | +| `ledgers`, `ctx_hint`, `[anchors]` | in-flight plan ledgers, `/clear` hint, required Ref anchors | +| `cloud`, `cloud_include`, `cloud_note` | cloud lane opt-in (below) | + +Add `out/` to the project's `.gitignore`: logs, batch summaries and scratch go there. + +### Resource ledger (`wf res`) +Optional, but needed once several agents build or test at the same time. + +```sh +wf res shell-init # prints an alias that starts claude inside agents.slice: add it to ~/.bashrc +wf res timer on # 1-minute timer: starts queued jobs, ends game mode, cleans scratch, adopts sessions +wf res status # budget, reservations, warnings +``` +Reserves and limits go in `~/.config/wf/resources.toml` (all keys optional: `user_reserve_gb`, +`user_reserve_cpus`, `game_reserve_gb`, `game_reserve_cpus`, `game_hours`, `small_headroom_gb`, +`scratch_hours`; see resource-ledger.md §3.3). Optional hook that hides the display from agent +commands while game mode is on, in `~/.claude/settings.json`: + +```json +"hooks": {"PreToolUse": [{"matcher": "Bash", "hooks": [{"type": "command", + "command": "python3 /projects/public/workflow/wf_res.py hook"}]}]} +``` + +`wf batch` (and `/wf-pilot`) need one allow rule in `~/.claude/settings.json`, because auto mode +does not let an agent start another unattended agent: `Bash(python3 /projects/public/workflow/wf.py batch:*)`. + +### Cloud lane (optional) +Tasks can also run as `claude --cloud` sessions, which are billed to your cloud budget, not your +machine. You need a Claude Code version with `--cloud` / `--teleport` and a claude.ai login. +Per project: `cloud = true` in `workflow.toml`, plus `cloud_include` for git-ignored inputs the +task needs. Set the spending cap once with `wf cloud ledger --budget 100` and check it with +`wf cloud ledger`. Only opus tasks marked `Cloud: yes` (or passing the fit rules) are sent. See +orchestrator.md "Cloud lane". + +## 2. Daily use + +There are four kinds of session. Start each one in the project folder with `claude` (the +`agents.slice` alias if you set it up). + +### Worker session: "continue" +Type `continue` (or "next task"). The agent runs `wf next --as <its model>`. It shows open +questions, items that need you, plans in flight and the next task, then works on that task: +test red → implement → green → `wf done` → one commit. It does not ask before it starts. +Several sessions in one project: give each a lane ("continue in the slow lane"). `wf next` then +switches them to worktree mode (one branch per task, `wf finish` merges it). One session per lane. + +### Orchestrator: `/wf-orchestrate` +After `/clear`, type `/wf-orchestrate` (`go` = start the proposed run without asking, +`batch N` = prepare an unattended batch). The orchestrator holds no lane and never reads code. +It spawns one fresh `wf-worker` per task (`wf orch pick` / `wf orch post`), one per lane in +parallel, on the task's model, and stops a lane on the first awaiting, handback or red +post-check. Use it when the backlog is runner-ready (section 3) and you want to steer. + +### Overnight: `/wf-pilot` +`/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]` in an idle session, which can be one you opened +from the phone. The pilot starts headless `wf batch` runs one after another until the deadline. +The session stays small, and you get push notifications on new questions, on stops and at the +end. Steering: `wf batch --stop` (graceful, running workers finish), `wf move <id> deferred` +(skip a task), `wf batch --status` (newest batch and its summary). A one-off short run without a +pilot is `wf batch N [--lanes …]`. `wf batch N --prep` makes tasks without a Done line +runner-ready instead. + +### Design: `/wf-design` +`/wf-design [topic | id]` is for brainstorming, specs and rulings with you. Flow: brainstorm → +spec (committed) → your approval → `wf add` slices ≤ 1h with Done, Model and Ref → the +orchestrator picks them up. It answers open questions with you and records each ruling where it +binds. + +### Your side +- **Awaiting** (`wf list -s awaiting`): questions agents could not decide. Answer them in a + design or orchestrator session. The session records the ruling and runs `wf done <a-id>`, + which unblocks the waiting tasks. +- **Needs human** (`wf list -s human`): things only you can do (play-test, hardware, accounts). + Check boxes with `wf tick <id> <n>`. Done → tell any session ("t-x is done") or run `wf done`. + Tasks marked `Sessions: owner` are picked only when you say you are present (`wf next --owner`). +- **Stop file**: `wf batch --stop` writes `out/wf-batch.stop`. Batches, pilots and + `wf lanes --wait` stop before the next spawn. +- **Game mode**: `wf res game on [--for 4h]` before you play. Agents keep running but get less + CPU and IO, and big jobs that do not fit wait. `wf res game off` when you are done. +- **Overview**: `wf projects` (all projects), `wf lanes`, `wf log -n 10`, `wf res status`. +- **Workflow inbox**: agents file problems with `wf report`. Triage `/projects/public/workflow/inbox.md` + now and then: fix the tool, or turn entries into tasks. + +## 3. Writing good tasks + +```sh +wf add "Cave seams. Close the slit beside the lintel." -p 1 -e '<1h' --model sonnet \ + --done "tests/test_cave.py::test_seam passes; render of fixture shows no gap" --ref DESIGN.md#terrain +``` + +- **Title. Goal.** One line: the title, then the outcome in one sentence. +- **Done line** (`--done`): a check a worker can run without you, such as a file, a test result + or an artifact. Never "owner confirms". Only tasks with a Done line are runner-ready: the + orchestrator, batch and pilot skip the others. (`wf batch --prep` drafts them.) +- **Model line** (`--model`): the cheapest model that fits. haiku = mechanical, exact steps. + sonnet = exact Done, an existing pattern to copy, and a real-data pass/fail. opus (the default + when the line is missing) = design, debugging, reverse engineering, anything unclear. A worker + that finds the task too hard raises it itself and hands it back. +- **Effort** (`-e`): `<1h`, `1h`, `5h`, `10h`, `100h`. Anything above `slice_above` (1h) becomes a + slice job: an agent splits it into `--parent` slices, each with its own Done/Model, and never + implements it whole. Slice big work yourself when you already know the steps. +- **Code anchors**: when you know them, add body lines like + `Code: src/terrain.py build_cave; test to copy: tests/test_cave.py::test_floor`. Write anchors + (function or string names), not line numbers. They save the worker most of its exploring. +- **Refs and dependencies**: `--ref doc.md#heading` makes `wf ctx` print that section; + `--after t-other` keeps it unpicked until `t-other` is done. +- **Areas**: keep each area's code map and test recipe in the project `CLAUDE.md`. `wf ctx` + prints the areas a task names. When code drifts, `wf done` adds a `t-map-…` refresh task + automatically. +- **Cloud** (cloud projects): `--cloud yes` only when the Done line can be checked from the repo + alone (no GUI, live server, LAN, local installs). When unsure, `no`. +- **Sessions**: `--sessions solo` for repo-wide changes (nothing else runs meanwhile), + `--sessions owner` when you must be there. + +## 4. Best practices and pitfalls + +- **One lane per session**, at most one worker per lane worktree. Parallel lanes are safe + (project lock, branch per task). Two sessions in one lane pick against each other. +- **Fresh context per task.** Cost is mostly cache reads, which grow with context. Short-lived + workers on the cheapest fitting model are cheaper than one long session. `/clear` the + orchestrator after a big design talk, but only when no worker is running (a `/clear` kills its + running command). +- **Big jobs through `wf res`**: anything over 2 GB, 4 cores or 10 minutes goes through + `wf res run --mem … --for … --title … -- cmd`. Exit 3 = busy: do other work and retry later, + or use `--queue`. `wf res hist` suggests sizes from past runs. +- **Never hand-edit TASKS.md** while agents run. Use `wf add/set/note/move/done`. If you must + edit by hand, run `wf check` afterwards. +- **Commit ≠ merge ≠ push.** Agents commit their own task. Workers ff-merge only their own task + branch. Pushing goes only to a remote named `home` (`git push home --all`), and nothing else is + pushed unless you ask. +- **Close the loop on problems.** Agents work around a workflow problem and `wf report` it. + Triaging the inbox is how the rules get better. Put project-only needs in the project's + `CLAUDE.md` or `workflow.toml`, not in the shared rules. +- **Watch cost**: `wf usage` (this session per agent), `wf usage --report` (per lane, model and + effort across every project's `out/wf-cost.log`). Expensive lanes usually mean vague Done lines + or a model that is too large. +- **Model choice**: run the orchestrator and design sessions on opus. Tasks get the cheapest + model that fits. A sonnet task without an exact Done line comes back as a handback. +- **Updating the tool**: `git -C /projects/public/workflow pull`. Read the top of `CHANGES.md`: each line + ends with "Projects: …", which says what a project must do (usually nothing). +- **Publishing a public copy**: keep the tool repo private (its history names your projects) and + share a fresh-history snapshot instead. Put your private words (project names, home paths), one + per line, in `publish-denylist.local` in the tool repo (git-ignored), then run + `python3 scripts/publish_snapshot.py [DEST]` (default `~/src/wf-public`). It copies master's + tree without `inbox.md`, `.worktrees/`, `__pycache__/`, `out/`, refuses if any denylist word is + in a path or file, and commits once per run (first: "initial public snapshot"; later: the new + `CHANGES.md` lines). It never pushes and never adds a remote: review DEST, then add the public + remote and `git push` there yourself. + +## 5. Troubleshooting + +| Symptom | Cause / fix | +|---|---| +| `busy: …` and exit 3 from `wf res run/note` | Not enough memory or CPUs beyond your reserve. Retry after the printed time, add `--queue`, or check `wf res status` for stale entries (`wf res release r-N`). If memory really is free (unused reservations), `--force` fits against real free memory. | +| `wf orch post` says post-check red | The worker said done, but the archive line is missing, the branch is still there, the worktree is dirty, or `wf check` has errors. The message names the problem. Fix it in the worktree (`wf merge`, commit or clean), then post again. | +| Task stuck `in progress` though no one works on it | A session or worker died. `wf list --stale` shows them, and `wf status --clear-stale` clears every claim without a live session (or `wf status <id> clear`). Re-run the worker with `wf orch pick <lane> --id <id> --recovery "<why>"` to keep its WIP. | +| `wf start` refuses: worktree has uncommitted changes | A dead worker left WIP. `--recovery` keeps it and shows it. Otherwise commit or reset it in that worktree. | +| `branch <lane>/<id> exists (local WIP?)` (cloud pull) | A local branch for that task already exists. Merge or delete it, then `wf cloud pull <id>` again. The cloud session is kept. | +| `wf cloud pull` exit 4 | The session is still running. Pull again later; after 24 h without a result it counts as lost. | +| `wf cloud pull` handback / bad patch | The task goes back to its local lane with a `Recovery: cloud attempt …` note. The patch is kept in `out/cloud/<id>.patch`. | +| `wf merge`: rebase conflict | The branch is untouched. `git rebase master`, resolve, run verify, `wf merge` again. | +| `wf finish` refused | The quick gate is red, files outside the given paths are uncommitted, or a path is outside the repo, missing or unchanged. Fix it and rerun. If it fails after `done:`, a rerun resumes. | +| A worker or batch does nothing | `wf list --runner` is empty: tasks lack a Done line (`wf batch --prep`), are blocked on an awaiting item, or wait on `After:`. `wf lanes` shows counts per lane. | +| `wf check` errors after a hand edit | Read the message (duplicate id, broken `[[link]]`, bad section) and fix the item. Ids are never reused. | +| `project format N is newer than this wf` | Update the tool: `git -C /projects/public/workflow pull`. | diff --git a/docs/multi-session.md b/docs/multi-session.md new file mode 100644 index 0000000..628aad9 --- /dev/null +++ b/docs/multi-session.md @@ -0,0 +1,46 @@ +# Multi-session: several sessions in one project + +Approved 2026-10-04. Problem: one work tree + `master` shared by sessions → a commit sweeps the +other's edits (`git add -A`), half-done code breaks the other's tests, two sessions of a lane pick +the same task. + +## 1. Claims +- `wf status ID progress …` writes `.wf/claims/ID.json` (who: socket, pid, lane, model); details in + `lanes.md` §4. `wf next` skips a task claimed by another live session. +- A second live session of a lane: `wf next` warns on its first line; the registry keeps the first one. + +- Task line `Sessions: solo` → picked only while no other session is live; while it is in progress + (claimed) every other session's `wf next` picks nothing (exit 1, names the holder); its `wf done` + prints `notify <lane> uds:…: solo <id> done, run wf next` for each other live session. + `Sessions: owner` → picked only by `wf next --owner` (owner present). + +## 2. Worktree mode (only while > 1 session is live, git projects) +- `wf next` prints `===== Multi-session =====`: from the main tree + `git worktree add .worktrees/<lane> -b <lane>/<task> master` (or `cd .worktrees/<lane> && git switch -c <lane>/<task> master` + when it exists), `&& wf setup` appended when workflow.toml has `worktree_setup`; inside a worktree a + reminder of where you are. +- `wf setup` (in a worktree) runs `worktree_setup` (cwd = worktree, env `WF_MAIN` = main tree): links / + copies git-ignored inputs (`out/…`). Idempotent; run after every worktree create/reuse. +- wf run inside a linked git worktree uses the main tree's project (same relative folder, when it + has `workflow.toml`): TASKS.md, archive, docs, `.wf/` are the main tree's. Bookkeeping never + conflicts; branches never touch TASKS.md. `wf check` there checks the main tree. +- One session alone: as before, on `master`. + +## 3. Done in worktree mode +`wf done` run inside a linked worktree prints, after verify/checklist: commit code (explicit paths), then +`wf merge`. `wf merge` (under the project lock, `.wf/lock`): worktree must be clean → `git rebase master` +(conflict → abort, branch untouched: rebase by hand, resolve, verify, `wf merge` again) → `merge --ff-only` +into the main tree → commit TASKS.md + archive there (`<id> done`, id from the branch `<lane>/<id>`) → +detach the worktree, delete the branch → push home if the remote exists (`--no-push` skips). +The ff-merge of one's own task branch is a standing owner permission (shared CLAUDE.md); never other branches. +`wf finish <id> -m ENTRY [--commit MSG PATH…]` does it in one call (workers): refuses first if files outside +PATHS are uncommitted or quick_gate is red (task stays open) → `wf done` → commit PATHS → `wf merge`. In a main tree +(no worktree) TASKS.md + archive join that commit and nothing is merged. Verify runs before it (`wf ctx` prints it). +Next task: new branch from master in the same worktree. + +## 3a. Project lock +Every wf write command and `wf merge` run under an exclusive flock on `.wf/lock` (main tree), so parallel +sessions/workers never lose a TASKS.md edit or race a merge-back. + +## 4. Commits +Shared rule: stage explicit paths, never `git add -A` / `git commit -a`. diff --git a/docs/orchestrator.md b/docs/orchestrator.md new file mode 100644 index 0000000..845c527 --- /dev/null +++ b/docs/orchestrator.md @@ -0,0 +1,152 @@ +# Orchestrator: one session talks to the owner, fresh workers do the tasks + +Why: cost is dominated by cache reads, which grow with context size. A fresh worker per task costs about +as much less as a cheaper model does; researching in one context and implementing in another is NOT cheaper +(rediscovery, handbacks). So: the orchestrator slices and rules, and one fresh worker per task researches +and implements. + +## Sessions +- **Orchestrator** (interactive, opus, `claude --autocompact 120k`): owner talk about runs, Awaiting rulings, + slicing, spawning workers, post-checks, unattended batches. Reads `wf` output and worker reports, never + code. Holds no lane: never `wf next --as` (use `wf lanes`, `wf list --runner --lane <lane>`). +- **Design session** (optional, big window): brainstorming, specs, hard rulings. Ends each topic with a + committed spec/task; the orchestrator picks it up from TASKS.md. +- **Workers**: the `wf-worker` subagent (`shared/agents/wf-worker.md`; install: + `ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md`), one task each. + +## Starting a session +Skills (install once: `ln -s /projects/public/workflow/shared/skills/wf-orchestrate ~/.claude/skills/` and the same for +`wf-design` and `wf-pilot`): after `/clear`, type `/wf-orchestrate [go | batch N]`, `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]` or `/wf-design [topic | id]` (not +"continue": that is the normal worker cold start). Agents and skills installed while a session runs appear +only in sessions started afterwards: restart it. + +## Runner-ready tasks (slicing rule) +- `Model:` line set to the cheapest lane that fits (sonnet: exact Done + a pattern to copy; haiku: mechanical). +- Done checkable headless: a file, test result or artifact, plus `wf add` for follow-ups. Never "report to + owner" / "owner confirms". +- `After:` satisfied, no `Sessions: owner`. `Sessions: solo` → run it alone. + +## One worker +Two commands do steps 1-3 and 5-7 (one call each): `wf orch pick <lane> [--id ID] [--recovery WHY]` (stop file, +pick, claim, free worktree, recorded in `.wf/orch/<id>.json`, prints the spawn line and the prompt below) and +`wf orch post <id> <lane> --result "<report result line>" [--commit SHA] [--agent A] [--duration S] [--no-pick]` +(post-check, merge if needed, leftover commit, orch + cost log lines, then the next pick or `stop lane <lane>: <why>`). +The steps they automate: + +1. Pick: `wf list --runner --lane <lane>` → first id (lanes: `wf lanes`). Claim: `wf status <id> progress "worker"`. No commit of + its own: `wf merge` or any later bookkeeping commit takes TASKS.md (and the claim) along; harmless. +2. Worktree: `<main>/.worktrees/<lane>`; if a live session or another worker uses it, `<lane>-2`, `-3`, … + (live interactive sessions: `~/.claude/sessions/*.json` field `cwd`). Branch `<lane>/<id>`. +3. Spawn: Agent tool, `subagent_type: wf-worker`, `model: <task Model, none = opus>`, background, no worktree isolation + (it blocks `wf merge` on the main checkout). Prompt: + + ``` + Task: <id> Lane: <lane> Model: <model> + Main tree: <main> Worktree: <path> Branch: <lane>/<id> + Final message: the 4 report lines only. + ``` + (the last line is needed: the agent rule alone did not stop prose before the report.) + Slice job (row starts `slice:` / `wf show` effort > slice_above): same spawn; the worker only slices. + (no `wf ctx` output in the prompt: the worker fetches it; the prompt stays in your context forever.) +4. At most one worker per lane at a time; lanes in parallel. `wf` writes and `wf merge` take the project + lock (`.wf/lock`), so parallel workers are safe. +5. On completion read only the 4-line report (`followups` = ids the worker added). Post-check: + - `done`: `wf log -n 3` shows the id; `git -C <main> branch --list <lane>/<id>` empty; `wf check` 0 errors; + `git -C <path> status --porcelain` empty; `git -C <path> merge-base --is-ancestor HEAD <master>` exit 0 + (worktree HEAD in master: no impl commit left on a detached HEAD; else `cd <path> && wf merge`). + - Everything else: commit leftover bookkeeping + `git -C <main> commit -m "<id> <outcome> (orchestrator)" -- TASKS.md <archive>` if they changed. +6. Outcome: + | Report / state | Action | + |---|---| + | done + post-check ok | next task in the lane | + | model raised by the worker | continue; next pick spawns it on the new model | + | sliced | next task in the lane | + | done+gate-red <culprit> <fix-id> (async gate red, earlier task's break) | post-check as done; next pick in the lane = the P0 fix (no Done/Model → add them first); never stop the lane | + | awaiting / handback / post-check red / no report | stop the lane, tell the owner | + | wip (after a wrap-up) | stop the lane; reslice or rerun later | + | crashed / killed (no report) | one fresh worker, prompt plus `Recovery: <why>` (it inspects git status, git log master..HEAD, `git diff -- TASKS.md` — often just the claim — and `wf show <id>`), then stop | +7. Log one line per task to `out/wf-orch.log` (git-ignored): time, lane, model, id, outcome, commit, duration. + Cost: `wf usage --agent <agent id> --log <id> <outcome> --effort <task estimate: <1h|1h|5h|10h|100h> --lane <lane>` appends one key=value line to + `out/wf-cost.log` (tokens + API-price $ from the subagent transcript). `--duration <s>` adds dur=; `wf usage --report` (med_dur) = per lane/model and + effort: n, done, median/total $, $ per done task, median turns, over every project's log. + The completion notice's token count is NOT the cost (one worker: notice 67k, transcript 5.0M cache reads); + real usage is in the subagent transcript `~/.claude/projects/<project>/<session>/subagents/agent-*.jsonl` + (`wf usage` per session/agent; one request is several transcript entries: counted once; + subagent transcripts lack final output counts → out estimated from content, shown `~`, log `est=N`). + +## Cloud lane +Virtual lane `cloud` (projects with `cloud = true`; design: cloud-lane spec §4.6): tasks run as +`claude --cloud` sessions billed to the cloud budget, not as local workers. +- `wf orch pick cloud [--id ID]` = stop file → ledger (`wf cloud ledger`) → first fitting task across all lanes + (runner-ready, not in progress, `wf cloud` fit rules; opus only, `Cloud: yes` first, then prio, then larger effort) → `wf cloud send` + (claim `in progress: cloud:<sid>`, record `.wf/orch/<id>.json` lane cloud). Nothing to spawn: the session runs + remotely. Otherwise `stop lane cloud: ledger (…)` (balance < reserve, or send refused) | `none fit (…)` | + `max parallel (…)`. A stop of the cloud lane never stops the local lanes. +- Local lanes skip tasks claimed `cloud:` (every `in progress` task is skipped by `wf orch pick <lane>`). +- On each wake-up, at most every 10 min: `wf cloud pull --all` (teleport ~30 s, no tokens). Per ended task it prints + `<id>: <state> (<sid>, $… )` (state done | awaiting | handback | lost; done also `report: commit <sha>`) → + `wf orch post <id> cloud --result <state> [--commit <sha>]` = the worker-report path (done: archive line, branch + `<task lane>/<id>` gone, `wf check`; else leftover commit of the note/claim pull left), log line, next cloud + pick. `running` (exit 4) → nothing; pull again next wake. awaiting / handback / lost → `stop lane cloud`, tell the + owner; the task is back in its local lane (handback note `Recovery: cloud attempt …`). +- Fill: one `wf orch pick cloud` per free slot (up to `max_parallel`), each in its own call. +- Batch (`wf batch` on a `cloud = true` project) pulls itself: the job runs `wf_res.py batch-sidecar` around + `claude -p`; every 10 min while it lives (until `out/wf-batch.stop`) it runs `wf cloud pull --all` and per ended + task `wf orch post <id> cloud --result <state> [--commit <sha>] --no-pick` + a `cloud <id> <state> <sha> (sidecar pull)` + line in the summary, so pulls never starve while the batch waits on foreground workers. The batch orchestrator + only picks cloud, never pulls. Two pulls of one id at once: the second prints `<id>: pulled by another wf cloud + pull, skipped` (rc 0; flock on `.wf/cloud/<id>.json`). + Tail: the orchestrator exited with `.wf/cloud/*.json` records left (sent at batch end) → the sidecar keeps + pulling every 10 min (no picks, no sends) until no record is left, the stop file appears or 24 h passed; the job's + ledger reservation shrinks to 0.2 GB / 1 cpu for it (ledger only, the unit's MemoryMax stays); the summary ends + with `cloud tail ended: <why>; <n> pulls; left: <ids|none> (sidecar)`. A next batch's sidecar may overlap it (record + flock). + +## Unattended batch (nightly, long runs) +The orchestrator does not babysit long runs; a headless batch orchestrator (separate `claude -p` process, +fresh context per batch) does. The owner starts it (auto mode denies an agent launching another unattended agent): `!` + the command, or +once an allow rule `Bash(python3 /projects/public/workflow/wf.py batch:*)` in `~/.claude/settings.json`. + +``` +wf batch N [--lanes sonnet,haiku] [--model opus] [--for 6h] [--mem 4G] # --dry-run prints the command; default mem/for: wf res hist +wf batch --status # newest job + out/wf-batch-*.md +wf batch K --prep [--lanes fast] # prep batch (below) +``` += `wf res run --mem 4G --for 6h --title wf-batch -- claude -p --model opus --permission-mode auto +--permission-prompts none "<templates/batch-prompt.md>"` with `CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS` = `--for`. +The prompt: follow 'One worker' for up to N tasks, never implement; each round one wf-worker per lane in ONE +message, foreground, post-check, one line per task to `out/wf-batch-<stamp>.md`; stop a lane on its first stop +outcome; never ask (follow-ups → `wf add -p 2`, decisions → `wf add -s awaiting`); end with the ids added. + +Why foreground: `claude -p` ends when the orchestrator's turn ends and kills background workers after a +wait ceiling (default 10 min; the env above raises it as a backstop). Workers spawned in one message run in +parallel and the turn waits for all. Background subagents cannot spawn subagents, so the batch +orchestrator is never a subagent. The interactive orchestrator checks `wf batch --status`. + +Headless batch = short fire-and-forget runs (a few tasks, no steering). Long runs: pilot below. + +Prep batch (`--prep`): makes the backlog runner-ready, never implements. Targets = up to K Pending tasks without +Done, not `Sessions: owner`, no status, no open slices (pickable first, then priority); none → prints `prep: … +nothing started`, starts nothing. Prompt `templates/prep-prompt.md`: sonnet subagents (≤ 3 tasks each) read the +task text, `wf set <id> --done "…"` (+ `--model` if missing); unclear → `wf add -s awaiting` + `wf status <id> +blocked <a-id>`. Summary lines `<id> → Done: …` | `<id> → awaiting <a-id>`; then it commits the task file. +Pilot runs one when idle with `not runner-ready` tasks (once per idle spell). + +## Pilot (nested) +`/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]`, typed into an idle project session (phone via Remote Control), +keeps that session tiny: subagents cannot spawn subagents, so it pilots headless batch orchestrators instead of +workers. Loop: `wf batch K --left <time to deadline>` (a wf res job, returns at once) → background Bash `wf res wait <rid>` → its exit +re-invokes the pilot → `wf batch --status` → one line to its context and `out/wf-orch.log` → next batch. A batch = +up to K tasks, rounds of one worker per lane in parallel (2 lanes, K=4 → two rounds). Nothing pickable → +prep batch if tasks are not runner-ready (once per idle spell), then background `wf lanes --wait 1800`, until `--for`. Deadline fit from history: task p90 = p90 of done/handback durations in `out/wf-orch.log` (local lanes, > 0; < 3 runs → 30 min); `--left` launches K' = min(K, floor(left / p90)) tasks with `--for` = left (K' = 0 → `fit: 0 …; nothing started`, the pilot stops); the batch prompt runs `wf batch --time-left <deadline>` before each round and spawns nothing more on `: stop` (left < p90). Steering: the owner talks to the pilot only; stop → +`wf batch --stop` (stop file, checked before every spawn: latency ≤ the running tasks), skip → `wf move <id> deferred`. Batch rc≠0 or no summary → +retry once, then stop. PushNotification on new awaiting items, on a stop and at run end; each alert also an `ALERT <text>` line in `out/wf-orch.log` and in the summary (push may be off). Phone push needs `agentPushNotifEnabled` (settings) / app notifications on. Needs the allow rule +`Bash(python3 /projects/public/workflow/wf.py batch:*)`. + +## Context +Orchestrator large (after design talk) → finish the topic, commit, ask the owner to `/clear`. State lives in +TASKS.md, the log and the batch summaries. +Safe to clear only with no background worker running. Tested (2026-10-05): `/clear` leaves a background +subagent alive and its hand-back + completion notice reach the new context, but SIGKILLs its running Bash +command (exit 137): a worker mid-test or mid-commit loses that step and reports a false failure or half state. diff --git a/docs/resource-ledger.md b/docs/resource-ledger.md new file mode 100644 index 0000000..17a0443 --- /dev/null +++ b/docs/resource-ledger.md @@ -0,0 +1,304 @@ +# wf res — shared resource ledger for agents — design + +Status: approved and implemented 2026-10-01. + +## 1. Goal + +1. No agent job is OOM-killed because another agent started a big job at the same time. +2. The user keeps a guaranteed share of RAM and CPU (gaming), switchable on demand. +3. An agent that cannot get resources now gets one clear line (who holds them, until when) and works on + something else instead of waiting or retrying blindly. +4. Agent-made waste (tmpfs scratch, failed units) is cleaned routinely. +5. Python ≥ 3.11 stdlib, Linux systemd user manager, no daemon. Works in any directory, wf project or not. + +Machine facts (2026-10-01): 16 cores, 30 GB RAM, 7 GB swap, `/tmp` is tmpfs (16 GB), user manager +delegates `cpu io memory pids`; `systemd-run --user --scope --slice=agents.slice` and transient units with +`MemoryMax`/`CPUWeight` verified working. + +## 2. User decisions (2026-10-01) + +| # | Decision | +|---|----------| +| D1 | Default user reserve 6 GB RAM, 4 CPUs. | +| D2 | Gaming reserve 12 GB RAM, 8 CPUs. | +| D3 | Gaming mode slows agents; never freezes or kills them. | +| D4 | Scratch under `/tmp/claude-<uid>/` untouched for 2 h (newest mtime in the entry) may be deleted automatically by any agent, whoever made it. Except a live session's own dir (`<project>/<session-id>`, live = `~/.claude/sessions/<pid>.json` with that `sessionId`, pid alive, `procStart` matching; added 2026-10-02 after an idle session lost its scratchpad). Same automatic clean sweeps top-level litter in `/tmp` and `/var/tmp` (own entries only): empty dirs idle 2 h (`scratch_hours`) and `clr-debug-pipe-<pid>-*` / `dotnet-diagnostic-<pid>-*` of dead pids; entries with content are never swept. Per-project `out/` patterns only list unless `--yes`. | +| D5 | v1 = everything in the proposal: run, status, wait, release, note, queue, game, clean, timer, shared rule. | +| D6 | `game on [--for D]`, default 4 h, turns itself off. | +| D7 | Install a 1-minute user timer via an explicit `wf res timer on`. | +| D8 | `game on` never fails and never squeezes running jobs; it switches budget and weights at once and prints the shortfall and the estimated time until the full reserve is free. | +| D9 | Small unreserved work: headroom (2 GB) in the budget **and** all agent processes inside `agents.slice`: running Claude sessions are moved there live by `wf res adopt` (run by every timer tick; no restart), new ones optionally start there via the `shell-init` alias. | + +## 3. Architecture + +### 3.1 Cgroup layout + +``` +user@<uid>.service +└── agents.slice ← MemoryHigh, CPUWeight, IOWeight (normal / gaming) + ├── run-…scope ← each Claude session (started via the alias), its shells, small jobs + └── agents-jobs.slice + └── wf-r-N.service ← each `wf res run` job: MemoryMax, MemoryHigh, MemorySwapMax=0, Nice=10 +``` + +- `agents.slice` caps the sum of everything agents do. The ledger decides who may start big work; the slice + is the kernel backstop for wrong estimates and unreserved small work. +- Slice properties are set with `systemctl --user set-property --runtime agents.slice …`. The slice unit file + `~/.config/systemd/user/agents.slice` is written on first use (then `daemon-reload`), so the slice exists + before any set-property. +- Normal: `CPUWeight=20`, `IOWeight=20`, `MemoryHigh = RAM − user_reserve_gb`. +- Gaming: `CPUWeight=5`, `IOWeight=5`, `MemoryHigh = max(RAM − game_reserve_gb, agents.slice MemoryCurrent)`; + re-lowered towards the target on every `wf res` call / tick (D8: no squeezing below current use). +- Weights only matter under contention; on an idle machine agents get full speed. + +### 3.2 Code layout + +- `wflib/res.py` — pure, text in / data out: parse `/proc/meminfo`, parse `systemctl show` output, durations + and sizes (`10G`, `40m`, `4h`), prune, capacity rule, queue start order, game slice values and shortfall, + stale-scratch selection (from a given list of (path, newest mtime)), all output lines. +- `wf_res.py` — I/O: lock, atomic ledger write, `/proc` and `/sys/fs/cgroup` reads, filesystem walk, the + runner (`subprocess.run` wrapper, injectable for tests), unit/timer file writing, printing. +- `wf.py` — only registers `wf res …` and dispatches to `wf_res.main(argv)`. `wf res` needs no + `workflow.toml`. + +### 3.3 Files + +- State dir `~/.local/state/wf/` (`$XDG_STATE_HOME/wf` if set): `resources.json`, `resources.lock`, + `logs/r-N.log`, `logs/r-N.rc`, `logs/r-N.peak`, `resources-history.jsonl` (finished rc=0 runs with a peak, + appended when pruned from the ledger after 24 h; newest 2000 kept; read by `wf res hist`). +- Config `~/.config/wf/resources.toml` (`$XDG_CONFIG_HOME`), all keys optional: + ```toml + user_reserve_gb = 6 + user_reserve_cpus = 4 + game_reserve_gb = 12 + game_reserve_cpus = 8 + game_hours = 4 + small_headroom_gb = 2 + scratch_hours = 2 + ``` +- Unit files under `~/.config/systemd/user/`: `agents.slice`, and with `timer on`: `wf-res.service`, + `wf-res.timer`. + +### 3.4 Ledger + +`resources.json`: `{"next": N, "game_until": iso|null, "last_clean": iso|null, "entries": [...]}`. + +Entry fields: `id` (`r-N`, never reused), `project` (git toplevel basename, else cwd basename), `owner` (pid of +the nearest ancestor process named `claude`, else the parent pid of `wf`), `title`, `mem_gb`, `cpus` +(default 1), `est_min`, `state` (`queued` | `running` | `note` | `done`), `unit` (`wf-r-N.service`, run only), +`cmd` (argv list), `cwd`, `log`, `queued` (iso), `started` (iso), `expires` (iso = started + 2 × est_min), +and for `done`: `ended`, `rc` (int or null), `peak_gb`, `why` (`exited` | `expired` | `owner gone` | `released` | a kill reason, e.g. +`killed: oom-kill by systemd-oomd, limit 4.0 GB; raise --mem`). +`env` (caller variables the user manager lacks; `WF_RES_ID` and secrets never), `by` (who to tell, run/note; +keys dropped when unknown): `name` (`--by`, else `WF_SESSION_NAME`, else `<lane> session` from the wf lane +record `.wf/sessions/*.json` with this `CLAUDE_PID`, main tree or toplevel), `task` (`WF_TASK`, else the +`<lane>/<id>` branch of cwd), `batch` (`WF_RES_ID`: every job unit gets `WF_RES_ID=r-N`, so a `wf batch` +worker's entries name the batch), `address` (`uds:$CLAUDE_CODE_MESSAGING_SOCKET`; a batch worker's = its +orchestrator's, reachable via SendMessage). Status lines of live entries end `[by <name> <task> batch r-N, +message uds:…]`; the throttle warning ends `; started by …`. + +Every command takes the lock (`fcntl.flock` exclusive on `resources.lock`), reads, then **prunes**, acts, +writes atomically (temp file in the same dir + `os.replace`), releases. + +Prune (pure, given unit states, live pids, now): +- `running` whose unit is inactive/failed/unknown → `done` (`rc` from `logs/r-N.rc`, `peak_gb` from + `logs/r-N.peak`). No `.rc` (the sh wrapper died with the job) → `rc` null and `why`/`peak_gb` from + `journalctl --user -u wf-r-N.service --since @started` (`Failed with result '…'`, `systemd-oomd killed`, + `… memory peak`), shown in `status` and `wait`; journal silent → `why` points at that journalctl command. +- `running` past `expires` → stays running (job not killed; `run` says so on start, status shows + `ETA overdue (still running, not killed)`), flagged `overdue` in status, and counted at + `max(reserved, used)`. +- `note` whose owner pid is gone or past `expires` → `done`. +- `done` older than 24 h → removed. +- `game_until` in the past → game off (slice back to normal values). + +### 3.5 Capacity rule + +``` +budget_mem = MemAvailable − reserve_gb − small_headroom_gb + − Σ over running: max(0, mem_gb − MemoryCurrent(unit)) + − Σ over notes: mem_gb +budget_cpus = nproc − reserve_cpus − Σ over running and notes: cpus +fits(req) = req.mem_gb ≤ budget_mem and req.cpus ≤ budget_cpus +``` + +- `reserve_*` is the user or the gaming value depending on game mode. +- A note counts its full reservation (its memory cannot be measured; overcounting briefly is safe). +- A running job's memory already in use is inside `MemAvailable`, so only its unused part is subtracted. + +## 4. Commands + +All print short lines; `--json` where stated. Exit: 0 ok · 3 busy (refused) · 1 `wf: …` one line (unknown +id, systemd failure) · 2 usage. + +### 4.1 `wf res run --mem 10G [--cpus N] --for 40m --title "…" [--queue] -- <cmd …>` + +- Fits → `systemd-run --user --slice=agents-jobs.slice --unit=wf-r-N --collect + -p MemoryMax=<mem> -p MemoryHigh=<0.9·mem> -p MemorySwapMax=0 -p Nice=10 + -p WorkingDirectory=<cwd> -p StandardOutput=append:<log> -p StandardError=append:<log> + /bin/sh -c '"$@"; rc=$?; cat /sys/fs/cgroup$(cut -d: -f3 /proc/self/cgroup)/memory.peak > <peak>; echo $rc > <rc>' sh <cmd …>` + (the unit is collected on exit, so the wrapper records exit code and peak bytes itself). Prints `r-N started; log <path>; ETA ~HH:MM`. Exit 0. +- Environment: a user unit starts from the user manager's environment, so the caller's variables that differ + from `systemctl --user show-environment` go along as `--setenv=NAME=VALUE` (stored in the entry, so a queued + job gets them too); skipped: `PWD OLDPWD SHLVL _` and names with TOKEN/SECRET/PASSWORD/PASSWD/CREDENTIAL. +- Does not fit, no `--queue` → exit 3, one line: + `busy: 7.5 GB held by proj-a "dotnet e2e" (r-4) until ~14:40; 2.1 GB free for agents; retry after ~14:40 or work on something else` + (names the holders whose release would make it fit, earliest ETA first; CPU shortage named the same way). +- Does not fit, `--queue` → `queued`, prints `r-N queued, position P; est. start ~HH:MM; cancel: wf res release r-N`. Exit 0. +- A request larger than the whole agent budget on an empty machine → exit 1 `wf: 20 GB can never fit (max …)`. +- `--force` (memory really free though the ledger says busy: unused reservations, stale notes): fits against + MemAvailable − user reserve (game reserve in game mode) − headroom only; ledger claims and the queue are ignored; + the entry is still recorded. Does not fit even so → exit 3 `busy even with --force: only X GB really free beyond + the reserve; …`. The normal busy line ends with `; X GB really free beyond the reserve: --force starts it past + the ledger (only if the holders will not use what they reserved)` when `--force` would fit. +- `systemd-run` failure → entry removed, exit 1 with its first stderr line. +- `--lock KEY`: one queued/running job per KEY and main tree (git common dir's parent: lane worktrees share it) + at a time, e.g. jobs sharing one checkout (a-game-project-builder `../.<project>-gate`). A title starting with the word + `gate` locks `gate` unasked (2026-10-06: two hand-started gates overlapped in one gate checkout). Held → exit 3 + `busy: lock 'KEY' held by r-N "title" (ETA ~HH:MM|queued): …; --queue waits for it (--force does not override + a lock)`; `--queue` → queued `(lock held by r-N)`. Ledger field `lock` = `KEY@<main tree>`. +- The command reaches the job verbatim (no systemd `%` specifier expansion: `--format='%h %s'` is safe). + +### 4.2 Queue + +Strict FIFO by `queued`: the head starts when it fits; entries behind a waiting head wait (a big job is never +starved); an entry whose lock a running job holds is skipped (it neither starts nor blocks the rest). Starting happens inside every `wf res` call (after prune) and every timer tick. A started queued job +runs with the cwd and argv it was queued with; its ETA counts from the actual start. + +### 4.3 `wf res status [r-N] [--json]` + +Without id: one line per entry (id, project, title, state, reserved / used / peak GB, cpus, started, ETA or +`overdue`), then the budget line (`agents may use X GB, Y cpus now; reserve 6 GB/4 cpus`), game mode with +time left and shortfall line (§4.6), unreserved agent memory (`agents.slice` MemoryCurrent − jobs; `(X GB of it file cache, reclaimable)` = +slice `memory.stat` `file` − `shmem`, capped at that figure), warnings: +- running job throttled at its own limit: used ≥ 0.85 × `--mem` **and** its cgroup `memory.pressure` + `some avg60` ≥ 20 % → `r-N throttled at its memory limit (…): likely too small; release --stop, re-run + with a bigger --mem` (page cache alone at the limit does not stall, so no false alarm; the timer tick + has no reader, so it does not warn); +- tmpfs `/tmp` above 25 % of RAM; +- `N claude sessions outside agents.slice (wf res adopt)`. +- `N jobs outside wf res (bare systemd-run): <unit> <used> GB, …`: running transient user services outside + `agents*` slices, except desktop ones (`app-*`, `dbus-*`). Warning only: a service cannot change slice live. + +With id: that entry in full, including `rc`, `peak_gb`, `why` for done entries. + +### 4.4 `wf res wait r-N [--timeout 2h]` + +Polls every 15 s (each poll is a normal locked call, so it also prunes and starts queued work) until the entry +is `done`; prints `r-N done rc=0 peak 9.4 GB in 37 min`. The throttle warning (§4.3) goes to stderr once. Exit 0 when done (whatever `rc`), 1 on timeout. If +reaped, `status r-N` answers from the ledger. + +### 4.5 `wf res release r-N [--stop]` + +- `queued` / `note` → removed. Running unit → refused with `wf: r-N is running; --stop to kill it` unless + `--stop`, then `systemctl --user stop`, entry → done (`why=released`). + +### 4.6 `wf res game on [--for 4h] | off` + +- `on`: `game_until = now + for`; set gaming slice values (§3.1); prints + ``` + game on until 22:40 (4h); CPU/IO now yours + short 3.1 GB of 12 GB: r-4 proj-a "dotnet e2e" 2.5 GB ~21:05, r-7 proj-b "r5 rebuild" 9 GB ~21:30 + full reserve free ~21:30 (est.); free now: wf res release r-7 --stop + ``` + The shortfall/estimate lines appear only when the reserve is not free now. Estimate = ETAs of the running + jobs that must end, in order (jobs are never stopped by game mode). New requests that do not fit are refused + or queued as usual. `on` while on → extends `game_until`. +- `off` or expiry: normal slice values, `game_until = null`. +- No windows over the game: `wf res hook` (Claude Code `PreToolUse` hook, matcher `Bash`, in + `~/.claude/settings.json`; the user adds it, wf never edits dotfiles) prefixes every agent Bash command with + `unset DISPLAY WAYLAND_DISPLAY;` while `game_until > now`, and adds context telling the agent why GUI apps + fail (run headless/offscreen or later). Reads the ledger only (no lock, no systemctl, ~60 ms); off, expired, + other tools or any error → no output, command unchanged. Live for running sessions: state is read per call. + ```json + "hooks": {"PreToolUse": [{"matcher": "Bash", "hooks": [{"type": "command", + "command": "python3 /projects/public/workflow/wf_res.py hook"}]}]} + ``` + +### 4.7 `wf res note --mem 3G [--cpus N] --for 20m [--force] "title"` + +Reserves for foreground work (no unit). Freed when `owner` exits, at `expires`, or by `release`. Refused like +`run` (exit 3) when it does not fit; `--force` as in `run`. + +### 4.8 `wf res clean [--yes]` + +- Automatic part (no `--yes` needed; also run by any `wf res` call when `last_clean` > 10 min ago, and by every + tick): + - every direct and second-level entry under `/tmp/claude-<uid>/` whose newest mtime (recursive) is older than + `scratch_hours` → deleted; lines `freed 1.2 GB: /tmp/claude-1000/-projects-x/<session>` (only printed by + `clean` itself; silent from other commands); + - `systemctl --user reset-failed 'wf-r-*'`. +- Listed only, deleted with `--yes`: per-project patterns from the current project's `workflow.toml` + `cleanup = ["out/prof", "out/history-logs/*.log:30d"]` (glob relative to the project root, optional `:Nd` + minimum age by mtime). Never follows symlinks; never leaves the project root. + +### 4.9 `wf res timer on | off`, `wf res tick` + +- `timer on` writes `wf-res.service` (`ExecStart=python3 /projects/public/workflow/wf.py res tick`) and + `wf-res.timer` (`OnBootSec=1min`, `OnUnitActiveSec=1min`), `daemon-reload`, `enable --now wf-res.timer`. + `off` disables and removes them. +- `tick` = lock, prune, start queued, game expiry and slice re-lowering, automatic clean; then adopt., then + session caps: every running scope in `agents.slice` (a Claude session) gets `MemoryHigh = session_mem_gb` + (config, default 6, 0 = off → `infinity`) via `systemctl --user set-property --runtime`, only where it differs. + MemoryHigh only: a session over it is throttled, never killed (a MemoryMax OOM kill could pick `claude`, whose + oom_score_adj is 200). `status` warns `claude session N at its memory cap (…)` at ≥ 0.9 × cap and ≥ 20 % stall. + `wf res run` jobs live in `agents-jobs.slice`, not in a session, so the cap never limits them. Silent. + +### 4.10 `wf res adopt` + +Moves running Claude sessions into `agents.slice` without a restart (verified 2026-10-01: a live process moved +by `StartTransientUnit` with `PIDs`). Session roots = processes named `claude` whose parent is not `claude`, +outside the slice; each root and its descendants outside the slice go into a new scope +`wf-claude-<root pid>.scope` via +`busctl --user call org.freedesktop.systemd1 /org/freedesktop/systemd1 org.freedesktop.systemd1.Manager +StartTransientUnit 'ssa(sv)a(sa(sv))' wf-claude-<pid>.scope fail 2 PIDs au <n> <pids…> Slice s agents.slice 0`. +Children forked later inherit the scope. Prints one line per session; errors (process gone) are reported, and +the timer retries next minute. + +### 4.11 `wf res shell-init` + +Prints the alias line for `~/.bashrc` (the user adds it; wf never edits dotfiles): +`alias claude='systemd-run --user --scope --quiet --slice=agents.slice claude'`. + +### 4.12 `wf res hist [--project P]` + +Reservation sizes from history: history file + the ledger's finished runs (rc=0, peak known). Grouped by +project and title kind = title minus trailing `(…)` and trailing sha/hex words (7–40 hex chars, ≥ 1 digit): +`gate 2c91cac7 (fix x)` → `gate`. One line per kind: `n`, median request, peak p50/p95, median estimate, +duration p50/p90, suggestion. Suggestion (≥ 3 runs): mem = p95 peak ×1.15 (up to 0.1 GB, ≥ 0.2 GB), for = p90 +duration ×1.5 (≥ 5 min); percentiles nearest-rank. `run`/`note` asking > 2× the suggestion (mem or for) print +`hint: history says ~X GB / Y min (n runs)` on stderr; nothing is resized. `wf batch` without `--mem`/`--for` +uses the `wf-batch` suggestion of the project (not `--prep`; the bg-wait ceiling stays 6h unless `--for`). + +## 5. Shared rule (`shared/CLAUDE.md`, ~4 lines) + +- Job expected > 2 GB RAM, > 4 cores or > 10 min → `wf res run` (never bare `systemd-run`, never a long + foreground command); big foreground step → `wf res note`. +- Exit 3 = busy: note it on the task, do other work, retry after the printed time. No tight polling. + The line names `--force` when the memory is really free (claims unused): rerun with it if the holders will not grow. +- No big scratch in `/tmp` (RAM); use the project's git-ignored `out/`. +- Session end: release your own entries (`wf res status`). + +## 6. Errors + +- Not Linux / no systemd user manager / no cgroup v2 → `wf: wf res needs a systemd user session` exit 1. +- Corrupt ledger → renamed to `resources.json.bad-<ts>`, start empty (id counter salvaged from its text: ids never reused), one warning line. Unknown entry keys (written by a newer version) are ignored, never corrupt: a `wf res wait` started before a release must not wipe the ledger. Starting a job deletes stale `logs/<id>.rc/.peak`. +- Lock is held at most for one command's critical section (no subprocess wait longer than a `systemctl` + call inside the lock; `wait` polls with the lock released between polls). + +## 7. Tests + +- Pure (`tests/test_res.py`, hand-written literals): meminfo parse; size/duration parse; capacity with mixed + running/note entries, used figures, headroom and game reserve; prune (unit gone, pid gone, expired note, + overdue run, 24 h done removal, game expiry); FIFO start order incl. blocked head; busy line holder choice and + wording; game slice values and shortfall/estimate lines; stale scratch selection; cleanup pattern + age. +- I/O (`tests/test_res_io.py`): fake runner records argv for run/stop/set-property/timer; atomic write; corrupt + ledger recovery; two real processes reserving at once under the real lock never overbook (budget faked). +- Smoke (manual, release step): real `wf res run --mem 100M --for 1m -- true`, `wait`, `status`. + +## 8. Release + +Slices (≈1 h each): pure core · ledger+lock+run/status/release · wait/rc/peak · queue · note · game+slice · +clean · timer+tick+shell-init · shared rule + CHANGES + release. Release rules of the workflow repo apply; +after release the workflow session runs `wf res timer on` once and tells the user to add the +`shell-init` alias. diff --git a/scripts/publish_snapshot.py b/scripts/publish_snapshot.py new file mode 100755 index 0000000..bcbb002 --- /dev/null +++ b/scripts/publish_snapshot.py @@ -0,0 +1,159 @@ +#!/usr/bin/env python3 +"""Export a scrubbed snapshot of the tool repo into a separate local git repo (fresh history). + +Usage: publish_snapshot.py [DEST] [--src REPO] [--ref REF] [--denylist FILE] + +- DEST (default ~/src/wf-public): created and `git init`ed on the first run. +- The tree of REF (default master) is copied; inbox.md, .worktrees/, __pycache__/, out/ are dropped. +- Denylist (default <src>/publish-denylist.local, git-ignored): one private word per line, `#` comments; + matched case-insensitively against every path and file content. Any match → nothing is written, exit 1. +- First run → one commit 'initial public snapshot'; later runs → one commit whose message is the + CHANGES.md lines new since the last snapshot; no change → no commit. +- Never pushes, never adds a remote: pushing is a manual step (docs/manual.md). +""" +import argparse +import io +import os +import shutil +import subprocess +import sys +import tarfile +from pathlib import Path + +EXCLUDE_NAMES = {"inbox.md", ".worktrees", "__pycache__", "out"} +DENYLIST_NAME = "publish-denylist.local" + + +class Fail(Exception): + pass + + +def git(cwd, *args, data=None): + r = subprocess.run(["git", *args], cwd=cwd, input=data, capture_output=True) + if r.returncode: + raise Fail(f"git {' '.join(args)}: {r.stderr.decode(errors='replace').strip()}") + return r.stdout + + +def excluded(path): + parts = path.split("/") + return any(p in EXCLUDE_NAMES for p in parts) + + +def read_tree(src, ref): + """{relative path: (bytes, mode)} of REF's tree, excluded paths dropped.""" + tar = tarfile.open(fileobj=io.BytesIO(git(src, "archive", "--format=tar", ref))) + files = {} + for m in tar.getmembers(): + if not (m.isfile() or m.issym()) or excluded(m.name): + continue + if m.issym(): + files[m.name] = (m.linkname.encode(), "link") + else: + files[m.name] = (tar.extractfile(m).read(), m.mode) + return files + + +def load_denylist(path): + if not path.is_file(): + raise Fail(f"no denylist at {path} (one private word per line; git-ignored) — refusing to publish") + words = [w.strip() for w in path.read_text().splitlines()] + words = [w for w in words if w and not w.startswith("#")] + if not words: + raise Fail(f"denylist {path} is empty — refusing to publish") + return words + + +def scan(files, words): + """Lines 'path[:line]: word' for every denylist hit.""" + low = [w.lower() for w in words] + hits = [] + for path in sorted(files): + for w, wl in zip(words, low): + if wl in path.lower(): + hits.append(f"{path}: {w} (path)") + text = files[path][0].decode("utf-8", errors="ignore").lower() + if not any(wl in text for wl in low): + continue + for n, line in enumerate(text.splitlines(), 1): + for w, wl in zip(words, low): + if wl in line: + hits.append(f"{path}:{n}: {w}") + return hits + + +def changes_lines(data): + return [l for l in data.decode("utf-8", errors="replace").splitlines() if l.startswith("- ")] + + +def write_tree(dest, files): + for p in dest.iterdir(): + if p.name == ".git": + continue + shutil.rmtree(p) if p.is_dir() and not p.is_symlink() else p.unlink() + for path, (data, mode) in files.items(): + f = dest / path + f.parent.mkdir(parents=True, exist_ok=True) + if mode == "link": + os.symlink(data.decode(), f) + else: + f.write_bytes(data) + os.chmod(f, 0o755 if mode & 0o111 else 0o644) + + +def publish(src, dest, ref, denylist): + words = load_denylist(denylist) + files = read_tree(src, ref) + hits = scan(files, words) + if hits: + raise Fail("denylist matches, nothing published:\n" + "\n".join(hits)) + first = not (dest / ".git").exists() + if first: + if dest.exists() and any(dest.iterdir()): + raise Fail(f"{dest} exists, is not empty and not a git repo") + dest.mkdir(parents=True, exist_ok=True) + git(dest, "init", "-q", "-b", "master") + old_changes = [] + else: + old = dest / "CHANGES.md" + old_changes = changes_lines(old.read_bytes()) if old.is_file() else [] + write_tree(dest, files) + git(dest, "add", "-A") + if not first and not git(dest, "status", "--porcelain").strip(): + return "nothing new: no commit" + if first: + msg = "initial public snapshot" + else: + new = [l for l in changes_lines(files.get("CHANGES.md", (b"", 0))[0]) if l not in set(old_changes)] + msg = "public snapshot\n\n" + ("\n".join(new) if new else "- (no new CHANGES.md lines)") + ident = [] # a fresh dest has no identity of its own: commit as the source repo's user + for key in ("user.name", "user.email"): + r = subprocess.run(["git", "config", key], cwd=src, capture_output=True, text=True) + if r.stdout.strip(): + ident += ["-c", f"{key}={r.stdout.strip()}"] + git(dest, *ident, "commit", "-q", "-F", "-", data=msg.encode()) + sha = git(dest, "rev-parse", "--short", "HEAD").decode().strip() + return f"committed {sha} in {dest} ({'first' if first else 'update'}; not pushed — see docs/manual.md)" + + +def main(argv=None): + here = Path(__file__).resolve().parent.parent + ap = argparse.ArgumentParser(prog="publish_snapshot.py", description=__doc__.splitlines()[0]) + ap.add_argument("dest", nargs="?", default=str(Path.home() / "src" / "wf-public"), + help="target repo (default ~/src/wf-public)") + ap.add_argument("--src", default=str(here), help="tool repo (default: this script's repo)") + ap.add_argument("--ref", default="master", help="source ref (default master)") + ap.add_argument("--denylist", help=f"private-word file (default <src>/{DENYLIST_NAME})") + a = ap.parse_args(argv) + src = Path(a.src).expanduser().resolve() + deny = Path(a.denylist).expanduser() if a.denylist else src / DENYLIST_NAME + try: + print(publish(src, Path(a.dest).expanduser().resolve(), a.ref, deny)) + except Fail as e: + print(f"wf: {e}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/shared/CLAUDE.md b/shared/CLAUDE.md new file mode 100644 index 0000000..8e293cc --- /dev/null +++ b/shared/CLAUDE.md @@ -0,0 +1,70 @@ +# CLAUDE.md — /projects (shared workflow) + +Project CLAUDE.md wins on conflict. `wf` = `python3 /projects/public/workflow/wf.py` (run in the project; `wf -h`). +No `workflow.toml` → rules still apply, without wf. + +## Style +Extremely concise, sacrifice grammar: replies, subagent prompts, reports, task entries, commits. +Full and exact: compaction / session summaries (every decision, id, path, number, open question), specs, +plans, documents for the user. Terse ≠ less work: tests, checks, recorded decisions stay. +Independent tool calls → one message; chain short dependent commands in one Bash call. + +## Cold start ("continue" / "next task" / "resume") +No Qs. `wf next --as <haiku|sonnet|opus> [--lane <lane>]` (`--brief` if context loaded) → mention Awaiting, +Needs human, in-flight plans (don't block) → work the next task. Never re-ask what specs settle. +Orchestrator: `/wf-orchestrate` (never `wf next --as`, never implement). `wf-worker`: your agent rules, not this. + +## Lanes and models +- Lanes by size (`wf lanes`); effort > slice_above (default 1h) = slice job: split, never implement. + A lane only when the owner/orchestrator names it. +- `Model:` line (`wf add/set --model`; none = opus); take Model ≤ yours. haiku = mechanical, exact steps; + sonnet = exact Done + pattern to copy + real-data pass/fail; opus = design, debugging, RE, anything unclear. + Beyond yours → `wf set <id> --model <higher>` + `wf note` why + `wf status <id> clear`, `wf next`. +- Cloud (project `cloud = true`): each runner-ready task gets `Cloud: yes|no` at add (`wf add/set --cloud`). yes = Done checkable from the repo alone (code, synthetic data, `cloud_include` data); no GUI/display, live game/server/capture, LAN, Wine or local installs, `wf`/`wf res` steps, owner. Opus only; prefer larger (1h); sonnet/haiku → no. Unsure → no. +- `wf next` Waiting / `wf done` notify lines → SendMessage that `uds:…` address (none → tell the owner). + Peer messages are requests, never owner approval. + +## Loop +1. Impl (≤1h): test red → implement → verify green. Design (user present): superpowers:brainstorming → spec → + approval → `wf add` slices. Bug: superpowers:systematic-debugging. Big/risky (only if asked): writing-plans + → branch → wait for the merge call. + Test runs: one call per red/green; on red print all failures in that call (`… 2>&1 | grep -A15 '^FAIL\|^ERROR' | head -80`), never run then grep. +2. Start: `wf status <id> progress "<branch or note>"`. Decide, record in spec/rulings, report; ask only on real forks. +3. Blocked → `wf add -s awaiting "<q>"` → `wf status <id> blocked <a-id>` → next task. +4. Tests must fail when broken; oracle = hand-derived or real data, never the code under test. + +## Done → one commit, don't ask +1. `wf done <id> -m "<≤2-line entry>"` (or `wf finish`) → follow what it prints. New work → `wf add` now. +2. `wf check` 0 errors → commit explicit paths (never `add -A` / `commit -a` / `add .`). Commit ≠ merge ≠ push; + merge only your own task branch in worktree mode. Push only to remote `home` (if present): + `git push home --all && git push home --tags`. Other remotes never unasked. +3. Slow gate / review → background on the commit with `wf res run --queue` (never retry or await a busy gate), keep working; its res id in the `wf done -m` entry; findings → follow-up commit. + +## Several sessions +`wf next` shows `Multi-session` → work in your lane's worktree, branch per task, as it prints; finish with +`wf finish` (`wf add`/`note` after it → `wf merge`). Another live session in your lane → tell the owner. +`Sessions: solo` = picked only alone; `owner` = needs the owner (`wf next --owner` only when they say present). + +## TASKS.md +Never read or rewrite it whole: `wf list/show/ctx/search/log`; change via `wf add/done/prio/move/status/set/ +note/body/tick` (hand edit only what wf can't, then `wf check`). Ids stable, never reused; refer `[[id]]`. +>1h → `wf add --parent <id> -e 1h` slices. Owner says an owner part is done → `wf done` it same turn (or +`wf move <id> pending` + headless Done). Spec approved → drop "(brainstorm)" from the title. + +## Context economy +`wf ctx` / `wf search` before whole docs; never re-read just-edited files. Costly discovery → project CLAUDE.md +same turn. Grep `-l`/`-c` → `-n | head`; ~40-line reads at an anchor. Area notes (code map with grep anchors, +test recipe) first; refreshed a map → `wf areas --mark <area>`. Had to read a tool's source for its usage → +fix its `-h`. + +## Memory / CPU (`wf res`) +Job > 2 GB, > 4 cores or > 10 min → `wf res run --mem … --for … --title … -- cmd` (never bare `systemd-run`, +nor in scripts; never a long foreground command); big foreground step → `wf res note`. Named session → +`WF_SESSION_NAME=<name>` / `--by`. Exit 3 = busy: note it, other work, retry after the printed time, no tight +polling; memory really free → `--force`. No big scratch in `/tmp`: git-ignored `out/`; delete temp dirs you +create; never delete what you did not create (only the consuming task deletes, noted in `wf note`). Session end: release your `wf res` entries. `game on` → no GUI, headless or later. + +## Workflow problems +wf bug, unclear rule, missing command, repeated manual work → `wf report "<what>" --kind bug|idea|friction`, +work around (hand edit + `wf check`), keep going. Never edit `/projects/public/workflow` from a project session. +Project-only need → project CLAUDE.md / workflow.toml. diff --git a/shared/agents/wf-worker.md b/shared/agents/wf-worker.md new file mode 100644 index 0000000..555b20c --- /dev/null +++ b/shared/agents/wf-worker.md @@ -0,0 +1,53 @@ +--- +name: wf-worker +description: Works exactly one wf task (id, lane worktree and branch given in the prompt) to done, awaiting or handback, then ends with a 4-line report. Spawned by a wf orchestrator; never picks tasks itself. +tools: Bash, Read, Edit, Write +permissionMode: auto +--- + +# wf worker + +You are a worker spawned by a wf orchestrator. Nobody answers questions; a denied permission is final. +The prompt gives: task id, lane, project main tree, worktree path, branch. First run `wf start` (below): +it prints the task, its refs and relations. `wf` = `python3 /projects/public/workflow/wf.py` (no alias in your shell). + +## Start +- One call, in the main tree: `cd <main> && wf start <id> --worktree <path> --branch <branch>` (prompt has + `Recovery: <why>` → add `--recovery`: a dead worker's dirty WIP is kept and shown; keep or reset it) = + worktree + branch, `wf setup`, status progress, `wf ctx`, a ready verify && finish line. Refused/failed → hand back. +- Run every command inside the worktree (`cd <path> && …`); never edit files in the main tree. + +## Work +- Explore cheap: area notes (code map, test recipe) first; grep -l/-c → -n | head; ~40-line reads; never page whole files. +- Only the given id. Never `wf next`. The Loop rules of the shared CLAUDE.md apply (test red → green, verify). +- Tests: run the target test file first; on red print all failures in the same call + (`unittest … 2>&1 | grep -A15 '^FAIL\|^ERROR' | head -80`), never run then grep in two calls. +- No background jobs (no `run_in_background`, no Monitor): long commands run in the foreground, or + `wf res run` and wait for it in the foreground. The end of your turn is the end of your work. +- Blocked on a decision → `wf add -s awaiting "<question>"`, `wf status <id> blocked <a-id>`, stop. +- Slice job (`wf next`/`wf show` says effort > slice_above): no code; `wf add --parent <id>` slices ≤ 1h, each with Steps/Done/Ref + Model; `wf note`; `wf status <id> clear`; report outcome `sliced`. +- Beyond your model → `wf set <id> --model <higher>`, `wf note <id> "<why>"`, `wf status <id> clear`, stop. +- Told to wrap up → one call in the worktree: `wf wip <id> -m "<state + next step>" --commit "<msg + footer>" <paths>` (commits WIP, notes, status clear), stop. +- Hygiene: `wf add` titles without the id; refs only `[[id]]`, never `#N`; one Co-Authored-By footer. + +## Done +Run the `Verify` commands `wf ctx` printed (long ones via +`wf res run`, waited for in the foreground; red → fix or hand back), then ONE call: +`wf finish <id> -m "<entry>" --commit "<msg + footer>" <explicit paths>` (in the worktree; no code → no `--commit`) += quick gate → done → commit → `wf merge`; a refusal (gate red, stray uncommitted files) leaves the task open: fix, rerun. +Paths = this worktree's files only (another repo's: commit there first, finish without them); never hand-edit or commit +TASKS.md/archive (wf writes the main tree's, `wf merge` commits them). Failed after `done:` → fix, rerun the same `wf finish` (it resumes). +Chain it after the verify in the same Bash call when they are short (`wf start` printed that line). A verify step that says background / gate on a commit → run it after `wf finish`, in the +foreground, from the main tree (`cd <main>`, so logs land in its `out/`), on the merged sha +(`git -C <main> rev-parse master`); red → `wf finish`/`wf gate` red prints the procedure (culprit, P0 fix task, `done+gate-red` report); same steps for a red post-merge gate. Rebase conflict → `git rebase master`, resolve, verify, `wf merge` again; still red → hand back. +Every loose end, open item or follow-up you would mention → `wf add -p 2 --model <model> -e <effort> --done "<check>" "<title>. <goal>"` +now (cloud project: add `--cloud yes|no` per the Cloud rule); prose is lost. Its id goes in the report's `followups` line. + +## Report +Your final message is exactly these 4 lines, no text before or after (details belong in the archive +entry, `wf note` or new tasks — the orchestrator never reads prose): + + id: <id> + result: done | done+gate-red <culprit> <fix-id> | awaiting <a-id> | handback <why> | wip + commit: <merged sha or -> (copy `wf finish`'s last line `report: commit <sha> [tool <sha>]`) + followups: <ids from wf add, or none> diff --git a/shared/skills/wf-design/SKILL.md b/shared/skills/wf-design/SKILL.md new file mode 100644 index 0000000..e9f02be --- /dev/null +++ b/shared/skills/wf-design/SKILL.md @@ -0,0 +1,30 @@ +--- +name: wf-design +description: Use when the owner starts or resumes this session as the project's wf design / rulings / planning session (after /clear, "design session", "rulings"). Optional args - a topic or task id to start with. +disable-model-invocation: true +--- + +# wf design session + +You turn questions into decisions and decisions into runner-ready tasks. A separate orchestrator session +runs the tasks (guide: /projects/public/workflow/docs/orchestrator.md); you never spawn workers. + +## Cold start +1. `wf list -s awaiting` (with `wf show <a-id>` for each), `wf list -s human`, `wf list --model opus` + (tasks needing design), new reports in `out/wf-batch-*.md` / `out/wf-orch.log` that ask for rulings. +2. Show the owner in ≤ 12 lines: open questions (id + one line), design tasks, your proposed order. +3. Args: topic or id → start there. None → the owner picks. + +## Work +- Rulings: decide with the owner (recommend; ask only on real forks) → record where it binds (spec + rulings section or the task body) → `wf done <a-id>` / `wf status <id> clear` to unblock. +- New design: REQUIRED SUB-SKILL superpowers:brainstorming → spec committed. Multi-step build: + superpowers:writing-plans. +- Output is always runner-ready tasks for the orchestrator: `wf add` slices ≤ 1h, `Model:` the cheapest lane + that fits (sonnet = exact Done + pattern to copy + real-data pass/fail; haiku = mechanical; opus = judgement), + Done checkable headless (file/test + `wf add` follow-ups, never "report to owner"), `Ref:` to the spec. +- Task bodies: add `Code: <file> <anchor>; test to copy: <file>::<name>` only when already in your context (no extra research). +- Stale owner state: owner part already done → `wf done` or `wf move <id> pending` + headless Done; title + still "(brainstorm)" after spec approval → `wf set <id> --title` without it. Check notes before asking. +- Implement only when the owner says so. Commit bookkeeping and specs as you go. +- Topic finished and context > ~200k → tell the owner "/clear, then /wf-design". diff --git a/shared/skills/wf-orchestrate-loop/SKILL.md b/shared/skills/wf-orchestrate-loop/SKILL.md new file mode 100644 index 0000000..dcdfde8 --- /dev/null +++ b/shared/skills/wf-orchestrate-loop/SKILL.md @@ -0,0 +1,8 @@ +--- +name: wf-orchestrate-loop +description: Retired - use /wf-pilot. Only when the owner types /wf-orchestrate-loop. +disable-model-invocation: true +--- + +Retired 2026-10-05. Tell the owner in one line: use `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]` +(N tasks → `--batch`), then run /projects/public/workflow/shared/skills/wf-pilot/SKILL.md with their args. diff --git a/shared/skills/wf-orchestrate/SKILL.md b/shared/skills/wf-orchestrate/SKILL.md new file mode 100644 index 0000000..53a4785 --- /dev/null +++ b/shared/skills/wf-orchestrate/SKILL.md @@ -0,0 +1,42 @@ +--- +name: wf-orchestrate +description: Use when the owner starts or resumes this session as the wf orchestrator of the project (after /clear, "orchestrate", "be the orchestrator"). Optional args - "go" (start the proposed run without asking), "batch N" (prepare an unattended batch of N tasks). +disable-model-invocation: true +--- + +# wf orchestrator + +`wf` = `python3 /projects/public/workflow/wf.py`. You are this project's orchestrator: you pick, spawn `wf-worker` subagents, post-check, report. You never +implement, never read code, hold no lane (never `wf next --as`). Background and the unattended batch +command: /projects/public/workflow/docs/orchestrator.md — read it only when a step below says so. + +## Cold start (status first, no questions) +1. `wf lanes --unregister` · `wf list -s awaiting` · `wf list --runner` · `wf res status` · `tail -n 20 out/wf-orch.log` · + `wf batch --status`. +2. Report in ≤ 10 lines: awaiting ids, ready per lane, batches running/finished + stop reasons, stale + "in progress" tasks (no live worker). +3. Propose the run (lanes, ids, interactive or batch). `go` → start; `batch N` → print + `wf batch N --lanes <lanes>` for the owner to start with `!` (or their allow rule); status: `wf batch --status`; none → wait for a yes. + +## One worker (repeat per lane, one per lane at a time, lanes in parallel) +1. `wf orch pick <lane>` (lanes: `wf lanes`) = stop-file check, pick, claim, free worktree, prompt. `none:` / `stop:` line → spawn nothing. + Slice job (it says so): same spawn; the worker only slices. +2. Agent: `subagent_type: wf-worker`, `model: <printed model>`, background, no isolation. Prompt = the printed lines, verbatim. +3. Report (4 lines) → `wf orch post <id> <lane> --result "<result line>" --commit <sha> --agent <agent id> --duration <duration_ms/1000>` + = post-check (done: archive line, branch gone, worktree clean + in master, else it runs `wf merge`; `wf check`), leftover + bookkeeping commit otherwise, `out/wf-orch.log` + `out/wf-cost.log` lines, then the lane's next pick (spawn it) or `stop lane …`. +4. `stop lane` → tell the owner (awaiting/handback/post-check-red/wip). No report (crash) → it prints the one recovery pick + (`wf orch pick <lane> --id <id> --recovery "<why>"`), then stop. done+gate-red with a not-runner-ready fix → add its Done/Model, run the printed pick. + +Cloud lane (project `cloud = true`): `wf orch pick cloud` → sends the first fitting task (`wf cloud send`), no agent to spawn; +repeat until `stop lane cloud: ledger|none fit|max parallel`. Each wake (≥ 10 min apart): `wf cloud pull --all`; each +`<id>: <state> …` line (not `running`) → `wf orch post <id> cloud --result <state> [--commit <sha from report: commit>]`. + +Unattended run (overnight, steerable, never asks; session stays tiny, runs headless batches): `/wf-pilot [--batch 4] [--lanes a,b] [--for 8h]`. + +## Keep small +Task bodies: add `Code: <file> <anchor>; test to copy: <file>::<name>` only when already in your context (no extra research). +Don't paste task bodies into prompts, don't re-read the guide, don't run /context. Not runner-ready or a +design question → `wf add -s awaiting "<question>"` for the design session (tasks: `wf add -p N …`). Topic done and context large → +tell the owner "/clear, then /wf-orchestrate" — only when no background worker runs (/clear kills its +running Bash command; notice still arrives). diff --git a/shared/skills/wf-pilot/SKILL.md b/shared/skills/wf-pilot/SKILL.md new file mode 100644 index 0000000..1f01f2b --- /dev/null +++ b/shared/skills/wf-pilot/SKILL.md @@ -0,0 +1,44 @@ +--- +name: wf-pilot +description: Use when the owner types /wf-pilot in a project session (often an idle session opened from the phone app via Remote Control) to pilot an overnight run of headless `wf batch` runs; the session itself stays tiny. Args - "[--batch 4] [--lanes a,b] [--for 8h]" (tasks per batch default 4; lanes default all; deadline default 8h). +disable-model-invocation: true +--- + +# wf pilot (nested overnight run, never asks) + +`wf` = `python3 /projects/public/workflow/wf.py`. You pilot headless batch orchestrators (`wf batch K`, each a fresh +`claude -p` that spawns the workers) — never pick, spawn workers, implement or read code; never `wf next --as`. +Background: /projects/public/workflow/docs/orchestrator.md 'Pilot (nested)'. +Cloud lane (project `cloud = true`) runs inside each batch: its prompt makes the batch run `wf orch pick cloud` until stop, `wf cloud pull --all` at most every 10 min, `wf orch post <id> cloud` per ended task; a cloud stop never stops local lanes (pilot needs no extra step). Needs the owner allow rule +`Bash(python3 /projects/public/workflow/wf.py batch:*)`; denied → tell the owner, end. + +## Start (no proposal, no questions) +K = `--batch` (4); deadline = now + `--for` (8h); lanes = `--lanes` or all. `wf lanes` · `wf list -s awaiting` +(remember the awaiting ids). One status line to the owner, then go. + +## Loop +1. Launch: `wf batch K [--lanes …] --left <time left>` (K shrinks to the tasks that fit by history, `--for` = left) → `<rid> started` → background Bash + `wf res wait <rid> --timeout <same>`; end the turn. `fit: 0 of K …; nothing started` → stop (step 6, why: deadline). Exit 3 (busy) → background Bash `sleep 600`, then retry. +2. Wait exit → `wf batch --status` (newest summary) → one line to the owner and `out/wf-orch.log`: + `<HH:MM> pilot batch <n> rc=<rc> done=<ids> stopped=<lanes: why> added=<ids>`. No task bodies, no logs. + A lane the batch stopped (handback/awaiting) is not relaunched in the next batch (`--lanes` without it) until the task/fix that stopped it is done, or the owner says so. +3. Failure (rc≠0, or no new summary file) → retry once; second failure → stop (step 6). + Memory: `wf res` throttled warning on a running batch → never kill it (workers mid-task); log it (`<HH:MM> pilot batch <n> throttled`); + next batch `--mem` = 1.5 × last (cap per `wf res free`). Batch oom-killed → re-run the same N with 1.5 × `--mem`. +4. `wf list -s awaiting` has new ids → alert `wf-pilot <project>: awaiting <ids>`; keep going. +5. Next: deadline passed → stop. Else `wf lanes`: a lane pickable → step 1. None, and a lane shows + `not runner-ready` → prep batch once per idle spell (until a batch ran again): `wf batch 10 --prep [--lanes …] --for 1h` + (`prep: … nothing started` → skip it) → background `wf res wait <rid>` → `wf batch --status` → log line + `<HH:MM> pilot prep rc=<rc> done=<ids> awaiting=<ids>` (new awaiting → alert, step 4) → step 5 again. Else background Bash + `wf lanes --wait <min(1800, seconds to deadline)>` (0 → step 1; 1 → poll again; 2 = stop file → delete it, stop). Seconds left < 60 → stop, no wait. +6. Stop: append a summary line to `out/wf-orch.log` (batches, done ids, added ids, why stopped), tell the owner + in ≤ 5 lines, alert `wf-pilot <project>: ended (<why>)`, end. Never kill a running batch. + +Alert = PushNotification AND an `ALERT <text>` line in `out/wf-orch.log` AND in the summary line/owner message (push may be off; the alert must survive). + +## Owner messages (win over the loop) +- stop → `wf batch --stop` (the running batch spawns nothing more: stops once its running workers finish), then stop after its exit. +- skip <id> → `wf move <id> deferred` + `wf note <id> "pilot: skipped by owner"` (a running batch no longer + picks it; one already on it finishes). +- pause / change lanes / batch size → apply from the next launch. Status? → `wf batch --status`, one line. +- Never /clear; context stays one line per batch. diff --git a/templates/CLAUDE.md b/templates/CLAUDE.md new file mode 100644 index 0000000..88f70b4 --- /dev/null +++ b/templates/CLAUDE.md @@ -0,0 +1,21 @@ +# CLAUDE.md — {name} + +Project rules; win over the shared /projects/CLAUDE.md on conflict. Terse. Update when you re-learn something. + +## Layout +- <dirs / entry points, one line each> + +## Loop additions +- Verify: <command> + +## Gotchas +- <env / layout / tool surprise, found at cost> + +## Areas +### <area> +- Code map: <grep anchors: function / class / string names, no line numbers> +- Test recipe: <command for this area, fixture / data notes> +- Tools: <scripts / CLIs, one line each: what, how to run> +- Worktree setup: <steps a fresh worktree needs: venv, submodules, data links> +- Paths: <globs the area's code lives in; none = whole repo> +- Checked: <short sha: wf areas --mark <area> after a refresh> diff --git a/templates/TASKS.md b/templates/TASKS.md new file mode 100644 index 0000000..2148242 --- /dev/null +++ b/templates/TASKS.md @@ -0,0 +1,12 @@ +# Tasks — {name} + +Format and rules: /projects/CLAUDE.md. Done history: [`tasks/archive.md`](tasks/archive.md). +Look and change with `wf` (`wf -h`), not by hand. + +## Awaiting your decision + +## Pending + +## Needs human + +## Deferred diff --git a/templates/archive.md b/templates/archive.md new file mode 100644 index 0000000..14e4146 --- /dev/null +++ b/templates/archive.md @@ -0,0 +1 @@ +# Task archive (newest first) diff --git a/templates/batch-prompt.md b/templates/batch-prompt.md new file mode 100644 index 0000000..1a0f338 --- /dev/null +++ b/templates/batch-prompt.md @@ -0,0 +1 @@ +You are a batch orchestrator. Follow {here}/docs/orchestrator.md 'One worker' for up to {n} tasks (lanes: {lanes}), never implement. Before EVERY spawn (each round, a lane's P0 fix, a recovery worker) run `test -e {stop}`: it exists → spawn nothing, let running workers finish (never kill one), post-check them, append 'stopped: stop file' to {out}, delete the stop file, exit. Else, before each round run `wf batch --time-left {deadline}`: its line ends in ': stop' (less time left than a task takes, p90) → spawn nothing new, let running workers finish, post-check them, append 'stopped: deadline (<that line>)' to {out}, exit. Else each round: `wf orch pick <lane>` per lane, spawn one wf-worker per picked lane in ONE message (model and prompt: as printed), foreground (not background), wait for all, post-check each with `wf orch post <id> <lane> --result "<result line>" --commit <sha> --agent <agent id> --duration <s> --no-pick`, append one line per task to {out} (it exists with its '# wf-batch' header: only append, never rewrite) (time, lane, id, outcome, commit), repeat. If the project has `cloud = true`, the virtual lane cloud also runs each round (nothing to spawn): `wf orch pick cloud` until a `stop lane cloud` / `none fit` / `max parallel` line; never run `wf cloud pull` or `wf orch post <id> cloud` yourself: the batch job's sidecar pulls every 10 min (`wf cloud pull --all` → `wf orch post <id> cloud … --no-pick`, a 'cloud <id> <state> … (sidecar pull)' line in {out}), also while you wait on workers; a cloud stop never stops the local lanes. Stop a lane on its first stop outcome, and keep it stopped (never resume because some other task is pickable) until the task/fix that stopped it is done or the owner says so; done+gate-red <culprit> <fix-id> is not one (post-check as done, the lane picks the P0 fix next). Never ask the owner: follow-ups → wf add -p 2 --model <model>, decisions → wf add -s awaiting. At batch start note `git branch --list 'fast/*' 'slow/*'` (and worktrees): a branch already there is pre-existing; report it ONCE in {out} as 'pre-existing branch <name>' (+ the awaiting id when a task/awaiting names it), never per round or per task. End with the ids added in that file, then exit. diff --git a/templates/cloud-prompt.md b/templates/cloud-prompt.md new file mode 100644 index 0000000..6020e90 --- /dev/null +++ b/templates/cloud-prompt.md @@ -0,0 +1,38 @@ +You are a remote worker for one task in a snapshot repo (branch main, base commit {{base}}). +Rules (they override CLAUDE.md where they differ): +- The task tool `wf` is absent: ignore wf / TASKS / lane / worktree / gate rules in CLAUDE.md; never create or edit TASKS.md or the archive. +- Never ask questions. Blocked on a decision you cannot make: end with result awaiting and the one-line question as the report. +- Work: test red -> implement -> green (test recipe below); commit on main, one-line messages; never rewrite {{base}}. +- A needed network source is blocked: say so in the report, continue with what the repo has. +- When done, measure your usage with the usage script at the end of this prompt (it sums your own transcript). +- Commit only what the task needs: `git add <paths>`, never -A or `.`; no build output (__pycache__, bin, obj). +- Then build the patch block: git format-patch --binary {{base}}..HEAD --stdout | gzip -9 > /tmp/wf.patch.gz; echo "WF-PATCH-BEGIN sha256=$(sha256sum < /tmp/wf.patch.gz | cut -c1-64) bytes=$(wc -c < /tmp/wf.patch.gz)"; base64 -w 76 /tmp/wf.patch.gz; echo WF-PATCH-END +- Your final message is plain text (not a tool call, no code fences, no other text), first line = the result line. + Every line keeps its key word (WF-RESULT, WF-REPORT, WF-USAGE, then the patch block) literally; replace only <...>: + WF-RESULT <done, awaiting or handback> + WF-REPORT <one or two lines: what changed, tests run + result, open questions> + <the WF-USAGE line the usage script printed> + <the patch block: its WF-PATCH-BEGIN line, the base64 lines, WF-PATCH-END> + No commits (awaiting / handback without work): still write the patch block, for the empty patch. + +Task {{id}}: +{{task}} + +Test recipe (area notes): +{{recipe}} +{{note}} +Usage script (run as is): +python3 - <<'PY' +import json, glob, os +f = max(glob.glob(os.path.expanduser("~/.claude/projects/*/*.jsonl")), key=os.path.getmtime) +m = {} +for line in open(f): + try: + j = json.loads(line); x = j.get("message") or {} + if x.get("usage"): m[x.get("id") or j.get("uuid")] = x + except Exception: pass +t = [sum(x["usage"].get(k) or 0 for x in m.values()) for k in + ("input_tokens", "cache_creation_input_tokens", "cache_read_input_tokens", "output_tokens")] +model = next((x["model"] for x in m.values() if x.get("model")), "?") +print("WF-USAGE in=%d cw=%d cr=%d out=%d model=%s" % (*t, model)) +PY diff --git a/templates/prep-prompt.md b/templates/prep-prompt.md new file mode 100644 index 0000000..a503e7a --- /dev/null +++ b/templates/prep-prompt.md @@ -0,0 +1 @@ +You are a prep batch orchestrator: make not-runner-ready tasks runner-ready; never implement, never edit files yourself. wf = python3 {here}/wf.py, project {root}. Tasks: {ids}. First, if out/wf-batch.stop exists → append 'stopped: stop file' to {out}, delete the stop file, exit. Else spawn sonnet workers (Agent tool, general-purpose, model sonnet, foreground, not background), up to 3 tasks each, all in ONE message, and wait for all. Worker prompt (fill in its ids): "Prep wf tasks <ids> in project {root} (cd there; wf = python3 {here}/wf.py). Read only task text: wf show <id> (wf ctx <id> for its refs, ~40-line reads). Never implement, never edit files, never commit. Per task: goal clear → wf set <id> --done "<one checkable line: command or observable result that proves it>". Done = observable check of the task text; never invert recorded original behaviour unless the task is a bug fix. Done names concrete files/runs, not 'each listed X'. Effort > slice_above (default 1h) → no Done on it: wf add --parent <id> slices (each with Done + Model), parent Done = all slices done. Also --model haiku|sonnet|opus if it has no Model line (haiku = mechanical, exact steps; sonnet = exact Done + existing pattern to copy; opus = design forks, debugging, anything unclear; any model above sonnet must be justified in the reply). Unclear (real fork, info only the owner has) → wf add -s awaiting "<id>: <question>", then wf status <id> blocked <a-id>. Reply one line per task: <id> → Done: <text> [Model: <m>] (model above sonnet: [opus: <why>]) | <id> → awaiting <a-id>." Then check each task with wf show <id> and append one line per task to {out} (it exists with its '# wf-batch' header: only append, never rewrite): <id> → Done: <text> | <id> → awaiting <a-id> | <id> → unchanged <why>. Finally commit only the task file: git -C {root} commit -m "prep batch: <ids changed>" -- {tasks}. Never ask the owner: questions → wf add -s awaiting. Then exit. diff --git a/templates/workflow.toml b/templates/workflow.toml new file mode 100644 index 0000000..b8b05cd --- /dev/null +++ b/templates/workflow.toml @@ -0,0 +1,36 @@ +# wf project config. Design: docs/design.md. Unknown keys are errors. +format = 1 +tasks = "TASKS.md" +archive = "tasks/archive.md" +docs = [] # files/folders whose [[id]] links `wf check` validates, e.g. ["DESIGN.md", "docs/"] +verify = [] # commands `wf done` reminds of (never runs), e.g. ["make test"] +done = [] # extra checklist lines `wf done` prints +# ledgers = ".superpowers/sdd" # in-flight plan ledgers shown by `wf next` +# ctx_hint = 100000 # tokens: `wf next`/`done` suggest /clear above it (0 = off) +# worktree_setup = ["mkdir -p out && ln -sfn \"$WF_MAIN/out/data\" out/data"] +# # `wf setup` in a lane worktree: cwd = worktree, env WF_MAIN = main tree; +# # git-ignored inputs; idempotent (workers run it every start) +# quick_gate = ["make quick-test"] # `wf gate`: fast regression check workers run before `wf done` +# # (cwd = this tree's project folder, env WF_MAIN; red = exit 1) + +# [anchors] # optional: every task Ref anchor into `index` must be a heading there +# index = "DESIGN.md" +# index_section = "Subsystems" +# specs = "docs/superpowers/specs" + +# slice_above = "1h" # effort above this → slice job on the slices lane (never implemented whole) +# [lanes.fast] # default lanes = fast + slow below; defining any [lanes.*] replaces both +# efforts = ["<1h"] # efforts this lane implements (disjoint, cover every effort ≤ slice_above) +# order = "unblock" # unblock = tasks others wait on first · priority = inherited priority first +# slices = true # takes slice jobs (one lane; default: the lane with "<1h") +# [lanes.slow] +# efforts = ["1h"] +# order = "priority" +# fallback = "fast" # nothing pickable here → pick from that lane +# areas = "CLAUDE.md" # area notes (## Areas / ### <area>): wf areas checks their code maps +# code_root = "../tool" # repo the areas' anchors / Paths / Checked refer to (default: this project); no uncovered nudge then +# area_stale_commits = 20 # commits touching an area's Paths since Checked → stale +# area_ignore = ["tests", "test", "docs", "doc"] # folders never nudged as uncovered (default shown) +# cloud = true # cloud lane opt-in (wf cloud send) +# cloud_include = ["out/pack"] # git-ignored inputs force-added to the cloud snapshot +# cloud_note = "Game data: out/pack" # appended to the cloud prompt diff --git a/tests/test_areas.py b/tests/test_areas.py new file mode 100644 index 0000000..558230d --- /dev/null +++ b/tests/test_areas.py @@ -0,0 +1,319 @@ +import subprocess +import tempfile +import unittest +from pathlib import Path + +from test_cli import Cli +from test_claims import git +from test_merge import IDENT +from wflib import areas # noqa: E402 (tests/ run with the repo root on sys.path) + +NOTES = """\ +# CLAUDE.md + +## Areas +### Parser +- Code map: `parse_item`, `NOPE_GONE` +- Test recipe: python3 -m unittest tests.test_p +- Paths: src/ +- Checked: {sha} + +### Docs +- Code map: `intro` + +## Gotchas +- x +""" + + +def build(root: Path) -> tuple[str, str]: + """Repo with src/p.py + docs.md, then 3 commits touching src/; returns (sha1, notes).""" + git(root, "init", "-q", "-b", "master") + (root / "src").mkdir(exist_ok=True) + (root / "src" / "p.py").write_text("def parse_item(): pass\n") + (root / "docs.md").write_text("intro\n") + (root / ".gitignore").write_text(".wf/\n") + git(root, "add", "-A") + git(root, "-c", "user.name=t", "-c", "user.email=t@t", "commit", "-qm", "c1") + sha1 = subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=root, capture_output=True, text=True).stdout.strip() + for n in "abc": + (root / "src" / f"{n}.txt").write_text(n) + git(root, "add", "-A") + git(root, "-c", "user.name=t", "-c", "user.email=t@t", "commit", "-qm", n) + return sha1, NOTES.format(sha=sha1) + + +class AreasTest(unittest.TestCase): + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.root = Path(self.tmp.name).resolve() + self.sha1, self.notes = build(self.root) + + def tearDown(self): + self.tmp.cleanup() + + def test_parse(self): + a = areas.parse(NOTES.format(sha="abc1234")) + self.assertEqual([(x.name, x.slug, x.anchors, x.paths, x.checked) for x in a], [ + ("Parser", "parser", ["parse_item", "NOPE_GONE"], ["src/"], "abc1234"), + ("Docs", "docs", ["intro"], [], None)]) + + def test_missing_and_commits(self): + parser, docs = areas.parse(self.notes) + self.assertEqual(areas.missing(self.root, parser), ["NOPE_GONE"]) + self.assertEqual(areas.commits_since(self.root, parser), 3) + self.assertEqual(areas.missing(self.root, docs), []) + self.assertIsNone(areas.commits_since(self.root, docs)) + + def test_stale(self): + parser, docs = areas.parse(self.notes) + self.assertEqual(areas.status(self.root, parser, 20), (True, ["NOPE_GONE"], 3)) + self.assertEqual(areas.status(self.root, docs, 20), (False, [], None)) + fixed = areas.parse(self.notes.replace(", `NOPE_GONE`", ""))[0] + self.assertEqual(areas.status(self.root, fixed, 3), (True, [], 3)) + self.assertEqual(areas.status(self.root, fixed, 4), (False, [], 3)) + + def test_unknown_checked_sha(self): + a = areas.parse(NOTES.format(sha="deadbee"))[0] + self.assertIsNone(areas.commits_since(self.root, a)) + + def test_mark(self): + out = areas.mark(NOTES.format(sha="abc1234"), "Docs", "f00ba44") + self.assertIn("### Docs\n- Code map: `intro`\n- Checked: f00ba44\n\n## Gotchas", out) + out = areas.mark(out, "Parser", "1111111") + self.assertIn("- Paths: src/\n- Checked: 1111111\n", out) + + +class MatchTest(unittest.TestCase): + def test_matches(self): + parser, docs = areas.parse(NOTES.format(sha="abc1234")) + self.assertTrue(areas.matches(parser, "Fix src/p.py crash")) + self.assertTrue(areas.matches(parser, "see x.py/src/p.py")) + self.assertTrue(areas.matches(parser, "parse_item() loops")) + self.assertTrue(areas.matches(parser, "the parser area")) + self.assertFalse(areas.matches(parser, "parse_items and srcs")) + self.assertFalse(areas.matches(docs, "introduction")) + self.assertTrue(areas.matches(docs, "docs intro")) + self.assertFalse(areas.matches(parser, "Fix src/p.py crash", {"src"})) + + def test_shared_paths(self): + a = areas.parse("## Areas\n### A\n- Paths: x.py lib/\n### B\n- Paths: ./x.py y.py\n### C\n- Paths: lib\n") + self.assertEqual(areas.shared_paths(a), {"x.py", "lib"}) + + def test_block(self): + text = NOTES.format(sha="abc1234") + self.assertEqual(areas.block(text, "Docs"), "### Docs\n- Code map: `intro`") + self.assertEqual(areas.block(text, "Parser").splitlines()[-1], "- Checked: abc1234") + self.assertEqual(areas.block(text, "X"), "") + + +class CodeRootTest(Cli): + """code_root: area anchors / Checked / staleness live in another repo than the project.""" + + def setUp(self): + super().setUp() + self.code = self.root.parent / "code" + self.code.mkdir() + self.sha1, notes = build(self.code) + (self.root / "CLAUDE.md").write_text(notes) + with open(self.root / "workflow.toml", "a") as f: + f.write('code_root = "../code"\n') + git(self.root, "init", "-q", "-b", "master") + (self.root / "engine").mkdir() + (self.root / "engine" / "x.py").write_text("x") + + def test_areas_use_code_root(self): + self.assertEqual(self.ok("areas"), f"Parser: stale · missing: NOPE_GONE · 3 commits since {self.sha1}\n" + "Docs: ok · no Checked\n") + + def test_mark_uses_code_head(self): + self.ok("areas", "--mark", "Docs") + head = subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=self.code, capture_output=True, + text=True).stdout.strip() + self.assertIn(f"- Checked: {head}", (self.root / "CLAUDE.md").read_text()) + + def test_check_uses_code_root(self): + out = self.ok("check") + self.assertIn("area Parser: anchor NOPE_GONE not found", out) + self.assertNotIn("Checked", out) + self.assertNotIn("anchor intro", out) + + def test_no_uncovered_nudge(self): + self.assertNotIn("uncovered", self.ok("areas")) + + +class CtxAreaTest(Cli): + tasks_text = ("# T\n\n## Pending\n\n- **t-a** [P1] (<1h): A. Fix src/p.py.\n\n" + "- **t-b** [P1] (<1h): B. Unrelated.\n") + + def setUp(self): + super().setUp() + (self.root / "CLAUDE.md").write_text(NOTES.format(sha="abc1234")) + + def test_ctx_prints_matching_area(self): + out = self.ok("ctx", "t-a") + self.assertIn("\nArea Parser (CLAUDE.md):\n### Parser\n- Code map: `parse_item`, `NOPE_GONE`\n" + "- Test recipe: python3 -m unittest tests.test_p\n", out) + self.assertNotIn("### Docs", out) + + def test_ctx_no_match(self): + self.assertNotIn("Area ", self.ok("ctx", "t-b")) + + def test_ctx_shared_path_no_match(self): + (self.root / "CLAUDE.md").write_text(NOTES.format(sha="abc1234").replace("- Code map: `intro`", + "- Code map: `intro`\n- Paths: src/")) + self.assertNotIn("Area ", self.ok("ctx", "t-a")) + + +class AreasCliTest(Cli): + def setUp(self): + super().setUp() + sha1, notes = build(self.root) + self.sha1 = sha1 + (self.root / "CLAUDE.md").write_text(notes) + + def test_areas_output(self): + out = self.ok("areas") + self.assertEqual(out, f"Parser: stale · missing: NOPE_GONE · 3 commits since {self.sha1}\n" + "Docs: ok · no Checked\n") + + def test_mark_then_show(self): + self.ok("areas", "--mark", "Docs") + head = subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=self.root, capture_output=True, + text=True).stdout.strip() + self.assertEqual(self.ok("areas", "Docs"), f"Docs: ok · 0 commits since {head}\n") + + def test_mark_unique_prefix(self): + self.ok("areas", "--mark", "Do") + head = subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=self.root, capture_output=True, + text=True).stdout.strip() + self.assertEqual(self.ok("areas", "Doc"), f"Docs: ok · 0 commits since {head}\n") + + def test_ambiguous_prefix(self): + txt = (self.root / "CLAUDE.md").read_text().replace("### Docs", "### Parser2 (x)") + (self.root / "CLAUDE.md").write_text(txt) + err = self.fails("areas", "--mark", "Par", code=2) + self.assertIn("ambiguous area 'Par'", err) + self.assertIn("Parser, Parser2 (x)", err) + self.ok("areas", "--mark", "Parser") + + def test_unknown_area(self): + err = self.fails("areas", "X", code=2) + self.assertIn("wf: no area 'X' in CLAUDE.md (Parser, Docs)", err) + + def test_check_warnings(self): + out = self.ok("check") + self.assertIn("area Parser: anchor NOPE_GONE not found", out) + (self.root / "CLAUDE.md").write_text(NOTES.format(sha="deadbee").replace("### Parser", "### X")) + self.assertIn("area X: Checked deadbee unknown", self.ok("check")) + + +class AutoTaskTest(Cli): + tasks_text = "# T\n\n## Pending\n\n- **t-a** [P1] (<1h): A. Do.\n\n- **t-b** [P1] (<1h): B. Do.\n" + + def setUp(self): + super().setUp() + sha1, notes = build(self.root) + (self.root / "CLAUDE.md").write_text(notes) + + def test_auto_task(self): + out = self.ok("done", "t-a", "-m", "x") + self.assertIn("added t-map-parser (area map stale)", out) + shown = self.ok("show", "t-map-parser") + self.assertIn("- **t-map-parser** [P1] (<1h): Refresh Parser area map.", shown) + self.assertIn("Steps: wf areas Parser; fix missing anchors + test recipe from git log --stat " + "<Checked>..HEAD -- <paths>; wf areas --mark Parser.", shown) + self.assertIn("Done: wf areas Parser shows ok.", shown) + self.assertIn("Model: sonnet", shown) + + def test_auto_task_once(self): + self.ok("done", "t-a", "-m", "x") + out = self.ok("done", "t-b", "-m", "y") + self.assertNotIn("added", out) + self.assertEqual(self.ok("list").count("t-map-parser"), 1) + + def test_archived_gets_suffix(self): + self.ok("done", "t-a", "-m", "x") + out = self.ok("done", "t-map-parser", "-m", "z") + self.assertIn("added t-map-parser-2 (area map stale)", out) + + def test_no_areas_file(self): + (self.root / "CLAUDE.md").unlink() + self.assertNotIn("added", self.ok("done", "t-a", "-m", "x")) + + +class UncoveredTest(unittest.TestCase): + def test_uncovered_groups_two_levels(self): + a = areas.parse(NOTES.format(sha="abc1234")) + files = ["src/p.py", "lib/x/y/z.py", "lib/x/w.py", "lib/q.py", "README.md", "tests/t.py", + ".github/ci.yml", "TASKS.md"] + self.assertEqual(areas.uncovered(files, a, ["tests"]), [("lib/x", 2), ("lib", 1)]) + + def test_paths_globs_and_files(self): + a = areas.parse("## Areas\n### A\n- Paths: app/*.py tools/run.sh\n") + self.assertEqual(areas.uncovered(["app/m.py", "tools/run.sh", "tools/other.sh"], a, []), [("tools", 1)]) + + def test_no_paths_anywhere_no_nudge(self): + a = areas.parse("## Areas\n### A\n- Code map: `x`\n") + self.assertEqual(areas.uncovered(["lib/q.py"], a, []), []) + + +class AnchorPathsTest(unittest.TestCase): + def test_anchor_folders_stand_in(self): + with tempfile.TemporaryDirectory() as d: + root = Path(d) + build(root) + a = areas.parse("## Areas\n### A\n- Code map: `parse_item`, `intro`\n### B\n- Paths: lib/\n") + self.assertEqual(areas.with_anchor_paths(root, a[0]).paths, ["docs.md", "src"]) + self.assertEqual(areas.with_anchor_paths(root, a[1]).paths, ["lib/"]) + + +class UncoveredCliTest(Cli): + tasks_text = "# T\n\n## Pending\n\n- **t-a** [P1] (<1h): A. Do.\n\n- **t-b** [P1] (<1h): B. Do.\n" + + def setUp(self): + super().setUp() + build(self.root) + notes = NOTES.format(sha="x").replace(", `NOPE_GONE`", "").replace("- Checked: x\n", "") + (self.root / "CLAUDE.md").write_text(notes) + git(self.root, "add", "-A") + git(self.root, "-c", "user.name=t", "-c", "user.email=t@t", "commit", "-qm", "notes") + (self.root / "engine" / "net").mkdir(parents=True) + (self.root / "engine" / "net" / "sock.py").write_text("x") + (self.root / "docs.md").write_text("changed") + + def test_done_adds_map_task(self): + out = self.ok("done", "t-a", "-m", "x") + self.assertIn("added t-map-engine-net (uncovered area: engine/net, 1 file)", out) + shown = self.ok("show", "t-map-engine-net") + self.assertIn("- **t-map-engine-net** [P1] (<1h): Map engine/net area. " + "No area's Paths covers it.", shown) + self.assertIn("Done: wf areas lists the area ok and no longer names engine/net uncovered.", shown) + self.assertIn("Model: sonnet", shown) + self.assertNotIn("added", self.ok("done", "t-b", "-m", "y")) + + def test_areas_names_uncovered(self): + self.assertIn("uncovered: engine/net (1 file)", self.ok("areas")) + + def test_branch_diff_counts(self): + git(self.root, "switch", "-qc", "topic") + git(self.root, "add", "-A") + git(self.root, "-c", "user.name=t", "-c", "user.email=t@t", "commit", "-qm", "work") + self.assertIn("uncovered: engine/net (1 file)", self.ok("areas")) + + def test_no_paths_uses_anchor_folders(self): + (self.root / "CLAUDE.md").write_text("## Areas\n### P\n- Code map: `parse_item`\n") + (self.root / "src" / "n.py").write_text("x") + out = self.ok("areas") + self.assertIn("uncovered: engine/net (1 file)", out) + self.assertNotIn("uncovered: src", out) + + def test_ignore_config(self): + with open(self.root / "workflow.toml", "a") as f: + f.write('area_ignore = ["engine"]\n') + self.assertNotIn("uncovered", self.ok("areas")) + self.assertNotIn("added", self.ok("done", "t-a", "-m", "x")) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_batch.py b/tests/test_batch.py new file mode 100644 index 0000000..2f0e834 --- /dev/null +++ b/tests/test_batch.py @@ -0,0 +1,534 @@ +import io +import shlex +import subprocess +import sys +import tempfile +import threading +import unittest +from contextlib import redirect_stderr, redirect_stdout +from pathlib import Path + +HERE = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(HERE)) +sys.path.insert(0, str(HERE / "tests")) + +import wf_res # noqa: E402 +from test_res_io import Box, hist_rec, seed_history # noqa: E402 + +OUT = "out/wf-batch-2026-10-01-1400.md" + + +class BatchBase(unittest.TestCase): + def setUp(self): + self._tmp = tempfile.TemporaryDirectory() + self.box = Box(Path(self._tmp.name)) + bin_ = self.box.tmp / "bin" + bin_.mkdir() + self.claude = bin_ / "claude" + self.claude.write_text("#!/bin/sh\n") + self.claude.chmod(0o755) + self.box.caller["PATH"] = f"{bin_}:/usr/bin" + + def tearDown(self): + self._tmp.cleanup() + + def batch(self, *argv): + out, err = io.StringIO(), io.StringIO() + with redirect_stdout(out), redirect_stderr(err): + code = wf_res.batch_main(list(argv), self.box.env) + return code, out.getvalue(), err.getvalue() + + def unit_argv(self): + return self.box.fake.ran("systemd-run")[0] + + +class Start(BatchBase): + def test_starts_claude_print_unit(self): + code, out, err = self.batch("2", "--lanes", "fast,slow") + self.assertEqual((code, err), (0, "")) + log = self.box.env.state / "logs" / "r-1.log" + self.assertEqual(out, f"r-1 started; log {log}; ETA ~20:00 (estimate: not killed when over)\n" + f"summary: {OUT} · status: wf batch --status\n") + argv = self.unit_argv() + i = argv.index(str(self.claude)) + self.assertEqual(argv[i - 1], "sh") + self.assertEqual(argv[i + 1:-1], ["-p", "--model", "opus", "--permission-mode", "auto", + "--permission-prompts", "none"]) + self.assertIn("--setenv=CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=21600000", argv) + self.assertIn(f"MemoryMax={4 * 1024 ** 3}", argv) + e = self.box.ledger().get("r-1") + self.assertEqual((e.title, e.est_min, e.project), ("wf-batch", 360, "proj")) + + def test_prompt_from_template(self): + self.batch("2", "--lanes", "fast,slow") + prompt = self.unit_argv()[-1] + self.assertIn("for up to 2 tasks (lanes: fast, slow)", prompt) + self.assertIn(f"append one line per task to {OUT}", prompt) + self.assertIn("/projects/public/workflow/docs/orchestrator.md", prompt.replace(str(HERE), "/projects/public/workflow")) + self.assertIn("foreground", prompt) + self.assertNotIn("{", prompt) + + def test_dry_run_writes_no_summary(self): + self.batch("2", "--dry-run") + self.assertFalse((self.box.env.root() / OUT).exists()) + + def test_existing_summary_not_overwritten(self): + f = self.box.env.root() / OUT + f.parent.mkdir(exist_ok=True) + f.write_text("live\n") + self.batch("2") + self.assertTrue(f.read_text().startswith("live\n# wf-batch")) + + def test_size_lane(self): + code, out, err = self.batch("2", "--lanes", "fast", "--dry-run") + self.assertEqual((code, err), (0, "")) + self.assertIn("lanes: fast", out) + + def test_default_lanes(self): + self.batch("3") + self.assertIn("for up to 3 tasks (lanes: every lane with ready tasks, see wf lanes)", self.unit_argv()[-1]) + + def test_for_sets_ceiling_and_model(self): + self.batch("1", "--for", "2h", "--model", "sonnet", "--mem", "2G") + argv = self.unit_argv() + self.assertIn("--setenv=CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=7200000", argv) + self.assertEqual(argv[argv.index("--model") + 1], "sonnet") + self.assertIn(f"MemoryMax={2 * 1024 ** 3}", argv) + + def test_no_claude(self): + self.box.caller["PATH"] = "/nonexistent" + code, out, err = self.batch("2") + self.assertEqual((code, out, err), (1, "", "wf: claude not found on PATH\n")) + self.assertEqual(self.box.fake.ran("systemd-run"), []) + + def test_bad_count(self): + code, _, err = self.batch("0") + self.assertEqual(code, 2) + self.assertIn("N must be ≥ 1", err) + + def test_summary_header_written(self): + code, out, err = self.batch("2", "--lanes", "fast") + self.assertEqual(code, 0) + text = (self.box.env.root() / OUT).read_text() + self.assertEqual(text, "# wf-batch proj 2026-10-01 14:00 (N=2, lanes: fast)\n") + + def test_dry_run(self): + code, out, err = self.batch("2", "--dry-run") + self.assertEqual((code, err), (0, "")) + self.assertIn("CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=21600000 wf res run --mem 4G --for 6h --title wf-batch -- ", out) + self.assertIn(f"{self.claude} -p --model opus", out) + self.assertEqual(self.box.fake.ran("systemd-run"), []) + + def test_mem_scales_with_n(self): + for args, mem in ((("4",), "6G"), (("1",), "4G"), (("9",), "8G"), (("4", "--mem", "3G"), "3G")): + code, out, err = self.batch(*args, "--dry-run") + self.assertIn(f"wf res run --mem {mem} --for", out, args) + + def test_defaults_from_history(self): + # peaks 2,3,4 → p95 4 ×1.15 = 4.6 GB; durations 20,30,40 → p90 40 ×1.5 = 60 min; ceiling stays 6h + seed_history(self.box, [hist_rec(f"r-{i}", title="wf-batch", peak=p, minutes=m) + for i, (p, m) in enumerate(((2.0, 20.0), (3.0, 30.0), (4.0, 40.0)))]) + code, out, err = self.batch("4", "--dry-run") + self.assertEqual((code, err), (0, "")) + self.assertIn("CEILING_MS=21600000 wf res run --mem 4.6G --for 60m --title wf-batch -- ", out) + out = self.batch("4", "--mem", "3G", "--for", "2h", "--dry-run")[1] + self.assertIn("CEILING_MS=7200000 wf res run --mem 3G --for 2h --title wf-batch -- ", out) + + def test_busy_exit_3(self): + self.box.available_gb = 1 + code, out, _ = self.batch("2") + self.assertEqual(code, 3) + self.assertIn("busy", out) + + def test_force_past_unused_claim(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + code, out, _ = self.batch("2") + self.assertEqual(code, 3) + code, out, err = self.batch("2", "--force") + self.assertEqual(code, 0, (out, err)) + self.assertIn("r-2 started", out) + + +class Status(BatchBase): + def test_none(self): + code, out, err = self.batch("--status") + self.assertEqual((code, out, err), (0, "no batch in this project\n", "")) + + def test_running_and_summary(self): + self.batch("2") + md = self.box.tmp / "proj" / OUT + with md.open("a") as f: + f.write("t-a sonnet done abc123\n") + (md.parent / "wf-batch-2026-09-30-0100.md").write_text("old\n") + code, out, _ = self.batch("--status") + self.assertEqual(code, 0) + lines = out.splitlines() + self.assertTrue(lines[0].startswith('r-1 proj "wf-batch" running'), lines[0]) + self.assertIn(f"log {self.box.env.state / 'logs' / 'r-1.log'}", lines[1]) + self.assertEqual(lines[2:], [f"{OUT}:", "# wf-batch proj 2026-10-01 14:00 (N=2, lanes: every lane with ready tasks, see wf lanes)", + "t-a sonnet done abc123"]) + + def test_orch_log_progress(self): + self.batch("2") + out_dir = self.box.tmp / "proj" / "out" + (out_dir / "wf-orch.log").write_text( + "2026-10-01 13:00 fast sonnet t-old done aaa\n2026-10-01 14:05 fast sonnet t-new done bbb\n") + _, out, _ = self.batch("--status") + self.assertIn("finished since batch start", out) + self.assertIn("t-new done bbb", out) + self.assertNotIn("t-old", out) + + def test_stop_creates_file_and_status_reports(self): + code, out, err = self.batch("--stop") + self.assertEqual((code, err), (0, "")) + f = self.box.tmp / "proj" / "out" / "wf-batch.stop" + self.assertTrue(f.exists()) + self.assertIn("out/wf-batch.stop", out) + code, out, _ = self.batch("--status") + self.assertIn("stop requested: out/wf-batch.stop", out) + + def test_stale_stop_cleared_on_start(self): + f = self.box.tmp / "proj" / "out" / "wf-batch.stop" + f.parent.mkdir(exist_ok=True) + f.touch() + code, out, err = self.batch("2") + self.assertEqual(code, 0) + self.assertIn("stale out/wf-batch.stop cleared", err) + + def test_stop_consumed_on_exit(self): + f = self.box.tmp / "proj" / "out" / "wf-batch.stop" + f.parent.mkdir(exist_ok=True) + orig = wf_res.cmd_run + + def fake(env, cfg, args): + f.touch() + return 0 + wf_res.cmd_run = fake + try: + code, _, _ = self.batch("2") + finally: + wf_res.cmd_run = orig + self.assertEqual(code, 0) + self.assertFalse(f.exists()) + + def test_prompt_has_stop_check(self): + code, out, _ = self.batch("2", "--dry-run") + self.assertIn("out/wf-batch.stop", out) + self.assertIn("stopped: stop file", out) + + def test_prompt_cloud_lane(self): + code, out, _ = self.batch("2", "--dry-run") + self.assertIn("wf orch pick cloud", out) + self.assertIn("never run `wf cloud pull`", out) + self.assertIn("sidecar pulls every 10 min", out) + self.assertIn(f"-- {self.claude} -p", out) # not a cloud project: plain claude -p job + + def test_prompt_stale_branch_once(self): + code, out, _ = self.batch("2", "--dry-run") + self.assertIn("pre-existing branch", out) + self.assertIn("ONCE", out) + + def test_prompt_checks_stop_before_every_spawn(self): + code, out, _ = self.batch("2", "--dry-run") + self.assertIn("Before EVERY spawn", out) + self.assertIn(f"test -e {self.box.tmp / 'proj' / 'out' / 'wf-batch.stop'}", out) + self.assertNotIn("Each round: first", out) + code, out, _ = self.batch("--stop") + self.assertIn("before its next spawn", out) + + def test_n_required_without_status(self): + code, _, err = self.batch() + self.assertEqual(code, 2) + + +PREP_TASKS = """\ +## Awaiting your decision + +## Pending + +- **t-ready** [P1] (<1h): Ready. + Done: works. + +- **t-low** [P3] (<1h): Low, no Done. + +- **t-hi** [P1] (1h): High, no Done. + +- **t-mid** [P2] (1h): Mid, no Done. + +## Needs human + +## Deferred +""" + + +class Prep(BatchBase): + def setUp(self): + super().setUp() + root = self.box.env.root() + (root / "workflow.toml").write_text('format = 1\ntasks = "TASKS.md"\narchive = "archive.md"\n') + (root / "TASKS.md").write_text(PREP_TASKS) + (root / "archive.md").write_text("# Archive\n") + + def test_prompt_lists_targets(self): + code, out, err = self.batch("2", "--prep") + self.assertEqual((code, err), (0, "")) + prompt = self.unit_argv()[-1] + self.assertIn("Tasks: t-hi t-mid.", prompt) + self.assertIn("never implement", prompt) + self.assertIn("wf set <id> --done", prompt) + self.assertIn("wf add -s awaiting", prompt) + self.assertIn(f"to {OUT}", prompt) + for ph in ("never invert recorded original behaviour", "wf add --parent <id> slices", + "concrete files/runs", "[opus: <why>]", "[Model: <m>]"): + self.assertIn(ph, prompt) + self.assertNotIn("{", prompt) + self.assertEqual((self.box.env.root() / OUT).read_text(), + "# wf-batch proj 2026-10-01 14:00 (prep K=2: t-hi t-mid)\n") + self.assertEqual(self.box.ledger().get("r-1").title, "wf-batch") + + def test_lanes_filter(self): + self.batch("5", "--prep", "--lanes", "fast") + self.assertIn("Tasks: t-low.", self.unit_argv()[-1]) + + def test_prep_alias_all(self): + out, err = io.StringIO(), io.StringIO() + with redirect_stdout(out), redirect_stderr(err): + code = wf_res.prep_main(["all"], self.box.env) + self.assertEqual((code, err.getvalue()), (0, "")) + self.assertIn("Tasks: t-hi t-mid", self.unit_argv()[-1]) + + def test_nothing_to_prep_starts_nothing(self): + code, out, err = self.batch("3", "--prep", "--lanes", "nolane") + self.assertEqual((code, out, err), (0, "prep: no pending task without Done (lanes: nolane); nothing started\n", "")) + self.assertEqual(self.box.fake.ran("systemd-run"), []) + self.assertFalse((self.box.env.root() / OUT).exists()) + + def test_prep_mem_default_small(self): + for args, mem in ((("3", "--prep"), "1G"), (("3", "--prep", "--mem", "2G"), "2G")): + code, out, err = self.batch(*args, "--dry-run") + self.assertIn(f"wf res run --mem {mem} --for", out, args) + + def test_dry_run(self): + code, out, err = self.batch("1", "--prep", "--dry-run") + self.assertEqual((code, err), (0, "")) + self.assertIn("--title wf-batch -- ", out) + self.assertIn("t-hi", out) + self.assertEqual(self.box.fake.ran("systemd-run"), []) + + def test_no_project(self): + (self.box.env.root() / "workflow.toml").unlink() + code, out, err = self.batch("1", "--prep") + self.assertEqual(code, 1) + self.assertTrue(err.startswith("wf: "), err) + + +class Dispatch(unittest.TestCase): + def test_wf_forwards_batch(self): + r = subprocess.run([sys.executable, str(HERE / "wf.py"), "batch", "-h"], capture_output=True, text=True) + self.assertEqual(r.returncode, 0) + self.assertIn("usage: wf batch", r.stdout) + + def test_listed_in_wf_help(self): + r = subprocess.run([sys.executable, str(HERE / "wf.py"), "-h"], capture_output=True, text=True) + self.assertIn("batch", r.stdout) + + +if __name__ == "__main__": + unittest.main() + + +FAKE_WF = """import sys, pathlib +d = pathlib.Path(sys.argv[0]).parent +with open(d / "calls.log", "a") as f: + f.write(" ".join(sys.argv[1:]) + "\\n") +if sys.argv[1:3] == ["cloud", "pull"] and (d / "clear").exists(): + for rec in (d / ".wf" / "cloud").glob("*.json"): + rec.unlink() +if sys.argv[1:3] == ["cloud", "pull"] and not (d / "pulled").exists(): + (d / "pulled").touch() + print("t-a: done (session_1, $0.40 usage): merged abc1234") + print("archive session_1 failed: x; archive by hand (wf cloud archive --ended)") + print("report: commit abc1234") + print("t-b: running (session_2, 12 min)") + print("t-c: handback (session_3, $0.10 usage): cloud handback: stuck") + print("t-d: pulled by another wf cloud pull, skipped") + sys.exit(1) +""" + + +class Sidecar(unittest.TestCase): + """wf batch on a cloud = true project: the job runs the orchestrator under the pull sidecar.""" + + def setUp(self): + self._tmp = tempfile.TemporaryDirectory() + self.root = Path(self._tmp.name) + (self.root / "out").mkdir() + self.summary = self.root / OUT + self.summary.write_text("# wf-batch\n") + self.fake = self.root / "fakewf.py" + self.fake.write_text(FAKE_WF) + + def tearDown(self): + self._tmp.cleanup() + + def run_sidecar(self, child_s, every=0.2, rc=7, **kw): + child = [sys.executable, "-c", f"import time, sys; time.sleep({child_s}); sys.exit({rc})"] + out = io.StringIO() + with redirect_stdout(out): + code = wf_res.sidecar(child, self.root, self.summary, every, [sys.executable, str(self.fake)], **kw) + return code, out.getvalue() + + def records(self, *ids): + d = self.root / ".wf" / "cloud" + d.mkdir(parents=True) + for i in ids: + (d / f"{i}.json").write_text("{}") + + def tail_line(self): + return [ln[6:] for ln in self.summary.read_text().splitlines()[1:] if "tail" in ln] + + def test_tail_ends_on_empty_records(self): + self.records("t-b") + (self.root / "clear").touch() + code, out = self.run_sidecar(0.05, every=0.1) + self.assertEqual(code, 7) + self.assertEqual(len([c for c in self.calls() if c.startswith("cloud pull")]), 1) + self.assertIn("orch post t-a cloud --result done --no-pick --commit abc1234", self.calls()) + self.assertEqual(self.tail_line(), ["cloud tail ended: no cloud records left; 1 pulls; left: none (sidecar)"]) + + def test_tail_ends_on_stop_file(self): + self.records("t-b", "t-e") + threading.Timer(0.5, (self.root / wf_res.STOP_FILE).touch).start() + code, _ = self.run_sidecar(0.05, every=0.2) + pulls = [c for c in self.calls() if c.startswith("cloud pull")] + self.assertGreaterEqual(len(pulls), 1) # pulled in the tail until the stop file + self.assertEqual(self.tail_line(), [f"cloud tail ended: stop file; {len(pulls)} pulls; left: t-b t-e (sidecar)"]) + self.assertEqual(code, 7) + + def test_tail_ends_on_cap(self): + self.records("t-b") + shrunk = [] + code, out = self.run_sidecar(0.05, every=0.1, cap=0.35, shrink=lambda: shrunk.append(1) or "shrunk") + self.assertEqual((code, shrunk), (7, [1])) # reservation shrunk once, entering the tail + self.assertIn("shrunk", out) + pulls = [c for c in self.calls() if c.startswith("cloud pull")] + self.assertGreaterEqual(len(pulls), 2) + self.assertTrue(all(c.startswith("cloud pull") or c.startswith("orch post") for c in self.calls())) # no picks + self.assertEqual(self.tail_line(), + [f"cloud tail ended: {0.35 / 3600:g}h cap; {len(pulls)} pulls; left: t-b (sidecar)"]) + + def test_no_tail_without_records(self): + shrunk = [] + code, _ = self.run_sidecar(0.05, every=0.1, shrink=lambda: shrunk.append(1) or "") + self.assertEqual((code, shrunk, self.calls(), self.tail_line()), (7, [], [], [])) + + def test_shrink_reservation(self): + b = BatchBase("run") + b.setUp() + try: + b.batch("4", "--mem", "6G") + msg = wf_res.shrink_reservation(b.box.env, "r-1") + e = b.box.ledger().get("r-1") + self.assertEqual((e.mem_gb, e.cpus, e.state), (0.2, 1, "running")) + self.assertIn("r-1 reservation 6", msg) + self.assertIn("not running", wf_res.shrink_reservation(b.box.env, "r-9")) + finally: + b.tearDown() + + def calls(self): + f = self.root / "calls.log" + return f.read_text().splitlines() if f.exists() else [] + + def test_pulls_and_posts_while_child_lives(self): + code, out = self.run_sidecar(1.0) + self.assertEqual(code, 7) # the orchestrator's exit code + calls = self.calls() + pulls = [c for c in calls if c.startswith("cloud pull")] + self.assertGreaterEqual(len(pulls), 2) # repeats every interval while the child lives + self.assertEqual(pulls[0], f"cloud pull --all --project {self.root}") + self.assertEqual([c for c in calls if c.startswith("orch post")], + ["orch post t-a cloud --result done --no-pick --commit abc1234", + "orch post t-c cloud --result handback --no-pick"]) + lines = self.summary.read_text().splitlines()[1:] + self.assertEqual([ln[6:] for ln in lines], ["cloud t-a done abc1234 (sidecar pull)", + "cloud t-c handback - (sidecar pull)"]) + self.assertIn("t-a: done", out) # pull output lands in the job log + + def test_stops_with_child(self): + code, _ = self.run_sidecar(0.1, every=5, rc=0) + self.assertEqual((code, self.calls()), (0, [])) + + def test_stop_file_ends_pulls(self): + (self.root / wf_res.STOP_FILE).touch() + code, out = self.run_sidecar(0.7) + self.assertEqual((code, self.calls()), (7, [])) + self.assertIn("sidecar: stop file, no more pulls", out) + + def test_cloud_project_job_wraps_claude(self): + b = BatchBase("run") + b.setUp() + try: + root = b.box.env.root() + (root / "workflow.toml").write_text('format = 1\ntasks = "TASKS.md"\narchive = "a.md"\ncloud = true\n') + code, out, _ = b.batch("2", "--dry-run") + self.assertEqual(code, 0) + argv = shlex.split(out.split(" -- ", 1)[1]) + self.assertEqual(argv[:9], [sys.executable, str(HERE / "wf_res.py"), "batch-sidecar", "--root", str(root), + "--summary", OUT, "--every", "600"]) + self.assertEqual(argv[9:12], ["--", str(b.claude), "-p"]) + finally: + b.tearDown() + + def test_main_entry(self): + r = subprocess.run([sys.executable, str(HERE / "wf_res.py"), "batch-sidecar", "--root", str(self.root), + "--summary", OUT, "--every", "5", "--wf", f"{sys.executable} {self.fake}", "--", + sys.executable, "-c", "import sys; sys.exit(3)"], capture_output=True, text=True) + self.assertEqual((r.returncode, r.stderr), (3, "")) + + +LOG_ROWS = "".join(f"2026-10-01T1{i}:00:00 slow opus t-{i} done abc {m}m00s\n" for i, m in enumerate((20, 25, 31))) + + +class Fit(BatchBase): + def log(self, text): + (self.box.env.root() / "out").mkdir(exist_ok=True) + (self.box.env.root() / "out" / "wf-orch.log").write_text(text) + + def test_left_shrinks_k_and_sets_for(self): + self.log(LOG_ROWS) # p90 of 20,25,31 = 31 min + code, out, err = self.batch("4", "--left", "1h40m", "--dry-run") + self.assertEqual((code, err), (0, "")) + self.assertIn("fit: 3 of 4 tasks (left 1h40m, task p90 31m from 3 runs)\n", out) + self.assertIn("CEILING_MS=6000000 wf res run --mem 4.5G --for 1h40m --title wf-batch -- ", out) + self.assertIn("for up to 3 tasks", out) + + def test_left_default_30m(self): + code, out, err = self.batch("4", "--left", "45m", "--dry-run") + self.assertIn("fit: 1 of 4 tasks (left 45m, task p90 30m default (0 runs))\n", out) + self.assertIn("--for 45m", out) + + def test_left_too_short_starts_nothing(self): + self.log(LOG_ROWS) + code, out, err = self.batch("4", "--left", "30m") + self.assertEqual((code, err), (0, "")) + self.assertEqual(out, "fit: 0 of 4 tasks (left 30m, task p90 31m from 3 runs); nothing started\n") + self.assertEqual(self.box.fake.ran("systemd-run"), []) + + def test_prompt_has_deadline_check(self): + code, out, _ = self.batch("2", "--for", "2h", "--dry-run") + self.assertIn("wf batch --time-left 2026-10-01T16:00+00:00", out) + self.assertIn("stopped: deadline", out) + + def test_time_left(self): + self.log(LOG_ROWS) + code, out, err = self.batch("--time-left", "2026-10-01T15:00+00:00") + self.assertEqual((code, out, err), (0, "time left 1h, task p90 31m (3 runs): spawn\n", "")) + code, out, err = self.batch("--time-left", "2026-10-01T14:30+00:00") + self.assertEqual((code, out), (0, "time left 30m, task p90 31m (3 runs): stop\n")) + code, out, err = self.batch("--time-left", "2026-10-01T13:00+00:00") + self.assertEqual((code, out), (0, "time left 0m, task p90 31m (3 runs): stop\n")) + + def test_time_left_bad(self): + code, out, err = self.batch("--time-left", "soon") + self.assertEqual(code, 1) + self.assertIn("wf: --time-left", err) diff --git a/tests/test_check.py b/tests/test_check.py new file mode 100644 index 0000000..e4f3d3f --- /dev/null +++ b/tests/test_check.py @@ -0,0 +1,304 @@ +import os +import sys +import time +import tempfile +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import check as K +from wflib import config as C + +TOML = 'format = 1\ntasks = "TASKS.md"\narchive = "tasks/archive.md"\ndocs = ["DESIGN.md", "docs/"]\n' +ANCHORS = '[anchors]\nindex = "DESIGN.md"\nindex_section = "Subsystems"\nspecs = "docs/specs"\n' +ARCHIVE = "# Archive (newest first)\n\n- 2026-09-01 **t-done** Done thing — ok\n" +DESIGN = """\ +# Design + +## Subsystems + +### terrain + +Heightmap. [spec](docs/specs/terrain.md#terrain) + +### input + +Keys. + +## Notes + +Free text. +""" +SPEC = '# Terrain spec\n\n<a id="terrain"></a>\n## Terrain\n\nText.\n' +CLEAN = """\ +# Tasks + +## Awaiting your decision + +- **a-key**: Key needed. Which one? + +## Pending + +- **t-one** [P1] (1h): One. + - After: [[t-done]] + Ref: DESIGN.md#terrain, docs/specs/terrain.md + +- **t-two** [P2] (5h) (blocked: [[a-key]]): Two. + - After: [[t-one]] + +## Needs human + +## Deferred +""" + + +class Base(unittest.TestCase): + toml = TOML + + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.root = Path(self.tmp.name).resolve() + (self.root / "tasks").mkdir() + (self.root / "docs" / "specs").mkdir(parents=True) + (self.root / "workflow.toml").write_text(self.toml) + (self.root / "tasks" / "archive.md").write_text(ARCHIVE) + (self.root / "DESIGN.md").write_text(DESIGN) + (self.root / "docs" / "specs" / "terrain.md").write_text(SPEC) + self.tasks(CLEAN) + + def tearDown(self): + self.tmp.cleanup() + + def tasks(self, text): + (self.root / "TASKS.md").write_text(text) + + def pending(self, items): + self.tasks("## Awaiting your decision\n\n- **a-key**: Key. Q?\n\n## Pending\n\n" + items + + "\n## Needs human\n\n## Deferred\n") + + def run_check(self): + errors, warnings = K.check(C.load(self.root)) + return [str(e) for e in errors], [str(w) for w in warnings] + + def errors(self): + return self.run_check()[0] + + +class TasksCheckTest(Base): + def test_clean(self): + self.assertEqual(self.run_check(), ([], [])) + + def test_after_prose_and_deferred(self): + self.tasks("## Awaiting your decision\n\n## Pending\n\n" + "- **t-a** [P1] (1h): A.\n After: when ready [[t-d]]\n\n" + "## Needs human\n\n## Deferred\n\n- **t-d** [P3] (1h): D.\n") + w = self.run_check()[1] + self.assertEqual(len(w), 2, w) + self.assertTrue(any("has prose" in x and ":6:" in x for x in w), w) + self.assertTrue(any("'t-d' is deferred" in x for x in w), w) + + def test_after_clean_links(self): + self.pending("- **t-a** [P1] (1h): A.\n After: [[t-done]], [[t-b]]\n\n- **t-b** [P1] (1h): B.\n") + self.assertEqual(self.run_check()[1], []) + + def test_duplicate_id(self): + self.pending("- **t-a** [P1] (1h): A.\n\n- **t-a** [P1] (1h): Again.\n") + self.assertEqual(self.errors(), ["TASKS.md:9: t-a: duplicate id (also line 7)"]) + + def test_id_reused_from_archive(self): + self.pending("- **t-done** [P1] (1h): A.\n") + self.assertEqual(self.errors(), ["TASKS.md:7: t-done: id already used in the archive (ids are never reused)"]) + + def test_bad_grammar(self): + self.pending("- **T-Foo** [P1] (1h): A.\n") + self.assertEqual(self.errors(), ["TASKS.md:7: T-Foo: bad id (want t-… or a-…, lowercase a-z 0-9 -)"]) + + def test_wrong_kind_for_section(self): + self.pending("- **a-x**: A. Q?\n") + self.assertEqual(self.errors(), ["TASKS.md:7: a-x: a- item in Pending (a- items belong in Awaiting, t- elsewhere)"]) + + def test_task_in_awaiting(self): + self.tasks("## Awaiting your decision\n\n- **t-x** [P1] (1h): A.\n\n## Pending\n\n## Needs human\n\n## Deferred\n") + self.assertEqual(self.errors(), + ["TASKS.md:3: t-x: t- item in Awaiting your decision (a- items belong in Awaiting, t- elsewhere)"]) + + def test_missing_priority_and_effort(self): + self.pending("- **t-a**: A.\n") + self.assertEqual(self.errors(), ["TASKS.md:7: t-a: no priority [P0]-[P3]", + "TASKS.md:7: t-a: no effort (want <1h, 1h, 5h, 10h, 100h)"]) + + def test_bad_priority_and_effort(self): + self.pending("- **t-a** [P7] (2h): A.\n") + self.assertEqual(self.errors(), ["TASKS.md:7: t-a: priority P7 (want P0-P3)", + "TASKS.md:7: t-a: effort '2h' (want <1h, 1h, 5h, 10h, 100h)"]) + + def test_dangling_link(self): + self.pending("- **t-a** [P1] (1h): A.\n - see [[t-none]]\n") + self.assertEqual(self.errors(), ["TASKS.md:8: [[t-none]]: no such id in TASKS or archive"]) + + def test_dangling_link_in_prose_section(self): + self.tasks(CLEAN + "\n## Notes\n\nSee [[a-gone]] and `[[t-example]]`.\n") + self.assertEqual(self.errors(), ["TASKS.md:22: [[a-gone]]: no such id in TASKS or archive"]) + + def test_blocked_on_missing_awaiting_item(self): + self.pending("- **t-a** [P1] (1h) (blocked: [[a-none]]): A.\n") + self.assertEqual(self.errors(), ["TASKS.md:7: [[a-none]]: no such id in TASKS or archive", + "TASKS.md:7: t-a: blocked on 'a-none', which is not an open Awaiting item"]) + + def test_after_cycle(self): + self.pending("- **t-a** [P1] (1h): A.\n - After: [[t-b]]\n\n- **t-b** [P1] (1h): B.\n - After: [[t-a]]\n") + self.assertIn("TASKS.md:7: t-a: After: cycle t-a → t-b → t-a", self.errors()) + + def test_item_before_its_open_dependency(self): + self.pending("- **t-a** [P1] (1h): A.\n - After: [[t-b]]\n\n- **t-b** [P1] (1h): B.\n") + self.assertEqual(self.errors(), ["TASKS.md:7: t-a: placed before 't-b', which it is After:"]) + + def test_after_bare_id_is_flagged(self): + self.pending("- **t-a** [P1] (1h): A.\n After: t-b\n\n- **t-b** [P1] (1h): B.\n") + self.assertEqual(self.errors(), ["TASKS.md:8: t-a: After: 't-b' is not a link (want [[t-b]])"]) + + def test_slices_bare_id_is_flagged(self): + self.pending("- **t-a** [P1] (1h): A.\n - Slices: [[t-b]], t-c\n\n- **t-b** [P1] (1h): B.\n") + self.assertEqual(self.errors(), ["TASKS.md:8: t-a: Slices: 't-c' is not a link (want [[t-c]])"]) + + def test_ref_path_missing(self): + self.pending("- **t-a** [P1] (1h): A.\n Ref: docs/none.md\n") + self.assertEqual(self.errors(), ["TASKS.md:8: t-a: Ref 'docs/none.md' does not exist"]) + + def test_ref_anchor_missing(self): + self.pending("- **t-a** [P1] (1h): A.\n Ref: DESIGN.md#nope\n") + self.assertEqual(self.errors(), ["TASKS.md:8: t-a: Ref 'DESIGN.md#nope': no such anchor"]) + + def test_old_numbered_items(self): + self.pending("1. **[P1] Old** (Effort: 1h) — goal.\n - Steps\n") + self.assertEqual(self.errors(), ["TASKS.md:7: old numbered item (wf migrate)"]) + + def test_malformed_header_has_line(self): + self.pending("- **t-ok** [P1] (1h): Fine.\n\n- **t-x** [P1] (1h) no colon\n") + self.assertEqual(self.errors(), + ["TASKS.md:9: t-x: bad header: want '- **id** [Pn] (effort) [(status)]: Title. Goal.'"]) + + def test_item_after_prose_is_flagged(self): + self.pending("- **t-a** [P1] (1h): A.\nstray\n- **t-b** [P1] (1h): B.\n") + self.assertEqual(self.errors(), ["TASKS.md:9: t-b: item after a flush-left prose line (line 8): indent or move the prose"]) + + def test_missing_section(self): + self.tasks("## Pending\n\n## Needs human\n\n## Deferred\n") + self.assertEqual(self.errors(), ["TASKS.md: no '## Awaiting your decision' section"]) + + def test_number_refs_warn(self): + self.pending("- **t-a** [P1] (1h): A, see #12.\n") + self.assertEqual(self.run_check(), ([], ["TASKS.md:7: '#12': number ref (tasks have ids: [[t-…]])"])) + + def test_empty_progress_note_warns(self): + self.pending("- **t-a** [P1] (1h) (in progress: ): A.\n") + self.assertEqual(self.run_check(), ([], ["TASKS.md:7: t-a: in progress without a branch or note"])) + + def test_doc_link_to_unknown_id(self): + (self.root / "docs" / "note.md").write_text("# N\n\nSee [[t-one]], [[t-done]], [[t-lost]], [[wiki-page]].\n") + self.assertEqual(self.errors(), ["docs/note.md:3: [[t-lost]]: no such id in TASKS or archive"]) + + +class StaleAwaitingTest(Base): + def commit(self, date): + import os + import subprocess + env = {**os.environ, "GIT_AUTHOR_DATE": date, "GIT_COMMITTER_DATE": date, + "GIT_AUTHOR_NAME": "t", "GIT_AUTHOR_EMAIL": "t@t", "GIT_COMMITTER_NAME": "t", "GIT_COMMITTER_EMAIL": "t@t"} + for cmd in (["init", "-q"], ["add", "-A"], ["commit", "-q", "-m", "x"]): + subprocess.run(["git", "-C", str(self.root), *cmd], check=True, env=env, capture_output=True) + + def test_old_unreferenced_awaiting_item_warns(self): + self.tasks(CLEAN.replace("- **a-key**: Key needed. Which one?", + "- **a-key**: Key needed. Which one?\n\n- **a-old**: Old question. Still open?")) + self.commit("2020-01-01T00:00:00") + errors, warnings = self.run_check() + self.assertEqual(errors, []) + self.assertEqual(len(warnings), 1) + self.assertRegex(warnings[0], r"^TASKS.md:7: a-old: waiting \d+ days, no task references it$") + + def test_fresh_item_does_not_warn(self): + self.tasks(CLEAN.replace("- **a-key**: Key needed. Which one?", + "- **a-key**: Key needed. Which one?\n\n- **a-new**: New question. Open?")) + import datetime + self.commit(datetime.datetime.now().isoformat(timespec="seconds")) + self.assertEqual(self.run_check(), ([], [])) + + +class ConfigCheckTest(Base): + def test_missing_tasks_file(self): + (self.root / "TASKS.md").unlink() + self.assertEqual(self.errors(), ["workflow.toml: tasks 'TASKS.md' does not exist"]) + + def test_missing_archive_and_docs(self): + (self.root / "tasks" / "archive.md").unlink() + (self.root / "DESIGN.md").unlink() + self.pending("- **t-a** [P1] (1h): A.\n") + self.assertEqual(self.errors(), ["workflow.toml: archive 'tasks/archive.md' does not exist", + "workflow.toml: docs 'DESIGN.md' does not exist"]) + + +class AnchorsCheckTest(Base): + toml = TOML + ANCHORS + + def test_clean(self): + self.assertEqual(self.run_check(), ([], [])) + + def test_task_anchor_not_in_index_section(self): + self.pending("- **t-a** [P1] (1h): A.\n Ref: DESIGN.md#notes\n") + self.assertEqual(self.errors(), ["TASKS.md:8: t-a: Ref 'DESIGN.md#notes' is not a heading under 'Subsystems'"]) + + def test_duplicate_index_slug(self): + (self.root / "DESIGN.md").write_text(DESIGN.replace("### input", "### Terrain")) + self.assertEqual(self.errors(), ["DESIGN.md:9: two 'Subsystems' headings give anchor 'terrain' (also line 5)"]) + + def test_index_link_to_missing_file_and_anchor(self): + (self.root / "DESIGN.md").write_text(DESIGN.replace("Keys.", "Keys. [a](docs/none.md) [b](docs/specs/terrain.md#nope) [c](https://x.y/z)")) + self.assertEqual(self.errors(), ["DESIGN.md:11: link 'docs/none.md': file does not exist", + "DESIGN.md:11: link 'docs/specs/terrain.md#nope': no such anchor"]) + + def test_explicit_spec_anchor_without_index_entry(self): + (self.root / "docs" / "specs" / "combat.md").write_text('# Combat\n\n<a id="combat"></a>\n## Combat\n') + self.assertEqual(self.errors(), ["docs/specs/combat.md:3: explicit anchor 'combat' has no heading under 'Subsystems' in DESIGN.md"]) + + +class ProblemTest(unittest.TestCase): + def test_key_ignores_line(self): + a = K.Problem("TASKS.md", 7, "t-a", "x") + b = K.Problem("TASKS.md", 9, "t-a", "x") + self.assertEqual(a.key, b.key) + self.assertEqual(str(a), "TASKS.md:7: t-a: x") + self.assertEqual(str(K.Problem("workflow.toml", None, "", "y")), "workflow.toml: y") + + +if __name__ == "__main__": + unittest.main() + + +class MergedToolWorktreeTest(unittest.TestCase): + def git(self, root, *a): + import subprocess + subprocess.run(["git", "-C", str(root), "-c", "user.name=t", "-c", "user.email=t@t", *a], + check=True, capture_output=True) + + def test_flags_only_merged_worktrees(self): + with tempfile.TemporaryDirectory() as d: + root = Path(d) / "tool" + root.mkdir() + self.git(root, "init", "-b", "master") + self.git(root, "commit", "--allow-empty", "-m", "a") + self.git(root, "worktree", "add", str(Path(d) / "fresh"), "-b", "fresh") + self.git(root, "worktree", "add", str(Path(d) / "done"), "-b", "done") + self.git(root, "worktree", "add", str(Path(d) / "open"), "-b", "open") + self.git(Path(d) / "open", "commit", "--allow-empty", "-m", "b") + self.git(root, "commit", "--allow-empty", "-m", "c") + old = time.time() - 3 * 3600 + for n in ("done", "open"): + os.utime(Path(d) / n / ".git", (old, old)) + got = K.merged_tool_worktrees(root) + names = sorted(p.message.split("branch ")[1].split(")")[0] for p in got) + self.assertEqual(names, ["done"]) # fresh (<2h) skipped + + def test_not_git(self): + with tempfile.TemporaryDirectory() as d: + self.assertEqual(K.merged_tool_worktrees(Path(d)), []) diff --git a/tests/test_claims.py b/tests/test_claims.py new file mode 100644 index 0000000..e6b5c5d --- /dev/null +++ b/tests/test_claims.py @@ -0,0 +1,182 @@ +import json +import os +import subprocess +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import tasks as T +from wflib import lanes as L +from test_cli import Cli +from test_model import LANES + +FOUR = LANES.replace("## Needs human", "- **t-four** [P3] (1h): Four.\n\n## Needs human") + + +class HeldPickTest(unittest.TestCase): + def test_pick_skips_held(self): + item, skipped = L.pick(T.parse(FOUR), set(), L.DEFAULT_LANES, "1h", None, "opus", + held={"t-one": "opus session uds:/a"}) + self.assertEqual((item.id, [(i.id, why) for i, why in skipped]), + ("t-two", [("t-one", "in progress by opus session uds:/a")])) + + +class ClaimCliTest(Cli): + tasks_text = FOUR + + def env(self, name, pid): + sock = self.root / f"{name}.sock" + sock.write_text("") + return {"CLAUDE_CODE_MESSAGING_SOCKET": str(sock), "CLAUDE_PID": str(pid), "CLAUDE_CODE_SESSION_ID": name} + + def dead_pid(self): + p = subprocess.Popen(["true"]) + p.wait() + return p.pid + + def claim(self, id): + return json.loads((self.root / ".wf" / "claims" / f"{id}.json").read_text()) + + def test_progress_claims_and_other_session_skips(self): + a, b = self.env("a", os.getpid()), self.env("b", os.getppid()) + self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=a) + self.ok("status", "t-three", "progress", "x", env=a) + c = self.claim("t-three") + self.assertEqual({k: c[k] for k in ("pid", "socket", "lane", "model")}, + {"pid": os.getpid(), "socket": str(self.root / "a.sock"), "lane": "fast", "model": "opus"}) + self.assertEqual(self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=a).splitlines()[0], + "- **t-three** [P2] (<1h) (in progress: x): Three.") + code, out, err = self.wf("next", "--lane", "fast", "--as", "opus", env=b) + self.assertEqual((code, err), (1, "wf: nothing pickable for fast (opus) in Pending\n")) + self.assertIn(f"===== Skipped =====\n- t-three: in progress by opus session uds:{self.root / 'a.sock'}\n", out) + out = self.ok("next", "--lane", "slow", "--as", "opus", env=b) + self.assertIn("===== Next task =====\n- **t-one**", out) + + def test_dead_claim_ignored(self): + self.ok("status", "t-three", "progress", "x", env=self.env("a", self.dead_pid())) + self.assertEqual(self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=self.env("b", os.getpid())).splitlines()[0], + "- **t-three** [P2] (<1h) (in progress: x): Three.") + + def test_claim_without_in_progress_status_ignored(self): + self.ok("status", "t-three", "progress", "x", env=self.env("a", os.getppid())) + self.ok("status", "t-three", "clear", env=self.env("b", os.getpid())) + self.assertFalse((self.root / ".wf" / "claims" / "t-three.json").exists()) + self.assertEqual(self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=self.env("b", os.getpid())).splitlines()[0], + "- **t-three** [P2] (<1h): Three.") + + def test_done_and_blocked_release(self): + a = self.env("a", os.getpid()) + self.ok("status", "t-three", "progress", "x", env=a) + self.ok("done", "t-three", "-m", "ok", env=a) + self.assertFalse((self.root / ".wf" / "claims" / "t-three.json").exists()) + self.ok("status", "t-four", "progress", "x", env=a) + self.ok("add", "-s", "awaiting", "Q?", env=a) + self.ok("status", "t-four", "blocked", "a-q", env=a) + self.assertFalse((self.root / ".wf" / "claims" / "t-four.json").exists()) + + def test_no_env_no_claim(self): + self.ok("status", "t-three", "progress", "x", env={"CLAUDE_PID": "", "CLAUDE_CODE_MESSAGING_SOCKET": ""}) + self.assertFalse((self.root / ".wf" / "claims").exists()) + self.assertFalse((self.root / ".wf" / "sessions").exists()) + + def test_same_lane_live_session_warns_and_keeps_registry(self): + a, b = self.env("a", os.getppid()), self.env("b", os.getpid()) + self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=a) + out = self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=b) + self.assertEqual(out.splitlines()[0], + f"another live fast session holds this lane: uds:{self.root / 'a.sock'} " + "(claims keep tasks apart; tell the owner if unintended)") + reg = json.loads((self.root / ".wf" / "sessions" / "fast.json").read_text()) + self.assertEqual(reg["socket"], str(self.root / "a.sock")) + self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=a) # own re-register: no warning + self.assertNotIn("another live", self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=a)) + + def test_lanes_unregister_drops_own_record(self): + a, b = self.env("a", os.getpid()), self.env("b", os.getppid()) + self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=a) + self.ok("next", "--lane", "slow", "--as", "sonnet", "--brief", env=b) + out = self.ok("lanes", "--unregister", env=a) + self.assertFalse((self.root / ".wf" / "sessions" / "fast.json").exists()) + self.assertTrue((self.root / ".wf" / "sessions" / "slow.json").exists()) + self.assertNotIn(str(self.root / "a.sock"), out) + + +def git(cwd, *args): + subprocess.run(["git", "-C", str(cwd), *args], check=True, capture_output=True, + env={**os.environ, "GIT_AUTHOR_NAME": "t", "GIT_AUTHOR_EMAIL": "t@t", "GIT_COMMITTER_NAME": "t", + "GIT_COMMITTER_EMAIL": "t@t"}) + + +class WorktreeTest(Cli): + tasks_text = FOUR + + def setUp(self): + super().setUp() + git(self.root, "init", "-q", "-b", "master") + (self.root / ".gitignore").write_text(".worktrees/\n") + git(self.root, "add", "-A") + git(self.root, "commit", "-qm", "init") + self.wt = self.root / ".worktrees" / "fast" + git(self.root, "worktree", "add", "-q", str(self.wt), "-b", "fast/t-three") + + def test_worktree_writes_main_tree(self): + none = {"CLAUDE_CODE_MESSAGING_SOCKET": ""} + self.wf("status", "t-three", "progress", "x", project=False, cwd=self.wt, env=none) + self.assertIn("(in progress: x)", (self.root / "TASKS.md").read_text()) + self.assertNotIn("in progress", (self.wt / "TASKS.md").read_text()) + code, out, err = self.wf("done", "t-four", "-m", "ok", project=False, cwd=self.wt / "docs", env=none) + self.assertEqual(code, 0, err) + self.assertIn("**t-four**", self.archive()) + self.assertNotIn("t-four", (self.wt / "tasks" / "archive.md").read_text()) + + def test_done_in_worktree_prints_merge_steps(self): + code, out, err = self.wf("done", "t-four", "-m", "ok", project=False, cwd=self.wt, + env={"CLAUDE_CODE_MESSAGING_SOCKET": ""}) + self.assertEqual(code, 0, err) + self.assertTrue(out.endswith( + "worktree mode (branch fast/t-three), after verify: commit your code here (explicit paths), then:\n" + " wf merge (rebase, ff-merge into master, commit TASKS.md tasks/archive.md, push home; " + "conflict → git rebase master, resolve, verify, wf merge again)\n"), out) + self.assertNotIn("worktree mode", self.ok("done", "t-three", "-m", "ok")) + + def env(self, name, pid): + sock = self.root / f"{name}.sock" + sock.write_text("") + return {"CLAUDE_CODE_MESSAGING_SOCKET": str(sock), "CLAUDE_PID": str(pid), "CLAUDE_CODE_SESSION_ID": name} + + def test_next_says_worktree_mode_when_two_sessions_live(self): + son, me = self.env("s", os.getppid()), self.env("o", os.getpid()) + self.assertNotIn("Multi-session", self.ok("next", "--lane", "fast", "--as", "opus", env=me)) + self.ok("next", "--lane", "slow", "--as", "sonnet", "--brief", env=son) + out = self.ok("next", "--lane", "fast", "--as", "opus", env=me) + self.assertIn("===== Multi-session =====\n" + "2 live sessions here: work in your lane's worktree, never on master:\n" + " cd .worktrees/fast && git switch -c fast/<task> master (wf there writes this TASKS.md)\n\n", out) + self.ok("next", "--lane", "slow", "--as", "sonnet", "--brief", env=son) + out = self.ok("next", "--lane", "slow", "--as", "sonnet", env=son) + self.assertIn(" git worktree add .worktrees/slow -b slow/<task> master (wf there writes this TASKS.md)\n", out) + code, out, err = self.wf("next", "--lane", "fast", "--as", "opus", project=False, cwd=self.wt, env=me) + self.assertIn("===== Multi-session =====\n2 live sessions here: you are in worktree .worktrees/fast " + "(branch fast/t-three); wf writes the main tree's TASKS.md\n", out) + + +if __name__ == "__main__": + unittest.main() + + +class StaleTest(ClaimCliTest): + def test_stale_list_and_clear(self): + self.ok("status", "t-three", "progress", "x", env=self.env("a", self.dead_pid())) + self.ok("status", "t-four", "progress", "y", env=self.env("b", os.getpid())) + self.assertEqual(self.ok("list", "--stale").splitlines()[0].split()[0], "t-three") + self.assertNotIn("t-four", self.ok("list", "--stale")) + self.assertIn("cleared t-three", self.ok("status", "--clear-stale")) + self.assertEqual(self.ok("list", "--progress").count("prog"), 1) + self.assertIn("t-four", self.ok("list", "--progress")) + self.assertFalse((self.root / ".wf" / "claims" / "t-three.json").exists()) + + def test_no_claim_listed(self): + self.ok("status", "t-three", "progress", "x") + (self.root / ".wf" / "claims" / "t-three.json").unlink(missing_ok=True) + self.assertIn("t-three", self.ok("list", "--stale")) diff --git a/tests/test_cli.py b/tests/test_cli.py new file mode 100644 index 0000000..43135c9 --- /dev/null +++ b/tests/test_cli.py @@ -0,0 +1,752 @@ +import datetime +import os +import subprocess +import sys +import tempfile +import unittest +from pathlib import Path + +HERE = Path(__file__).resolve().parent.parent +WF = HERE / "wf.py" +sys.path.insert(0, str(HERE)) + +TOML = ('format = 1\ntasks = "TASKS.md"\narchive = "tasks/archive.md"\ndocs = ["DESIGN.md", "docs/"]\n' + 'verify = ["make test"]\ndone = ["update CATALOG"]\nledgers = ".superpowers/sdd"\n') +TASKS = """\ +# Tasks — demo + +## Awaiting your decision + +- **a-key**: Key needed. Which one? + +## Pending + +- **t-one** [P1] (1h) (in progress: master): First thing. Do it. + - Steps: a + Ref: DESIGN.md#terrain, docs/plan.md + +- **t-two** [P2] (5h) (blocked: [[a-key]]): Second thing. + +- **t-three** [P2] (<1h): Third thing. Later. + - After: [[t-one]] + +## Needs human + +- **t-play** [P1] (<1h): Play-test feel. + - [ ] keyboard works + +## Deferred + +- **t-far** [P3] (100h): Far future. +""" +ARCHIVE = "# Archive (newest first)\n\n- 2026-09-01 **t-done** Done thing — ok\n- 2026-08-01 **t-older** Older thing — fine\n" +TODAY = datetime.date.today().isoformat() + + +class Cli(unittest.TestCase): + tasks_text = TASKS + toml = TOML + + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.root = Path(self.tmp.name).resolve() / "demo" + (self.root / "tasks").mkdir(parents=True) + (self.root / "docs").mkdir() + (self.root / "workflow.toml").write_text(self.toml) + (self.root / "TASKS.md").write_bytes(self.tasks_text.encode()) + (self.root / "tasks" / "archive.md").write_text(ARCHIVE) + (self.root / "DESIGN.md").write_text("# Design\n\n## terrain\n\nHeightmap.\n") + (self.root / "docs" / "plan.md").write_text("# The Plan\n\n**Goal:** Ship it.\n") + self.inbox = self.root.parent / "inbox.md" + + def tearDown(self): + self.tmp.cleanup() + + def wf(self, *args, stdin="", cwd=None, project=True, env=None, hints=False): + cmd = [sys.executable, str(WF), *(["--project", str(self.root)] if project else []), *args] + full = {**os.environ, "WF_INBOX": str(self.inbox), "WF_ROOT": str(self.root.parent), + "WF_TOOL_ROOT": str(self.root.parent / "no-tool"), "CLAUDE_CONFIG_DIR": str(self.root.parent / "no-claude"), **(env or {})} + r = subprocess.run(cmd, input=stdin, capture_output=True, text=True, cwd=cwd or self.root, env=full, timeout=30) + err = r.stderr if hints else "".join(l for l in r.stderr.splitlines(True) if not l.startswith("hint:")) + return r.returncode, r.stdout, err + + def ok(self, *args, **kw): + code, out, err = self.wf(*args, **kw) + self.assertEqual((code, err), (0, ""), out + err) + return out + + def fails(self, *args, code=1, **kw): + got, out, err = self.wf(*args, **kw) + self.assertEqual(got, code, out + err) + self.assertNotIn("Traceback", err) + return err + + def tasks(self): + return (self.root / "TASKS.md").read_text() + + def archive(self): + return (self.root / "tasks" / "archive.md").read_text() + + def item(self, id): + return self.ok("show", id) + + +class ReadTest(Cli): + def test_list_default_is_pending(self): + self.assertEqual(self.ok("list"), + "t-one P1 1h prog opus slow First thing\n" + "t-two P2 5h blkd opus fast Second thing\n" + "t-three P2 <1h - opus fast Third thing\n" + "pending 3 · human 1 · awaiting 1 · deferred 1\n") + + def test_list_all(self): + self.assertEqual(self.ok("list", "-s", "all"), + "== Awaiting your decision\n" + "a-key -- - - - - Key needed\n" + "== Pending\n" + "t-one P1 1h prog opus slow First thing\n" + "t-two P2 5h blkd opus fast Second thing\n" + "t-three P2 <1h - opus fast Third thing\n" + "== Needs human\n" + "t-play P1 <1h - opus fast Play-test feel\n" + "== Deferred\n" + "t-far P3 100h - opus fast Far future\n" + "pending 3 · human 1 · awaiting 1 · deferred 1\n") + + def first_column(self, *args): + return [l.split()[0] for l in self.ok("list", *args).splitlines()[:-1]] + + def test_list_filters(self): + self.assertEqual(self.first_column("-p", "1"), ["t-one"]) + self.assertEqual(self.first_column("--ready"), ["t-one"]) + self.assertEqual(self.first_column("--blocked"), ["t-two"]) + self.assertEqual(self.first_column("--progress"), ["t-one"]) + self.assertEqual(self.first_column("--ref", "DESIGN.md"), ["t-one"]) + self.assertEqual(self.first_column("--ref", "DESIGN.md#terrain"), ["t-one"]) + self.assertEqual(self.first_column("--ref", "DESIGN.md#other"), []) + self.assertEqual(self.first_column("-n", "2"), ["t-one", "t-two"]) + self.assertEqual(self.first_column("-s", "deferred"), ["t-far"]) + + def test_list_cuts_long_titles_to_100_columns(self): + self.ok("set", "t-far", "--title", "x" * 150) + line = self.ok("list", "-s", "deferred").splitlines()[0] + self.assertEqual(len(line), 100) + self.assertTrue(line.endswith("xxx…")) + + def test_show(self): + self.assertEqual(self.ok("show", "t-three", "a-key"), + "- **t-three** [P2] (<1h): Third thing. Later.\n - After: [[t-one]]\n\n" + "- **a-key**: Key needed. Which one?\n") + + def test_show_by_unique_prefix(self): + self.assertEqual(self.ok("show", "t-thr"), "- **t-three** [P2] (<1h): Third thing. Later.\n - After: [[t-one]]\n") + + def test_show_ambiguous_prefix(self): + self.assertEqual(self.fails("show", "t-t"), "wf: 't-t' matches t-three, t-two\n") + + def test_show_unknown(self): + self.assertRegex(self.fails("show", "t-zzz"), r"^wf: unknown id 't-zzz' \(nearest: .*\)\n$") + + def test_next(self): + self.assertEqual(self.ok("next", "--as", "opus"), + "===== Awaiting your decision (mention, don't block) =====\n" + "- a-key: Key needed. Which one?\n\n" + "===== Needs human (not picked) =====\n" + "- t-play: Play-test feel\n\n" + "===== Lanes =====\n" + "fast: 0 pickable · 2 waiting · no session\n" + "slow: 1 pickable · 0 waiting · no session → orchestrator or owner: start one (wf next --lane slow)\n\n" + "===== Next task =====\n" + "- **t-one** [P1] (1h) (in progress: master): First thing. Do it.\n" + " - Steps: a\n" + " Ref: DESIGN.md#terrain, docs/plan.md\n\n" + "===== DESIGN.md#terrain (line 3) =====\n## terrain\n\nHeightmap.\n\n" + "docs/plan.md: The Plan — Ship it.\n") + + def test_next_brief(self): + self.assertEqual(self.ok("next", "--as", "opus", "--brief"), + "- **t-one** [P1] (1h) (in progress: master): First thing. Do it.\n" + " - Steps: a\n Ref: DESIGN.md#terrain, docs/plan.md\n") + + def test_next_lists_what_it_skipped(self): + self.ok("move", "t-one", "deferred") + self.ok("prio", "t-far", "3") + code, out, err = self.wf("next", "--as", "opus") + self.assertEqual(code, 1) + self.assertIn("===== Skipped =====\n- t-two: blocked: a-key\n- t-three: after: t-one\n", out) + self.assertEqual(err, "wf: nothing pickable for all lanes (opus) in Pending\n") + + def test_next_shows_ledgers_and_inbox(self): + sdd = self.root / ".superpowers" / "sdd" / "p" + sdd.mkdir(parents=True) + (self.root / "docs" / "p.md").write_text("# P\n**Goal:** g\n### Task 1: A\n### Task 2: B\n### Task 3: C\n") + (sdd / "progress.md").write_text("# SDD ledger — plan: docs/p.md\nTask 1: complete (x)\nTask 2: Ruling: y\nTask 2: complete (z)\n") + self.inbox.write_text("- 2026-09-29 demo bug: x @abc\n- 2026-09-29 demo idea: y @abc\n") + out = self.ok("next", "--as", "opus") + self.assertIn("===== In-flight plans =====\n- docs/p.md: done 1, 2; resume at Task 3 (task-start PLAN 3) " + "(1 rulings; ledger .superpowers/sdd/p/progress.md)\n\n", out) + self.assertTrue(out.endswith("\nworkflow inbox: 2 reports (triage: workflow session)\n"), out) + + def test_ctx_item(self): + self.assertEqual(self.ok("ctx", "t-one"), + "- **t-one** [P1] (1h) (in progress: master): First thing. Do it.\n" + " - Steps: a\n" + " Ref: DESIGN.md#terrain, docs/plan.md\n\n" + "Section: Pending\n" + "Needed by: t-three\n" + "Verify (run before wf finish): make test\n\n" + "===== DESIGN.md#terrain (line 3) =====\n## terrain\n\nHeightmap.\n\n" + "docs/plan.md: The Plan — Ship it.\n") + + def test_ctx_dependencies_and_links(self): + out = self.ok("ctx", "t-three") + self.assertIn("Section: Pending\nAfter: t-one (open, Pending)\n", out) + self.assertEqual(self.ok("ctx", "a-key").split("\n\n", 1)[1], "Section: Awaiting your decision\nBlocks: t-two\n") + + def test_ctx_archived_id(self): + self.assertEqual(self.ok("ctx", "t-done"), "done: - 2026-09-01 **t-done** Done thing — ok\n") + + def test_ctx_anchor(self): + self.assertEqual(self.ok("ctx", "DESIGN.md#terrain"), + "===== DESIGN.md#terrain (line 3) =====\n## terrain\n\nHeightmap.\n\n" + "Tasks: t-one\n") + + def test_search(self): + self.assertEqual(self.ok("search", "thing", "later"), + "t-three Third thing. Later.\n" + "t-one First thing. Do it.\n" + "t-two Second thing.\n" + "tasks/archive.md:3 t-done Done thing — ok\n" + "tasks/archive.md:4 t-older Older thing — fine\n") + + def test_search_docs(self): + self.assertEqual(self.ok("search", "--docs", "height"), "DESIGN.md:3 terrain — Heightmap.\n") + + def test_search_nothing(self): + self.assertEqual(self.fails("search", "zzz"), "wf: no hits\n") + + def test_log(self): + self.assertEqual(self.ok("log", "-n", "1"), "- 2026-09-01 **t-done** Done thing — ok\n") + self.assertEqual(self.ok("log", "older"), "- 2026-08-01 **t-older** Older thing — fine\n") + + def test_check_clean(self): + self.assertEqual(self.ok("check"), "OK: 0 errors · 0 warnings\n") + + def test_check_errors(self): + (self.root / "TASKS.md").write_text(TASKS.replace("[[t-one]]", "[[t-gone]]")) + code, out, err = self.wf("check") + self.assertEqual((code, err), (1, "")) + self.assertEqual(out, "ERROR: TASKS.md:16: [[t-gone]]: no such id in TASKS or archive\n1 errors · 0 warnings\n") + + def test_projects(self): + other = self.root.parent / "group" / "other" + (other / "tasks").mkdir(parents=True) + (other / "workflow.toml").write_text('format = 1\ntasks = "TASKS.md"\narchive = "tasks/archive.md"\n') + (other / "TASKS.md").write_text("## Awaiting your decision\n\n## Pending\n\n- **t-x** [P9] (1h): X.\n\n## Needs human\n\n## Deferred\n") + (other / "tasks" / "archive.md").write_text("# Archive\n") + hidden = self.root.parent / ".gate" + hidden.mkdir() + (hidden / "workflow.toml").write_text(TOML) + tool = self.root.parent / "workflow" # a wf checkout: its templates/ is no project + (tool / "wflib").mkdir(parents=True) + (tool / "templates").mkdir() + (tool / "wf.py").write_text("") + (tool / "templates" / "workflow.toml").write_text(TOML) + self.assertEqual(self.ok("projects", project=False), + "demo pending 3 human 1 awaiting 1 errors 0 next: t-one First thing\n" + "group/other pending 1 human 0 awaiting 0 errors 1 next: t-x X\n") + + def test_outside_a_project(self): + err = self.fails("list", project=False, cwd=self.root.parent) + self.assertRegex(err, r"^wf: no workflow.toml in .* or above: not a wf project \(wf init makes one\)\n$") + + def test_broken_config(self): + (self.root / "workflow.toml").write_text(TOML + "verfy = []\n") + self.assertEqual(self.fails("list"), "wf: workflow.toml: unknown key 'verfy'\n") + + def test_finds_project_from_subfolder(self): + self.assertIn("t-one", self.ok("list", project=False, cwd=self.root / "docs")) + + def test_no_command_is_usage_error(self): + self.fails(code=2) + + +class AddTest(Cli): + def test_add_places_by_priority_and_prints_header(self): + out = self.ok("add", "Cave seams. Close the slit.", "-p", "1", "-e", "<1h") + self.assertEqual(out, "- **t-cave-seams** [P1] (<1h): Cave seams. Close the slit.\n") + self.assertEqual(self.first_ids(), ["t-one", "t-cave-seams", "t-two", "t-three"]) + + def first_ids(self, section="pending"): + return [l.split()[0] for l in self.ok("list", "-s", section).splitlines()[:-1]] + + def test_add_with_options(self): + self.ok("add", "Docs pass", "-p", "2", "-e", "5h", "--sessions", "owner", "--id", "t-docs", + "--after", "t-one,t-done", "--ref", "docs/plan.md,DESIGN.md#terrain") + self.assertEqual(self.item("t-docs"), + "- **t-docs** [P2] (5h): Docs pass.\n" + " Sessions: owner\n" + " - After: [[t-one]], [[t-done]]\n" + " Ref: docs/plan.md, DESIGN.md#terrain\n") + + def test_add_cloud(self): + self.ok("add", "Code fix", "-p", "1", "-e", "1h", "--done", "tests pass", "--model", "sonnet", "--cloud", "yes", + "--id", "t-cf") + self.assertEqual(self.item("t-cf"), "- **t-cf** [P1] (1h): Code fix.\n Done: tests pass\n" + " Model: sonnet\n Cloud: yes\n") + self.ok("add", "Gui fix", "-p", "1", "-e", "1h", "--cloud", "no", "--id", "t-gf") + self.assertIn(" Cloud: no\n", self.item("t-gf")) + + def test_add_body_from_stdin(self): + self.ok("add", "With body", "-p", "3", "-e", "1h", "-b", stdin="- Steps: a\n - sub\n- Done: b\n") + self.assertEqual(self.item("t-with-body"), + "- **t-with-body** [P3] (1h): With body.\n - Steps: a\n - sub\n - Done: b\n") + + def test_add_to_other_sections(self): + self.assertEqual(self.ok("add", "Which colour. Red or blue?", "-s", "awaiting"), + "- **a-which-colour** : Which colour. Red or blue?\n".replace(" :", ":")) + self.ok("add", "Feel check", "-p", "2", "-e", "<1h", "-s", "human") + self.assertEqual(self.first_ids("human"), ["t-play", "t-feel-check"]) + + def test_add_slice(self): + self.ok("add", "Map scaffold", "-e", "1h", "--parent", "t-far") + self.ok("add", "Combat rows", "-e", "1h", "--parent", "t-far") + self.assertEqual(self.ok("show", "t-far", "t-far-1", "t-far-2"), + "- **t-far** [P3] (100h): Far future.\n - Slices: [[t-far-1]], [[t-far-2]]\n\n" + "- **t-far-1** [P3] (1h): Map scaffold.\n\n" + "- **t-far-2** [P3] (1h): Combat rows.\n - After: [[t-far-1]]\n") + self.assertEqual(self.first_ids("deferred"), ["t-far", "t-far-1", "t-far-2"]) + + def test_add_slice_skips_deferred_previous_slice(self): + self.ok("add", "Map scaffold", "-e", "1h", "--parent", "t-far") + self.ok("add", "Combat rows", "-e", "1h", "--parent", "t-far", "-s", "pending") + self.ok("add", "Loot rows", "-e", "1h", "--parent", "t-far", "-s", "pending") + self.assertEqual(self.ok("show", "t-far-2", "t-far-3"), + "- **t-far-2** [P3] (1h): Combat rows.\n\n" + "- **t-far-3** [P3] (1h): Loot rows.\n - After: [[t-far-2]]\n") + + def test_add_named_slice_chains_after_previous(self): + self.ok("add", "Map scaffold", "-e", "1h", "--parent", "t-far", "--id", "t-map") + self.ok("add", "Combat rows", "-e", "1h", "--parent", "t-far", "--id", "t-combat") + self.assertEqual(self.ok("show", "t-far", "t-combat"), + "- **t-far** [P3] (100h): Far future.\n - Slices: [[t-map]], [[t-combat]]\n\n" + "- **t-combat** [P3] (1h): Combat rows.\n - After: [[t-map]]\n") + self.assertEqual(self.fails("done", "t-far", "-m", "x"), "wf: 't-far' has open slices: t-map, t-combat\n") + self.ok("done", "t-map", "-m", "x") + self.assertIn("last slice of t-far: finish it", self.ok("done", "t-combat", "-m", "x")) + + def test_add_leading_id_in_text(self): + self.assertEqual(self.ok("add", "a-colour: Which colour? Red or blue?", "-s", "awaiting"), + "- **a-colour**: Which colour? Red or blue?\n") + self.assertEqual(self.ok("add", "t-seams: Cave seams. Close them", "-p", "1", "-e", "1h"), + "- **t-seams** [P1] (1h): Cave seams. Close them.\n") + + def test_add_help_says_id_works_with_parent(self): + self.assertIn("--parent", self.ok("add", "-h").split("--id ID", 2)[2].split("--after")[0]) + + def test_add_block(self): + self.ok("add", "-", stdin="- **t-blk** [P0] (1h): Block. Goal.\n - Steps: a\n Done: x\n") + self.assertEqual(self.item("t-blk"), "- **t-blk** [P0] (1h): Block. Goal.\n - Steps: a\n Done: x\n") + self.assertEqual(self.first_ids()[0], "t-blk") + + def test_add_block_in_old_format_is_refused(self): + self.assertEqual(self.fails("add", "-", stdin="3. **[P2] Old** (Effort: 1h) — goal.\n"), + "wf: old numbered format: write '- **id** [Pn] (effort): Title. Goal.'\n") + + def test_add_needs_priority_and_effort(self): + self.assertIn("-p", self.fails("add", "No prio", "-e", "1h", code=2)) + self.assertIn("-e", self.fails("add", "No effort", "-p", "1", code=2)) + self.assertEqual(self.tasks(), TASKS) + + def test_add_title_without_letters(self): + self.assertEqual(self.fails("add", "???", "-p", "1", "-e", "1h"), + "wf: no id can be made from title '???': give one with --id\n") + + def test_add_same_title_twice(self): + self.ok("add", "Cave seams", "-p", "1", "-e", "1h") + self.assertEqual(self.ok("add", "Cave seams", "-p", "1", "-e", "1h"), "- **t-cave-seams-2** [P1] (1h): Cave seams.\n") + + def test_add_refused_when_result_fails_check(self): + self.assertEqual(self.fails("add", "Bad ref", "-p", "1", "-e", "1h", "--ref", "docs/none.md"), + "wf: refused: nothing written, the change adds problems:\n TASKS.md:14: t-bad-ref: Ref 'docs/none.md' does not exist\n") + self.assertEqual(self.tasks(), TASKS) + + def test_dry_run(self): + out = self.ok("add", "Cave seams", "-p", "1", "-e", "1h", "--dry-run") + self.assertIn("+- **t-cave-seams** [P1] (1h): Cave seams.\n", out) + self.assertIn("--- TASKS.md\n+++ TASKS.md (new)\n", out) + self.assertEqual(self.tasks(), TASKS) + + +class DoneTest(Cli): + def test_done_builtin_checklist_empty_cfg(self): + self.toml = TOML.replace('verify = ["make test"]\ndone = ["update CATALOG"]\n', '') + self.setUp() + out = self.ok("done", "t-one", "-m", "x") + self.assertEqual(out, "done: t-one → tasks/archive.md\nfast work now pickable: t-three (no session: tell the owner)\nchecklist:\n - re-learned anything (>3 greps to find)? one anchor line → that area's code map\n") + + def test_done(self): + out = self.ok("done", "t-one", "-m", "shipped\nwith care") + self.assertEqual(out, "done: t-one → tasks/archive.md\nfast work now pickable: t-three (no session: tell the owner)\nverify:\n make test\nchecklist:\n - re-learned anything (>3 greps to find)? one anchor line → that area's code map\n - update CATALOG\n") + self.assertNotIn("t-one**", self.tasks()) + self.assertEqual(self.archive().splitlines()[2], f"- {TODAY} **t-one** First thing — shipped with care") + self.assertEqual(self.ok("check"), "OK: 0 errors · 0 warnings\n") + self.assertEqual(self.ok("next", "--as", "opus", "--brief").splitlines()[0], "- **t-three** [P2] (<1h): Third thing. Later.") + + def test_done_needs_entry(self): + self.assertEqual(self.fails("done", "t-one", code=2), "wf: done: -m \"<entry>\" is required for a single task\n") + + def test_done_several_uses_goals(self): + self.ok("done", "t-one", "t-three") + self.assertEqual(self.archive().splitlines()[2:4], + [f"- {TODAY} **t-three** Third thing — Later.", f"- {TODAY} **t-one** First thing — Do it."]) + + def test_done_refuses_parent_with_open_slices(self): + self.ok("add", "Map scaffold", "-e", "1h", "--parent", "t-far") + self.assertEqual(self.fails("done", "t-far", "-m", "x"), "wf: 't-far' has open slices: t-far-1\n") + + def test_done_awaiting_item_unblocks(self): + self.assertEqual(self.ok("done", "a-key"), "removed: a-key\nunblocked: t-two\n") + self.assertEqual(self.item("t-two"), "- **t-two** [P2] (5h): Second thing.\n") + self.assertEqual(self.archive(), ARCHIVE) + + def test_done_unknown(self): + self.assertRegex(self.fails("done", "t-on", "-m", "x"), r"^wf: unknown id 't-on' \(nearest: t-one") + self.assertEqual(self.archive(), ARCHIVE) + + +class EditTest(Cli): + def order(self): + return [l.split()[0] for l in self.ok("list").splitlines()[:-1]] + + def test_prio(self): + self.assertEqual(self.ok("prio", "t-three", "0"), "- **t-three** [P0] (<1h): Third thing. Later.\n") + + def test_prio_refused_when_it_jumps_a_dependency(self): + # t-three is After: t-one, the insert rule keeps it behind t-one + self.ok("prio", "t-three", "0") + self.assertEqual(self.order(), ["t-one", "t-three", "t-two"]) + + def test_move_section(self): + self.assertEqual(self.ok("move", "t-two", "deferred"), "- **t-two** [P2] (5h) (blocked: [[a-key]]): Second thing.\n") + self.assertEqual(self.order(), ["t-one", "t-three"]) + + def test_move_relative(self): + self.ok("move", "t-three", "--before", "t-two") + self.assertEqual(self.order(), ["t-one", "t-three", "t-two"]) + self.assertEqual(self.fails("move", "t-two", "--before", "t-one"), + "wf: moving 't-two' there breaks priority order (wf prio, or --force)\n") + self.ok("move", "t-two", "--before", "t-one", "--force") + self.assertEqual(self.order(), ["t-two", "t-one", "t-three"]) + + def test_move_needs_one_target(self): + self.fails("move", "t-two", code=2) + self.fails("move", "t-two", "deferred", "--before", "t-one", code=2) + + def test_status(self): + self.assertEqual(self.ok("status", "t-three", "progress", "feature/x"), + "- **t-three** [P2] (<1h) (in progress: feature/x): Third thing. Later.\n") + self.assertEqual(self.ok("status", "t-three", "blocked", "a-key"), + "- **t-three** [P2] (<1h) (blocked: [[a-key]]): Third thing. Later.\n") + self.assertEqual(self.ok("status", "t-three", "clear"), "- **t-three** [P2] (<1h): Third thing. Later.\n") + self.assertEqual(self.fails("status", "t-three", "blocked", "a-none"), "wf: 'a-none' is not an open Awaiting item\n") + self.fails("status", "t-three", "progress", code=2) + + def test_set(self): + self.ok("set", "t-two", "--title", "Second: renamed", "--effort", "10h", "--sessions", "owner", + "--after", "t-one", "--ref", "docs/plan.md (goal)") + self.assertEqual(self.item("t-two"), + "- **t-two** [P2] (10h) (blocked: [[a-key]]): Second: renamed.\n" + " Sessions: owner\n - After: [[t-one]]\n Ref: docs/plan.md (goal)\n") + self.ok("set", "t-two", "--after", "", "--ref", "", "--sessions", "") + self.assertEqual(self.item("t-two"), "- **t-two** [P2] (10h) (blocked: [[a-key]]): Second: renamed.\n") + + def test_set_done(self): + self.ok("set", "t-two", "--done", "two works") + self.assertEqual(self.item("t-two"), "- **t-two** [P2] (5h) (blocked: [[a-key]]): Second thing.\n Done: two works\n") + self.ok("set", "t-two", "--done", "") + self.assertEqual(self.item("t-two"), "- **t-two** [P2] (5h) (blocked: [[a-key]]): Second thing.\n") + + def test_rename(self): + self.assertEqual(self.ok("rename", "a-key", "a-which-key"), "renamed: a-key → a-which-key (1 link)\n") + self.assertEqual(self.item("t-two"), "- **t-two** [P2] (5h) (blocked: [[a-which-key]]): Second thing.\n") + self.assertEqual(self.fails("rename", "t-one", "t-done"), "wf: id 't-done' already used in the archive (ids are never reused)\n") + + def test_set_refused_on_bad_ref(self): + self.assertIn("Ref 'missing.md' does not exist", self.fails("set", "t-two", "--ref", "missing.md")) + self.assertEqual(self.tasks(), TASKS) + + def test_set_without_fields(self): + self.fails("set", "t-two", code=2) + + def test_note(self): + self.ok("note", "t-one", "cause found") + self.assertEqual(self.item("t-one"), + "- **t-one** [P1] (1h) (in progress: master): First thing. Do it.\n" + " - Steps: a\n - cause found\n Ref: DESIGN.md#terrain, docs/plan.md\n") + + def test_body(self): + self.ok("body", "t-one", stdin="- Steps: b\n- Done: c\n") + self.assertEqual(self.item("t-one"), + "- **t-one** [P1] (1h) (in progress: master): First thing. Do it.\n" + " - Steps: b\n - Done: c\n Ref: DESIGN.md#terrain, docs/plan.md\n") + + def test_body_refusal_says_refused_and_writes_nothing(self): + before = self.tasks() + for _ in range(2): + err = self.fails("body", "t-one", stdin="- After: [[t-three]]\n") + self.assertTrue(err.startswith("wf: refused: nothing written"), err) + self.assertIn("placed before 't-three'", err) + self.assertEqual(self.tasks(), before) + + def test_body_positional_text_hints_stdin(self): + self.assertIn("stdin", self.fails("body", "t-one", "some", "text")) + self.assertIn("stdin", self.ok("body", "-h")) + + def test_tick(self): + self.assertEqual(self.ok("tick", "t-play", "keyboard"), "ticked: keyboard works\n") + self.assertIn(" - [x] keyboard works\n", self.tasks()) + self.assertEqual(self.fails("tick", "t-play", "1"), "wf: box already ticked: keyboard works\n") + + def test_write_needs_exact_id(self): + self.assertRegex(self.fails("prio", "t-thr", "1"), r"^wf: unknown id 't-thr' \(nearest: t-three") + + def test_untouched_text_survives_a_write(self): + self.ok("note", "t-far", "x") + self.assertEqual(self.tasks(), TASKS.replace("Far future.\n", "Far future.\n - x\n")) + + +class OldProblemsTest(Cli): + tasks_text = TASKS + "\n## Notes\n\nSee [[t-lost]].\n" + + def test_existing_problems_do_not_block_other_edits(self): + self.ok("note", "t-far", "x") + self.assertIn(" - x\n", self.tasks()) + + +class CrlfTest(Cli): + tasks_text = TASKS.replace("\n", "\r\n") + + def test_write_keeps_crlf(self): + self.ok("add", "Cave seams", "-p", "1", "-e", "1h") + raw = (self.root / "TASKS.md").read_bytes() + self.assertIn(b"- **t-cave-seams** [P1] (1h): Cave seams.\r\n", raw) + self.assertEqual(raw.count(b"\n"), raw.count(b"\r\n")) + self.assertTrue(raw.endswith(b"Far future.\r\n")) + + def test_read_commands_work(self): + self.assertEqual(self.ok("show", "a-key"), "- **a-key**: Key needed. Which one?\n") + + +class FormatGateTest(Cli): + toml = TOML.replace("format = 1", "format = 0") + + def test_write_refused(self): + self.assertEqual(self.fails("note", "t-one", "x"), + "wf: TASKS format 0, wf needs 1: run wf migrate --write (idle project, one commit)\n") + self.assertEqual(self.tasks(), TASKS) + + def test_read_works(self): + self.assertIn("t-one", self.ok("list")) + + +class NewerFormatTest(Cli): + toml = TOML.replace("format = 1", "format = 2") + + def test_refused(self): + self.assertEqual(self.fails("list"), "wf: project format 2 is newer than this wf (1): update /projects/public/workflow\n") + + +class ReportTest(Cli): + def test_report_line(self): + self.assertEqual(self.ok("report", "prio change\nneeds two commands", "--kind", "friction", "--cmd", "wf prio x 1"), + "reported (workflow inbox); carry on\n") + self.assertRegex(self.inbox.read_text(), + rf"^- {TODAY} demo friction: prio change needs two commands \(cmd: wf prio x 1\) @[0-9a-f]{{7}}\n$") + + def test_default_kind_and_outside_project(self): + self.ok("report", "something", project=False, cwd=self.root.parent) + self.assertRegex(self.inbox.read_text(), rf"^- {TODAY} {self.root.parent.name} bug: something @") + + def test_parallel_reports_stay_whole(self): + env = {**os.environ, "WF_INBOX": str(self.inbox)} + procs = [subprocess.Popen([sys.executable, str(WF), "--project", str(self.root), "report", f"report {n} " + "x" * 3000], + env=env, stdout=subprocess.DEVNULL) for n in range(8)] + for p in procs: + self.assertEqual(p.wait(timeout=30), 0) + lines = self.inbox.read_text().splitlines() + self.assertEqual(len(lines), 8) + for line in lines: + self.assertRegex(line, rf"^- {TODAY} demo bug: report \d x{{3000}} @[0-9a-f]{{7}}$") + + def test_empty_report(self): + self.fails("report", " ", code=2) + + +class InitTest(Cli): + def test_init(self): + new = self.root.parent / "fresh" + new.mkdir() + self.assertEqual(self.ok("init", project=False, cwd=new), + "wrote workflow.toml\nwrote TASKS.md\nwrote tasks/archive.md\nwrote CLAUDE.md\n") + self.assertTrue((new / "TASKS.md").read_text().startswith("# Tasks — fresh\n")) + self.assertTrue((new / "CLAUDE.md").read_text().startswith("# CLAUDE.md — fresh\n")) + self.assertEqual(self.ok("check", project=False, cwd=new), "OK: 0 errors · 0 warnings\n") + self.assertEqual(self.ok("add", "First", "-p", "1", "-e", "1h", project=False, cwd=new), "- **t-first** [P1] (1h): First.\n") + + def test_init_twice(self): + self.assertEqual(self.fails("init"), f"wf: {self.root}/workflow.toml exists already\n") + + def test_init_keeps_existing_tasks_file(self): + new = self.root.parent / "old" + new.mkdir() + (new / "TASKS.md").write_text("mine\n") + self.assertEqual(self.ok("init", project=False, cwd=new), + "wrote workflow.toml\nkept TASKS.md (exists; old format → wf migrate)\nwrote tasks/archive.md\nwrote CLAUDE.md\n") + self.assertEqual((new / "TASKS.md").read_text(), "mine\n") + + def test_init_keeps_existing_claude_md(self): + new = self.root.parent / "oldc" + new.mkdir() + (new / "CLAUDE.md").write_text("mine\n") + out = self.ok("init", project=False, cwd=new) + self.assertIn("kept CLAUDE.md (exists)\n", out) + self.assertEqual((new / "CLAUDE.md").read_text(), "mine\n") + + def test_init_inside_a_project_subfolder_makes_a_new_project_there(self): + self.assertEqual(self.ok("init", project=False, cwd=self.root / "docs").splitlines()[0], "wrote workflow.toml") + self.assertTrue((self.root / "docs" / "workflow.toml").is_file()) + + +class ProjectsRootTest(unittest.TestCase): + def setUp(self): + import wf + self.wf = wf + self.tmp = tempfile.TemporaryDirectory() + self.top = Path(self.tmp.name) + self.tool = self.top / "public" / "workflow" + self.tool.mkdir(parents=True) + + def tearDown(self): + self.tmp.cleanup() + + def project(self, rel): + (self.top / rel).mkdir(parents=True) + (self.top / rel / "workflow.toml").write_text("") + + def test_parent_with_projects_wins(self): + self.project("public/a") + self.project("games/b") + self.assertEqual(self.wf.default_root(self.tool), self.top / "public") + + def test_parent_without_projects_falls_back_to_grandparent(self): + self.project("games/b") + self.assertEqual(self.wf.default_root(self.tool), self.top) + + def test_no_projects_anywhere_keeps_parent(self): + self.assertEqual(self.wf.default_root(self.tool), self.top / "public") + + +class WriteSafetyTest(unittest.TestCase): + def setUp(self): + import wf + self.wf = wf + self.tmp = tempfile.TemporaryDirectory() + self.path = Path(self.tmp.name) / "TASKS.md" + self.path.write_text("one\n") + + def tearDown(self): + self.tmp.cleanup() + + def test_write(self): + stamp = self.wf.stamp(self.path) + self.wf.write_if_unchanged(self.path, "two\n", stamp) + self.assertEqual(self.path.read_text(), "two\n") + self.assertEqual(sorted(p.name for p in self.path.parent.iterdir()), ["TASKS.md"]) + + def test_stops_when_the_file_changed_since_read(self): + stamp = self.wf.stamp(self.path) + self.path.write_text("someone else wrote this\n") + with self.assertRaisesRegex(self.wf.Failure, "TASKS.md changed on disk since it was read: nothing written, run again"): + self.wf.write_if_unchanged(self.path, "two\n", stamp) + self.assertEqual(self.path.read_text(), "someone else wrote this\n") + + def test_keeps_file_mode(self): + self.path.chmod(0o640) + self.wf.write_if_unchanged(self.path, "two\n", self.wf.stamp(self.path)) + self.assertEqual(self.path.stat().st_mode & 0o777, 0o640) + + +if __name__ == "__main__": + unittest.main() + + +OLD_TASKS = """\ +# Tasks — demo + +## Pending + +1. **[P1] First thing** (Effort: 1h) — do it. + - Steps: a + Reference: [DESIGN.md#terrain](DESIGN.md#terrain) + +2. **[P2] Second thing** (Effort: 5h) — later. + - After: First thing + +## Needs human + +## Awaiting your decision + +- Which key to use. +""" +MIGRATED = """\ +# Tasks — demo + +## Pending + +- **t-first-thing** [P1] (1h): First thing. do it. + - Steps: a + Ref: DESIGN.md#terrain + +- **t-second-thing** [P2] (5h): Second thing. later. + - After: [[t-first-thing]] + +## Needs human + +## Awaiting your decision + +- **a-which-key-to-use**: Which key to use. + +## Deferred +""" + + +class MigrateCliTest(Cli): + tasks_text = OLD_TASKS + toml = TOML.replace("format = 1", "format = 0") + + def test_dry_run_prints_diff_and_map_and_writes_nothing(self): + out = self.ok("migrate") + self.assertIn("+- **t-first-thing** [P1] (1h): First thing. do it.\n", out) + self.assertIn("\nids:\n t-first-thing First thing\n t-second-thing Second thing\n a-which-key-to-use Which key to use\n", out) + self.assertTrue(out.endswith("check after migrate: 0 errors\ndry run: nothing written (wf migrate --write)\n"), out) + self.assertEqual(self.tasks(), OLD_TASKS) + self.assertIn("format = 0", (self.root / "workflow.toml").read_text()) + + def test_write(self): + out = self.ok("migrate", "--write") + self.assertTrue(out.endswith("check after migrate: 0 errors\nwrote TASKS.md, workflow.toml (format = 1)\n"), out) + self.assertEqual(self.tasks(), MIGRATED) + self.assertEqual((self.root / "workflow.toml").read_text(), TOML) + self.assertEqual(self.ok("check"), "OK: 0 errors · 0 warnings\n") + self.assertEqual(self.ok("next", "--as", "opus", "--brief").splitlines()[0], "- **t-first-thing** [P1] (1h): First thing. do it.") + + def test_second_run_changes_nothing(self): + self.ok("migrate", "--write") + self.assertEqual(self.ok("migrate", "--write"), "nothing to migrate\n") + self.assertEqual(self.tasks(), MIGRATED) + + def test_ids_avoid_archived_ones(self): + (self.root / "tasks" / "archive.md").write_text("# A\n\n- 2026-09-01 **t-first-thing** First thing — old\n") + self.ok("migrate", "--write") + self.assertIn("- **t-first-thing-2** [P1] (1h): First thing. do it.\n", self.tasks()) diff --git a/tests/test_cloud.py b/tests/test_cloud.py new file mode 100644 index 0000000..e922f9c --- /dev/null +++ b/tests/test_cloud.py @@ -0,0 +1,1174 @@ +import datetime as dt +import io +import json +import os +import shutil +import sys +import tempfile +import unittest +import subprocess +import textwrap +from contextlib import redirect_stderr, redirect_stdout +from pathlib import Path + +HERE = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(HERE)) +import wf_cloud # noqa: E402 +from wflib import cloud as C, usage as U # noqa: E402 + +T0 = dt.datetime(2026, 10, 6, 12, 0, tzinfo=dt.timezone.utc) + + +def led_with(n_running, **kw): + led = {**C.new(), **kw} + for i in range(n_running): + C.add(led, f"t{i}", "p", f"s{i}", "claude-sonnet-5-5", T0) + return led + + +class Pure(unittest.TestCase): + def test_default_budget(self): + self.assertEqual(C.new()["budget"], 240.0) + + def test_balance(self): + # 240 - 10 - 4*2 = 222 + self.assertEqual(C.balance(led_with(2, spent=10.0)), 222.0) + self.assertEqual(C.balance(C.new()), 240.0) + + def test_refuse_low_balance(self): + # 20 - 17 - 0 = 3 < 4 + self.assertIn("balance", C.refusal(led_with(0, budget=20.0, spent=17.0))) + # exactly 4 left: ok + self.assertIsNone(C.refusal(led_with(0, budget=20.0, spent=16.0))) + + def test_refuse_max_parallel(self): + self.assertIn("max_parallel", C.refusal(led_with(3))) + self.assertIsNone(C.refusal(led_with(2))) + + def test_balance_counts_reserve_toward_refusal(self): + # 10 - 0 - 4*2 = 2 < 4 + self.assertIn("balance", C.refusal(led_with(2, budget=10.0))) + + def test_set_balance(self): + led = led_with(0, spent=50.0) + row = C.set_balance(led, 200.0, T0) + self.assertEqual(led["spent"], 40.0) + self.assertEqual((row["usd_source"], row["usd"]), ("owner", -10.0)) + + def test_charge_from_usage(self): + led = led_with(1) + u = U.Usage(turns=1, inp=1_000_000, out=1_000_000) # sonnet-5-5: 2 + 10 = 12 ; x1.15 = 13.8 + e = C.end(led, "s0", "done", u) + self.assertAlmostEqual(e["usd"], 13.8) + self.assertEqual(e["usd_source"], "self") + self.assertAlmostEqual(led["spent"], 13.8) + self.assertEqual(C.running(led), 0) + + def test_charge_missing_usage_is_reserve(self): + led = led_with(1) + e = C.end(led, "s0", "handback", None) + self.assertEqual((e["usd"], e["usd_source"]), (4.0, "est")) + + def test_end_twice_refused(self): + led = led_with(1) + C.end(led, "s0", "done", None) + with self.assertRaises(C.CloudError): + C.end(led, "s0", "done", None) + + def test_expire_lost_after_24h(self): + led = led_with(2) + led["entries"][1]["sent"] = (T0 + dt.timedelta(hours=20)).isoformat() + lost = C.expire(led, T0 + dt.timedelta(hours=25)) + self.assertEqual([e["sid"] for e in lost], ["s0"]) + self.assertEqual(led["entries"][0]["state"], "lost") + self.assertEqual(led["spent"], 4.0) + self.assertEqual(C.running(led), 1) + + def test_roundtrip_and_bad(self): + led = led_with(1) + self.assertEqual(C.loads(C.dumps(led)), led) + with self.assertRaises(C.CloudError): + C.loads("{nope") + + +class ArchivePure(unittest.TestCase): + CREDS = '{"claudeAiOauth": {"accessToken": "tok-1", "refreshToken": "r"}, "other": 1}' + + def test_request_url_and_headers(self): + url, h = C.archive_request("session_01AbCdEf", self.CREDS, "2.1.291") + self.assertEqual(url, "https://api.anthropic.com/v1/code/sessions/session_01AbCdEf/archive") + self.assertEqual(h, {"Authorization": "Bearer tok-1", "Content-Type": "application/json", + "anthropic-version": "2023-06-01", "User-Agent": "claude-code/2.1.291"}) + + def test_trusted_device_token_sent(self): + creds = '{"claudeAiOauth": {"accessToken": "tok-1"}, "trustedDeviceToken": "dev-9"}' + _, h = C.archive_request("session_01AbCdEf", creds, "2.1.291") + self.assertEqual(h["X-Trusted-Device-Token"], "dev-9") + + def test_no_token_or_bad_sid(self): + for creds in ("", "{}", '{"claudeAiOauth": {}}', "not json"): + with self.assertRaisesRegex(C.CloudError, "no claude.ai login token"): + C.archive_request("session_01AbCdEf", creds, "1") + for sid in ("pending:x", "-", "session_01/../x"): + with self.assertRaisesRegex(C.CloudError, "not a cloud session id"): + C.archive_request(sid, self.CREDS, "1") + + def test_problem(self): + self.assertIsNone(C.archive_problem(200, "")) + self.assertIsNone(C.archive_problem(409, "already")) + self.assertEqual(C.archive_problem(401, "x"), "HTTP 401 (login expired? run claude once)") + self.assertEqual(C.archive_problem(404, "no\nsuch " + "y" * 200), "HTTP 404: no such " + "y" * 102) + + def test_unarchived(self): + led = led_with(3) + C.end(led, "s0", "done") + C.end(led, "s1", "lost") + led["entries"][1]["archived"] = True + for e in led["entries"]: + e["sid"] = "session_" + e["sid"] + C.set_balance(led, 200, T0) + self.assertEqual(C.unarchived(led), ["session_s0"]) + + +class Cli(unittest.TestCase): + def run_wf(self, st, *argv): + out = io.StringIO() + with redirect_stdout(out): + code = wf_cloud.main(list(argv), state=st, now=T0) + return code, out.getvalue() + + def test_ledger_flow(self): + with tempfile.TemporaryDirectory() as d: + st = Path(d) + code, out = self.run_wf(st, "ledger") + self.assertEqual(code, 0) + self.assertIn("budget $240.00 spent $0.00 balance $240.00 running 0/3", out) + _, out = self.run_wf(st, "ledger", "--budget", "100", "--set-balance", "70") + self.assertIn("budget $100.00 spent $30.00 balance $70.00", out) + _, out = self.run_wf(st, "ledger") # persisted + self.assertIn("spent $30.00", out) + + +# Synthetic, same shape as a real `claude --cloud` run in a pty (escape codes, CRLF, title line). +SEND_OUT = ("\x1b[?2004l\x1b]3008;start=x;type=command\x1b\\\x1b7\x1b[r\x1b8\x1b[?25h\x1b[>4;2m\x1b[c" + "Created cloud session: Some title\r\n" + "View: https://claude.ai/code/session_01AbCdEfGhIjKlMnOpQrStUv?from=cli&m=0\r\n" + "Resume with: claude --teleport session_01AbCdEfGhIjKlMnOpQrStUv\r\n") + +FAKE = r"""#!/usr/bin/env python3 +import os, sys, time, tty +mode, st = os.environ.get("FAKE_MODE", ""), os.environ["FAKE_STATE"] +def say(s): sys.stdout.write(s); sys.stdout.flush() +def line(): + return sys.stdin.readline().strip() +n = int(open(st).read()) if os.path.exists(st) else 0 +open(st, "w").write(str(n + 1)) +if "trust" in mode and n == 0: + say("\x1b[2J Do\x1b[4Gyou trust the files in this folder?\r\n 1. No, exit\r\n 2. Yes, proceed\r\n") + if line() != "\x1b[B": + say("refused\r\n"); sys.exit(1) +if sys.argv[1] == "--cloud": + if os.environ.get("FAKE_ARGS"): + open(os.environ["FAKE_ARGS"], "w").write(os.getcwd() + "\n" + sys.argv[2]) + if "nosid" in mode: + say("Error: something\r\nline2\r\n"); sys.exit(1) + say(%r); sys.exit(0) +assert sys.argv[1] == "--teleport", sys.argv +say("\x1b[3G◯\x1b[5GChecking\x1b[14Gout\x1b[18Gbranch\r\n") +if "noresume" in mode or ("flaky" in mode and n == 0): + sys.exit(0) +say("●\x1b[3GSession\x1b[11Gresumed\r\n❯ ") +while True: + cmd = line() + if cmd.startswith("/export "): + if line() == "": # 2nd Enter (autocomplete) needed + open(cmd.split(" ", 1)[1], "w").write("● WF-RESULT done\n") + say("Conversation exported\r\n") + elif cmd == "/exit": + sys.exit(0) +""" % SEND_OUT + + +class Parse(unittest.TestCase): + def test_sid_from_send_output(self): + self.assertEqual(C.parse_sid(SEND_OUT), "session_01AbCdEfGhIjKlMnOpQrStUv") + + def test_sid_from_view_only(self): + self.assertEqual(C.parse_sid("View: https://claude.ai/code/session_01Zz9Zz9Zz9Zz9?from=cli\n"), + "session_01Zz9Zz9Zz9Zz9") + + def test_sid_absent(self): + self.assertIsNone(C.parse_sid("Error: not logged in\nclaude --teleport\n")) + + def test_strip_ansi_gaps(self): + # CSI n G (column) and CSI n C (forward) are word gaps in the TUI + self.assertEqual(C.strip_ansi("\x1b[39m●\x1b[3GSession\x1b[11Gresumed\r\n"), "● Session resumed\n") + self.assertEqual(C.strip_ansi("a\x1b[2Cb"), "a b") + + def test_last_lines(self): + self.assertEqual(C.last_lines("a\r\n\r\nb\nc\n", 2), ["b", "c"]) + + def test_trust_keeps_other_keys(self): + out = json.loads(C.trust('{"x": 1, "projects": {"/a": {"k": 2}}}', "/b")) + self.assertEqual(out, {"x": 1, "projects": {"/a": {"k": 2}, "/b": {"hasTrustDialogAccepted": True}}}) + out = json.loads(C.trust('{"projects": {"/a": {"k": 2}}}', "/a")) + self.assertEqual(out["projects"]["/a"], {"k": 2, "hasTrustDialogAccepted": True}) + self.assertEqual(json.loads(C.trust("", "/c")), {"projects": {"/c": {"hasTrustDialogAccepted": True}}}) + with self.assertRaises(C.CloudError): + C.trust("[1]", "/c") + + +class PtyDriver(unittest.TestCase): + """send/export against a fake claude on a real pty.""" + + def setUp(self): + self.d = Path(tempfile.mkdtemp()) + fake = self.d / "claude" + fake.write_text(FAKE) + fake.chmod(0o755) + self.repo = self.d / "repo" + self.repo.mkdir() + self.cj = self.d / "claude.json" + self.cj.write_text('{"keep": true}') + self.env = {"WF_CLAUDE_JSON": str(self.cj), "FAKE_STATE": str(self.d / "n")} + self.old = (wf_cloud.CLAUDE, wf_cloud.SETTLE, dict(os.environ)) + wf_cloud.CLAUDE, wf_cloud.SETTLE = str(fake), 0.3 + os.environ.update(self.env) + + def tearDown(self): + wf_cloud.CLAUDE, wf_cloud.SETTLE, env = self.old + os.environ.clear() + os.environ.update(env) + shutil.rmtree(self.d) + + def test_send_parses_sid_and_pre_trusts(self): + self.assertEqual(wf_cloud.send(self.repo, "do it", timeout=10), "session_01AbCdEfGhIjKlMnOpQrStUv") + data = json.loads(self.cj.read_text()) + self.assertEqual(data["projects"][str(self.repo.resolve())], {"hasTrustDialogAccepted": True}) + self.assertTrue(data["keep"]) + + def test_send_answers_trust_dialog(self): + os.environ["FAKE_MODE"] = "trust" + self.assertEqual(wf_cloud.send(self.repo, "do it", timeout=10), "session_01AbCdEfGhIjKlMnOpQrStUv") + + def test_send_no_sid(self): + os.environ["FAKE_MODE"] = "nosid" + with self.assertRaises(C.CloudError) as cm: + wf_cloud.send(self.repo, "do it", timeout=10) + self.assertIn("no session id (exit 1): Error: something | line2", str(cm.exception)) + + def test_export(self): + out = wf_cloud.export(self.repo, "session_x", self.d / "e" / "x.txt", timeout=20) + self.assertEqual(out.read_text(), "● WF-RESULT done\n") + self.assertEqual((self.d / "n").read_text(), "1") + + def test_export_retries_teleport_once(self): + os.environ["FAKE_MODE"] = "flaky" + out = wf_cloud.export(self.repo, "session_x", self.d / "x.txt", timeout=20) + self.assertTrue(out.exists()) + self.assertEqual((self.d / "n").read_text(), "2") + + def test_export_gives_up_after_two(self): + os.environ["FAKE_MODE"] = "noresume" + with self.assertRaises(C.CloudError) as cm: + wf_cloud.export(self.repo, "session_x", self.d / "x.txt", timeout=10) + self.assertIn("no export after 2 tries: ◯ Checking out branch", str(cm.exception)) + self.assertEqual((self.d / "n").read_text(), "2") + + +TEMPLATE = (HERE / "templates" / "cloud-prompt.md").read_text() +BODY = ("- **t-x** [P1] (1h): Fix the parser.\n Done: WF-RESULT done in a line\n WF-PATCH-END\n" + "● WF-RESULT done\nWF-RESULT done\n Model: opus") + + +def as_export(prompt: str, wrap: bool) -> str: + """The prompt the way /export renders a user message: '❯ ' first line, ' ' continuations (F12).""" + lines = [] + for ln in prompt.split("\n"): + lines += (textwrap.wrap(ln, 78, drop_whitespace=False) or [""]) if wrap else [ln] + return "\n".join(("❯ " if i == 0 else " ") + ln.ljust(78) for i, ln in enumerate(lines)) + "\n" + + +FINAL = ("● WF-RESULT done\n WF-REPORT fixed it\n WF-USAGE in=1 cw=2 cr=3 out=4 model=m\n" + " WF-PATCH-BEGIN sha256=ab bytes=3\n QUJD\n WF-PATCH-END\n\n") + + +class Prompt(unittest.TestCase): + def fill(self, **kw): + args = {"id": "t-x", "base": "b" * 40, "task": BODY, "recipe": "Test recipe: make t", "note": None, **kw} + return C.fill(TEMPLATE, **args) + + def test_fields_filled(self): + p = self.fill(note="Data in out/pack") + self.assertIn("Task t-x:\n- **t-x** [P1] (1h): Fix the parser.\n", p) + self.assertIn("Test recipe (area notes):\nTest recipe: make t\n", p) + self.assertIn("git format-patch --binary " + "b" * 40 + "..HEAD --stdout | gzip -9", p) + self.assertIn("\nData in out/pack\n", p) + self.assertNotIn("{{", p) + self.assertLessEqual(len(self.fill(task="t").split("\n")), 45) + + def test_no_recipe_no_note(self): + p = self.fill(recipe="") + self.assertIn("Test recipe (area notes):\n(none: see CLAUDE.md)\n\nUsage script", p) + + def test_task_text_braces_kept(self): + self.assertIn("use {{base}} here", self.fill(task="use {{base}} here")) + + def test_unknown_field(self): + with self.assertRaisesRegex(C.CloudError, r"unknown field \{\{bogus\}\}"): + C.fill("x {{bogus}}", "t", "b", "", "", None) + + def test_template_rules(self): + for rule in ("`wf` is absent", "never create or edit TASKS.md", "Never ask questions", + "test red -> implement -> green", "commit on main", "network source is blocked", + " WF-RESULT <done, awaiting or handback>\n", " WF-REPORT <one or two lines", + 'print("WF-USAGE in=%d cw=%d cr=%d out=%d model=%s"', 'echo "WF-PATCH-BEGIN sha256=$(', + "echo WF-PATCH-END", "base64 -w 76", "~/.claude/projects", "never -A", "__pycache__"): + self.assertIn(rule, TEMPLATE) + # live run 2026-10-06: '[WF-RESULT R]' placeholders made the session drop every key word + self.assertNotIn("[WF-", TEMPLATE) + self.assertIn("key word", TEMPLATE) + + def test_parser_never_matches_prompt(self): + p = self.fill() + for wrap in (False, True): + self.assertIsNone(C.final_message(as_export(p, wrap)), wrap) + + def test_parser_takes_final_message_after_prompt(self): + exp = as_export(self.fill(), True) + "\n● Working on it.\n Ran 3 shell commands\n\n" + FINAL + self.assertEqual(C.final_message(exp), ["WF-RESULT done", "WF-REPORT fixed it", + "WF-USAGE in=1 cw=2 cr=3 out=4 model=m", + "WF-PATCH-BEGIN sha256=ab bytes=3", "QUJD", "WF-PATCH-END"]) + + def test_parser_last_result_wins_and_stops_at_next_message(self): + exp = "● WF-RESULT handback\n WF-REPORT old\n\n● WF-RESULT awaiting\n WF-REPORT q?\n\n❯ more\n" + self.assertEqual(C.final_message(exp), ["WF-RESULT awaiting", "WF-REPORT q?"]) + + def test_include_problem(self): + for bad in ("/abs", "../x", "a/../../b", "", ".git", ".git/config"): + self.assertIsNotNone(C.include_problem(bad), bad) + for ok in ("out/pack", "data.bin", "out/a/b/"): + self.assertIsNone(C.include_problem(ok), ok) + + def test_size_refusal(self): + self.assertIsNone(C.size_refusal(90_000_000)) + self.assertEqual(C.size_refusal(91_400_000), "snapshot 91 MB > 90 MB; trim cloud_include") + + +from test_claims import FOUR, git # noqa: E402 +from test_setup import GitCli # noqa: E402 + +AREAS = "# Demo\n\n## Areas\n\n### app\n- Test recipe: RECIPE-MARK python3 -m unittest\n- Paths: src/app.py\n" + + + +def stub_archive(tc, d: Path, status=200): + """No network: wf_cloud.post records (url, headers) and answers `status`; fake credentials file.""" + creds = d / "credentials.json" + creds.write_text('{"claudeAiOauth": {"accessToken": "tok-1"}}') + tc.posts, tc.answer = [], status + saved = (wf_cloud.post, wf_cloud.cli_version, os.environ.get("WF_CLAUDE_CREDENTIALS")) + + def post(url, headers, timeout=10): + tc.posts.append((url, headers)) + if isinstance(tc.answer, Exception): + raise tc.answer + return tc.answer, "body" + + def restore(): + wf_cloud.post, wf_cloud.cli_version, env = saved + os.environ.pop("WF_CLAUDE_CREDENTIALS", None) + if env is not None: + os.environ["WF_CLAUDE_CREDENTIALS"] = env + wf_cloud.post, wf_cloud.cli_version = post, lambda: "9.9.9" + os.environ["WF_CLAUDE_CREDENTIALS"] = str(creds) + tc.addCleanup(restore) + + +class CloudProject(GitCli): + """A cloud-opted git project + fake claude; wf cloud in-process.""" + tasks_text = FOUR.replace("Four.", "Four: fix src/app.py.") + toml = GitCli.toml + 'cloud = true\ncloud_include = ["out/pack"]\ncloud_note = "NOTE-MARK data in out/pack"\n' + + def setUp(self): + super().setUp() + (self.root / "src").mkdir() + (self.root / "src" / "app.py").write_text("v1\n") + (self.root / "CLAUDE.md").write_text(AREAS) + (self.root / "out").mkdir() + (self.root / "out" / "tracked.txt").write_text("tracked under out/\n") + git(self.root, "add", "-A") + git(self.root, "add", "-f", "out/tracked.txt") + git(self.root, "commit", "-qm", "app") + (self.root / "src" / "app.py").write_text("dirty\n") # working tree: not in the snapshot + (self.root / "out" / "pack").mkdir(parents=True) + (self.root / "out" / "pack" / "data.bin").write_text("pack\n") + (self.root / "out" / "other.txt").write_text("not included\n") + (self.root / ".wf").mkdir(exist_ok=True) + (self.root / ".wf" / "x").write_text("state\n") + self.d = Path(tempfile.mkdtemp()) + fake = self.d / "claude" + fake.write_text(FAKE) + fake.chmod(0o755) + self.st = self.d / "state" + self.old = (wf_cloud.CLAUDE, wf_cloud.CAP, dict(os.environ)) + wf_cloud.CLAUDE = str(fake) + os.environ.update({"WF_CLAUDE_JSON": str(self.d / "claude.json"), "FAKE_STATE": str(self.d / "n"), + "FAKE_ARGS": str(self.d / "args"), "CLAUDE_CODE_MESSAGING_SOCKET": "", + "WF_INBOX": str(self.d / "inbox.md")}) + self.snap = self.root / "out" / "cloud" / "t-four" + self.rec = self.root / ".wf" / "cloud" / "t-four.json" + stub_archive(self, self.d) + + def tearDown(self): + wf_cloud.CLAUDE, wf_cloud.CAP, env = self.old + os.environ.clear() + os.environ.update(env) + shutil.rmtree(self.d) + super().tearDown() + + def send(self, *extra): + out, err = io.StringIO(), io.StringIO() + with redirect_stdout(out), redirect_stderr(err): + code = wf_cloud.main(["send", "t-four", "--project", str(self.root), *extra], state=self.st, now=T0) + return code, out.getvalue(), err.getvalue() + + def ledger(self): + return json.loads((self.st / "cloud.json").read_text()) if (self.st / "cloud.json").exists() else None + + def sh(self, cwd, *args): + return subprocess.run(["git", "-C", str(cwd), *args], capture_output=True, text=True).stdout.strip() + +class Send(CloudProject): + """wf cloud send against the fake claude.""" + + def test_send_snapshot_prompt_record_claim(self): + master = self.sh(self.root, "rev-parse", "master") + code, out, err = self.send() + self.assertEqual((code, err), (0, ""), out) + sid = "session_01AbCdEfGhIjKlMnOpQrStUv" + files = sorted(str(p.relative_to(self.snap)) for p in self.snap.rglob("*") + if p.is_file() and ".git" not in p.relative_to(self.snap).parts) + self.assertEqual(files, [".gitignore", "CLAUDE.md", "DESIGN.md", "docs/plan.md", "out/pack/data.bin", + "src/app.py", "workflow.toml"]) + self.assertEqual((self.snap / "src" / "app.py").read_text(), "v1\n") + self.assertEqual(self.sh(self.snap, "log", "--format=%s"), f"base {master}") + self.assertEqual(self.sh(self.snap, "ls-files").split("\n"), files) # all committed (-f) + self.assertEqual(self.sh(self.snap, "status", "--porcelain"), "") + self.assertEqual(self.sh(self.snap, "branch", "--show-current"), "main") + base = self.sh(self.snap, "rev-parse", "HEAD") + cwd, prompt = (self.d / "args").read_text().split("\n", 1) + self.assertEqual(Path(cwd), self.snap.resolve()) + self.assertIn("Task t-four:\n- **t-four** [P3] (1h): Four: fix src/app.py.\n", prompt) + self.assertIn("RECIPE-MARK", prompt) + self.assertIn("NOTE-MARK data in out/pack", prompt) + self.assertIn(f"format-patch --binary {base}..HEAD", prompt) + rec = json.loads(self.rec.read_text()) + self.assertEqual({k: rec[k] for k in ("id", "sid", "master", "base", "lane", "sent", "model")}, + {"id": "t-four", "sid": sid, "master": master, "base": base, "lane": "slow", + "sent": "2026-10-06T12:00:00+00:00", "model": "claude-opus-5-5"}) + self.assertIn(f"**t-four** [P3] (1h) (in progress: cloud:{sid}): Four", self.tasks()) + led = self.ledger() + self.assertEqual([(e["id"], e["sid"], e["state"]) for e in led["entries"]], [("t-four", sid, "running")]) + self.assertIn(f"sent t-four: {sid}", out) + code, _, err = self.send() # twice: refused + self.assertEqual(code, 1) + self.assertIn(f"t-four already sent ({sid})", err) + + def test_dry_run_sends_nothing(self): + code, out, err = self.send("--dry-run") + self.assertEqual((code, err), (0, "")) + self.assertRegex(out, r"^snapshot 0\.0 MB \(master [0-9a-f]{12}, base [0-9a-f]{12}\)\n\nYou are a remote") + self.assertIn("Task t-four:", out) + self.assertFalse(self.snap.exists()) + self.assertFalse(self.snap.parent.exists()) # out/cloud gone too, out/ kept + self.assertFalse(self.rec.exists()) + self.assertIsNone(self.ledger()) + self.assertFalse((self.d / "n").exists()) # claude never ran + self.assertNotIn("in progress", self.tasks()) + + def test_dry_run_keeps_live_snapshot(self): + self.assertEqual(self.send()[0], 0) + head = self.sh(self.snap, "rev-parse", "HEAD") + code, out, err = self.send("--dry-run") + self.assertEqual((code, err), (0, "")) + self.assertEqual(self.sh(self.snap, "rev-parse", "HEAD"), head) # live session's folder intact + self.assertEqual(sorted(p.name for p in self.snap.parent.iterdir()), ["t-four"]) # dry-run folder dropped + + def test_size_refused(self): + wf_cloud.CAP = 10 + code, _, err = self.send() + self.assertEqual(code, 1) + self.assertIn("MB > 0 MB; trim cloud_include", err) + self.assertFalse(self.snap.exists()) + self.assertIsNone(self.ledger()) + self.assertFalse((self.d / "n").exists()) + + def test_ledger_refused_exit_3(self): + with wf_cloud.locked(self.st) as led: + for i in range(3): + C.add(led, f"t{i}", "p", f"s{i}", "claude-opus-5-5", T0) + code, out, err = self.send() + self.assertEqual(code, 3) + self.assertEqual(err.count("\n"), 1) + self.assertIn("cloud refused: ", err) + self.assertIn("max_parallel", err) + self.assertEqual(len(self.ledger()["entries"]), 3) + self.assertFalse(self.snap.exists()) + self.assertFalse(self.rec.exists()) + self.assertFalse((self.d / "n").exists()) + + def test_send_failure_releases_reserve(self): + os.environ["FAKE_MODE"] = "nosid" + code, _, err = self.send() + self.assertEqual(code, 1) + self.assertIn("no session id", err) + self.assertEqual(self.ledger()["entries"], []) + self.assertFalse(self.snap.exists()) + self.assertFalse(self.rec.exists()) + + def test_not_opted_in(self): + (self.root / "workflow.toml").write_text(GitCli.toml) + code, _, err = self.send() + self.assertEqual(code, 1) + self.assertIn("not opted in", err) + + def test_include_outside_refused(self): + (self.root / "workflow.toml").write_text(GitCli.toml + 'cloud = true\ncloud_include = ["../x"]\n') + code, _, err = self.send("--dry-run") + self.assertEqual(code, 1) + self.assertIn("cloud_include '../x': must be a path inside the project", err) + + +if __name__ == "__main__": + unittest.main() + + +class PullPure(unittest.TestCase): + def result(self, raw: bytes, sha=None, n=None): + import base64, gzip, hashlib + gz = gzip.compress(raw) + b64 = base64.b64encode(gz).decode() + return ["WF-RESULT done", "WF-REPORT fixed the parser,", "tests green", + "WF-USAGE in=1000 cw=2000 cr=3000 out=4000 model=claude-opus-5-5", + f"WF-PATCH-BEGIN sha256={sha or hashlib.sha256(gz).hexdigest()} bytes={n or len(gz)}", + *[b64[i:i + 76] for i in range(0, len(b64), 76)], "WF-PATCH-END"] + + def test_parse_and_decode(self): + r = C.parse_result(self.result(b"diff --git a/x b/x\n")) + self.assertEqual((r.state, r.report, r.model), ("done", "fixed the parser, tests green", "claude-opus-5-5")) + self.assertEqual((r.usage.inp, r.usage.cw, r.usage.cr, r.usage.out), (1000, 2000, 3000, 4000)) + self.assertEqual(C.decode_patch(r), b"diff --git a/x b/x\n") + + def test_header_wrapped_anywhere(self): + lines = self.result(b"abc") + head = lines[4] + lines[4:5] = [head[:30], head[30:61], head[61:]] # sha split over three lines + self.assertEqual(C.decode_patch(C.parse_result(lines)), b"abc") + + def test_mismatches(self): + with self.assertRaisesRegex(C.CloudError, "sha256 mismatch"): + C.decode_patch(C.parse_result(self.result(b"abc", sha="0" * 64))) + with self.assertRaisesRegex(C.CloudError, "bytes .* != 7 announced"): + C.decode_patch(C.parse_result(self.result(b"abc", n=7))) + with self.assertRaisesRegex(C.CloudError, "bad WF-RESULT 'maybe'"): + C.parse_result(["WF-RESULT maybe"]) + with self.assertRaisesRegex(C.CloudError, "without WF-PATCH-END"): + C.parse_result(self.result(b"abc")[:-1]) + self.assertIsNone(C.parse_result(["WF-RESULT handback", "WF-REPORT no"]).usage) + + def keyless(self, raw: bytes) -> str: + """The keys-dropped shape seen live: '● done', report, usage, header split, base64, bare WF-PATCH-END.""" + import base64, gzip, hashlib + gz = gzip.compress(raw) + b64 = base64.b64encode(gz).decode() + msg = ["done", "Added sub and a test.", "usage in=4 cw=5 cr=6 out=7 model=claude-opus-5-5", + f"sha256={hashlib.sha256(gz).hexdigest()}", f"bytes={len(gz)}", + *[b64[i:i + 76] for i in range(0, len(b64), 76)], "WF-PATCH-END"] + return "\n".join(("● " if i == 0 else " ") + ln for i, ln in enumerate(msg)) + "\n\n" + + def test_keyless_final_message(self): + prompt = as_export("do it\nend with\n [WF-PATCH-END]\nWF-PATCH-END", wrap=False) + text = prompt + "● working\n\n" + self.keyless(b"diff --git a/x b/x\n") + "● Session resumed\n" + self.assertIsNone(C.final_message(text)) + lines = C.keyless_message(text) + self.assertEqual((lines[0], lines[-1]), ("done", "WF-PATCH-END")) + r = C.keyless_result(lines) + self.assertEqual((r.state, r.model, r.usage.inp, r.usage.out), ("handback", "claude-opus-5-5", 4, 7)) + self.assertEqual(C.decode_patch(r), b"diff --git a/x b/x\n") + self.assertIsNone(C.keyless_message(prompt)) # prompt only: running + self.assertIsNone(C.keyless_message(prompt + "● still working\n")) + self.assertIsNone(C.keyless_message(text + as_export("redo with the keys", wrap=False))) # redo sent + self.assertIsNotNone(C.keyless_message(text + "❯ \n")) # empty input line + r = C.keyless_result(["done", "no patch here", "WF-PATCH-END"]) + self.assertEqual((r.has_patch, r.usage), (False, None)) + + def test_patch_files_and_problem(self): + patch = ("diff --git a/src/a.py b/src/a.py\n--- a/src/a.py\n+++ b/src/a.py\n" + "diff --git a/old.txt b/new.txt\nrename from old.txt\nrename to new.txt\n") + self.assertEqual(C.patch_files(patch), ["src/a.py", "old.txt", "new.txt"]) + never = ["TASKS.md", "tasks/archive.md", ".wf", "out", "data/pack"] + self.assertIsNone(C.patch_problem(["src/a.py", "outish.txt"], "", never)) + self.assertIn("TASKS.md", C.patch_problem(["TASKS.md"], "", never)) + self.assertIn("data/pack", C.patch_problem(["data/pack/x"], "", never)) + self.assertIn("outside the tree", C.patch_problem(["../x"], "", never)) + self.assertIn("outside the tree", C.patch_problem([".git/hooks/x"], "", never)) + self.assertIn("outside the project folder", C.patch_problem(["other/x"], "proj", never)) + self.assertIsNone(C.patch_problem(["proj/src/a.py"], "proj", never)) + self.assertIn("out", C.patch_problem(["proj/out/x"], "proj", never)) + + +class Pull(CloudProject): + """wf cloud pull --export FIXTURE on a git project with a recorded send.""" + toml = None # set in setUp: no worktree_setup (it writes log.txt), a quick_gate + + def setUp(self): + from test_cli import TOML + self.toml = (TOML + 'cloud = true\ncloud_include = ["out/pack"]\n' + 'quick_gate = ["grep -q fixed src/app.py"]\n') + super().setUp() + os.environ.update({"GIT_AUTHOR_NAME": "t", "GIT_AUTHOR_EMAIL": "t@t", "GIT_COMMITTER_NAME": "t", + "GIT_COMMITTER_EMAIL": "t@t"}) + (self.root / "src" / "app.py").write_text("v1\n") + import wf + from wflib import config + cfg = config.load(self.root) + self.master, self.base, _ = wf_cloud.snapshot(cfg, self.snap) + self.sid = "session_01PullPullPullPull" + with wf_cloud.locked(self.st) as led: + C.add(led, "t-four", str(self.root), self.sid, C.MODEL, T0) + self.rec.parent.mkdir(parents=True, exist_ok=True) + self.rec.write_text(json.dumps({"id": "t-four", "sid": self.sid, "project": str(self.root), "lane": "slow", + "master": self.master, "base": self.base, "folder": str(self.snap), + "sent": T0.isoformat(), "model": C.MODEL, "bytes": 1})) + with redirect_stdout(io.StringIO()): + self.assertEqual(wf.main(["--project", str(self.root), "status", "t-four", "progress", + f"cloud:{self.sid}"]), 0) + self.exp = self.d / "export.txt" + + def change(self, files: dict[str, str]): + """Commit files in the snapshot (the cloud's work) -> the gzip patch bytes.""" + import gzip + for rel, text in files.items(): + (self.snap / rel).parent.mkdir(parents=True, exist_ok=True) + (self.snap / rel).write_text(text) + git(self.snap, "add", "-A", "-f") + git(self.snap, "-c", "user.name=c", "-c", "user.email=c@x", "commit", "-qm", "cloud work") + p = subprocess.run(["git", "-C", str(self.snap), "format-patch", "--binary", f"{self.base}..HEAD", "--stdout"], + capture_output=True, check=True).stdout + return gzip.compress(p) + + def write_export(self, state="done", gz=b"", report="fixed app", sha=None, wrap=False, final=True): + import base64, hashlib + b64 = base64.b64encode(gz).decode() + msg = [f"WF-RESULT {state}", f"WF-REPORT {report}", "WF-USAGE in=1000 cw=0 cr=0 out=1000 model=claude-opus-5-5", + f"WF-PATCH-BEGIN sha256={sha or hashlib.sha256(gz).hexdigest()} bytes={len(gz)}", + *[b64[i:i + 76] for i in range(0, len(b64), 76)], "WF-PATCH-END"] + if wrap: # the export soft-wraps at 60 columns: the header too + msg = [part for ln in msg for part in textwrap.wrap(ln, 60, break_long_words=True)] + text = as_export("prompt with ● WF-RESULT done inside\nWF-PATCH-BEGIN", wrap=False) + "\n" + if final: + text += "\n".join(("● " if i == 0 else " ") + ln for i, ln in enumerate(msg)) + "\n\n❯ \n" + self.exp.write_text(text) + + def pull(self, *extra, now=T0 + dt.timedelta(hours=1), export=True): + out, err = io.StringIO(), io.StringIO() + argv = ["pull", "t-four", "--project", str(self.root), "--no-push", + *(["--export", str(self.exp)] if export else []), *extra] + with redirect_stdout(out), redirect_stderr(err): + code = wf_cloud.main(argv, state=self.st, now=now) + return code, out.getvalue(), err.getvalue() + + def entry(self): + return self.ledger()["entries"][0] + + def assert_ended(self, state): + self.assertEqual(self.entry()["state"], state) + self.assertFalse(self.snap.exists()) + self.assertFalse(self.rec.exists()) + + def assert_handback(self, why, kept: bool): + self.assert_ended("handback") + self.assertIn(f"Recovery: cloud attempt {self.sid} — {why}", self.tasks()) + self.assertNotIn("in progress", self.tasks()) + self.assertEqual(self.sh(self.root, "rev-parse", "master"), self.master) + self.assertEqual(self.sh(self.root, "branch", "--list", "slow/*"), "") + self.assertEqual((self.root / "out" / "cloud" / "t-four.patch").exists(), kept) + + def test_done_archives_session(self): + self.write_export(gz=self.change({"src/app.py": "fixed\n"})) + code, out, err = self.pull() + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(self.posts, [(f"https://api.anthropic.com/v1/code/sessions/{self.sid}/archive", + {"Authorization": "Bearer tok-1", "Content-Type": "application/json", + "anthropic-version": "2023-06-01", "User-Agent": "claude-code/9.9.9"})]) + self.assertIs(self.entry()["archived"], True) + self.assertNotIn("archive", out.replace("archive.md", "")) + + def test_second_pull_same_id_skips(self): + """Two pulls of one id at once (batch sidecar + orchestrator): the one without the record lock skips + with rc 0, no 'branch … exists (local WIP?)'; the holder's pull still ends the task.""" + import fcntl + self.write_export(gz=self.change({"src/app.py": "fixed\n"})) + fd = os.open(self.rec, os.O_RDONLY) + try: + fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB) # the other pull, mid-flight + code, out, err = self.pull() + finally: + os.close(fd) + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(out, "t-four: pulled by another wf cloud pull, skipped\n") + self.assertTrue(self.rec.exists()) + self.assertEqual(self.entry()["state"], "running") + self.assertEqual(self.sh(self.root, "branch", "--list", "slow/*"), "") + code, out, err = self.pull() + self.assertEqual((code, err), (0, ""), out) + self.assert_ended("done") + + def test_pull_claim_record_gone(self): + with wf_cloud.pull_claim(self.d / "gone.json") as mine: + self.assertFalse(mine) + with wf_cloud.pull_claim(self.rec) as mine: + self.assertTrue(mine) + self.rec.unlink() # winner ended it while we waited: re-checked on the next claim + with wf_cloud.pull_claim(self.rec) as mine: + self.assertFalse(mine) + + def archive_fails(self, answer, why): + self.answer = answer + self.write_export(state="handback", report="cannot") + code, out, err = self.pull() + self.assertEqual((code, err), (0, ""), out) + self.assert_ended("handback") + self.assertIn(f"archive {self.sid} failed: {why}; archive by hand (wf cloud archive --ended)", out) + self.assertNotIn("archived", self.entry()) + + def test_archive_http_error_warns_pull_still_ends(self): + self.archive_fails(403, "HTTP 403: body") + + def test_archive_network_error_warns_pull_still_ends(self): + self.archive_fails(C.CloudError("timed out"), "timed out") + + def test_running_not_archived(self): + self.write_export(final=False) + code, out, err = self.pull() + self.assertEqual(code, 4, out + err) + self.assertEqual(self.posts, []) + + def test_archive_command_ended(self): + self.write_export(state="handback", report="cannot") + self.answer = 500 + self.pull() + self.answer = 200 + out = io.StringIO() + with redirect_stdout(out): + code = wf_cloud.main(["archive", "--ended"], state=self.st, now=T0) + self.assertEqual((code, out.getvalue()), (0, f"archived {self.sid}\n")) + self.assertIs(self.entry()["archived"], True) + with redirect_stdout(out := io.StringIO()): + code = wf_cloud.main(["archive", "--ended"], state=self.st, now=T0) + self.assertEqual((code, out.getvalue()), (0, "no ended cloud session left to archive\n")) + + def test_archive_command_one_sid_failure_exit_1(self): + self.answer = 401 + err = io.StringIO() + with redirect_stdout(io.StringIO()), redirect_stderr(err): + code = wf_cloud.main(["archive", "session_01Other"], state=self.st, now=T0) + self.assertEqual(code, 1) + self.assertEqual(err.getvalue(), "wf: archive session_01Other failed: HTTP 401 (login expired? run claude once)\n") + + def test_done_applies_gates_merges(self): + self.write_export(gz=self.change({"src/app.py": "fixed\n", "src/new.py": "new\n"})) + code, out, err = self.pull() + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(self.sh(self.root, "show", "master:src/app.py"), "fixed") + self.assertEqual(self.sh(self.root, "show", "master:src/new.py"), "new") + self.assertIn("cloud work", self.sh(self.root, "log", "--format=%s", "master")) + self.assertNotIn("t-four", self.tasks()) + self.assertIn(f"fixed app (cloud {self.sid})", self.archive()) + self.assertEqual(self.sh(self.root, "status", "--porcelain", "TASKS.md", "tasks"), "") # bookkeeping committed + self.assert_ended("done") + e = self.entry() + self.assertEqual(e["usd_source"], "self") + self.assertAlmostEqual(e["usd"], U.cost(C.MODEL, U.Usage(inp=1000, out=1000)) * C.OVERHEAD) + self.assertFalse((self.root / "out" / "cloud" / "t-four.patch").exists()) + self.assertRegex(out, r"report: commit [0-9a-f]{7,}\n$") + self.assertIn("$ grep -q fixed src/app.py", out) + + def test_wrapped_indented_base64_rejoined(self): + self.write_export(gz=self.change({"src/app.py": "fixed\n"}), wrap=True) + code, out, err = self.pull() + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(self.sh(self.root, "show", "master:src/app.py"), "fixed") + + def test_sha_mismatch_hands_back(self): + self.write_export(gz=self.change({"src/app.py": "fixed\n"}), sha="0" * 64) + code, out, err = self.pull() + self.assertEqual(code, 0, err) + self.assert_handback("patch sha256 mismatch", kept=False) + + def test_gate_red_hands_back_keeps_patch(self): + self.write_export(gz=self.change({"src/app.py": "broken\n"})) + code, out, err = self.pull() + self.assertEqual(code, 0, err) + self.assert_handback("quick_gate 'grep -q fixed src/app.py' red (exit 1)", kept=True) + wt = self.root / ".worktrees" / "slow" + self.assertEqual(self.sh(wt, "status", "--porcelain"), "") + self.assertEqual(self.sh(wt, "branch", "--show-current"), "") + + def test_patch_touching_cloud_include_refused(self): + self.write_export(gz=self.change({"src/app.py": "fixed\n", "out/pack/data.bin": "x\n"})) + code, out, err = self.pull() + self.assertEqual(code, 0, err) + self.assert_handback("patch touches out/pack/data.bin", kept=True) + + def test_patch_touching_tasks_refused(self): + self.write_export(gz=self.change({"src/app.py": "fixed\n", "TASKS.md": "x\n"})) + code, out, err = self.pull() + self.assertEqual(code, 0, err) + self.assert_handback("patch touches TASKS.md", kept=True) + + def test_awaiting_adds_question_and_blocks(self): + self.write_export(state="awaiting", report="Which format: csv or json?") + code, out, err = self.pull() + self.assertEqual(code, 0, err) + self.assert_ended("awaiting") + m = __import__("re").search(r"\*\*(a-[^*]+)\*\*: Which format: csv or json\? \[\[t-four\]\]", self.tasks()) + self.assertTrue(m, self.tasks()) + self.assertIn(f"(blocked: [[{m[1]}]])", self.tasks()) + self.assertEqual(self.sh(self.root, "rev-parse", "master"), self.master) + + def test_cloud_handback(self): + self.write_export(state="handback", report="needs the LAN host") + code, out, err = self.pull() + self.assertEqual(code, 0, err) + self.assert_handback("cloud handback: needs the LAN host", kept=False) + + def test_no_marker_running_exit_4(self): + self.write_export(final=False) + code, out, err = self.pull() + self.assertEqual((code, err), (4, "")) + self.assertIn(f"t-four: running ({self.sid}, sent 1h00m ago)", out) + self.assertTrue(self.rec.exists() and self.snap.exists()) + self.assertEqual(self.entry()["state"], "running") + self.assertIn(f"in progress: cloud:{self.sid}", self.tasks()) + + def test_keyless_final_message_hands_back_keeps_patch(self): + import gzip + self.write_export(final=False) + msg = PullPure().keyless(gzip.decompress(self.change({"src/app.py": "fixed\n"}))) + self.exp.write_text(self.exp.read_text() + msg + "● Session resumed\n") + code, out, err = self.pull() + self.assertEqual(code, 0, err) + self.assert_handback("bad result: no WF-RESULT key; patch out/cloud/t-four.patch", kept=True) + self.assertIn(b"+fixed", (self.root / "out" / "cloud" / "t-four.patch").read_bytes()) + self.assertEqual(self.entry()["usd_source"], "self") + + def test_teleport_export_used_and_removed(self): + self.write_export(final=False) + seen = [] + + def fake_export(folder, sid, out, timeout=300): + seen.append((Path(folder), sid)) + Path(out).parent.mkdir(parents=True, exist_ok=True) + Path(out).write_text(self.exp.read_text()) + return Path(out) + + old, wf_cloud.export = wf_cloud.export, fake_export + try: + code, out, err = self.pull(export=False) + finally: + wf_cloud.export = old + self.assertEqual((code, seen), (4, [(self.snap, self.sid)])) + self.assertFalse((self.root / "out" / "cloud" / "t-four.export.txt").exists()) + + def test_over_24h_lost(self): + self.write_export(final=False) + code, out, err = self.pull(now=T0 + dt.timedelta(hours=25)) + self.assertEqual(code, 0, err) + self.assert_ended("lost") + self.assertEqual((self.entry()["usd"], self.entry()["usd_source"]), (4.0, "est")) + self.assertIn(f"Recovery: cloud attempt {self.sid} — lost: no WF-RESULT after 25h00m", self.tasks()) + self.assertNotIn("in progress", self.tasks()) + + def test_all_and_not_sent(self): + os.environ["FAKE_MODE"] = "noresume" # teleport never resumes: export fails + out, err = io.StringIO(), io.StringIO() + with redirect_stdout(out), redirect_stderr(err): + code = wf_cloud.main(["pull", "--all", "--project", str(self.root)], state=self.st, + now=T0 + dt.timedelta(hours=25)) + self.assertEqual(code, 0, err.getvalue()) # no export, > 24h: lost + self.assert_ended("lost") + code, _, err = self.pull() + self.assertEqual(code, 1) + self.assertIn("t-four is not out in the cloud", err) + + +# ------------------------------------------------------------------ fit (spec §4.5) + +FIT_DOC = """\ +# Tasks — demo + +## Awaiting your decision + +## Pending + +- **t-ok** [P1] (1h): Plain code. Goal. + Done: tests pass +- **t-sonnet** [P1] (1h): Plain code. Goal. + Done: tests pass + Model: sonnet +- **t-slice** [P1] (5h): Big. Goal. + Done: tests pass +- **t-owner** [P1] (1h): Plain. Goal. + Done: tests pass + Sessions: owner +- **t-solo** [P1] (1h): Plain. Goal. + Done: tests pass + Sessions: solo +- **t-nodone** [P1] (1h): Plain. Goal. +- **t-no** [P1] (1h): Plain. Goal. + Done: tests pass + Cloud: no +- **t-wf** [P1] (1h): Plain. Goal. + Steps: run wf set x --model opus then check + Done: tests pass +- **t-res** [P1] (1h): Plain. Goal. + Steps: wf res run the build + Done: tests pass +- **t-gui** [P1] (1h): Plain. Goal. + Steps: click through the GUI + Done: tests pass +- **t-lan** [P1] (1h): Plain. Goal. + Steps: copy to 10.0.0.20 + Done: tests pass +- **t-srv** [P1] (1h): Plain. Goal. + Steps: hit the live server + Done: tests pass +- **t-yes-wf** [P1] (1h): Plain. Goal. + Steps: run wf set x --model opus then check + Done: tests pass + Cloud: yes +- **t-yes-sonnet** [P1] (1h): Plain. Goal. + Done: tests pass + Model: sonnet + Cloud: yes +- **t-yes-owner** [P1] (1h): Plain. Goal. + Done: tests pass + Sessions: owner + Cloud: yes +- **t-yes-haiku** [P1] (<1h): Plain. Goal. + Done: tests pass + Model: haiku + Cloud: yes +- **t-no-opus** [P1] (1h): Plain. Goal. + Done: tests pass + Model: opus + Cloud: no +- **t-yes-slice** [P1] (5h): Big. Goal. + Done: tests pass + Cloud: yes + +## Needs human + +## Deferred +""" + + +class FitTest(unittest.TestCase): + def setUp(self): + from wflib import tasks as TK + self.TK = TK + self.doc = TK.parse(FIT_DOC) + self.cfg = type("Cfg", (), {"cloud": True, "slice_above": "1h"})() + + def fit(self, id): + return C.fit(self.doc.item(id), self.cfg) + + def test_fits(self): + self.assertIsNone(self.fit("t-ok")) + + def test_exclusions(self): + for id, why in [("t-sonnet", "Model sonnet"), ("t-slice", "slice"), ("t-owner", "not runner-ready"), + ("t-solo", "Sessions: solo"), ("t-nodone", "not runner-ready"), ("t-no", "Cloud: no"), + ("t-wf", "`wf` commands"), ("t-res", "wf res"), ("t-gui", "GUI"), ("t-lan", "LAN"), + ("t-srv", "live server")]: + with self.subTest(id): + self.assertIn(why, self.fit(id) or "") + + def test_not_opted_in(self): + self.cfg.cloud = False + self.assertIn("not opted in", self.fit("t-ok")) + + def test_cloud_yes_skips_regex_only_opus_only(self): + self.assertIsNone(self.fit("t-yes-wf")) + self.assertEqual(self.fit("t-yes-sonnet"), "Model sonnet stays local") + self.assertEqual(self.fit("t-yes-haiku"), "Model haiku stays local") + self.assertEqual(self.fit("t-no-opus"), "Cloud: no") + self.assertEqual(self.fit("t-sonnet"), "Model sonnet stays local") # unset: old rule + self.assertIn("not runner-ready", self.fit("t-yes-owner")) + self.assertIn("slice", self.fit("t-yes-slice")) + + def test_cloud_property_and_set(self): + self.assertEqual(self.doc.item("t-no").cloud, "no") + self.assertIsNone(self.doc.item("t-ok").cloud) + self.TK.set_fields(self.doc, "t-ok", cloud="yes") + self.assertEqual(self.doc.item("t-ok").body, [" Done: tests pass", " Cloud: yes"]) + self.TK.set_fields(self.doc, "t-ok", cloud="no", model="sonnet") + self.assertEqual(self.doc.item("t-ok").body, [" Done: tests pass", " Model: sonnet", " Cloud: no"]) + self.TK.set_fields(self.doc, "t-ok", cloud="") + self.assertEqual(self.doc.item("t-ok").body, [" Done: tests pass", " Model: sonnet"]) + with self.assertRaises(self.TK.TaskError): + self.TK.set_fields(self.doc, "t-ok", cloud="maybe") + + +class PickKeyTest(unittest.TestCase): + def test_order(self): + from wflib import tasks as TK + doc = TK.parse("""# Tasks + +## Pending + +- **t-p1-small** [P1] (<1h): A. G. + Done: x +- **t-p1-big** [P1] (1h): A. G. + Done: x +- **t-p2-yes** [P2] (<1h): A. G. + Done: x + Cloud: yes +- **t-p0** [P0] (<1h): A. G. + Done: x + +## Needs human + +## Deferred +""") + items = [doc.item(i) for i in ("t-p1-small", "t-p1-big", "t-p2-yes", "t-p0")] + got = [i.id for i in sorted(items, key=C.pick_key)] + # yes first; then prio; then 1h before <1h + self.assertEqual(got, ["t-p2-yes", "t-p0", "t-p1-big", "t-p1-small"]) + + +class CloudLineCheck(unittest.TestCase): + def test_check(self): + sys.path.insert(0, str(HERE / "tests")) + from test_check import Base + + class B(Base): + def runTest(self): + self.pending("- **t-a** [P1] (1h): A. G.\n Done: x\n Cloud: yes\n") + self.assertEqual(self.run_check(), ([], [])) + self.pending("- **t-a** [P1] (1h): A. G.\n Done: x\n Cloud: maybe\n") + self.assertTrue(any("Cloud 'maybe'" in e for e in self.errors())) + self.pending("- **t-a** [P1] (1h): A. G.\n Done: x\n Cloud: yes\n Cloud: no\n") + self.assertTrue(any("two Cloud lines" in e for e in self.errors())) + r = unittest.TextTestRunner(stream=io.StringIO()).run(B()) + self.assertTrue(r.wasSuccessful(), r.failures + r.errors) + + +ORCH_TASKS = """\ +# Tasks — demo + +## Awaiting your decision + +## Pending + +- **t-fast** [P1] (<1h): Fast sonnet task. + - Done: fast works + - Model: sonnet + +- **t-gui** [P1] (<1h): Check the GUI screenshot. + - Done: gui works + +- **t-hai** [P1] (<1h): Haiku forced to the cloud. + - Done: hai works + - Model: haiku + - Cloud: yes + +- **t-slow** [P2] (1h): Slow opus task. + - Done: slow works + +- **t-yes** [P3] (<1h): Opus task marked for the cloud. + - Done: yes works + - Cloud: yes + +## Needs human + +## Deferred +""" + + +class OrchCloud(CloudProject): + """wf orch pick cloud (subprocess, fake claude via WF_CLAUDE, ledger via XDG_STATE_HOME).""" + tasks_text = ORCH_TASKS + + def setUp(self): + super().setUp() + from test_merge import IDENT + self.env = {**IDENT, "WF_CLAUDE": wf_cloud.CLAUDE, "XDG_STATE_HOME": str(self.d / "xdg")} + + def orch(self, *args): + out = self.ok("orch", *args, env=self.env) + return out + + def put_ledger(self, **kw): + n = kw.pop("running", 0) + led = {**C.new(), **kw} + for i in range(n): + C.add(led, f"t-x{i}", "/x", f"session_x{i}", "opus", T0) + (self.d / "xdg" / "wf").mkdir(parents=True, exist_ok=True) + (self.d / "xdg" / "wf" / "cloud.json").write_text(C.dumps(led)) + + def status(self, id): + return next(l for l in (self.root / "TASKS.md").read_text().splitlines() if f"**{id}**" in l) + + def test_pick_cloud_yes_first_then_opus_then_none(self): + out = self.orch("pick", "cloud") + self.assertIn("pick: t-yes (lane cloud, model opus", out) # Cloud: yes beats P2 1h; haiku never + out = self.orch("pick", "cloud") + self.assertIn("pick: t-slow (lane cloud, model opus", out) + self.assertIn("in progress: cloud:session_01AbCdEfGhIjKlMnOpQrStUv", self.status("t-slow")) + rec = json.loads((self.root / ".wf" / "orch" / "t-slow.json").read_text()) + self.assertEqual((rec["lane"], rec["task_lane"], rec["branch"]), ("cloud", "slow", "slow/t-slow")) + self.assertTrue((self.root / ".wf" / "cloud" / "t-slow.json").exists()) + out = self.orch("pick", "cloud") + self.assertEqual(out.strip().splitlines()[0][:26], "stop lane cloud: none fit ") + led = json.loads((self.d / "xdg" / "wf" / "cloud.json").read_text()) + self.assertEqual([e["id"] for e in led["entries"]], ["t-yes", "t-slow"]) + + def test_local_lane_skips_cloud_claim(self): + self.orch("pick", "cloud", "--id", "t-slow") + out = self.orch("pick", "slow") + self.assertNotIn("t-slow", out.splitlines()[0]) + self.assertIn("pick: t-fast", out) + + def test_stop_ledger(self): + self.put_ledger(budget=3.0) + out = self.orch("pick", "cloud") + self.assertIn("stop lane cloud: ledger (balance $3.00 < reserve $4.00)", out) + self.assertNotIn("in progress", self.status("t-slow")) + + def test_stop_max_parallel(self): + self.put_ledger(max_parallel=2, running=2) + out = self.orch("pick", "cloud") + self.assertIn("stop lane cloud: max parallel (2 running >= max_parallel 2)", out) + self.assertFalse((self.root / ".wf" / "cloud" / "t-slow.json").exists()) + + def test_id_unfit_refused(self): + out = self.orch("pick", "cloud", "--id", "t-gui") + self.assertIn("stop lane cloud: none fit (t-gui: needs a GUI)", out) + + def test_not_opted_in(self): + (self.root / "workflow.toml").write_text(GitCli.toml) + out = self.orch("pick", "cloud") + self.assertIn("stop lane cloud: none fit (project not opted in", out) + + def test_post_handback_commits_and_stops(self): + self.orch("pick", "cloud", "--id", "t-slow") + self.ok("note", "t-slow", "Recovery: cloud attempt x — red") # what wf cloud pull leaves behind + self.ok("status", "t-slow", "clear") + out = self.orch("post", "t-slow", "cloud", "--result", "handback") + self.assertIn("post: t-slow handback", out) + self.assertIn("committed leftover", out) + self.assertIn("stop lane cloud: handback", out) + self.assertIn(" cloud opus t-slow handback ", (self.root / "out" / "wf-orch.log").read_text()) diff --git a/tests/test_config.py b/tests/test_config.py new file mode 100644 index 0000000..e51d2bc --- /dev/null +++ b/tests/test_config.py @@ -0,0 +1,134 @@ +import sys +import tempfile +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import config, lanes as L +from wflib import config as C + + +class ConfigTest(unittest.TestCase): + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.root = Path(self.tmp.name).resolve() + + def tearDown(self): + self.tmp.cleanup() + + def write(self, text): + (self.root / "workflow.toml").write_text(text) + + def test_find_root_from_nested_dir(self): + self.write('format = 1\ntasks = "TASKS.md"\narchive = "a.md"\n') + deep = self.root / "a" / "b" + deep.mkdir(parents=True) + self.assertEqual(C.find_root(deep), self.root) + + def test_find_root_none(self): + with self.assertRaisesRegex(C.ConfigError, "no workflow.toml in .* or above"): + C.find_root(self.root) + + def test_minimal(self): + self.write('format = 1\ntasks = "TASKS.md"\narchive = "tasks/archive.md"\n') + cfg = C.load(self.root) + self.assertEqual((cfg.format, cfg.tasks, cfg.archive), (1, self.root / "TASKS.md", self.root / "tasks/archive.md")) + self.assertEqual((cfg.docs, cfg.verify, cfg.done, cfg.ledgers, cfg.anchors_index), ([], [], [], None, None)) + + def test_full(self): + self.write('format = 1\ntasks = "docs/TASKS.md"\narchive = "docs/DONE.md"\ndocs = ["DESIGN.md", "docs/"]\n' + 'verify = ["make test"]\ndone = ["update CATALOG"]\nledgers = ".superpowers/sdd"\n' + '[anchors]\nindex = "DESIGN.md"\nindex_section = "Subsystems"\nspecs = "docs/specs"\n') + cfg = C.load(self.root) + self.assertEqual(cfg.docs, [self.root / "DESIGN.md", self.root / "docs"]) + self.assertEqual((cfg.verify, cfg.done), (["make test"], ["update CATALOG"])) + self.assertEqual(cfg.ledgers, self.root / ".superpowers/sdd") + self.assertEqual((cfg.anchors_index, cfg.anchors_section, cfg.anchors_specs), + (self.root / "DESIGN.md", "Subsystems", self.root / "docs/specs")) + + def test_cloud_keys(self): + self.write('format = 1\ntasks = "T.md"\narchive = "a.md"\n') + cfg = C.load(self.root) + self.assertEqual((cfg.cloud, cfg.cloud_include, cfg.cloud_note), (False, [], None)) + self.write('format = 1\ntasks = "T.md"\narchive = "a.md"\ncloud = true\ncloud_include = ["out/p"]\n' + 'cloud_note = "data in out/p"\n') + cfg = C.load(self.root) + self.assertEqual((cfg.cloud, cfg.cloud_include, cfg.cloud_note), (True, ["out/p"], "data in out/p")) + self.write('format = 1\ntasks = "T.md"\narchive = "a.md"\ncloud = "yes"\n') + with self.assertRaisesRegex(C.ConfigError, "'cloud' must be true or false"): + C.load(self.root) + + def test_unknown_key(self): + self.write('format = 1\ntasks = "T.md"\narchive = "a.md"\nverfy = []\n') + with self.assertRaisesRegex(C.ConfigError, "workflow.toml: unknown key 'verfy'"): + C.load(self.root) + + def test_unknown_anchors_key(self): + self.write('format = 1\ntasks = "T.md"\narchive = "a.md"\n[anchors]\nindx = "D.md"\n') + with self.assertRaisesRegex(C.ConfigError, r"unknown key 'anchors.indx'"): + C.load(self.root) + + def test_missing_required(self): + self.write('format = 1\narchive = "a.md"\n') + with self.assertRaisesRegex(C.ConfigError, "workflow.toml: missing 'tasks'"): + C.load(self.root) + + def test_wrong_type(self): + self.write('format = 1\ntasks = "T.md"\narchive = "a.md"\nverify = "make"\n') + with self.assertRaisesRegex(C.ConfigError, "'verify' must be a list of strings"): + C.load(self.root) + + def test_bad_toml(self): + self.write('format = \n') + with self.assertRaisesRegex(C.ConfigError, "workflow.toml: "): + C.load(self.root) + + + + + +class LaneConfigTest(unittest.TestCase): + def load(self, extra: str): + d = Path(tempfile.mkdtemp()) + self.addCleanup(__import__("shutil").rmtree, d) + (d / "workflow.toml").write_text('format = 1\ntasks = "TASKS.md"\narchive = "a.md"\n' + extra) + return config.load(d) + + def test_default_lanes(self): + cfg = self.load("") + self.assertEqual(cfg.slice_above, "1h") + self.assertEqual(cfg.lanes, (L.Lane("fast", ("<1h",), "unblock", None, True), + L.Lane("slow", ("1h",), "priority", "fast", False))) + self.assertEqual(cfg.area_stale_commits, 20) + self.assertEqual(cfg.areas_file, cfg.root / "CLAUDE.md") + + def test_custom_lanes_replace_default(self): + cfg = self.load('slice_above = "5h"\n[lanes.quick]\nefforts = ["<1h", "1h"]\norder = "unblock"\n' + '[lanes.long]\nefforts = ["5h"]\nfallback = "quick"\n') + self.assertEqual(cfg.lanes, (L.Lane("quick", ("<1h", "1h"), "unblock", None, True), + L.Lane("long", ("5h",), "priority", "quick", False))) + + def test_lane_errors(self): + bad = { + '[lanes.a]\nefforts = ["<1h", "1h"]\n[lanes.b]\nefforts = ["1h"]\n': "effort '1h' in lanes a and b", + '[lanes.a]\nefforts = ["<1h"]\n': "effort '1h' (≤ slice_above) in no lane", + '[lanes.a]\nefforts = ["<1h"]\nslices = true\n[lanes.b]\nefforts = ["1h"]\nslices = true\n': + "slices = true on more than one lane", + '[lanes.a]\nefforts = ["<1h"]\nfallback = "b"\n[lanes.b]\nefforts = ["1h"]\nfallback = "a"\n': + "fallback cycle: a ↔ b", + '[lanes.a]\nefforts = ["<1h", "1h", "5h"]\n': "lane a: effort '5h' above slice_above '1h'", + '[lanes.a]\nefforts = ["<1h", "1h"]\norder = "fifo"\n': "lane a: order 'fifo' (want unblock, priority)", + '[lanes.all]\nefforts = ["<1h", "1h"]\n': "lane name 'all' is reserved", + '[lanes.A]\nefforts = ["<1h", "1h"]\n': "bad lane name 'A'", + '[lanes.a]\nefforts = ["<1h", "1h"]\nspeed = 1\n': "unknown key 'lanes.a.speed'", + 'slice_above = "2h"\n': "slice_above '2h' (want <1h, 1h, 5h, 10h, 100h)", + 'area_stale_commits = 0\n': "'area_stale_commits' must be a number ≥ 1", + } + for extra, msg in bad.items(): + with self.subTest(msg=msg), self.assertRaises(config.ConfigError) as cm: + self.load(extra) + self.assertIn(msg, str(cm.exception)) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_ctx_hint.py b/tests/test_ctx_hint.py new file mode 100644 index 0000000..7750284 --- /dev/null +++ b/tests/test_ctx_hint.py @@ -0,0 +1,98 @@ +import json +import sys +import unittest +from pathlib import Path + +HERE = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(HERE)) +sys.path.insert(0, str(HERE / "tests")) + +import test_cli # noqa: E402 +from test_usage import entry # noqa: E402 +from wflib import usage # noqa: E402 + +SID = "99999999-2222-3333-4444-555555555555" +USER = json.dumps({"type": "user", "timestamp": "2026-10-05T09:00:00Z", "message": {"role": "user", "content": "hi"}}) + + +def transcript(*sizes): + """One request per size (inp 10, cw 1000, rest cache read), each streamed as 2 entries.""" + lines = [USER] + for n, size in enumerate(sizes): + e = entry(f"m{n}", f"r{n}", "claude-opus-5-5", f"2026-10-05T09:0{n}:01Z", inp=10, cw5=1000, cr=size - 1010) + lines += [e, e, USER] + return "\n".join(lines) + "\n" + + +class ContextTokens(unittest.TestCase): + def test_last_request_prompt_size(self): + self.assertEqual(usage.context_tokens(transcript(50_000, 152_000)), 152_000) + + def test_skips_synthetic_and_sidechain(self): + text = transcript(40_000) + synth = json.loads(entry("s", "rs", "<synthetic>", "2026-10-05T10:00:00Z", inp=1)) + side = json.loads(entry("x", "rx", "claude-opus-5-5", "2026-10-05T10:00:01Z", cr=999_000)) + side["isSidechain"] = True + text += json.dumps(synth) + "\n" + json.dumps(side) + "\n" + self.assertEqual(usage.context_tokens(text), 40_000) + + def test_partial_first_line_and_empty(self): + self.assertEqual(usage.context_tokens('ens": 5}}}\n' + transcript(30_000)), 30_000) + self.assertIsNone(usage.context_tokens(USER + "\n")) + + def test_hint_line(self): + self.assertEqual(usage.ctx_hint(152_000, 100_000), + "context ~152k tokens (> 100k): ask the owner to /clear, then continue " + "(subagent: ignore, this is the main session)") + self.assertIsNone(usage.ctx_hint(99_999, 100_000)) + self.assertIsNone(usage.ctx_hint(None, 100_000)) + self.assertIsNone(usage.ctx_hint(500_000, 0)) + + +class HintCli(test_cli.Cli): + def setUp(self): + super().setUp() + self.home = self.root.parent / "claude-home" + (self.home / "projects" / "-work-demo").mkdir(parents=True) + self.jsonl = self.home / "projects" / "-work-demo" / f"{SID}.jsonl" + + def env(self, sid=SID): + return {"CLAUDE_CONFIG_DIR": str(self.home), "CLAUDE_CODE_SESSION_ID": sid} + + HINT = "context ~152k tokens (> 100k)" + + def test_done_prints_hint_when_big(self): + self.jsonl.write_text(transcript(152_000)) + out = self.ok("done", "t-one", "-m", "ok", env=self.env()) + self.assertIn(self.HINT, out.splitlines()[-1]) + + def test_next_prints_hint_when_big(self): + self.jsonl.write_text(transcript(152_000)) + out = self.ok("next", "--as", "opus", env=self.env()) + self.assertIn(self.HINT, out.splitlines()[-1]) + + def test_silent_when_small_or_missing(self): + self.jsonl.write_text(transcript(60_000)) + self.assertNotIn("context ~", self.ok("next", "--as", "opus", env=self.env())) + self.assertNotIn("context ~", self.ok("next", "--as", "opus", env=self.env(sid=""))) + self.assertNotIn("context ~", self.ok("next", "--as", "opus", env=self.env(sid="nope"))) + + def test_threshold_from_toml(self): + (self.root / "workflow.toml").write_text(self.toml + "ctx_hint = 200000\n") + self.jsonl.write_text(transcript(152_000)) + self.assertNotIn("context ~", self.ok("next", "--as", "opus", env=self.env())) + (self.root / "workflow.toml").write_text(self.toml + "ctx_hint = 0\n") + self.jsonl.write_text(transcript(900_000)) + self.assertNotIn("context ~", self.ok("next", "--as", "opus", env=self.env())) + + def test_bad_threshold(self): + (self.root / "workflow.toml").write_text(self.toml + 'ctx_hint = "big"\n') + self.assertIn("'ctx_hint' must be a number", self.fails("list")) + + def test_reads_only_the_tail(self): + self.jsonl.write_text(transcript(152_000) + (USER + "\n") * 40_000 + transcript(160_000)) + self.assertIn("context ~160k tokens", self.ok("next", "--as", "opus", env=self.env())) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_finish.py b/tests/test_finish.py new file mode 100644 index 0000000..38fceb1 --- /dev/null +++ b/tests/test_finish.py @@ -0,0 +1,210 @@ +import subprocess +import unittest + +from test_cli import TOML, Cli +from test_claims import git +from test_merge import IDENT + + +class FinishTest(Cli): + def setUp(self): + super().setUp() + git(self.root, "init", "-q", "-b", "master") + (self.root / ".gitignore").write_text(".worktrees/\n.wf/\n") + git(self.root, "add", "-A") + git(self.root, "commit", "-qm", "init") + self.wt = self.root / ".worktrees" / "sonnet" + git(self.root, "worktree", "add", "-q", str(self.wt), "-b", "sonnet/t-three") + + def out(self, *args, cwd=None): + return subprocess.run(["git", *args], cwd=cwd or self.root, capture_output=True, text=True).stdout + + def finish(self, *args, cwd=None): + return self.wf("finish", *args, project=False, cwd=cwd or self.wt, env=IDENT) + + def write_code(self): + (self.wt / "code.txt").write_text("x\n") + + def test_report_line_tool_commit(self): + self.write_code() + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", "--no-push", + "--tool-commit", "abc1234") + self.assertEqual(code, 0, err) + self.assertRegex(out, r"report: commit [0-9a-f]+ tool abc1234\n$") + + def test_report_line_no_worktree(self): + (self.root / "code.txt").write_text("x\n") + code, out, err = self.wf("finish", "t-three", "-m", "ok", "--commit", "impl", "code.txt", "--no-push", + project=False, cwd=self.root, env=IDENT) + self.assertEqual(code, 0, err) + sha = self.out("rev-parse", "--short", "HEAD").strip() + self.assertTrue(out.endswith(f"report: commit {sha}\n"), out) + + def test_commit_done_merge_in_one(self): + self.write_code() + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", "--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertIn("done: t-three → tasks/archive.md\n", out) + self.assertIn("committed code.txt\n", out) + self.assertIn("merged sonnet/t-three into master\n", out) + sha = self.out("log", "--format=%h", "-n1", "--grep=^impl$", "master").strip() + self.assertTrue(out.endswith(f"report: commit {sha}\n"), out) + self.assertNotIn("wf merge", out) # no merge-step reminder: finish merged + self.assertEqual(self.out("log", "--format=%s", "master"), "t-three done\nimpl\ninit\n") + self.assertEqual(self.out("status", "--porcelain"), "") + self.assertEqual((self.root / "code.txt").read_text(), "x\n") + self.assertIn("t-three", (self.root / "tasks" / "archive.md").read_text()) + + def test_wip_commits_notes_clears(self): + self.write_code() + self.wf("status", "t-three", "progress", "br", project=False, cwd=self.wt, env=IDENT) + code, out, err = self.wf("wip", "t-three", "-m", "half done, next: tests", "--commit", "wip", "code.txt", + project=False, cwd=self.wt, env=IDENT) + self.assertEqual(code, 0, err) + self.assertIn("committed code.txt", out) + self.assertEqual(self.out("log", "--format=%s", "-n1", cwd=self.wt), "wip\n") + self.assertEqual(self.out("status", "--porcelain", cwd=self.wt, ).count("code.txt"), 0) + code, out, err = self.wf("show", "t-three", project=False, cwd=self.wt) + self.assertIn("half done, next: tests", out) + self.assertNotIn("in progress", out) + + def test_no_commit_bookkeeping_only(self): + git(self.wt, "switch", "-q", "--detach", "master") + code, out, err = self.finish("t-three", "-m", "ok", "--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertTrue(out.endswith("committed TASKS.md tasks/archive.md\n"), out) + self.assertEqual(self.out("log", "--format=%s", "master"), "bookkeeping\ninit\n") + + def test_commit_tasks_only_no_path_lists_once(self): + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "bk", "--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(out.count("TASKS.md"), 1, out) + self.assertIn("committed TASKS.md tasks/archive.md\n", out) + self.assertEqual(self.out("log", "--format=%s", "master"), "t-three done\ninit\n") + + def test_commit_tasks_path_named_lists_once(self): + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "bk", "TASKS.md", "--no-push", cwd=self.root) + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(out.count("TASKS.md"), 1, out) + self.assertEqual(self.out("log", "--format=%s", "master"), "bk\ninit\n") + + def test_dirty_outside_paths_refused_before_done(self): + self.write_code() + (self.wt / "DESIGN.md").write_text("stray\n") + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", "--no-push") + self.assertEqual((code, err), (1, "wf: uncommitted changes outside the --commit paths: DESIGN.md\n")) + self.assertIn("t-three", (self.root / "TASKS.md").read_text()) + self.assertEqual(self.out("log", "--format=%s", "master"), "init\n") + + def test_gate_red_refused_before_done(self): + (self.root / "workflow.toml").write_text(TOML + 'quick_gate = ["exit 3"]\n') + git(self.root, "commit", "-qam", "gate") + git(self.wt, "rebase", "-q", "master") + code, out, err = self.finish("t-three", "-m", "ok", "--no-push") + self.assertEqual(code, 1, out + err) + self.assertIn("quick_gate 'exit 3' red", err) + self.assertIn("wf add -p 0", err) + self.assertIn("done+gate-red", err) + self.assertIn("t-three", (self.root / "TASKS.md").read_text()) + + def test_paths_need_commit_message(self): + code, out, err = self.finish("t-three", "-m", "ok", "code.txt") + self.assertEqual(code, 2, out + err) + self.assertIn("paths need --commit", err) + + def test_main_tree_commits_code_and_bookkeeping(self): + (self.root / "code.txt").write_text("y\n") + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", cwd=self.root) + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(self.out("log", "--format=%s", "master"), "impl\ninit\n") + self.assertEqual(self.out("show", "--stat", "--format=", "master").split("|")[0].strip(), "TASKS.md") + self.assertEqual(self.out("status", "--porcelain"), "") + + def test_main_tree_pushes_home(self): + bare = self.root.parent / "home-finish.git" + subprocess.run(["git", "init", "-q", "--bare", str(bare)], check=True) + git(self.root, "remote", "add", "home", str(bare)) + (self.root / "code.txt").write_text("y\n") + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", cwd=self.root) + self.assertEqual((code, err), (0, ""), out) + self.assertIn("pushed home\n", out) + self.assertEqual(self.out("log", "--format=%s", "master", cwd=bare), "impl\ninit\n") + + def test_main_tree_no_push_flag(self): + bare = self.root.parent / "home-finish2.git" + subprocess.run(["git", "init", "-q", "--bare", str(bare)], check=True) + git(self.root, "remote", "add", "home", str(bare)) + (self.root / "code.txt").write_text("y\n") + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", "--no-push", cwd=self.root) + self.assertEqual((code, err), (0, ""), out) + self.assertNotIn("pushed home", out) + + # books stay with the task: no half-done finish, no main-tree done of a worktree's task + + def test_path_outside_repo_refused_before_done(self): + other = self.root.parent / "other-repo-file.md" + other.write_text("x\n") + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", str(other), "--no-push") + self.assertEqual(code, 1, out + err) + self.assertIn("is outside this repo", err) + self.assertIn("t-three", (self.root / "TASKS.md").read_text()) + self.assertEqual(self.out("status", "--porcelain"), "") + + def test_missing_path_refused_before_done(self): + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "shared/CLAUDE.md", "--no-push") + self.assertEqual(code, 1, out + err) + self.assertIn("does not exist", err) + self.assertIn("t-three", (self.root / "TASKS.md").read_text()) + + def test_unchanged_paths_refused_before_done(self): + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "workflow.toml", "--no-push") + self.assertEqual(code, 1, out + err) + self.assertIn("nothing to commit", err) + self.assertIn("t-three", (self.root / "TASKS.md").read_text()) + + def test_worktree_books_copy_refused(self): + (self.wt / "TASKS.md").write_text((self.wt / "TASKS.md").read_text() + "\n") + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "bk", "TASKS.md", "--no-push") + self.assertEqual(code, 1, out + err) + self.assertIn("copy of the books", err) + self.assertIn("t-three", (self.root / "TASKS.md").read_text()) + + def test_already_done_resumes_commit_and_merge(self): + self.write_code() + code, out, err = self.wf("done", "t-three", "-m", "ok", project=False, cwd=self.wt, env=IDENT) + self.assertEqual(code, 0, err) + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", "--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertIn("t-three already done (tasks/archive.md): resuming", out) + self.assertEqual(self.out("log", "--format=%s", "master"), "t-three done\nimpl\ninit\n") + self.assertEqual(self.out("status", "--porcelain"), "") + self.assertEqual((self.root / "tasks" / "archive.md").read_text().count("**t-three**"), 1) + + def test_main_tree_done_of_worktree_task_refused(self): + self.wf("status", "t-three", "progress", "sonnet/t-three", project=False, cwd=self.wt, env=IDENT) + for cmd in (["done", "t-three", "-m", "ok"], ["finish", "t-three", "-m", "ok", "--no-push"]): + code, out, err = self.wf(*cmd, project=False, cwd=self.root, env=IDENT) + self.assertEqual(code, 1, out + err) + self.assertIn("in progress in worktree", err) + self.assertIn("branch sonnet/t-three", err) + self.assertIn("t-three", (self.root / "TASKS.md").read_text()) + # from the worktree it goes through, books committed on master by the merge + self.write_code() + code, out, err = self.finish("t-three", "-m", "ok", "--commit", "impl", "code.txt", "--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(self.out("log", "--format=%s", "master"), "t-three done\nimpl\ninit\n") + + def test_main_tree_done_after_status_clear(self): + self.wf("status", "t-three", "progress", "sonnet/t-three", project=False, cwd=self.wt, env=IDENT) + self.wf("status", "t-three", "clear", project=False, cwd=self.root, env=IDENT) + code, out, err = self.wf("done", "t-three", "-m", "ok", project=False, cwd=self.root, env=IDENT) + self.assertEqual(code, 0, out + err) + + def test_main_tree_done_branch_without_worktree_ok(self): + self.wf("status", "t-three", "progress", "fast/t-three", project=False, cwd=self.root, env=IDENT) + code, out, err = self.wf("done", "t-three", "-m", "ok", project=False, cwd=self.root, env=IDENT) + self.assertEqual(code, 0, out + err) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_gate.py b/tests/test_gate.py new file mode 100644 index 0000000..7a33c1f --- /dev/null +++ b/tests/test_gate.py @@ -0,0 +1,74 @@ +import sys +import tempfile +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import config as C +from test_claims import git +from test_cli import TOML +from test_setup import GitCli + +GATE = TOML + 'quick_gate = ["echo \\"$WF_MAIN\\" > gate.txt", "echo second >> gate.txt"]\n' + + +class GateConfigTest(unittest.TestCase): + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.root = Path(self.tmp.name).resolve() + + def tearDown(self): + self.tmp.cleanup() + + def load(self, extra): + (self.root / "workflow.toml").write_text('format = 1\ntasks = "T.md"\narchive = "a.md"\n' + extra) + return C.load(self.root) + + def test_default_empty(self): + self.assertEqual(self.load("").quick_gate, []) + + def test_list(self): + self.assertEqual(self.load('quick_gate = ["make quick"]\n').quick_gate, ["make quick"]) + + def test_wrong_type(self): + with self.assertRaisesRegex(C.ConfigError, "'quick_gate' must be a list of strings"): + self.load('quick_gate = "make quick"\n') + + +class GateCliTest(GitCli): + toml = GATE + + def setUp(self): + super().setUp() + self.wt = self.root / ".worktrees" / "opus" + git(self.root, "worktree", "add", "-q", str(self.wt), "-b", "opus/t-three") + + def test_runs_in_worktree_with_wf_main(self): + code, out, err = self.wf("gate", project=False, cwd=self.wt / "docs") + self.assertEqual((code, err), (0, ""), out) + self.assertEqual((self.wt / "gate.txt").read_text(), f"{self.root}\nsecond\n") + self.assertFalse((self.root / "gate.txt").exists()) + self.assertIn("$ echo second >> gate.txt\n", out) + self.assertTrue(out.endswith("quick_gate: 2 commands green\n"), out) + + def test_runs_in_main_tree(self): + code, out, err = self.wf("gate", project=False, cwd=self.root) + self.assertEqual((code, err), (0, ""), out) + self.assertEqual((self.root / "gate.txt").read_text(), f"{self.root}\nsecond\n") + + def test_red_exit_1_rest_skipped(self): + (self.wt / "workflow.toml").write_text(TOML + 'quick_gate = ["exit 4", "echo c > c.txt"]\n') + code, out, err = self.wf("gate", project=False, cwd=self.wt) + self.assertEqual(code, 1, out + err) + self.assertTrue(err.startswith("wf: quick_gate 'exit 4' red (exit 4): fix it before wf done; later commands skipped\n"), err) + self.assertIn("done+gate-red", err) + self.assertFalse((self.wt / "c.txt").exists()) + + def test_none_configured(self): + (self.wt / "workflow.toml").write_text(TOML) + code, out, err = self.wf("gate", project=False, cwd=self.wt) + self.assertEqual((code, out, err), (0, "no quick_gate in workflow.toml: nothing to do\n", "")) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_human_done.py b/tests/test_human_done.py new file mode 100644 index 0000000..e6affee --- /dev/null +++ b/tests/test_human_done.py @@ -0,0 +1,40 @@ +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import tasks as T + + +def ready(done): + doc = T.parse(f"# T\n\n## Pending\n\n- **t-x** [P2] (<1h): X.\n Done: {done}\n") + return doc.sections[0].items[0].runner_ready if hasattr(doc, "sections") else None + + +class HumanDoneTest(unittest.TestCase): + def test_owner_bound_subject(self): + self.assertTrue(ready("owner-bound task with Done not counted")) + + def test_owner_actions(self): + self.assertFalse(ready("owner confirms X")) + self.assertFalse(ready("owner approves X")) + self.assertFalse(ready("report to owner")) + self.assertFalse(ready("confirm it works")) + + def test_word_alone_is_fine(self): + self.assertTrue(ready("owner-facing docs updated; owner-private names scrubbed")) + self.assertTrue(ready("tests confirm the parser handles X")) + self.assertTrue(ready("the owner file is read")) + + def test_more_actions(self): + self.assertFalse(ready("owner says it is fine")) + self.assertFalse(ready("Owner tests it")) + self.assertFalse(ready("confirm with the owner")) + + def test_match_exposed(self): + doc = T.parse("# T\n\n## Pending\n\n- **t-x** [P2] (<1h): X.\n Done: owner approves X\n") + self.assertEqual(doc.sections[0].items[0].human_done_match, "owner approves") + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_lanes.py b/tests/test_lanes.py new file mode 100644 index 0000000..a9bcf93 --- /dev/null +++ b/tests/test_lanes.py @@ -0,0 +1,408 @@ +import json +import os +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import lanes as L +from wflib import tasks as T +from test_cli import Cli + +ITEMS = """\ +## Awaiting your decision + +## Pending + +- **t-a** [P1] (1h): A. + +- **t-b** [P1] (1h): B. + - After: [[t-s]] + +- **t-s** [P1] (1h): S. + Model: sonnet + +- **t-s2** [P2] (1h): S2. + Model: sonnet + - After: [[t-x]] + +- **t-c** [P2] (1h): C. + - After: [[t-s2]] + +## Needs human + +## Deferred +""" +SOCK = "/run/user/1000/cc-socks/1.sock" + + +class LanesTest(unittest.TestCase): + def setUp(self): + self.doc = T.parse(ITEMS) + + def test_counts(self): + self.assertEqual(L.counts(self.doc, set(), L.DEFAULT_LANES, "1h"), + {"fast": (0, 0, 0), "slow": (2, 3, 0)}) + + def test_counts_held(self): + self.assertEqual(L.counts(self.doc, set(), L.DEFAULT_LANES, "1h", {"t-a": "x"}), + {"fast": (0, 0, 0), "slow": (1, 4, 0)}) + + def test_counts_slice_jobs(self): + self.assertEqual(L.counts(T.parse(PICK), {"t-old1"}, L.DEFAULT_LANES, "1h"), + {"fast": (3, 2, 1), "slow": (1, 1, 0)}) + + def test_lanes_block(self): + counts = {"fast": (3, 1, 1), "slow": (0, 2, 0)} + sessions = {"slow": {"socket": SOCK, "alive": True, "model": "opus"}} + self.assertEqual(L.lanes_block(counts, sessions, "fast"), [ + "fast (you): 3 pickable (1 slice) · 1 waiting", + f"slow: 0 pickable · 2 waiting · session uds:{SOCK} (alive, opus)", + ]) + + def test_lanes_block_no_session(self): + self.assertEqual(L.lanes_block({"fast": (2, 0, 2), "slow": (0, 0, 0)}, {}, None), [ + "fast: 2 pickable (2 slices) · 0 waiting · no session → orchestrator or owner: start one (wf next --lane fast)", + "slow: 0 pickable · 0 waiting · no session", + ]) + + def test_cross_waits_by_lane(self): + doc = T.parse(PICK) + self.assertEqual([(m.id, b.id) for m, b in L.cross_waits(doc, {"t-old1"}, "slow", L.DEFAULT_LANES, "1h")], + [("t-w1", "t-blocker")]) + self.assertEqual(L.cross_waits(doc, {"t-old1"}, "fast", L.DEFAULT_LANES, "1h"), []) + + def test_waiting_block(self): + doc = T.parse(PICK) + waits = L.cross_waits(doc, {"t-old1"}, "slow", L.DEFAULT_LANES, "1h") + self.assertEqual(L.waiting_block(waits, {"fast": {"socket": SOCK, "alive": True}}, L.DEFAULT_LANES, "1h"), + [f'- t-w1 waits on t-blocker (fast lane) → message uds:{SOCK}: ' + '"t-blocker blocks my t-w1, please take it"']) + self.assertEqual(L.waiting_block(waits, {}, L.DEFAULT_LANES, "1h"), + ["- t-w1 waits on t-blocker (fast lane) → no fast session: tell the owner"]) + + def test_notify_block_by_lane(self): + doc = T.parse(PICK) + freed = [doc.item("t-w1"), doc.item("t-w2")] + self.assertEqual(L.notify_block(freed, {"slow": {"socket": SOCK, "alive": True}}, L.DEFAULT_LANES, "1h"), [ + "fast work now pickable: t-w2 (no session: tell the owner)", + f"notify slow uds:{SOCK}: now pickable t-w1", + ]) + + def test_solo_done_block_any_session_name(self): + sessions = {"fast": {"socket": SOCK, "alive": True, "pid": 1}, "all": {"socket": "/b", "alive": True, "pid": 2}, + "slow": {"socket": "/c", "alive": False, "pid": 3}} + self.assertEqual(L.solo_done_block(["t-s"], sessions, "2"), + [f"notify fast uds:{SOCK}: solo t-s done, run wf next"]) + + +PICK = """\ +## Awaiting your decision + +## Pending + +- **t-big** [P2] (5h): Big unsliced. + +- **t-p0** [P0] (<1h): Small P0. + +- **t-blocker** [P3] (<1h): Small, two wait on it. + +- **t-w1** [P2] (1h): Waits. + - After: [[t-blocker]] + +- **t-w2** [P3] (<1h): Waits too. + - After: [[t-blocker]] + Model: sonnet + +- **t-mid** [P1] (1h): Mid. + Model: sonnet + +- **t-done-parent** [P1] (10h): Sliced, slices archived. + - Slices: [[t-old1]] + +## Needs human + +## Deferred +""" + + +class PickTest(unittest.TestCase): + def setUp(self): + self.doc = T.parse(PICK) + + def pick(self, lane=None, model=None, archived=frozenset({"t-old1"})): + item, _ = L.pick(self.doc, set(archived), L.DEFAULT_LANES, "1h", lane, model) + return item.id if item else None + + def test_lane_of(self): + got = {i.id: L.lane_of(i, L.DEFAULT_LANES, "1h") for i in self.doc.section("pending").items} + self.assertEqual(got, {"t-big": "fast", "t-p0": "fast", "t-blocker": "fast", "t-w1": "slow", + "t-w2": "fast", "t-mid": "slow", "t-done-parent": "fast"}) + + def test_no_effort_is_fast(self): + item = T.Item(id="t-x", prio=1) + self.assertEqual(L.lane_of(item, L.DEFAULT_LANES, "1h"), "fast") + + def test_fast_unblock_order(self): + # unblock: waited-on first (t-blocker: 2 waiters), then the slice job t-big (0 waiters) vs t-p0 by prio + self.assertEqual([i.id for i in L.ranked(self.doc, {"t-old1"}, L.DEFAULT_LANES, "1h", "fast")], + ["t-blocker", "t-p0", "t-big"]) + + def test_slow_priority_order_then_fallback(self): + self.assertEqual(self.pick("slow"), "t-mid") + self.assertEqual([i.id for i in L.ranked(self.doc, {"t-old1"}, L.DEFAULT_LANES, "1h", "slow")], + ["t-mid", "t-blocker", "t-p0", "t-big"]) + + def test_fallback_when_lane_empty(self): + doc = T.parse(PICK.replace("(1h): Mid.", "(<1h): Mid.")) # slow has only t-w1 (blocked) + item, _ = L.pick(doc, {"t-old1"}, L.DEFAULT_LANES, "1h", "slow") + self.assertEqual(item.id, "t-blocker") + + def test_all_lanes_priority_order(self): + self.assertEqual(self.pick(None), "t-p0") + + def test_model_ceiling(self): + # haiku session: no task has Model haiku; opus (default) and sonnet are above it + self.assertIsNone(self.pick("fast", "haiku")) + self.assertEqual(self.pick("slow", "sonnet"), "t-mid") + self.assertEqual(self.pick("fast", "sonnet"), None) # t-w2 is blocked; others are opus + self.assertEqual(self.pick("fast", "opus"), "t-blocker") + + def test_slice_job(self): + big = self.doc.item("t-big") + self.assertTrue(L.is_slice_job(big, "1h")) + self.assertFalse(L.is_slice_job(self.doc.item("t-p0"), "1h")) + + def test_sliced_parent_not_slice_job(self): + parent = self.doc.item("t-done-parent") + self.assertFalse(L.is_slice_job(parent, "1h")) + self.assertTrue(L.sliced_out(parent, "1h")) + # skipped lists items ranked before the pick: drop the others so the parent is first + doc = T.parse(PICK[:PICK.index("- **t-big**")] + PICK[PICK.index("- **t-done-parent**"):]) + _, skipped = L.pick(doc, {"t-old1"}, L.DEFAULT_LANES, "1h", "fast") + self.assertIn(("t-done-parent", "sliced: all slices done → wf done t-done-parent or add slices"), + [(i.id, why) for i, why in skipped]) + + +PICK_DONE = PICK.replace("Small P0.\n", "Small P0.\n Done: x\n").replace( + "two wait on it.\n", "two wait on it.\n Done: x\n").replace("Big unsliced.\n", "Big unsliced.\n Done: x\n") + + +class LanesCliTest(Cli): + tasks_text = "# Tasks — demo\n\n" + PICK_DONE + + def env(self, name="me"): + sock = self.root / f"{name}.sock" + sock.write_text("") + return {"CLAUDE_CODE_MESSAGING_SOCKET": str(sock), "CLAUDE_PID": str(os.getpid()), "CLAUDE_CODE_SESSION_ID": name} + + def session(self, name): + return json.loads((self.root / ".wf" / "sessions" / f"{name}.json").read_text()) + + def test_next_lane_registers_lane(self): + out = self.ok("next", "--lane", "slow", "--as", "opus", env=self.env()) + self.assertIn("===== Next task =====\n- **t-mid**", out) + s = self.session("slow") + self.assertEqual((s["lane"], s["model"], s["socket"]), ("slow", "opus", str(self.root / "me.sock"))) + self.assertIn("===== Lanes =====\nfast: 3 pickable (1 slice) · 2 waiting · no session → " + "orchestrator or owner: start one (wf next --lane fast)\n" + "slow (you): 1 pickable · 1 waiting\n\n" + "===== Waiting on other lanes =====\n" + "- t-w1 waits on t-blocker (fast lane) → no fast session: tell the owner\n", out) + + def test_next_without_lane_registers_all(self): + out = self.ok("next", "--as", "opus", env=self.env()) + self.assertIn("===== Next task =====\n- **t-p0**", out) + self.assertEqual(self.session("all")["lane"], "all") + self.assertNotIn("Waiting on other lanes", out) + + def test_slice_job_output(self): + (self.root / "TASKS.md").write_text("# Tasks — demo\n\n## Awaiting your decision\n\n## Pending\n\n" + "- **t-big** [P2] (5h): Big unsliced.\n\n## Needs human\n\n## Deferred\n") + out = self.ok("next", "--lane", "fast", "--as", "opus", "--brief") + self.assertEqual(out, "- **t-big** [P2] (5h): Big unsliced.\n\n" + "===== Slice job (no code) =====\n" + "effort 5h > slice_above 1h: split it, don't implement.\n" + ' wf add "<title>. <goal>" -e <1h|1h> --parent t-big --model haiku|sonnet|opus' + " (then wf body: Steps/Done/Ref)\n" + ' wf note t-big "sliced into …" · wf status t-big clear · not wf done' + " (the parent waits on its slices)\n") + + def test_unknown_lane(self): + self.assertEqual(self.fails("next", "--lane", "nope", code=2), "wf: unknown lane 'nope' (fast, slow)\n") + self.assertEqual(self.fails("list", "--lane", "nope", code=2), "wf: unknown lane 'nope' (fast, slow)\n") + + def test_nothing_pickable_text(self): + self.assertEqual(self.fails("next", "--lane", "fast", "--as", "haiku", "--brief"), + "wf: nothing pickable for fast (haiku) in Pending\n") + self.assertEqual(self.fails("next", "--as", "haiku", "--brief").splitlines()[-1], + "wf: nothing pickable for all lanes (haiku) in Pending") + + def test_list_runner_lane_pick_order(self): + out = self.ok("list", "--runner", "--lane", "fast") + self.assertEqual([l.split()[0] for l in out.splitlines()[:-1]], ["t-blocker", "t-p0", "t-big"]) + + def test_list_lane_filter_and_column(self): + out = self.ok("list", "--lane", "slow") + self.assertEqual(out.splitlines()[:2], ["t-w1 P2 1h - opus slow Waits", + "t-mid P1 1h - sonnet slow Mid"]) + + def test_lanes_command(self): + out = self.ok("lanes", "--lane", "fast", "--as", "sonnet", env=self.env()) + self.assertEqual(out, "fast (you): 3 pickable (1 slice) · 2 waiting\n" + "slow: 0 pickable · 1 waiting · 1 not runner-ready (no Done) · no session\n") + self.assertEqual((self.session("fast")["lane"], self.session("fast")["model"]), ("fast", "sonnet")) + self.ok("lanes", "--unregister", env=self.env()) + self.assertFalse((self.root / ".wf" / "sessions" / "fast.json").exists()) + + def test_done_notifies_by_lane(self): + sock = self.root / "s.sock" + sock.write_text("") + d = self.root / ".wf" / "sessions" + d.mkdir(parents=True) + (d / "slow.json").write_text(json.dumps({"lane": "slow", "model": "opus", "socket": str(sock), + "pid": os.getpid()})) + out = self.ok("done", "t-blocker", "-m", "ok", env={"CLAUDE_CODE_MESSAGING_SOCKET": ""}) + self.assertIn(f"notify slow uds:{sock}: now pickable t-w1\n", out) + self.assertNotIn("t-w2", out) # same lane as t-blocker: the doer sees it + + + + +WAIT_NONE = """\ +## Awaiting your decision + +## Pending + +- **t-x** [P1] (1h): X. + - After: [[t-nope]] + +## Needs human + +## Deferred +""" + + +class LanesWaitTest(Cli): + tasks_text = WAIT_NONE + env = {"WF_LANES_POLL": "0.1"} + + def test_pickable_exits_0_at_once(self): + (self.root / "TASKS.md").write_text(DONE_ITEMS) + out = self.ok("lanes", "--wait", "5", env=self.env) + self.assertIn("slow: 2 pickable", out) + + def test_timeout_exits_1(self): + self.assertEqual(self.fails("lanes", "--wait", "1", env=self.env), "") + + def test_added_mid_wait(self): + import threading + import time + + def add(): + time.sleep(0.6) + (self.root / "TASKS.md").write_text(DONE_ITEMS) + th = threading.Thread(target=add) + th.start() + out = self.ok("lanes", "--wait", "10", env=self.env) + th.join() + self.assertIn("slow: 2 pickable", out) + + def test_stop_file_exits_2_at_once(self): + (self.root / "out").mkdir() + (self.root / "out" / "wf-batch.stop").touch() + code, out, err = self.wf("lanes", "--wait", "30", env=self.env) + self.assertEqual((code, err), (2, "")) + self.assertIn("stop requested: out/wf-batch.stop", out) + + def test_no_done_not_pickable(self): + (self.root / "TASKS.md").write_text(NO_DONE) + code, out, err = self.wf("lanes", "--wait", "1", env=self.env) + self.assertEqual((code, err), (1, "")) + self.assertEqual(out, + "fast: 0 pickable · 0 waiting · 2 not runner-ready (no Done) · no session\n" + "slow: 0 pickable · 0 waiting · no session\n") + + def test_lanes_plain_shows_not_ready(self): + (self.root / "TASKS.md").write_text(NO_DONE) + self.assertIn("fast: 0 pickable · 0 waiting · 2 not runner-ready (no Done)", self.ok("lanes")) + + +NO_DONE = """\ +## Awaiting your decision + +## Pending + +- **t-a** [P1] (<1h): A. + +- **t-b** [P1] (<1h): B. + +## Needs human + +## Deferred +""" + +DONE_ITEMS = ITEMS.replace("Model: sonnet\n", "Model: sonnet\n Done: ok.\n").replace("(1h): A.", "(1h): A.\n Done: ok.") \ + .replace("(1h): B.", "(1h): B.\n Done: ok.") + + +class LibRunnerTest(unittest.TestCase): + def test_counts_runner_and_not_ready(self): + doc = T.parse(NO_DONE.replace("A.", "A.\n Done: ok.")) + a = (doc, set(), L.DEFAULT_LANES, "1h") + self.assertEqual(L.counts(*a), {"fast": (2, 0, 0), "slow": (0, 0, 0)}) + self.assertEqual(L.counts(*a, runner=True), {"fast": (1, 0, 0), "slow": (0, 0, 0)}) + self.assertEqual(L.not_ready(*a), {"fast": 1, "slow": 0}) + + def test_not_ready_skips_owner_bound_with_done(self): + doc = T.parse(NO_DONE.replace("A.", "A.\n Done: ok.\n Sessions: owner")) + a = (doc, set(), L.DEFAULT_LANES, "1h") + self.assertEqual(L.not_ready(*a), {"fast": 1, "slow": 0}) + self.assertEqual(len(L.prep_targets(*a)), 1) + + +PREP = """\ +## Pending + +- **t-ok** [P1] (<1h): Has Done. + Done: x works. + +- **t-low** [P3] (<1h): Low, no Done. + +- **t-wait** [P0] (1h): Waits, no Done. + - After: [[t-low]] + +- **t-hi** [P1] (1h): High, no Done. + +- **t-own** [P0] (<1h): Owner, no Done. + Sessions: owner + +- **t-prog** [P0] (<1h) (in progress: w): Running. + +- **t-blk** [P0] (<1h) (blocked: a-q): Blocked. + +- **t-par** [P0] (5h): Parent. + +- **t-par-1** [P2] (<1h): Slice. + Done: y. + +## Needs human + +- **t-h** [P0] (<1h): Human. +""" + + +class PrepTargetsTest(unittest.TestCase): + def ids(self, **kw): + return [i.id for i in L.prep_targets(T.parse(PREP), set(), L.DEFAULT_LANES, "1h", **kw)] + + def test_pickable_first_then_prio(self): + self.assertEqual(self.ids(), ["t-hi", "t-low", "t-wait"]) + + def test_lane_filter(self): + self.assertEqual(self.ids(only={"fast"}), ["t-low"]) + self.assertEqual(self.ids(only={"slow"}), ["t-hi", "t-wait"]) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_ledgers.py b/tests/test_ledgers.py new file mode 100644 index 0000000..1a1add1 --- /dev/null +++ b/tests/test_ledgers.py @@ -0,0 +1,26 @@ +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import ledgers as L + +PLAN = "# P\n**Goal:** g\n### Task 1: A\nx\n### Task 2: B\n### Task 3: C\n" + + +class SummaryTest(unittest.TestCase): + def test_resumes_at_first_task_without_complete_line(self): + ledger = "# SDD ledger — plan: p.md\nTask 1: complete (x)\nTask 2: Ruling: y\nTask 2: complete (z)\n" + self.assertEqual(L.summary(PLAN, ledger), "done 1, 2; resume at Task 3 (task-start PLAN 3)") + + def test_ruling_line_is_not_completion(self): + self.assertEqual(L.summary(PLAN, "Task 1: Ruling: complete rewrite\n"), + "done none; resume at Task 1 (task-start PLAN 1)") + + def test_all_done_reports_last_final_review_line(self): + ledger = "Task 1: complete\nTask 2: complete\nTask 3: complete\nFinal review: running\nFinal review: clean\n" + self.assertEqual(L.summary(PLAN, ledger), "done 1, 2, 3; all tasks done; Final review: clean") + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_lock.py b/tests/test_lock.py new file mode 100644 index 0000000..15d5e88 --- /dev/null +++ b/tests/test_lock.py @@ -0,0 +1,30 @@ +import fcntl +import subprocess +import sys +import time +import unittest + +from test_cli import Cli, WF + + +class LockTest(Cli): + def test_write_waits_for_lock(self): + lock_path = self.root / ".wf" / "lock" + lock_path.parent.mkdir(exist_ok=True) + with open(lock_path, "w") as held: + fcntl.flock(held, fcntl.LOCK_EX) + p = subprocess.Popen([sys.executable, str(WF), "--project", str(self.root), "note", "t-three", "first"], + stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) + time.sleep(0.5) + self.assertIsNone(p.poll(), "write ran while the lock was held") + self.ok("show", "t-three") # read commands never wait + out, err = p.communicate(timeout=10) + self.assertEqual((p.returncode, err), (0, "")) + self.ok("note", "t-three", "second") + body = self.item("t-three") + self.assertIn("first", body) + self.assertIn("second", body) + + def test_lock_file_is_git_ignored_folder(self): + self.ok("note", "t-three", "x") + self.assertEqual((self.root / ".wf" / ".gitignore").read_text(), "*\n") diff --git a/tests/test_merge.py b/tests/test_merge.py new file mode 100644 index 0000000..cb14a20 --- /dev/null +++ b/tests/test_merge.py @@ -0,0 +1,159 @@ +import fcntl +import subprocess +import sys +import time +import unittest + +from test_cli import Cli, WF +from test_claims import git + +IDENT = {"GIT_AUTHOR_NAME": "t", "GIT_AUTHOR_EMAIL": "t@t", "GIT_COMMITTER_NAME": "t", "GIT_COMMITTER_EMAIL": "t@t", + "CLAUDE_CODE_MESSAGING_SOCKET": ""} + + +class MergeTest(Cli): + def setUp(self): + super().setUp() + git(self.root, "init", "-q", "-b", "master") + (self.root / ".gitignore").write_text(".worktrees/\n.wf/\n") + git(self.root, "add", "-A") + git(self.root, "commit", "-qm", "init") + self.wt = self.root / ".worktrees" / "sonnet" + git(self.root, "worktree", "add", "-q", str(self.wt), "-b", "sonnet/t-three") + + def merge(self, *args): + return self.wf("merge", *args, project=False, cwd=self.wt, env=IDENT) + + def out(self, *args, cwd=None): + return subprocess.run(["git", *args], cwd=cwd or self.root, capture_output=True, text=True).stdout + + def commit_code(self): + (self.wt / "code.txt").write_text("x\n") + git(self.wt, "add", "code.txt") + git(self.wt, "commit", "-qm", "code") + + def test_merge_ff_commits_bookkeeping_and_detaches(self): + self.commit_code() + self.wf("done", "t-three", "-m", "ok", project=False, cwd=self.wt, env=IDENT) + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(out, "rebased onto master\nfast-forwarded master\n" + "committed TASKS.md tasks/archive.md\nmerged sonnet/t-three into master\n") + self.assertEqual(self.out("log", "--format=%s", "master"), "t-three done\ncode\ninit\n") + self.assertEqual(self.out("branch", "--list", "sonnet/t-three"), "") + self.assertEqual(self.out("status", "--porcelain"), "") + self.assertEqual((self.root / "code.txt").read_text(), "x\n") + + def test_merge_commits_notes_and_followups_after_done(self): + self.commit_code() + self.wf("done", "t-three", "-m", "ok", project=False, cwd=self.wt, env=IDENT) + r = self.wf("add", "-p", "2", "-e", "1h", "follow up", project=False, cwd=self.wt, env=IDENT); self.assertEqual(r[0], 0, r) + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(self.out("status", "--porcelain"), "") + self.assertIn("follow up", self.out("show", "master:TASKS.md")) + + def test_merge_rebases_onto_moved_master(self): + self.commit_code() + (self.root / "other.txt").write_text("o\n") + git(self.root, "add", "other.txt") + git(self.root, "commit", "-qm", "other") + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(self.out("log", "--format=%s", "master"), "code\nother\ninit\n") + + def test_merge_conflict_aborts_and_keeps_branch(self): + (self.root / "DESIGN.md").write_text("main side\n") + git(self.root, "commit", "-qam", "main edit") + (self.wt / "DESIGN.md").write_text("branch side\n") + git(self.wt, "commit", "-qam", "branch edit") + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (1, "wf: rebase onto master conflicts: git rebase master, resolve, verify, " + "then wf merge again\n")) + self.assertEqual((self.wt / "DESIGN.md").read_text(), "branch side\n") + self.assertEqual(self.out("status", "--porcelain", cwd=self.wt), "") + self.assertEqual(self.out("branch", "--show-current", cwd=self.wt), "sonnet/t-three\n") + + def test_merge_refuses_dirty_worktree(self): + (self.wt / "DESIGN.md").write_text("dirty\n") + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (1, "wf: worktree has uncommitted changes: commit them first\n")) + + def test_merge_outside_worktree_fails(self): + self.assertEqual(self.fails("merge"), "wf: merge runs inside a linked git worktree (lane worktree)\n") + + def test_merge_detached_fails(self): + git(self.wt, "switch", "-q", "--detach", "master") + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (1, "wf: worktree is on a detached HEAD: nothing to merge\n")) + + def test_merge_detached_commits_leftover_bookkeeping(self): + self.commit_code() + self.wf("done", "t-three", "-m", "ok", project=False, cwd=self.wt, env=IDENT) + self.merge("--no-push") + self.wf("add", "-p", "2", "-e", "1h", "late follow up", project=False, cwd=self.wt, env=IDENT) + code, out, err = self.merge("--no-push") + self.assertEqual((code, err, out), (0, "", "committed TASKS.md tasks/archive.md\n")) + self.assertEqual(self.out("status", "--porcelain"), "") + self.assertIn("late follow up", self.out("show", "master:TASKS.md")) + + def test_merge_detached_with_impl_commit_merges_it(self): + git(self.wt, "switch", "-q", "--detach", "master") + (self.root / "other.txt").write_text("o\n") + git(self.root, "add", "other.txt") + git(self.root, "commit", "-qm", "other") + self.commit_code() + r = self.wf("done", "t-three", "-m", "ok", project=False, cwd=self.wt, env=IDENT) + self.assertIn("detached HEAD has 1 commit not in master: wf merge merges it", r[1]) + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (0, ""), out) + self.assertEqual(out, "rebased onto master\nfast-forwarded master\n" + "committed TASKS.md tasks/archive.md\nmerged detached HEAD into master\n") + self.assertEqual(self.out("log", "--format=%s", "master"), "bookkeeping\ncode\nother\ninit\n") + self.assertEqual((self.root / "code.txt").read_text(), "x\n") + self.assertEqual(self.out("rev-parse", "HEAD", cwd=self.wt), self.out("rev-parse", "master")) + self.assertEqual(self.out("status", "--porcelain"), "") + + def test_merge_detached_refuses_dirty_worktree(self): + git(self.wt, "switch", "-q", "--detach", "master") + (self.wt / "code.txt").write_text("x\n") + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (1, "wf: worktree has uncommitted changes: commit them first\n")) + + def test_merge_detached_conflict_aborts_and_keeps_commit(self): + git(self.wt, "switch", "-q", "--detach", "master") + (self.root / "DESIGN.md").write_text("main side\n") + git(self.root, "commit", "-qam", "main edit") + (self.wt / "DESIGN.md").write_text("branch side\n") + git(self.wt, "commit", "-qam", "branch edit") + code, out, err = self.merge("--no-push") + self.assertEqual((code, err), (1, "wf: rebase onto master conflicts: git rebase master, resolve, verify, " + "then wf merge again\n")) + self.assertEqual(self.out("log", "-1", "--format=%s", cwd=self.wt), "branch edit\n") + + def test_merge_pushes_to_home(self): + bare = self.root.parent / "home.git" + subprocess.run(["git", "init", "-q", "--bare", str(bare)], check=True) + git(self.root, "remote", "add", "home", str(bare)) + self.commit_code() + code, out, err = self.merge() + self.assertEqual((code, err), (0, ""), out) + self.assertIn("pushed home\n", out) + self.assertEqual(self.out("log", "--format=%s", "master", cwd=bare), "code\ninit\n") + + def test_merge_waits_for_lock(self): + self.commit_code() + (self.root / ".wf").mkdir(exist_ok=True) + with open(self.root / ".wf" / "lock", "w") as held: + fcntl.flock(held, fcntl.LOCK_EX) + import os + p = subprocess.Popen([sys.executable, str(WF), "merge", "--no-push"], cwd=self.wt, text=True, + stdout=subprocess.PIPE, stderr=subprocess.PIPE, env={**os.environ, **IDENT}) + time.sleep(0.5) + self.assertIsNone(p.poll(), "merge ran while the lock was held") + out, err = p.communicate(timeout=20) + self.assertEqual((p.returncode, err), (0, "")) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_migrate.py b/tests/test_migrate.py new file mode 100644 index 0000000..bd02f92 --- /dev/null +++ b/tests/test_migrate.py @@ -0,0 +1,160 @@ +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import migrate as M +from wflib import tasks as T + +OLD = """\ +# Tasks — demo + +Conventions: CLAUDE.md. + +## Awaiting your decision + +- **Smoke screenshots (need you present)**: windows open hidden, + so screenshots are black. +- Online mod updates (spec step 4) not built. + +## Pending + +1. **[P1] Terrain: cave seams** (Effort: <1h) — close the slit beside the lintel. + - Tests (red first): no edge left open. + - After: Terrain 6/6: cave rendering + sound + Reference: [DESIGN.md#terrain](DESIGN.md#terrain) + +2. N. **[P2] Reputation factions** (Effort: 5h, interactive) (in progress: master; slices 1–2/3 done (see plan)) — make factions friendly. + - Steps: use the table + - nested + - After: Terrain: cave seams + Reference: docs/notes/exe-map.md (ai.c rows), [spec](docs/spec.md#rep). + +3. **[P2] Spirit Powers 3/5: technique keys** (Effort: 1h) — keys 1/2/3. + +4. **[P3] Spirit Powers 4/5: passives. And more** (Effort: 1h) + - After: Spirit Powers 3/5: technique keys + +## Needs human + +Only what a script can't judge. + +1. **[P1] Play-test feel** (Effort: <1h) — last pass 2026-09-13. + - [ ] keyboard works + - [ ] motion smooth + +## Deferred (explicitly out of scope, tracked for later) + +**Final art**: later. +""" + +NEW = """\ +# Tasks — demo + +Conventions: CLAUDE.md. + +## Awaiting your decision + +- **a-smoke-screenshots**: Smoke screenshots (need you present). windows open hidden, + so screenshots are black. + +- **a-online-mod-updates-not-built**: Online mod updates (spec step 4) not built. + +## Pending + +- **t-terrain-cave-seams** [P1] (<1h): Terrain: cave seams. close the slit beside the lintel. + - Tests (red first): no edge left open. + Ref: DESIGN.md#terrain + +- **t-reputation-factions** [P2] (5h, interactive) (in progress: master; slices 1–2/3 done [see plan]): Reputation factions. make factions friendly. + - Steps: use the table + - nested + - After: [[t-terrain-cave-seams]] + Ref: docs/notes/exe-map.md (ai.c rows), docs/spec.md#rep + +- **t-spirit-powers-3** [P2] (1h): Spirit Powers 3/5: technique keys. keys 1/2/3. + +- **t-spirit-powers-4** [P3] (1h): Spirit Powers 4/5: passives, And more. + - After: [[t-spirit-powers-3]] + +## Needs human + +Only what a script can't judge. + +- **t-play-test-feel** [P1] (<1h): Play-test feel. last pass 2026-09-13. + - [ ] keyboard works + - [ ] motion smooth + +## Deferred + +(explicitly out of scope, tracked for later) + +**Final art**: later. +""" + + +class MigrateTest(unittest.TestCase): + def test_converts_all_oddities(self): + new, report = M.migrate(OLD) + self.assertEqual(new, NEW) + self.assertEqual(report.ids, { + "Smoke screenshots (need you present)": "a-smoke-screenshots", + "Online mod updates (spec step 4) not built": "a-online-mod-updates-not-built", + "Terrain: cave seams": "t-terrain-cave-seams", + "Reputation factions": "t-reputation-factions", + "Spirit Powers 3/5: technique keys": "t-spirit-powers-3", + "Spirit Powers 4/5: passives. And more": "t-spirit-powers-4", + "Play-test feel": "t-play-test-feel", + }) + self.assertEqual(report.notes, [ + "t-terrain-cave-seams: dropped 'After: Terrain 6/6: cave rendering + sound' (no open task with that title: done)"]) + + def test_result_parses_clean(self): + doc = T.parse(M.migrate(OLD)[0]) + self.assertEqual([i.error for i in doc.all_items()], [None] * 7) + self.assertEqual(T.render(doc), NEW) + + def test_idempotent(self): + new, report = M.migrate(NEW) + self.assertEqual(new, NEW) + self.assertEqual((report.ids, report.notes), ({}, [])) + + def test_missing_sections_are_added(self): + new, _ = M.migrate("# T\n\n## Pending\n\n1. **[P1] A** (Effort: 1h) — a.\n\n## Notes\n\nprose\n") + self.assertEqual(new, "# T\n\n## Pending\n\n- **t-a** [P1] (1h): A. a.\n\n## Notes\n\nprose\n\n" + "## Awaiting your decision\n\n## Needs human\n\n## Deferred\n") + + def test_same_title_twice_gets_distinct_ids(self): + new, _ = M.migrate("## Pending\n\n1. **[P1] A** (Effort: 1h) — x.\n\n2. **[P1] A** (Effort: 1h) — y.\n") + self.assertIn("- **t-a** [P1] (1h): A. x.", new) + self.assertIn("- **t-a-2** [P1] (1h): A. y.", new) + + def test_ids_avoid_the_archive(self): + new, _ = M.migrate("## Pending\n\n1. **[P1] A** (Effort: 1h) — x.\n", taken={"t-a"}) + self.assertIn("- **t-a-2** [P1] (1h): A. x.", new) + + def test_one_column_deeper_under_a_colon_line_is_a_child(self): + new, _ = M.migrate("## Pending\n\n1. **[P1] A** (Effort: 1h) — x.\n - Seed Qs:\n - first?\n - second?\n" + " - [ ] box\n - [ ] sibling typo\n") + self.assertIn("- **t-a** [P1] (1h): A. x.\n - Seed Qs:\n - first?\n - second?\n" + " - [ ] box\n - [ ] sibling typo\n", new) + + def test_reference_words_that_are_no_path_become_a_note(self): + new, _ = M.migrate("## Pending\n\n1. **[P1] A** (Effort: 1h) — x.\n" + " Reference: docs/a.md (fate rows), skill row notes, src/X.cs, DESIGN.md#terrain\n\n" + "2. **[P1] B** (Effort: 1h) — y.\n Reference: main spec §5\n") + self.assertIn(" Ref: docs/a.md (fate rows; skill row notes), src/X.cs, DESIGN.md#terrain\n", new) + self.assertIn("- **t-b** [P1] (1h): B. y.\n - Reference: main spec §5\n", new) + + def test_reference_note_keeps_the_body_indent(self): + new, _ = M.migrate("## Pending\n\n1. **[P1] A** (Effort: 1h) — x.\n - Steps: s\n Reference: main spec §5\n") + self.assertIn("- **t-a** [P1] (1h): A. x.\n - Steps: s\n - Reference: main spec §5\n", new) + + def test_unreadable_numbered_item_is_reported_and_kept(self): + new, report = M.migrate("## Pending\n\n1. something without the pattern\n body\n") + self.assertIn("1. something without the pattern\n body\n", new) + self.assertEqual(report.notes, ["line 3: numbered item not understood, left as it is: 1. something without the pattern"]) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_model.py b/tests/test_model.py new file mode 100644 index 0000000..b2e0878 --- /dev/null +++ b/tests/test_model.py @@ -0,0 +1,247 @@ +import json +import os +import subprocess +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import tasks as T +from wflib import lanes as L +from test_cli import Cli +from test_check import Base, CLEAN +from wflib import check as K +from wflib import config as C + +LANES = """\ +# Tasks — demo + +## Awaiting your decision + +## Pending + +- **t-one** [P1] (1h): One. + - Steps: a + Model: sonnet + - After: [[t-done]] + +- **t-two** [P2] (1h): Two. + - Model: haiku + +- **t-three** [P2] (<1h): Three. + +## Needs human + +## Deferred +""" + + +class ModelLineTest(unittest.TestCase): + def setUp(self): + self.doc = T.parse(LANES) + + def test_parsed_from_body_default_opus(self): + self.assertEqual([i.model for i in self.doc.section("pending").items], ["sonnet", "haiku", "opus"]) + + def test_extra_words_ignored(self): + item = T.parse_block("- **t-x** [P1] (1h): X.\n - Model: sonnet ok\n") + self.assertEqual(item.model, "sonnet") + + def test_set_replaces_in_place(self): + T.set_fields(self.doc, "t-one", model="opus") + self.assertEqual(self.doc.item("t-one").body, [" - Steps: a", " Model: opus", " - After: [[t-done]]"]) + + def test_set_adds_before_after_and_ref(self): + doc = T.parse(CLEAN) + T.set_fields(doc, "t-one", model="haiku") + self.assertEqual(doc.item("t-one").body, + [" Model: haiku", " - After: [[t-done]]", " Ref: DESIGN.md#terrain, docs/specs/terrain.md"]) + + def test_set_empty_removes(self): + T.set_fields(self.doc, "t-two", model="") + self.assertEqual(self.doc.item("t-two").body, []) + self.assertEqual(self.doc.item("t-two").model, "opus") + + def test_set_bad_value(self): + with self.assertRaisesRegex(T.TaskError, r"model 'gpt' \(want haiku, sonnet, opus\)"): + T.set_fields(self.doc, "t-one", model="gpt") + + def test_note_goes_before_model_line(self): + T.add_note(self.doc, "t-two", "why") + self.assertEqual(self.doc.item("t-two").body, [" - why", " - Model: haiku"]) + + +def doc(items): + return T.parse("## Awaiting your decision\n\n## Pending\n\n" + items + "\n## Needs human\n\n## Deferred\n") + + +class PickOrderTest(unittest.TestCase): + def pick(self, items, model="opus", archived=()): + item, _ = L.pick(doc(items), set(archived), L.DEFAULT_LANES, "1h", None, model) + return item.id if item else None + + def test_priority_inherited_from_waiting_task(self): + self.assertEqual(self.pick("- **t-c** [P1] (1h): C.\n\n- **t-b** [P3] (1h): B.\n\n" + "- **t-a** [P0] (1h): A.\n - After: [[t-b]]\n"), "t-b") + + def test_inherited_transitively(self): + self.assertEqual(self.pick("- **t-c** [P1] (1h): C.\n\n- **t-x** [P3] (1h): X.\n\n" + "- **t-b** [P3] (1h): B.\n - After: [[t-x]]\n\n" + "- **t-a** [P0] (1h): A.\n - After: [[t-b]]\n"), "t-x") + + def test_unblocking_only_low_work_does_not_jump(self): + self.assertEqual(self.pick("- **t-c** [P1] (1h): C.\n\n- **t-b** [P3] (1h): B.\n\n" + "- **t-a** [P3] (1h): A.\n - After: [[t-b]]\n"), "t-c") + + def test_other_lane_waiting_first(self): + self.assertEqual(self.pick("- **t-f** [P2] (1h): F.\n\n- **t-k** [P2] (1h): K.\n\n" + "- **t-l** [P2] (1h): L.\n - After: [[t-k]]\n\n- **t-g** [P2] (1h): G.\n\n" + "- **t-h** [P2] (<1h): H.\n - After: [[t-g]]\n"), "t-g") + + def test_more_waiting_first(self): + self.assertEqual(self.pick("- **t-f** [P2] (1h): F.\n - After: [[t-z]]\n\n- **t-g** [P2] (1h): G.\n\n" + "- **t-i** [P2] (1h): I.\n - After: [[t-g]]\n\n" + "- **t-j** [P2] (1h): J.\n - After: [[t-i]]\n\n" + "- **t-z** [P2] (1h): Z.\n", archived=()), "t-g") + + def test_parent_waits_on_its_slices(self): + self.assertEqual(self.pick("- **t-q** [P1] (1h): Q.\n\n- **t-p** [P0] (5h): P.\n - Slices: [[t-p-1]]\n\n" + "- **t-p-1** [P2] (1h): P1.\n"), "t-p-1") + + def test_lane_filter(self): + items = ("- **t-a** [P0] (1h): A.\n\n- **t-b** [P1] (1h): B.\n Model: sonnet\n\n" + "- **t-c** [P2] (1h): C.\n Model: haiku\n") + self.assertEqual([self.pick(items, m) for m in ("opus", "sonnet", "haiku", None)], + ["t-a", "t-b", "t-c", "t-a"]) + self.assertIsNone(self.pick("- **t-a** [P0] (1h): A.\n", "haiku")) + + def test_skipped_are_ranked_before_pick(self): + d = doc("- **t-a** [P0] (1h) (blocked: [[a-k]]): A.\n\n- **t-s** [P0] (1h): S.\n Model: sonnet\n\n" + "- **t-b** [P1] (1h): B.\n\n- **t-c** [P2] (1h): C.\n - After: [[t-b]]\n") + item, skipped = L.pick(d, set(), L.DEFAULT_LANES, "1h", "slow", "opus") + self.assertEqual((item.id, [(i.id, why) for i, why in skipped]), ("t-s", [("t-a", "blocked: a-k")])) + item, skipped = L.pick(d, set(), L.DEFAULT_LANES, "1h", "slow", "haiku") + self.assertEqual((item, skipped), (None, [])) + + +class ModelCheckTest(Base): + def test_values(self): + self.pending("- **t-a** [P1] (1h): A.\n Model: gpt\n\n" + "- **t-b** [P1] (1h): B.\n - Model: sonnet ok\n\n" + "- **t-c** [P1] (1h): C.\n Model: haiku\n Model: opus\n\n" + "- **t-d** [P1] (1h): D.\n") + errors, warnings = K.check(C.load(self.root)) + self.assertEqual([(p.id, p.message) for p in errors], + [("t-a", "Model 'gpt' (want haiku, sonnet, opus)"), ("t-c", "two Model lines")]) + self.assertEqual([(p.id, p.message) for p in warnings], + [("t-b", "Model line 'sonnet ok': write 'Model: sonnet' (wf set --model)")]) + + +class ModelCliTest(Cli): + tasks_text = LANES + + def test_list_column_and_filter(self): + self.assertEqual(self.ok("list"), + "t-one P1 1h - sonnet slow One\n" + "t-two P2 1h - haiku slow Two\n" + "t-three P2 <1h - opus fast Three\n" + "pending 3 · human 0 · awaiting 0 · deferred 0\n") + self.assertEqual(self.ok("list", "--model", "opus"), + "t-three P2 <1h - opus fast Three\n" + "pending 3 · human 0 · awaiting 0 · deferred 0\n") + + def test_add_and_set(self): + self.assertEqual(self.ok("add", "Four. Goal.", "-p", "3", "-e", "1h", "--model", "haiku", "--ref", "DESIGN.md"), + "- **t-four** [P3] (1h): Four. Goal.\n") + self.assertEqual(self.item("t-four"), "- **t-four** [P3] (1h): Four. Goal.\n Model: haiku\n Ref: DESIGN.md\n") + self.ok("set", "t-four", "--model", "sonnet") + self.assertEqual(self.item("t-four"), "- **t-four** [P3] (1h): Four. Goal.\n Model: sonnet\n Ref: DESIGN.md\n") + self.ok("set", "t-four", "--model", "") + self.assertEqual(self.item("t-four"), "- **t-four** [P3] (1h): Four. Goal.\n Ref: DESIGN.md\n") + + def test_next_model_ceiling(self): + # --as M takes tasks with Model ≤ M; no --lane = all lanes, priority order + self.assertEqual(self.ok("next", "--as", "haiku", "--brief"), "- **t-two** [P2] (1h): Two.\n - Model: haiku\n") + self.assertEqual(self.ok("next", "--as", "opus", "--brief").splitlines()[0], "- **t-one** [P1] (1h): One.") + self.assertEqual(self.ok("next", "--as", "sonnet", "--brief").splitlines()[0], "- **t-one** [P1] (1h): One.") + self.assertEqual(self.ok("next", "--lane", "fast", "--as", "opus", "--brief"), "- **t-three** [P2] (<1h): Three.\n") + self.ok("set", "t-two", "--model", "opus") + self.assertEqual(self.fails("next", "--as", "haiku", "--brief").splitlines()[-1], + "wf: nothing pickable for all lanes (haiku) in Pending") + + def test_next_without_as_is_haiku_with_hint(self): + out = self.ok("next", "--brief") + self.assertEqual(out, "no --as: treated as haiku; pass --as haiku|sonnet|opus\n" + "- **t-two** [P2] (1h): Two.\n - Model: haiku\n") + + def session_env(self, sock): + sock.write_text("") + return {"CLAUDE_CODE_MESSAGING_SOCKET": str(sock), "CLAUDE_PID": str(os.getpid()), + "CLAUDE_CODE_SESSION_ID": "sess-1"} + + def dead_pid(self): + p = subprocess.Popen(["true"]) + p.wait() + return p.pid + + def test_next_registers_session(self): + env = self.session_env(self.root / "me.sock") + self.ok("next", "--as", "opus", "--brief", env=env) + reg = json.loads((self.root / ".wf" / "sessions" / "all.json").read_text()) + self.assertEqual({k: reg[k] for k in ("lane", "model", "socket", "pid", "session")}, + {"lane": "all", "model": "opus", "socket": str(self.root / "me.sock"), "pid": os.getpid(), "session": "sess-1"}) + self.assertEqual((self.root / ".wf" / ".gitignore").read_text(), "*\n") + + def test_no_register_without_as_or_env(self): + none = {"CLAUDE_CODE_MESSAGING_SOCKET": "", "CLAUDE_PID": "", "CLAUDE_CODE_SESSION_ID": ""} + self.ok("next", "--as", "opus", env=none) + self.ok("next", env=self.session_env(self.root / "me.sock")) + self.assertFalse((self.root / ".wf").exists()) + + def register(self, lane, model, sock, pid): + d = self.root / ".wf" / "sessions" + d.mkdir(parents=True, exist_ok=True) + (d / f"{lane}.json").write_text(json.dumps({"lane": lane, "model": model, "socket": str(sock), "pid": pid, + "session": "s", "at": "2026-10-04T10:00"})) + + def test_next_shows_lanes_and_waiting(self): + self.ok("set", "t-three", "--after", "t-one") + self.ok("set", "t-two", "--model", "opus") + sock = self.root / "son.sock" + sock.write_text("") + self.register("slow", "sonnet", sock, os.getpid()) + self.register("fast", "haiku", self.root / "gone.sock", self.dead_pid()) + none = {"CLAUDE_CODE_MESSAGING_SOCKET": ""} + code, out, err = self.wf("next", "--lane", "fast", "--as", "opus", env=none) # fast: no fallback lane + self.assertEqual((code, err), (1, "wf: nothing pickable for fast (opus) in Pending\n")) + self.assertIn("===== Lanes =====\n" + "fast (you): 0 pickable · 1 waiting\n" + f"slow: 2 pickable · 0 waiting · session uds:{sock} (alive, sonnet)\n\n" + "===== Waiting on other lanes =====\n" + f'- t-three waits on t-one (slow lane) → message uds:{sock}: "t-one blocks my t-three, please take it"\n', out) + self.assertEqual(self.ok("lanes", env=none), + "fast: 0 pickable · 1 waiting · no session\n" + f"slow: 0 pickable · 0 waiting · 2 not runner-ready (no Done) · session uds:{sock} (alive, sonnet)\n") + + def test_next_single_lane_prints_no_lanes_block(self): + (self.root / "workflow.toml").write_text(self.toml + '[lanes.one]\nefforts = ["<1h", "1h"]\n') + self.assertNotIn("Lanes", self.ok("next", "--as", "opus", env={"CLAUDE_CODE_MESSAGING_SOCKET": ""})) + + def test_done_notifies_other_lanes(self): + self.ok("add", "Four.", "-p", "3", "-e", "<1h", "--model", "haiku", "--after", "t-three") + self.ok("add", "Five.", "-p", "3", "-e", "1h", "--after", "t-three") + self.ok("add", "Six.", "-p", "3", "-e", "1h", "--model", "sonnet", "--after", "t-three") + sock = self.root / "h.sock" + sock.write_text("") + self.register("slow", "opus", sock, os.getpid()) + out = self.ok("done", "t-three", "-m", "ok") + self.assertIn(f"notify slow uds:{sock}: now pickable t-five, t-six\n", out) + self.assertNotIn("t-four", out) # same lane as t-three (fast) + + def test_bad_model_is_usage_error(self): + self.fails("add", "Four.", "-p", "3", "-e", "1h", "--model", "gpt", code=2) + self.fails("set", "t-one", "--model", "gpt", code=2) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_orch.py b/tests/test_orch.py new file mode 100644 index 0000000..83216b2 --- /dev/null +++ b/tests/test_orch.py @@ -0,0 +1,227 @@ +import json +import os +import subprocess +import unittest + +from test_cli import TOML, Cli +from test_claims import git +from test_merge import IDENT + +TASKS = """\ +# Tasks — demo + +## Awaiting your decision + +## Pending + +- **t-one** [P1] (<1h): One. + - Done: one works + - Model: sonnet + +- **t-two** [P2] (<1h): Two. + - Done: two works + +- **t-raw** [P2] (<1h): Raw, no Done line. + +- **t-big** [P2] (1h): Big. + - Done: big works + +## Needs human + +## Deferred +""" + + +class OrchTest(Cli): + tasks_text = TASKS + toml = TOML.replace('verify = ["make test"]\n', "") + + def setUp(self): + super().setUp() + git(self.root, "init", "-q", "-b", "master") + (self.root / ".gitignore").write_text(".worktrees/\n.wf/\nout/\n") + git(self.root, "add", "-A") + git(self.root, "commit", "-qm", "init") + + def orch(self, *args, code=0): + got, out, err = self.wf("orch", *args, env=IDENT) + self.assertEqual(got, code, out + err) + self.assertEqual(err, "") + return out + + def run_wf(self, cwd, *args): + got, out, err = self.wf(*args, project=False, cwd=cwd, env=IDENT) + self.assertEqual(got, 0, out + err) + return out + + def record(self, id): + return json.loads((self.root / ".wf" / "orch" / f"{id}.json").read_text()) + + def worker_done(self, id="t-one", lane="fast", merge=True): + """What a worker does: wf start, code, wf done/finish (finish merges, done alone leaves the branch).""" + wt = self.root / ".worktrees" / lane + self.run_wf(self.root, "start", id, "--worktree", str(wt), "--branch", f"{lane}/{id}") + (wt / "code.txt").write_text("x\n") + if merge: + self.run_wf(wt, "finish", id, "-m", "ok", "--commit", "impl", "code.txt", "--no-push") + else: + git(wt, "add", "code.txt") + git(wt, "commit", "-qm", "impl") + self.run_wf(wt, "done", id, "-m", "ok") + return wt + + def log(self): + return (self.root / "out" / "wf-orch.log").read_text() + + def test_pick_claims_and_prints_prompt(self): + out = self.orch("pick", "fast") + wt = self.root / ".worktrees" / "fast" + self.assertEqual(out, "pick: t-one (lane fast, model sonnet, effort <1h) · claimed (in progress: worker)\n" + "agent: subagent_type wf-worker, model sonnet, no isolation; prompt:\n" + "Task: t-one Lane: fast Model: sonnet\n" + f"Main tree: {self.root} Worktree: {wt} Branch: fast/t-one\n" + "Final message: the 4 report lines only.\n") + self.assertIn("- **t-one** [P1] (<1h) (in progress: worker): One.", self.tasks()) + self.assertEqual(self.record("t-one")["worktree"], str(wt)) + + def test_second_pick_skips_in_progress_and_busy_worktree(self): + self.orch("pick", "fast") + out = self.orch("pick", "fast") + self.assertIn("pick: t-two (lane fast, model opus", out) + self.assertIn(f"Worktree: {self.root / '.worktrees' / 'fast-2'} Branch: fast/t-two\n", out) + + def test_worktree_on_other_branch_not_reused(self): + git(self.root, "worktree", "add", "-q", str(self.root / ".worktrees" / "fast"), "-b", "fast/t-old") + self.assertIn(".worktrees/fast-2 Branch: fast/t-one", self.orch("pick", "fast")) + + def test_clean_detached_worktree_reused(self): + git(self.root, "worktree", "add", "-q", "--detach", str(self.root / ".worktrees" / "fast"), "master") + self.assertIn(".worktrees/fast Branch: fast/t-one", self.orch("pick", "fast")) + + def test_worktree_with_live_session_not_reused(self): + wt = self.root / ".worktrees" / "fast" + git(self.root, "worktree", "add", "-q", "--detach", str(wt), "master") + folder = self.root.parent / "no-claude" / "sessions" + folder.mkdir(parents=True) + (folder / "1.json").write_text(json.dumps({"pid": os.getpid(), "cwd": str(wt / "docs")})) + self.assertIn(".worktrees/fast-2 Branch: fast/t-one", self.orch("pick", "fast")) + + def test_stop_file_spawns_nothing(self): + (self.root / "out").mkdir() + (self.root / "out" / "wf-batch.stop").write_text("") + self.assertEqual(self.orch("pick", "fast"), "stop: out/wf-batch.stop exists: spawn nothing (let running workers finish)\n") + self.assertNotIn("in progress", self.tasks()) + + def test_none_pickable(self): + self.orch("pick", "fast") + self.orch("pick", "fast") + self.assertTrue(self.orch("pick", "fast").startswith("none: lane fast has no runner-ready task")) + + def test_explicit_id_and_recovery_line(self): + out = self.orch("pick", "slow", "--id", "t-big", "--recovery", "crashed") + self.assertIn("Branch: slow/t-big\nRecovery: crashed\nFinal message", out) + + def test_post_done_logs_and_picks_next(self): + self.orch("pick", "fast") + self.worker_done() + out = self.orch("post", "t-one", "fast", "--result", "done", "--commit", "abc1234", "--duration", "75") + self.assertTrue(out.startswith("post: t-one done\n\npick: t-two"), out) + self.assertRegex(self.log(), r"^\S+ fast sonnet t-one done abc1234 1m15s\n$") + self.assertFalse((self.root / ".wf" / "orch" / "t-one.json").exists()) + + def test_post_no_next_claims_nothing(self): + self.orch("pick", "fast") + self.worker_done() + out = self.orch("post", "t-one", "fast", "--result", "done", "--no-next", "--duration", "5") + self.assertTrue(out.startswith("post: t-one done"), out) + self.assertNotIn("pick:", out) + self.assertEqual([f.name for f in (self.root / ".wf" / "orch").glob("*.json")], []) + + def test_post_done_without_record_takes_model_from_history(self): + self.orch("pick", "fast") + self.worker_done() + (self.root / ".wf" / "orch" / "t-one.json").unlink() + self.orch("post", "t-one", "fast", "--result", "done", "--no-pick", "--duration", "5") + self.assertRegex(self.log(), r"^\S+ fast sonnet t-one done ") + + def test_post_done_merges_unmerged_branch(self): + self.orch("pick", "fast") + wt = self.worker_done(merge=False) + out = self.orch("post", "t-one", "fast", "--result", "done", "--no-pick", "--no-push") + self.assertEqual(out, "post: t-one done\n merged worktree HEAD: merged fast/t-one into master\n") + self.assertEqual(subprocess.run(["git", "-C", str(self.root), "show", "master:code.txt"], + capture_output=True, text=True).stdout, "x\n") + self.assertEqual(subprocess.run(["git", "-C", str(wt), "branch", "--show-current"], + capture_output=True, text=True).stdout, "") + + def test_post_done_without_archive_is_red(self): + self.orch("pick", "fast") + out = self.orch("post", "t-one", "fast", "--result", "done") + self.assertEqual(out, "post: t-one post-check-red\n no archive line for t-one\n" + "stop lane fast: post-check-red → tell the owner\n") + self.assertIn(" post-check-red ", self.log()) + + def backdate_pick(self, id="t-one", secs=125): + f = self.root / ".wf" / "orch" / f"{id}.json" + rec = json.loads(f.read_text()) + rec["at"] -= secs + f.write_text(json.dumps(rec) + "\n") + + def test_post_zero_duration_falls_back_to_pick_time(self): + self.orch("pick", "fast") + self.backdate_pick() + self.worker_done() + self.orch("post", "t-one", "fast", "--result", "done", "--no-pick", "--duration", "0") + self.assertRegex(self.log(), r" t-one done \S+ 2m0[5-9]s\n$") + + def test_post_red_keeps_pick_time_for_the_repost(self): + self.orch("pick", "fast") + self.backdate_pick() + self.orch("post", "t-one", "fast", "--result", "done") + self.assertTrue((self.root / ".wf" / "orch" / "t-one.json").exists()) + self.worker_done() + self.orch("post", "t-one", "fast", "--result", "done", "--no-pick") + self.assertRegex(self.log().splitlines()[-1], r" t-one done \S+ 2m0[5-9]s$") + self.assertFalse((self.root / ".wf" / "orch" / "t-one.json").exists()) + + def test_post_handback_commits_leftovers_and_stops(self): + self.orch("pick", "fast") + out = self.orch("post", "t-one", "fast", "--result", "handback unclear spec") + self.assertEqual(out, "post: t-one handback\n committed leftover TASKS.md tasks/archive.md\n" + "stop lane fast: handback → tell the owner\n") + self.assertEqual(subprocess.run(["git", "-C", str(self.root), "log", "-1", "--format=%s"], + capture_output=True, text=True).stdout, "t-one handback (orchestrator)\n") + + def test_post_done_on_slice_job_counts_as_sliced(self): + self.orch("pick", "slow", "--id", "t-big") + self.run_wf(self.root, "add", "--parent", "t-big", "-e", "<1h", "--done", "x", "Slice one") + out = self.orch("post", "t-big", "slow", "--result", "done", "--no-next", "--duration", "5") + self.assertTrue(out.startswith("post: t-big sliced"), out) + self.assertNotIn("post-check-red", self.log()) + self.assertIn(" t-big sliced ", self.log()) + + def test_post_model_raised_continues(self): + self.orch("pick", "fast") + self.wf("set", "t-one", "--model", "opus") + self.wf("status", "t-one", "clear") + out = self.orch("post", "t-one", "fast", "--result", "handback beyond sonnet") + self.assertTrue(out.startswith("post: t-one model-raised\n"), out) + self.assertIn("pick: t-one (lane fast, model opus", out) + + def test_post_no_report_suggests_recovery(self): + self.orch("pick", "fast") + self.assertIn('wf orch pick fast --id t-one --recovery "<why>"', self.orch("post", "t-one", "fast")) + + def test_post_agent_without_transcript_still_logs(self): + self.orch("pick", "fast") + out = self.orch("post", "t-one", "fast", "--result", "wip", "--agent", "abc") + self.assertIn(" cost: not logged (no transcript for agent abc)\n", out) + + def test_refused_in_linked_worktree(self): + git(self.root, "worktree", "add", "-q", "--detach", str(self.root / ".worktrees" / "x"), "master") + got, _, err = self.wf("orch", "pick", "fast", project=False, cwd=self.root / ".worktrees" / "x") + self.assertEqual((got, err), (1, "wf: orch runs in the main tree (the orchestrator's)\n")) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_publish_snapshot.py b/tests/test_publish_snapshot.py new file mode 100644 index 0000000..3243d4b --- /dev/null +++ b/tests/test_publish_snapshot.py @@ -0,0 +1,122 @@ +import os +import subprocess +import sys +import tempfile +import unittest +from pathlib import Path + +SCRIPT = Path(__file__).resolve().parent.parent / "scripts" / "publish_snapshot.py" +ENV = {**os.environ, "GIT_AUTHOR_NAME": "t", "GIT_AUTHOR_EMAIL": "t@t", "GIT_COMMITTER_NAME": "t", + "GIT_COMMITTER_EMAIL": "t@t", "GIT_CONFIG_GLOBAL": "/dev/null"} + + +def git(cwd, *args): + return subprocess.run(["git", *args], cwd=cwd, capture_output=True, text=True, env=ENV, check=True).stdout + + +class PublishSnapshotTest(unittest.TestCase): + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + t = Path(self.tmp.name) + self.src, self.dest, self.deny = t / "tool", t / "pub", t / "deny.txt" + self.src.mkdir() + git(self.src, "init", "-q", "-b", "master") + self.put({"wf.py": "print('hi')\n", "CHANGES.md": "# Changes\n\n- 2026-01-01 first.\n", + "wflib/a.py": "x = 1\n", "docs/d.md": "doc\n", + "inbox.md": "tracked inbox\n", "out/log": "x\n", "sub/__pycache__/c.pyc": "bin\n"}) + (self.src / "wf.py").chmod(0o755) + self.commit("one") + self.deny.write_text("# private\nSecretProj\n\n") + + def tearDown(self): + self.tmp.cleanup() + + def put(self, files): + for p, text in files.items(): + f = self.src / p + f.parent.mkdir(parents=True, exist_ok=True) + f.write_text(text) + + def commit(self, msg): + git(self.src, "add", "-A") + git(self.src, "commit", "-qm", msg) + + def run_it(self, *extra): + return subprocess.run([sys.executable, str(SCRIPT), str(self.dest), "--src", str(self.src), + "--denylist", str(self.deny), *extra], capture_output=True, text=True, env=ENV) + + def files(self): + return sorted(git(self.dest, "ls-files").split()) + + def test_first_run_one_commit_excludes(self): + r = self.run_it() + self.assertEqual(r.returncode, 0, r.stderr) + self.assertEqual(self.files(), ["CHANGES.md", "docs/d.md", "wf.py", "wflib/a.py"]) + self.assertEqual(git(self.dest, "log", "--format=%s").splitlines(), ["initial public snapshot"]) + self.assertTrue(os.access(self.dest / "wf.py", os.X_OK)) + self.assertEqual(git(self.dest, "remote"), "") + + def test_denylist_hit_content_and_path_writes_nothing(self): + self.put({"docs/d.md": "doc\nsee secretproj here\n", "SecretProj.md": "x\n"}) + self.commit("leak") + r = self.run_it() + self.assertEqual(r.returncode, 1) + self.assertIn("docs/d.md:2: SecretProj", r.stderr) + self.assertIn("SecretProj.md: SecretProj (path)", r.stderr) + self.assertFalse(self.dest.exists()) + + def test_uncommitted_files_not_published(self): + (self.src / "wip.py").write_text("SecretProj\n") + r = self.run_it() + self.assertEqual(r.returncode, 0, r.stderr) + self.assertNotIn("wip.py", self.files()) + + def test_missing_or_empty_denylist_refuses(self): + self.deny.unlink() + r = self.run_it() + self.assertEqual(r.returncode, 1) + self.assertIn("no denylist", r.stderr) + self.deny.write_text("# only comments\n") + self.assertEqual(self.run_it().returncode, 1) + self.assertFalse(self.dest.exists()) + + def test_second_run_commit_message_new_changes_lines(self): + self.run_it() + self.put({"CHANGES.md": "# Changes\n\n- 2026-01-03 third.\n- 2026-01-02 second.\n- 2026-01-01 first.\n", + "wflib/b.py": "y = 2\n"}) + (self.src / "docs/d.md").unlink() + self.commit("two") + r = self.run_it() + self.assertEqual(r.returncode, 0, r.stderr) + self.assertEqual(git(self.dest, "log", "-1", "--format=%B").strip(), + "public snapshot\n\n- 2026-01-03 third.\n- 2026-01-02 second.") + self.assertEqual(self.files(), ["CHANGES.md", "wf.py", "wflib/a.py", "wflib/b.py"]) + self.assertEqual(len(git(self.dest, "log", "--format=%h").split()), 2) + + def test_no_change_no_commit(self): + self.run_it() + r = self.run_it() + self.assertEqual(r.returncode, 0, r.stderr) + self.assertIn("nothing new", r.stdout) + self.assertEqual(len(git(self.dest, "log", "--format=%h").split()), 1) + + def test_ref_option_and_never_touches_remote(self): + self.run_it() + git(self.dest, "remote", "add", "home", "/nonexistent") + self.put({"wflib/a.py": "x = 2\n"}) + self.commit("two") + self.assertEqual(self.run_it("--ref", "HEAD~1").stdout.strip(), "nothing new: no commit") + r = self.run_it() + self.assertIn("committed", r.stdout) + self.assertEqual(git(self.dest, "remote").split(), ["home"]) + + def test_nonempty_non_git_dest_refused(self): + self.dest.mkdir() + (self.dest / "keep.txt").write_text("x") + r = self.run_it() + self.assertEqual(r.returncode, 1) + self.assertIn("not a git repo", r.stderr) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_refs.py b/tests/test_refs.py new file mode 100644 index 0000000..4242a03 --- /dev/null +++ b/tests/test_refs.py @@ -0,0 +1,109 @@ +import sys +import tempfile +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import refs as R + +DOC = """\ +# Design + +Intro. + +## Subsystems + +### Character Model + +Bones and poses. + +#### Detail + +Deep. + +<a id="terrain"></a> +### Terrain & caves! + +Heightmap. + +## Other + +<a id="loose"></a> +Loose paragraph anchor. +""" + + +class LinksTest(unittest.TestCase): + def test_links(self): + self.assertEqual(R.links("see [[t-a]] and [[a-b|label]] and [[doc#part]]"), ["t-a", "a-b", "doc"]) + + def test_code_is_skipped(self): + self.assertEqual(R.links("`[[t-x]]` then [[t-y]]\n```\n[[t-z]]\n```\n"), ["t-y"]) + + +class SectionsTest(unittest.TestCase): + def test_slugify(self): + self.assertEqual(R.slugify("Character Model"), "character-model") + self.assertEqual(R.slugify("Terrain & caves!"), "terrain-caves") + + def test_sections(self): + got = [(s.level, s.heading, sorted(s.anchors), s.start, s.end) for s in R.sections(DOC)] + self.assertEqual(got, [ + (1, "Design", ["design"], 0, 23), + (2, "Subsystems", ["subsystems"], 4, 19), + (3, "Character Model", ["character-model"], 6, 14), + (4, "Detail", ["detail"], 10, 14), + (3, "Terrain & caves!", ["terrain", "terrain-caves"], 15, 19), + (2, "Other", ["other"], 19, 23), + (0, "", ["loose"], 21, 23), + ]) + + def test_anchors(self): + self.assertEqual(R.anchors(DOC), {"design", "subsystems", "character-model", "detail", + "terrain", "terrain-caves", "other", "loose"}) + + def test_headings_in_code_fences_are_not_sections(self): + self.assertEqual(R.anchors("# A\n```\n# not a heading\n```\n"), {"a"}) + + +class ResolveTest(unittest.TestCase): + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.root = Path(self.tmp.name) + (self.root / "docs").mkdir() + (self.root / "DESIGN.md").write_text(DOC) + long = "# Long\n\n## Part\n" + "".join(f"line {n}\n" for n in range(1, 100)) + (self.root / "docs" / "x.md").write_text(long) + (self.root / "docs" / "plan.md").write_text("# The Plan\n\n**Goal:** Ship it.\n\nMore.\n") + + def tearDown(self): + self.tmp.cleanup() + + def test_anchor_section(self): + self.assertEqual(R.resolve(self.root, "DESIGN.md", "terrain"), + "===== DESIGN.md#terrain (line 16) =====\n### Terrain & caves!\n\nHeightmap.") + + def test_long_section_is_cut(self): + out = R.resolve(self.root, "docs/x.md", "part").split("\n") + self.assertEqual(out[0], "===== docs/x.md#part (line 3) =====") + self.assertEqual(out[1], "## Part") + self.assertEqual(out[80], "line 79") + self.assertEqual(out[81], "… (20 more lines: docs/x.md:83)") + self.assertEqual(len(out), 82) + + def test_missing_anchor(self): + self.assertEqual(R.resolve(self.root, "DESIGN.md", "nope"), "(no anchor 'nope' in DESIGN.md)") + + def test_missing_file(self): + self.assertEqual(R.resolve(self.root, "docs/none.md", None), "(missing: docs/none.md)") + + def test_path_only_gives_heading_and_goal(self): + self.assertEqual(R.resolve(self.root, "docs/plan.md", None), "docs/plan.md: The Plan — Ship it.") + self.assertEqual(R.resolve(self.root, "DESIGN.md", None), "DESIGN.md: Design") + + def test_directory(self): + self.assertEqual(R.resolve(self.root, "docs", None), "docs/ (directory)") + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_res.py b/tests/test_res.py new file mode 100644 index 0000000..6c92549 --- /dev/null +++ b/tests/test_res.py @@ -0,0 +1,751 @@ +import datetime as dt +import json +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import res as R + +UTC = dt.timezone.utc + + +def T(h, m=0, day=1): + return dt.datetime(2026, 10, day, h, m, tzinfo=UTC) + + +NOW = T(14, 10) +GIB = 1024 ** 3 + + +def running(id="r-4", project="proj-a", title="dotnet e2e", mem=10.0, cpus=2, est=40, started=T(14, 0)): + return R.Entry(id=id, project=project, owner=1, title=title, mem_gb=mem, cpus=cpus, est_min=est, + state="running", unit=f"wf-{id}.service", started=started, + expires=started + dt.timedelta(minutes=2 * est)) + + +def note(id="r-8", owner=77, mem=3.0, cpus=1, expires=T(14, 30)): + return R.Entry(id=id, project="proj-b", owner=owner, title="dotnet test", mem_gb=mem, cpus=cpus, + est_min=20, state="note", started=T(14, 0), expires=expires) + + +def facts(available=20.0, units=None, live=(), results=None, slice_gb=0.0, now=NOW): + return R.Facts(now=now, total_gb=30.0, available_gb=available, nproc=16, units=units or {}, + live=set(live), results=results or {}, slice_gb=slice_gb) + + +class Parse(unittest.TestCase): + def test_size(self): + self.assertEqual(R.parse_size("10G"), 10.0) + self.assertEqual(R.parse_size("512M"), 0.5) + self.assertEqual(R.parse_size("1.5g"), 1.5) + with self.assertRaisesRegex(R.ResError, "size '10'"): + R.parse_size("10") + + def test_duration(self): + self.assertEqual(R.parse_duration("40m"), 40) + self.assertEqual(R.parse_duration("4h"), 240) + self.assertEqual(R.parse_duration("1h30m"), 90) + self.assertEqual(R.parse_duration("90"), 90) + for bad in ("0m", "x", "", "h"): + with self.assertRaises(R.ResError): + R.parse_duration(bad) + + def test_meminfo(self): + text = "MemTotal: 31457280 kB\nMemFree: 1 kB\nMemAvailable: 12582912 kB\n" + self.assertEqual(R.meminfo(text), (30.0, 12.0)) + with self.assertRaises(R.ResError): + R.meminfo("MemFree: 1 kB\n") + + def test_show_units(self): + text = ("Id=agents.slice\nActiveState=active\nMemoryCurrent=0\n\n" + "Id=wf-r-1.service\nActiveState=inactive\nMemoryCurrent=[not set]\n") + self.assertEqual(R.show_units(text), { + "agents.slice": {"Id": "agents.slice", "ActiveState": "active", "MemoryCurrent": "0"}, + "wf-r-1.service": {"Id": "wf-r-1.service", "ActiveState": "inactive", "MemoryCurrent": "[not set]"}}) + + def test_gb_or_none(self): + self.assertEqual(R.gb_or_none("1073741824"), 1.0) + self.assertIsNone(R.gb_or_none("[not set]")) + self.assertIsNone(R.gb_or_none("infinity")) + + def test_formats(self): + self.assertEqual(R.fmt_gb(2.25), "2.2 GB") + self.assertEqual(R.fmt_bytes(1288490188), "1.2 GB") + self.assertEqual(R.fmt_bytes(300 * 1024 ** 2), "300 MB") + self.assertEqual(R.fmt_dur(240), "4h") + self.assertEqual(R.fmt_dur(90), "1h30m") + self.assertEqual(R.fmt_dur(40), "40m") + self.assertEqual(R.hhmm(T(9, 5)), "09:05") + + +class ConfigTest(unittest.TestCase): + def test_defaults_and_override(self): + cfg = R.load_config("user_reserve_gb = 8\n") + self.assertEqual((cfg.user_reserve_gb, cfg.user_reserve_cpus, cfg.game_reserve_gb, cfg.game_reserve_cpus, + cfg.game_hours, cfg.small_headroom_gb, cfg.scratch_hours), (8.0, 4, 12.0, 8, 4.0, 2.0, 2.0)) + + def test_unknown_key(self): + self.assertEqual(R.load_config("session_mem_gb = 12").session_mem_gb, 12.0) + self.assertEqual(R.Config().session_mem_gb, 6.0) + with self.assertRaisesRegex(R.ResError, "unknown key\\(s\\) foo"): + R.load_config("foo = 1\n") + + def test_negative(self): + with self.assertRaisesRegex(R.ResError, "game_hours must be a number ≥ 0"): + R.load_config("game_hours = -1\n") + + +class LedgerJson(unittest.TestCase): + def test_roundtrip(self): + led = R.Ledger(next=9, game_until=T(18), entries=[running(), note()]) + text = R.dumps(led) + self.assertIn('"started": "2026-10-01T14:00:00+00:00"', text) + self.assertEqual(R.loads(text), led) + + def test_empty(self): + self.assertEqual(R.loads(""), R.Ledger()) + + def test_corrupt(self): + for bad in ("{", '{"entries": [{"id": "r-1"}]}', "[]"): + with self.assertRaisesRegex(R.ResError, "corrupt ledger"): + R.loads(bad) + + def test_unknown_entry_keys_ignored(self): + d = json.loads(R.dumps(R.Ledger(next=9, entries=[running()]))) + d["entries"][0]["from_a_newer_version"] = 1 + self.assertEqual(R.loads(json.dumps(d)), R.Ledger(next=9, entries=[running()])) + + def test_unknown_keys_survive_rewrite(self): + d = json.loads(R.dumps(R.Ledger(next=9, entries=[running()]))) + d["entries"][0]["field_x"] = {"a": 1} + d["top_x"] = 5 + out = json.loads(R.dumps(R.loads(json.dumps(d)))) + self.assertEqual(out["entries"][0]["field_x"], {"a": 1}) + self.assertEqual(out["top_x"], 5) + + def test_new_id_and_get(self): + led = R.Ledger(next=3) + self.assertEqual(led.new_id(), "r-3") + self.assertEqual(led.next, 4) + with self.assertRaisesRegex(R.ResError, "no entry 'r-9'"): + led.get("r-9") + + +class Prune(unittest.TestCase): + def test_prune(self): + r1, r2 = running("r-1"), running("r-2") + n3, n5 = note("r-3", owner=77), note("r-5", owner=78, expires=T(14, 5)) + d6 = R.Entry(id="r-6", project="p", owner=1, title="old", mem_gb=1, cpus=1, est_min=1, state="done", + ended=NOW - dt.timedelta(hours=25)) + d7 = R.Entry(id="r-7", project="p", owner=1, title="new", mem_gb=1, cpus=1, est_min=1, state="done", + ended=NOW - dt.timedelta(hours=23)) + led = R.Ledger(game_until=T(14, 9), entries=[r1, r2, n3, n5, d6, d7]) + f = facts(units={"wf-r-1.service": R.Unit(False)}, live={78}, results={"r-1": (0, 9.4, "")}) + lines = R.prune(led, f) + self.assertEqual(lines, ["r-1 exited rc=0", "r-2 exited rc=?", "r-3 freed (owner gone)", + "r-5 freed (expired)", "game off (expired)"]) + self.assertEqual([e.id for e in led.entries], ["r-1", "r-2", "r-3", "r-5", "r-7"]) + self.assertEqual((r1.state, r1.rc, r1.peak_gb, r1.why, r1.ended), ("done", 0, 9.4, "exited", NOW)) + self.assertEqual((n3.why, n5.why), ("owner gone", "expired")) + self.assertIsNone(led.game_until) + + def test_prune_missing_results(self): + r = running("r-1") + R.prune(R.Ledger(entries=[r]), facts(units={"wf-r-1.service": R.Unit(False)})) + self.assertEqual((r.state, r.rc, r.peak_gb), ("done", None, None)) + + def test_overdue_running_is_kept(self): + r = running(started=T(10)) + R.prune(R.Ledger(entries=[r]), facts(units={"wf-r-4.service": R.Unit(True, 4.0)})) + self.assertEqual(r.state, "running") + + +class Capacity(unittest.TestCase): + def setUp(self): + self.cfg = R.Config() + + def test_budget(self): + led = R.Ledger(entries=[running(), note()]) + f = facts(units={"wf-r-4.service": R.Unit(True, 4.0)}, live={77}) + # 20 − 6 reserve − 2 headroom − (10−4) unused − 3 note = 3; 16 − 4 − (2+1) = 9 + self.assertEqual(R.budget(self.cfg, led, f), (3.0, 9)) + + def test_budget_gaming(self): + led = R.Ledger(game_until=T(15), entries=[running(), note()]) + f = facts(units={"wf-r-4.service": R.Unit(True, 4.0)}, live={77}) + # 20 − 12 − 2 − 6 − 3 = −3; 16 − 8 − 3 = 5 + self.assertEqual(R.budget(self.cfg, led, f), (-3.0, 5)) + + def test_held_over_reservation_is_zero(self): + self.assertEqual(R.held_gb(running(), facts(units={"wf-r-4.service": R.Unit(True, 12.0)})), 0.0) + self.assertEqual(R.frees_gb(running(), facts(units={"wf-r-4.service": R.Unit(True, 12.0)})), 12.0) + + def test_fits(self): + self.assertTrue(R.fits(3.0, 9, (3.0, 9))) + self.assertFalse(R.fits(3.1, 1, (3.0, 9))) + self.assertFalse(R.fits(1.0, 10, (3.0, 9))) + + def test_never_fits(self): + f = facts() + self.assertEqual(R.never_fits(self.cfg, f, 25.0, 1), "25.0 GB can never fit (max 22.0 GB for agents)") + self.assertEqual(R.never_fits(self.cfg, f, 1.0, 13), "13 cpus can never fit (max 12 for agents)") + self.assertIsNone(R.never_fits(self.cfg, f, 22.0, 12)) + + def test_queued_total(self): + q = running("r-9") + q.state = "queued" + self.assertEqual(R.queued_total(R.Ledger(entries=[q, running()])), (10.0, 2)) + + +class BusyLine(unittest.TestCase): + def setUp(self): + self.cfg = R.Config() + self.led = R.Ledger(entries=[running()]) + self.f = facts(units={"wf-r-4.service": R.Unit(True, 4.0)}) # budget 6.0 GB, 10 cpus + + def test_memory(self): + self.assertEqual(R.busy_line(self.cfg, self.led, self.f, 8.0, 1), + 'busy: 10.0 GB held by proj-a "dotnet e2e" (r-4) until ~14:40; 6.0 GB free for agents; ' + 'retry after ~14:40 or work on something else') + + def test_cpus(self): + self.assertEqual(R.busy_line(self.cfg, self.led, self.f, 1.0, 12), + 'busy: 2 cpus held by proj-a "dotnet e2e" (r-4) until ~14:40; 10 cpus free for agents; ' + 'retry after ~14:40 or work on something else') + + def test_no_holder(self): + self.assertEqual(R.busy_line(self.cfg, R.Ledger(), facts(available=7.0), 1.0, 1), + "busy: only 0.0 GB free for agents and no agent job holds any; other programs use the rest; " + "retry later or work on something else") + + def test_busy_overdue_holder(self): + led = R.Ledger(entries=[running(started=T(10))]) + self.assertEqual(R.busy_line(self.cfg, led, self.f, 8.0, 1), + 'busy: 10.0 GB held by proj-a "dotnet e2e" (r-4) until overdue; 6.0 GB free for agents; ' + 'retry later or work on something else') + + def test_two_holders_earliest_first(self): + a = running("r-1", project="a", title="A", mem=4.0, cpus=1, est=60) # ends 15:00 + b = running("r-2", project="b", title="B", mem=4.0, cpus=1, est=30) # ends 14:30 + f = facts(available=14.0, units={}) # 14 − 6 − 2 − 4 − 4 = −2.0 budget + self.assertEqual(R.busy_line(self.cfg, R.Ledger(entries=[a, b]), f, 5.0, 1), + 'busy: 4.0 GB held by b "B" (r-2) until ~14:30, 4.0 GB held by a "A" (r-1) until ~15:00; ' + '0.0 GB free for agents; retry after ~15:00 or work on something else') + + +class Lines(unittest.TestCase): + def test_run_argv(self): + e = running("r-3", mem=10.0) + e.cwd, e.log, e.cmd = "/projects/x", "/s/logs/r-3.log", ["make", "it big", "--fast"] + self.assertEqual(R.run_argv(e, "/s/logs"), [ + "systemd-run", "--user", "--quiet", "--collect", "--slice=agents-jobs.slice", "--unit=wf-r-3.service", + "--working-directory=/projects/x", "--setenv=WF_RES_ID=r-3", + "-p", "MemoryMax=10737418240", "-p", "MemoryHigh=9663676416", "-p", "MemorySwapMax=0", "-p", "Nice=10", + "-p", "StandardOutput=append:/s/logs/r-3.log", "-p", "StandardError=append:/s/logs/r-3.log", + "/bin/sh", "-c", + '"$@"; rc=$?; cat /sys/fs/cgroup$(cut -d: -f3 /proc/self/cgroup)/memory.peak > /s/logs/r-3.peak ' + '2>/dev/null; echo $rc > /s/logs/r-3.rc', + "sh", "make", "it big", "--fast"]) + + def test_run_argv_env(self): + e = running("r-3", mem=1.0) + e.cwd, e.log, e.cmd, e.env = "/p", "/s/r-3.log", ["x"], {"B": "2", "A": "1 2"} + argv = R.run_argv(e, "/s") + self.assertEqual(argv[argv.index("--working-directory=/p") + 1:][:2], ["--setenv=A=1 2", "--setenv=B=2"]) + + def test_env_diff(self): + self.assertEqual(R.env_diff({"A": "1", "B": "2", "C": "3", "PWD": "/", "OLDPWD": "/", "SHLVL": "1", "_": "/x", + "MY_TOKEN": "t", "API_SECRET": "s", "PGPASSWORD": "p", "SSH_AUTH_SOCK": "/s"}, + {"A": "1", "B": "9", "SSH_AUTH_SOCK": "/s"}), + {"B": "2", "C": "3"}) + + def test_manager_env_parse(self): + self.assertEqual(R.parse_show_environment("A=1\nB=x=y\n\nbad\n"), {"A": "1", "B": "x=y"}) + + def test_entry_lines(self): + f = facts(units={"wf-r-4.service": R.Unit(True, 4.0)}, live={77}) + q = running("r-9", project="p", title="Q", mem=1.0, cpus=1) + q.state = "queued" + d = R.Entry(id="r-2", project="p", owner=1, title="D", mem_gb=1, cpus=1, est_min=1, state="done", + ended=T(13, 50), rc=None, peak_gb=None, why="exited") + led = R.Ledger(entries=[running(), note(), q, d]) + self.assertEqual([R.entry_line(e, led, f) for e in led.entries], [ + 'r-4 proj-a "dotnet e2e" running 10.0 GB used 4.0 GB 2 cpu since 14:00 ETA ~14:40', + 'r-8 proj-b "dotnet test" note 3.0 GB 1 cpu until ~14:30', + 'r-9 p "Q" queued #1 1.0 GB 1 cpu', + 'r-2 p "D" done (exited) rc=? peak ? at 13:50']) + + def test_status_lines(self): + f = facts(units={"wf-r-4.service": R.Unit(True, 4.0)}, slice_gb=5.5) + self.assertEqual(R.status_lines(R.Config(), R.Ledger(entries=[running()]), f), [ + 'r-4 proj-a "dotnet e2e" running 10.0 GB used 4.0 GB 2 cpu since 14:00 ETA ~14:40', + "agents may use 6.0 GB, 10 cpus now; reserve 6.0 GB/4 cpus", + "unreserved agent memory 1.5 GB"]) + f.slice_cache_gb = 0.6 + self.assertEqual(R.status_lines(R.Config(), R.Ledger(entries=[running()]), f)[-1], + "unreserved agent memory 1.5 GB (0.6 GB of it file cache, reclaimable)") + f.slice_cache_gb = 3.0 # cache partly inside the jobs' 4.0 GB → capped at the unreserved figure + self.assertEqual(R.status_lines(R.Config(), R.Ledger(entries=[running()]), f)[-1], + "unreserved agent memory 1.5 GB (1.5 GB of it file cache, reclaimable)") + + def test_cache_gb(self): + # agents.slice memory.stat excerpt from this machine: file 5602209792, shmem (tmpfs, not reclaimable) 3861041152 + self.assertEqual(round(R.cache_gb("anon 11335516160\nfile 5602209792\nkernel 1\nshmem 3861041152\n"), 3), 1.622) + self.assertEqual(R.cache_gb(""), 0.0) + + def test_status_lines_empty_overdue(self): + self.assertEqual(R.status_lines(R.Config(), R.Ledger(), facts(available=7.0)), [ + "no reservations", "agents may use 0.0 GB, 12 cpus now; reserve 6.0 GB/4 cpus"]) + r = running(started=T(10)) + self.assertIn("ETA overdue (still running, not killed)", R.entry_line(r, R.Ledger(entries=[r]), facts())) + + def test_status_json(self): + import json + d = json.loads(R.status_json(R.Config(), R.Ledger(entries=[running()]), + facts(units={"wf-r-4.service": R.Unit(True, 4.0)}))) + self.assertEqual((d["budget_gb"], d["budget_cpus"], d["gaming"], d["entries"][0]["id"]), + (6.0, 10, False, "r-4")) + + def test_done_line(self): + e = running("r-3", started=T(14, 0)) + e.state, e.ended, e.rc, e.peak_gb = "done", T(14, 37), 0, 9.4 + self.assertEqual(R.done_line(e), "r-3 done rc=0 peak 9.4 GB in 37 min") + e.rc, e.peak_gb = None, None + self.assertEqual(R.done_line(e), "r-3 done rc=? peak ? in 37 min") + e.why = "killed: timeout" + self.assertEqual(R.done_line(e), "r-3 done rc=? peak ? in 37 min (killed: timeout)") + + def test_prune_killed(self): + r = running("r-1") + f = facts(units={"wf-r-1.service": R.Unit(False)}, results={"r-1": (None, 3.9, "killed: timeout")}) + self.assertEqual(R.prune(R.Ledger(entries=[r]), f), ["r-1 killed: timeout"]) + self.assertEqual((r.state, r.rc, r.peak_gb, r.why), ("done", None, 3.9, "killed: timeout")) + + +# journalctl --user -u wf-r-12.service -o cat, verbatim from a real systemd-oomd kill (systemd 258) +OOMD = """Started wf-r-12.service - [systemd-run] /bin/sh -c "x" sh dotnet test. +wf-r-12.service: systemd-oomd killed some process(es) in this unit. +wf-r-12.service: Main process exited, code=killed, status=9/KILL +wf-r-12.service: Failed with result 'oom-kill'. +wf-r-12.service: Consumed 11min 2.837s CPU time, 3.9G memory peak. +""" + + +class JournalReason(unittest.TestCase): + def test_oomd(self): + self.assertEqual(R.journal_reason(OOMD, 4.0), + ("killed: oom-kill by systemd-oomd, limit 4.0 GB; raise --mem", 3.9)) + + def test_kernel_oom(self): + text = ("u: A process of this unit has been killed by the OOM killer.\n" + "u: Main process exited, code=killed, status=9/KILL\n" + "u: Failed with result 'oom-kill'.\nu: Consumed 2s CPU time, 512M memory peak.\n") + self.assertEqual(R.journal_reason(text, 0.5), + ("killed: oom-kill at MemoryMax, limit 0.5 GB; raise --mem", 0.5)) + + def test_signal(self): + text = "u: Main process exited, code=killed, status=15/TERM\nu: Failed with result 'signal'.\n" + self.assertEqual(R.journal_reason(text, 1.0), ("killed: signal 15/TERM", None)) + + def test_timeout_and_exit(self): + self.assertEqual(R.journal_reason("u: Failed with result 'timeout'.\n", 1.0), ("killed: timeout", None)) + text = "u: Main process exited, code=exited, status=3/NOTIMPLEMENTED\nu: Failed with result 'exit-code'.\n" + self.assertEqual(R.journal_reason(text, 1.0), ("failed: exit 3/NOTIMPLEMENTED", None)) + + def test_nothing(self): + self.assertEqual(R.journal_reason("", 1.0), ("", None)) + self.assertEqual(R.journal_reason("u: Deactivated successfully.\nu: Consumed 1s CPU time, 1.5G memory peak.\n", 1.0), + ("", 1.5)) + +def queued(id, mem, cpus=1, at=T(13, 0)): + e = R.Entry(id=id, project="p", owner=1, title=id, mem_gb=mem, cpus=cpus, est_min=10, state="queued", + cmd=["x"], queued=at) + return e + + +class Queue(unittest.TestCase): + def test_fifo_blocked_head(self): + # budget 20 − 6 − 2 = 12 GB + led = R.Ledger(entries=[queued("r-1", 5, at=T(13, 0)), queued("r-2", 9, at=T(13, 1)), + queued("r-3", 1, at=T(13, 2))]) + self.assertEqual([e.id for e in R.to_start(R.Config(), led, facts())], ["r-1"]) + + def test_all_fit(self): + led = R.Ledger(entries=[queued("r-2", 4, at=T(13, 1)), queued("r-1", 4, at=T(13, 0))]) + self.assertEqual([e.id for e in R.to_start(R.Config(), led, facts())], ["r-1", "r-2"]) + + def test_estimate(self): + # running r-4 holds 10 (used 4), ends 14:40; budget 6; queue r-1 (5) ahead of r-2 (4): need 9 − 6 = 3 + led = R.Ledger(entries=[running(), queued("r-1", 5), queued("r-2", 4, at=T(13, 5))]) + f = facts(units={"wf-r-4.service": R.Unit(True, 4.0)}) + self.assertEqual(R.queue_estimate(R.Config(), led, f, led.get("r-2")), T(14, 40)) + + def test_estimate_unknown(self): + led = R.Ledger(entries=[queued("r-1", 20)]) + self.assertIsNone(R.queue_estimate(R.Config(), led, facts(), led.get("r-1"))) + +class Game(unittest.TestCase): + def setUp(self): + self.cfg = R.Config() + + def test_slice_props(self): + self.assertEqual(R.slice_props(self.cfg, R.Ledger(), facts()), + {"CPUWeight": "20", "IOWeight": "20", "MemoryHigh": str(24 * GIB)}) + on = R.Ledger(game_until=T(18)) + self.assertEqual(R.slice_props(self.cfg, on, facts(slice_gb=10.0))["MemoryHigh"], str(18 * GIB)) + self.assertEqual(R.slice_props(self.cfg, on, facts(slice_gb=20.0)), + {"CPUWeight": "5", "IOWeight": "5", "MemoryHigh": str(20 * GIB)}) + + def test_shortfall(self): + a = running("r-4", project="proj-a", title="dotnet e2e", mem=2.5, est=55, started=T(14, 10)) # ~15:05 + b = running("r-7", project="proj-b", title="r5 rebuild", mem=9.0, est=80, started=T(14, 10)) # ~15:30 + f = facts(available=11.5, units={"wf-r-4.service": R.Unit(True, 2.5), "wf-r-7.service": R.Unit(True, 9.0)}) + led = R.Ledger(game_until=T(18), entries=[a, b]) + # free for you 11.5 − 0 unused = 11.5 → short 0.5: r-4 alone frees 2.5 + self.assertEqual(R.shortfall_lines(self.cfg, led, f), [ + 'short 0.5 GB of 12.0 GB: r-4 proj-a "dotnet e2e" 2.5 GB ~15:05', + "full reserve free ~15:05 (est.); free now: wf res release r-4 --stop"]) + f.available_gb = 8.9 # short 3.1: needs both + self.assertEqual(R.shortfall_lines(self.cfg, led, f), [ + 'short 3.1 GB of 12.0 GB: r-4 proj-a "dotnet e2e" 2.5 GB ~15:05, r-7 proj-b "r5 rebuild" 9.0 GB ~15:30', + "full reserve free ~15:30 (est.); free now: wf res release r-7 --stop"]) + f.available_gb = 13.0 + self.assertEqual(R.shortfall_lines(self.cfg, led, f), []) + + def test_shortfall_not_agents(self): + self.assertEqual(R.shortfall_lines(self.cfg, R.Ledger(game_until=T(18)), facts(available=5.0)), [ + "short 7.0 GB of 12.0 GB: other programs, not agent jobs", + "agent jobs alone cannot free it; close other programs"]) + + def test_game_on_lines(self): + led = R.Ledger(game_until=T(18, 10)) + self.assertEqual(R.game_on_lines(self.cfg, led, facts(available=20.0), 240), + ["game on until 18:10 (4h); CPU/IO now yours"]) + + def test_status_shows_game(self): + lines = R.status_lines(self.cfg, R.Ledger(game_until=T(18, 10)), facts(available=5.0)) + self.assertEqual(lines[1:], ["agents may use 0.0 GB, 8 cpus now; reserve 12.0 GB/8 cpus", + "game on until 18:10 (4h left)", + "short 7.0 GB of 12.0 GB: other programs, not agent jobs", + "agent jobs alone cannot free it; close other programs"]) + +class Clean(unittest.TestCase): + def test_scratch_victims(self): + now = 10_000_000.0 + h = 3600 + tree = { + "/s/old-file.png": (now - 3 * h, {}), + "/s/-projects-a": (now - 3 * h, {"/s/-projects-a/s1": now - 3 * h}), + "/s/-projects-b": (now - 60, {"/s/-projects-b/s1": now - 5 * h, "/s/-projects-b/s2": now - 60}), + "/s/new.txt": (now - 60, {}), + } + self.assertEqual(R.scratch_victims(tree, now, 2.0), + ["/s/-projects-a", "/s/-projects-b/s1", "/s/old-file.png"]) + # live sessions keep their dir however idle; a stale top holding one is pruned per child + self.assertEqual(R.scratch_victims(tree, now, 2.0, keep={"s1"}), ["/s/old-file.png"]) + tree["/s/-projects-a"] = (now - 3 * h, {"/s/-projects-a/s1": now - 3 * h, "/s/-projects-a/s9": now - 3 * h}) + self.assertEqual(R.scratch_victims(tree, now, 2.0, keep={"s1"}), ["/s/-projects-a/s9", "/s/old-file.png"]) + + def test_live_session_id(self): + text = '{"pid": 817151, "sessionId": "4ad140ab-69ca", "procStart": "31379851", "kind": "interactive"}' + self.assertEqual(R.live_session_id(text, "31379851"), "4ad140ab-69ca") + self.assertIsNone(R.live_session_id(text, "999")) # pid reused by another process + self.assertEqual(R.live_session_id('{"sessionId": "x"}', "5"), "x") # no procStart recorded → trust pid + self.assertIsNone(R.live_session_id("{oops", "5")) + self.assertIsNone(R.live_session_id('{"pid": 1}', "5")) + + def test_cleanup_rule(self): + self.assertEqual(R.cleanup_rule("out/prof"), ("out/prof", None)) + self.assertEqual(R.cleanup_rule("out/history-logs/*.log:30d"), ("out/history-logs/*.log", 30)) + for bad in ("/etc", "../x", "out/../../x"): + with self.assertRaisesRegex(R.ResError, "cleanup pattern"): + R.cleanup_rule(bad) + + def test_proc_name(self): + # background sessions run the versioned binary: comm is the version, argv0 the path or `claude` + self.assertEqual(R.proc_name("2.1.283", "claude"), "claude") + self.assertEqual(R.proc_name("2.1.283", "claude bg-pty-host --bg-pty-host /tmp/x.sock"), "claude") # rewritten title + self.assertEqual(R.proc_name("2.1.283", "/h/u/.local/share/claude/versions/2.1.283"), "claude") + self.assertEqual(R.proc_name("claude", "/h/u/.local/bin/claude"), "claude") + self.assertEqual(R.proc_name("2.1.284", "ugrep"), "2.1.284") # tool re-exec of the binary: not a session + self.assertEqual(R.proc_name("bash", "/bin/bash"), "bash") + self.assertEqual(R.proc_name("x", ""), "x") # kernel thread / unreadable cmdline + + def test_adopt_groups(self): + procs = [ + (100, 1, "bash", "/u/app.slice/tab1.scope"), + (101, 100, "claude", "/u/app.slice/tab1.scope"), # root, outside + (102, 101, "bash", "/u/app.slice/tab1.scope"), # its shell + (103, 102, "make", "/u/app.slice/tab1.scope"), + (104, 101, "claude", "/u/app.slice/tab1.scope"), # child claude: not a root, but a descendant + (105, 101, "job", "/u/agents.slice/agents-jobs.slice/wf-r-1.service"), # already inside: skipped + (200, 1, "claude", "/u/agents.slice/run-9.scope"), # root already inside + (300, 1, "vim", "/u/app.slice/tab2.scope"), + ] + self.assertEqual(R.adopt_groups(procs), {101: [101, 102, 103, 104]}) + + def test_warning_lines(self): + self.assertEqual(R.warning_lines(30.0, 7.6, 2), [ + "warning: /tmp (RAM) holds 7.6 GB; wf res clean", + "warning: 2 claude sessions outside agents.slice (wf res adopt)"]) + self.assertEqual(R.warning_lines(30.0, 7.4, 0), []) + +class Units(unittest.TestCase): + def test_adopt_argv(self): + self.assertEqual(R.adopt_argv(101, [101, 102]), [ + "busctl", "--user", "call", "org.freedesktop.systemd1", "/org/freedesktop/systemd1", + "org.freedesktop.systemd1.Manager", "StartTransientUnit", "ssa(sv)a(sa(sv))", + "wf-claude-101.scope", "fail", "2", "PIDs", "au", "2", "101", "102", "Slice", "s", "agents.slice", "0"]) + + def test_texts(self): + self.assertEqual(R.service_text("/usr/bin/python3", "/projects/public/workflow/wf.py"), + "[Unit]\nDescription=wf res tick\n\n[Service]\nType=oneshot\n" + "ExecStart=/usr/bin/python3 /projects/public/workflow/wf.py res tick\n") + self.assertEqual(R.timer_text(), + "[Unit]\nDescription=wf res tick every minute\n\n[Timer]\nOnBootSec=1min\n" + "OnUnitActiveSec=1min\n\n[Install]\nWantedBy=timers.target\n") + self.assertEqual(R.shell_init_line(), + "alias claude='systemd-run --user --scope --quiet --slice=agents.slice claude'") + + +if __name__ == "__main__": + unittest.main() + + +class Throttle(unittest.TestCase): + def test_psi(self): + text = "some avg10=41.50 avg60=33.20 avg300=12.00 total=99\nfull avg10=30.00 avg60=25.00 avg300=9.00 total=88\n" + self.assertEqual(R.psi_some_avg60(text), 33.2) + self.assertIsNone(R.psi_some_avg60("")) + + def test_lines(self): + hot = running("r-1", mem=4.0) # 3.5 of 4.0 GB (≥ 0.85×), stall 33% → warn + cache = running("r-2", mem=4.0) # full of page cache, no stall → quiet + small = running("r-3", mem=10.0) # stalled but far below its limit (machine pressure) → quiet + f = facts(units={"wf-r-1.service": R.Unit(True, 3.5, 33.2), "wf-r-2.service": R.Unit(True, 3.9, 0.0), + "wf-r-3.service": R.Unit(True, 2.0, 50.0)}) + self.assertEqual(R.throttle_lines(R.Ledger(entries=[hot, cache, small]), f), [ + "r-1 throttled at its memory limit (3.5 of 4.0 GB, stalled 33% of the last minute): " + "likely too small; `wf res release r-1 --stop` and re-run with a bigger --mem"]) + + +class Owner(unittest.TestCase): + REC = [{"lane": "fast", "pid": 11, "socket": "/s/a"}, {"lane": "slow", "pid": 22, "socket": "/s/b"}] + + def test_from_env_branch_and_session_record(self): + env = {"CLAUDE_PID": "22", "CLAUDE_CODE_MESSAGING_SOCKET": "/s/b", "WF_RES_ID": "r-7"} + self.assertEqual(R.owner_by(env, "slow/t-res-owner", self.REC), + {"name": "slow session", "task": "t-res-owner", "batch": "r-7", "address": "uds:/s/b"}) + + def test_explicit_name_wins_then_env(self): + env = {"CLAUDE_PID": "22", "WF_SESSION_NAME": "pilot", "WF_TASK": "t-x"} + self.assertEqual(R.owner_by(env, "slow/t-y", self.REC, "worker-3"), {"name": "worker-3", "task": "t-x"}) + self.assertEqual(R.owner_by(env, "", self.REC), {"name": "pilot", "task": "t-x"}) + + def test_nothing_known(self): + self.assertEqual(R.owner_by({"CLAUDE_PID": "99"}, "master", self.REC), {}) + self.assertEqual(R.owner_by({}, "feature-x", [{"pid": 1}]), {}) + + def test_status_and_throttle_name_owner(self): + e = running("r-1", mem=4.0) + e.by = {"name": "slow session", "task": "t-a", "batch": "r-7", "address": "uds:/s/b"} + f = facts(units={"wf-r-1.service": R.Unit(True, 3.5, 33.2)}) + led = R.Ledger(entries=[e]) + self.assertTrue(R.status_lines(R.Config(), led, f)[0].endswith( + " [by slow session t-a batch r-7, message uds:/s/b]")) + self.assertTrue(R.throttle_lines(led, f)[0].endswith( + "bigger --mem; started by slow session t-a batch r-7, message uds:/s/b")) + + def test_ledger_roundtrip_and_old_entries(self): + e = running("r-1") + e.by = {"task": "t-a"} + self.assertEqual(R.loads(R.dumps(R.Ledger(entries=[e]))).get("r-1").by, {"task": "t-a"}) + d = json.loads(R.dumps(R.Ledger(entries=[running("r-2")]))) + for k in ("by", "lock"): # entries written before these fields + del d["entries"][0][k] + old = R.loads(json.dumps(d)).get("r-2") + self.assertEqual((old.by, old.lock), ({}, "")) + + +class BareUnits(unittest.TestCase): + def test_bare_units(self): + blocks = R.show_units( + "Id=wh-t27.service\nTransient=yes\nSlice=app.slice\nMemoryCurrent=2147483648\n\n" + "Id=run-u42.service\nTransient=yes\nSlice=app.slice\nMemoryCurrent=[not set]\n\n" + "Id=app-org.kde.konsole@65c0.service\nTransient=yes\nSlice=app.slice\nMemoryCurrent=1\n\n" + "Id=dbus-:1.1-org.kde.kwalletd6@0.service\nTransient=yes\nSlice=app.slice\nMemoryCurrent=1\n\n" + "Id=wf-r-3.service\nTransient=yes\nSlice=agents-jobs.slice\nMemoryCurrent=1\n\n" + "Id=pipewire.service\nTransient=no\nSlice=session.slice\nMemoryCurrent=1\n") + self.assertEqual(R.bare_units(blocks), [("run-u42.service", None), ("wh-t27.service", 2.0)]) + + def test_warning(self): + self.assertEqual(R.warning_lines(30.0, 0.0, 0, [("wh-t27.service", 2.0), ("run-u42.service", None)]), [ + "warning: 2 jobs outside wf res (bare systemd-run): wh-t27.service 2.0 GB, run-u42.service ? GB; " + "start jobs with wf res run, also from project scripts"]) + + +class Hook(unittest.TestCase): + BASH = {"tool_name": "Bash", "tool_input": {"command": "dotnet run --project Desktop", "timeout": 5000}} + + def test_gaming_hides_display(self): + out = R.hook_output(self.BASH, T(16, 0), NOW) + self.assertEqual(out, {"hookSpecificOutput": { + "hookEventName": "PreToolUse", + "updatedInput": {"command": "unset DISPLAY WAYLAND_DISPLAY; dotnet run --project Desktop", "timeout": 5000}, + "additionalContext": "wf res game mode until 16:00: this command has no display, so GUI windows fail. " + "Run headless/offscreen, or do other work until game off."}}) + + def test_quiet_when_not_gaming_or_not_bash(self): + self.assertIsNone(R.hook_output(self.BASH, None, NOW)) + self.assertIsNone(R.hook_output(self.BASH, T(14, 0), NOW)) # expired + self.assertIsNone(R.hook_output({"tool_name": "Read", "tool_input": {"file_path": "x"}}, T(16, 0), NOW)) + self.assertIsNone(R.hook_output({"tool_name": "Bash", "tool_input": {}}, T(16, 0), NOW)) + + +class Litter(unittest.TestCase): + def test_victims(self): + now, h = 10_000_000.0, 3600 + entries = [("/t/aB3xYz", "emptydir", now - 3 * h), # old empty dir → go + ("/t/MSBuildTemp1000", "emptydir", now - 5 * h), # → go + ("/t/fresh", "emptydir", now - 60), # too new + ("/t/full", "dir", now - 9 * h), # has content → never + ("/t/clr-debug-pipe-111-222-in", "other", now - 60), # pid 111 dead → go + ("/t/dotnet-diagnostic-333-444-socket", "other", now), # pid 333 alive + ("/t/notes.txt", "file", now - 9 * h)] # file → never + alive = {333}.__contains__ + self.assertEqual(R.litter_victims(entries, now, 2.0, alive), + ["/t/MSBuildTemp1000", "/t/aB3xYz", "/t/clr-debug-pipe-111-222-in"]) +class SessionCap(unittest.TestCase): + BLOCKS = { + "wf-claude-1.scope": {"Slice": "agents.slice", "MemoryHigh": "infinity", "MemoryCurrent": str(GIB)}, + "wf-claude-2.scope": {"Slice": "agents.slice", "MemoryHigh": str(10 * GIB), "MemoryCurrent": str(GIB)}, + "run-u7.scope": {"Slice": "agents.slice", "MemoryHigh": str(4 * GIB), "MemoryCurrent": str(GIB)}, + "podman-pause.scope": {"Slice": "user.slice", "MemoryHigh": "infinity", "MemoryCurrent": str(GIB)}, + } + + def test_caps_to_set(self): + self.assertEqual(R.session_caps(self.BLOCKS, 10.0), + [("run-u7.scope", str(10 * GIB)), ("wf-claude-1.scope", str(10 * GIB))]) + self.assertEqual(R.session_caps(self.BLOCKS, 0.0), # 0 = off → lift caps + [("run-u7.scope", "infinity"), ("wf-claude-2.scope", "infinity")]) + + def test_warn_at_cap(self): + blocks = {"wf-claude-5.scope": {"Slice": "agents.slice", "MemoryCurrent": str(int(9.6 * GIB))}, + "wf-claude-6.scope": {"Slice": "agents.slice", "MemoryCurrent": str(int(9.8 * GIB))}, + "wf-claude-7.scope": {"Slice": "agents.slice", "MemoryCurrent": str(int(2 * GIB))}} + stall = {"wf-claude-5.scope": 45.0, "wf-claude-6.scope": 1.0, "wf-claude-7.scope": 90.0} + self.assertEqual(R.session_cap_lines(blocks, stall, 10.0), [ + "warning: claude session 5 at its memory cap (9.6 of 10.0 GB, stalled 45% of the last minute): " + "run big work with wf res run"]) + self.assertEqual(R.session_cap_lines(blocks, stall, 0.0), []) + + +def run(title="gate 30735e0", peak=1.0, minutes=10.0, project="proj", mem=12.0, est=60, rc=0, id="r-1"): + return {"id": id, "project": project, "title": title, "mem_gb": mem, "est_min": est, "rc": rc, + "peak_gb": peak, "min": minutes} + + +TEN = [run(id=f"r-{i}", title=f"gate {i:07x}a", peak=float(i), minutes=6.0 * i) for i in range(1, 11)] + + +class History(unittest.TestCase): + def test_title_kind(self): + for title, kind in (("gate 30735e0", "gate"), ("gate 2c91cac7 (fix optimizer test)", "gate"), + ("mem-profile 16x5M (t-optimizer-x)", "mem-profile 16x5M"), ("wf-batch", "wf-batch"), + ("gate 97e6044b direct", "gate 97e6044b direct"), ("build deadbeef", "build deadbeef"), + ("gate a3cbde1 0cf9a79", "gate"), ("abc1234", "abc1234"), (" x ", "x")): + self.assertEqual(R.title_kind(title), kind, title) + + def test_pct_nearest_rank(self): + xs = [5.0, 1.0, 3.0, 2.0, 4.0] + self.assertEqual([R.pct(xs, p) for p in (0.5, 0.9, 0.95, 0.2)], [3.0, 5.0, 5.0, 1.0]) + + def test_history_record(self): + e = running(project="p", title="gate x", mem=10.0, est=40) + e.state, e.ended, e.rc, e.peak_gb = "done", T(14, 13, ), 0, 9.1 + e.ended = e.started + dt.timedelta(minutes=13, seconds=30) + self.assertEqual(R.history_record(e), {"id": "r-4", "project": "p", "title": "gate x", "mem_gb": 10.0, + "est_min": 40, "rc": 0, "peak_gb": 9.1, "min": 13.5}) + for change in ({"rc": 1}, {"peak_gb": None}, {"state": "running"}, {"ended": None}): + bad = R.Entry(**{**{f.name: getattr(e, f.name) for f in R.fields(R.Entry)}, **change}) + self.assertIsNone(R.history_record(bad), change) + + def test_suggest_p95_p90(self): + # peaks 1..10 → p95 = 10 → ×1.15 = 11.5; durations 6..60 → p90 = 54 → ×1.5 = 81 + self.assertEqual(R.suggest(TEN, "proj", "gate"), (11.5, 81, 10)) + + def test_suggest_floors_and_rounding(self): + tiny = [run(id=f"r-{i}", peak=0.01, minutes=1.0) for i in range(3)] + self.assertEqual(R.suggest(tiny, "proj", "gate"), (0.2, 5, 3)) + odd = [run(id=f"r-{i}", peak=1.01, minutes=7.1) for i in range(3)] + self.assertEqual(R.suggest(odd, "proj", "gate"), (1.2, 11, 3)) # 1.1615 → 1.2 up; 10.65 → 11 up + + def test_suggest_needs_three_matching(self): + two = TEN[:2] + [run(id="r-a", project="other"), run(id="r-b", title="build"), + run(id="r-c", rc=1), run(id="r-d", peak=None)] + self.assertIsNone(R.suggest(two, "proj", "gate")) + self.assertIsNone(R.suggest([], "proj", "gate")) + + def test_hint_line(self): + sug = (11.5, 81, 10) + line = "hint: history says ~11.5 GB / 81 min (10 runs)" + self.assertEqual(R.hint_line(23.1, 60, sug), line) + self.assertEqual(R.hint_line(10.0, 163, sug), line) + self.assertIsNone(R.hint_line(23.0, 162, sug)) + self.assertIsNone(R.hint_line(99.0, 999, None)) + + def test_merge_runs(self): + led = R.Ledger(entries=[running(id="r-4", project="proj", title="gate 1234567"), running(id="r-5")]) + led.entries[0].state, led.entries[0].rc, led.entries[0].peak_gb = "done", 0, 2.0 + led.entries[0].ended = led.entries[0].started + dt.timedelta(minutes=3) + merged = R.all_runs([run(id="r-1"), run(id="r-4", peak=7.0)], led) + self.assertEqual([(r["id"], r["peak_gb"]) for r in merged], [("r-1", 1.0), ("r-4", 7.0)]) + merged = R.all_runs([run(id="r-1")], led) + self.assertEqual([(r["id"], r["peak_gb"], r["min"]) for r in merged], [("r-1", 1.0, 10.0), ("r-4", 2.0, 3.0)]) + + def test_hist_lines(self): + runs = TEN + [run(id="r-x", title="wf-batch", peak=3.0, minutes=20.0, mem=6.0, est=400), + run(id="r-y", project="other", title="build")] + self.assertEqual(R.hist_lines(runs, "proj"), [ + "proj · gate · n 10 · req 12.0 GB · peak 5.0/10.0 GB · est 1h · dur 30m/54m · suggest 11.5 GB 1h21m", + "proj · wf-batch · n 1 · req 6.0 GB · peak 3.0/3.0 GB · est 6h40m · dur 20m/20m · suggest - (< 3 runs)"]) + self.assertEqual(len(R.hist_lines(runs, None)), 3) + self.assertEqual(R.hist_lines([], None), ["no finished runs with a peak yet"]) + + +ORCH_LOG = """\ +2026-10-06T10:00:00 slow opus t-a done abc1234 10m00s +2026-10-06T10:10:00 fast haiku t-b handback - 20m30s +2026-10-06T10:20:00 slow opus t-c done+gate-red abc1234 1h05m +2026-10-06T10:30:00 fast haiku t-d post-check-red abc1234 50m00s (branch x still there) +2026-10-06T10:40:00 fast haiku t-e done abc1234 0m00s +2026-10-06T10:50:00 slow opus t-f done abc1234 - +2026-10-06T11:00:00 cloud opus t-g done abc1234 3h00m +2026-10-06T11:10:00 slow sonnet t-h awaiting - 40m00s +21:30 pilot batch 1 rc=0 done=t-a stopped= added= +ALERT wf-pilot proj: awaiting a-1 +2026-10-06T11:20:00 slow opus t-i done abc1234 30m00s +""" + + +class TaskFit(unittest.TestCase): + def test_durations_from_log(self): + # done/handback rows of local lanes with a real duration only (minutes) + self.assertEqual(R.orch_durations(ORCH_LOG), [10.0, 20.5, 65.0, 30.0]) + + def test_p90(self): + # 4 rows: nearest rank ceil(0.9*4)=4 → 65 min + self.assertEqual(R.task_p90(ORCH_LOG), (65, 4)) + self.assertEqual(R.task_p90(""), (30, 0)) + two = "2026-10-06T10:00:00 slow opus t-a done abc 5m00s\n" * 2 + self.assertEqual(R.task_p90(two), (30, 2)) # < 3 runs → default 30 min + odd = "2026-10-06T10:00:00 slow opus t-a done abc 4m10s\n" * 3 + self.assertEqual(R.task_p90(odd), (5, 3)) # rounded up to whole minutes + + def test_batch_fit(self): + self.assertEqual(R.batch_fit(4, 150, 30), 4) + self.assertEqual(R.batch_fit(4, 90, 30), 3) + self.assertEqual(R.batch_fit(4, 59, 30), 1) + self.assertEqual(R.batch_fit(4, 29, 30), 0) + self.assertEqual(R.batch_fit(4, 0, 30), 0) diff --git a/tests/test_res_io.py b/tests/test_res_io.py new file mode 100644 index 0000000..d91046a --- /dev/null +++ b/tests/test_res_io.py @@ -0,0 +1,815 @@ +import datetime as dt +import io +import json +import os +import subprocess +import sys +import tempfile +import unittest +from contextlib import redirect_stderr, redirect_stdout +from pathlib import Path + +HERE = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(HERE)) +import wf_res # noqa: E402 +from wflib import res as R # noqa: E402 + +UTC = dt.timezone.utc +GIB_KB = 1024 * 1024 + + +class Fake: + """Records argv; answers systemctl show from self.units {name: (ActiveState, MemoryCurrent bytes|None)}.""" + + def __init__(self): + self.calls, self.units, self.fail, self.journal, self.cgroups, self.bare = [], {}, {}, {}, {}, {} + self.scopes = {} # name → {prop: value} for agent session scopes + self.manager_env = "PATH=/usr/bin\nHOME=/h/u\n" + self.git = {} # git subcommand (branch | rev-parse) → stdout + + def __call__(self, argv): + self.calls.append(list(argv)) + key = argv[0] if argv[0] != "systemctl" else argv[2] + if key in self.fail: + return subprocess.CompletedProcess(argv, 1, "", self.fail[key] + "\n") + out = "" + if argv[:3] == ["systemctl", "--user", "show"]: + blocks = [] + for name in argv[5:]: + if name in self.scopes: + blocks.append("\n".join([f"Id={name}"] + [f"{k}={v}" for k, v in self.scopes[name].items()])) + continue + if name in self.bare: + blocks.append(f"Id={name}\nTransient=yes\nSlice=app.slice\nMemoryCurrent={self.bare[name]}") + continue + state, cur = self.units.get(name, ("inactive", None)) + blocks.append(f"Id={name}\nActiveState={state}\nMemoryCurrent={'[not set]' if cur is None else cur}" + f"\nControlGroup={self.cgroups.get(name, '')}") + out = "\n\n".join(blocks) + "\n" + elif argv[:3] == ["systemctl", "--user", "list-units"]: + names = self.scopes if "--type=scope" in argv else self.bare + out = "".join(f"{n} loaded active running x\n" for n in names) + elif argv[:3] == ["systemctl", "--user", "set-property"]: + if argv[4] in self.scopes: + self.scopes[argv[4]].update(a.split("=", 1) for a in argv[5:]) + elif argv[:3] == ["systemctl", "--user", "show-environment"]: + out = self.manager_env + elif argv[0] == "git": + out = self.git.get(argv[3], "") + elif argv[0] == "journalctl": + out = self.journal.get(argv[argv.index("-u") + 1], "") + elif argv[0] == "systemd-run": + unit = next(a for a in argv if a.startswith("--unit="))[len("--unit="):] + self.units[unit] = ("active", 0) + elif argv[:3] == ["systemctl", "--user", "stop"]: + self.units[argv[3]] = ("inactive", None) + return subprocess.CompletedProcess(argv, 0, out, "") + + def ran(self, word): + return [c for c in self.calls if word in c[:7]] + + +class Box: + """Temp dirs + fake machine; .env is a wf_res.Env.""" + + def __init__(self, tmp: Path, available_gb=20, total_gb=30, now=dt.datetime(2026, 10, 1, 14, 0, tzinfo=UTC)): + self.tmp, self.fake, self.clock, self.alive = tmp, Fake(), [now], {4242} + self.available_gb, self.total_gb = available_gb, total_gb + (tmp / "proj").mkdir(exist_ok=True) + self.env = wf_res.Env( + state=tmp / "state", config=tmp / "cfg" / "resources.toml", units=tmp / "units", + scratch=tmp / "scratch", run=self.fake, + meminfo=lambda: f"MemTotal: {int(self.total_gb * GIB_KB)} kB\nMemAvailable: {int(self.available_gb * GIB_KB)} kB\n", + nproc=16, now=lambda: self.clock[0], pid_alive=lambda p: p in self.alive, owner=lambda: 4242, + root=lambda: tmp / "proj", cwd=lambda: str(tmp / "proj"), procs=lambda: [], + sleep=self.tick, tmp_used_gb=lambda: 0.0, cgread=lambda cg, name: self.psi.get(cg, "") if name == "memory.pressure" + else self.stat.get(cg, "")) + self.psi, self.stat = {}, {} + self.env.live_sessions = lambda: set() + self.caller = {"PATH": "/usr/bin", "HOME": "/h/u"} + self.env.environ = lambda: dict(self.caller) + self.env.tmp_dirs = [] + + def tick(self, seconds): + self.clock[0] += dt.timedelta(seconds=seconds) + + def wf(self, *argv): + out, err = io.StringIO(), io.StringIO() + with redirect_stdout(out), redirect_stderr(err): + code = wf_res.main(list(argv), self.env) + return code, out.getvalue(), err.getvalue() + + def ledger(self): + return R.loads((self.env.state / "resources.json").read_text()) + + +class IOBase(unittest.TestCase): + def setUp(self): + self._tmp = tempfile.TemporaryDirectory() + self.box = Box(Path(self._tmp.name)) + + def tearDown(self): + self._tmp.cleanup() + + +class Run(IOBase): + def test_run_starts_unit(self): + code, out, err = self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build", "--", "make", "-j8") + self.assertEqual((code, err), (0, "")) + log = self.box.env.state / "logs" / "r-1.log" + self.assertEqual(out, f"r-1 started; log {log}; ETA ~14:40 (estimate: not killed when over)\n") + argv = self.box.fake.ran("systemd-run")[0] + self.assertEqual(argv[-3:], ["sh", "make", "-j8"]) + self.assertIn("--unit=wf-r-1.service", argv) + e = self.box.ledger().get("r-1") + self.assertEqual((e.state, e.project, e.owner, e.cwd), ("running", "proj", 4242, str(self.box.tmp / "proj"))) + self.assertTrue((self.box.env.units / "agents.slice").exists()) + + def test_run_keeps_caller_env(self): + self.box.caller.update({"APP_DIR": "/data/arc", "PATH": "/venv/bin:/usr/bin", "PWD": "/x", "SHLVL": "2", + "CLAUDE_CODE_MESSAGING_TOKEN": "t", "GH_TOKEN": "t", "DB_PASSWORD": "p"}) + self.box.wf("run", "--mem", "1G", "--for", "1m", "--title", "x", "--", "x") + argv = self.box.fake.ran("systemd-run")[0] + self.assertEqual([a for a in argv if a.startswith("--setenv=")], + ["--setenv=APP_DIR=/data/arc", "--setenv=PATH=/venv/bin:/usr/bin", "--setenv=WF_RES_ID=r-1"]) + self.assertEqual(self.box.ledger().get("r-1").env, {"APP_DIR": "/data/arc", "PATH": "/venv/bin:/usr/bin"}) + + def test_run_records_owner(self): + main = self.box.tmp / "main" + (main / ".wf" / "sessions").mkdir(parents=True) + (main / ".wf" / "sessions" / "slow.json").write_text('{"lane": "slow", "pid": 77, "socket": "/s/o"}\n') + self.box.fake.git = {"branch": "slow/t-big\n", "rev-parse": f"{main}/.git\n"} + self.box.caller.update({"CLAUDE_PID": "77", "CLAUDE_CODE_MESSAGING_SOCKET": "/s/o", "WF_RES_ID": "r-9"}) + self.box.wf("run", "--mem", "1G", "--for", "1m", "--title", "x", "--", "x") + self.assertEqual(self.box.ledger().get("r-1").by, + {"name": "slow session", "task": "t-big", "batch": "r-9", "address": "uds:/s/o"}) + argv = self.box.fake.ran("systemd-run")[0] + self.assertEqual([a for a in argv if a.startswith("--setenv=WF_RES_ID")], ["--setenv=WF_RES_ID=r-1"]) + self.assertIn('"x" running', self.box.wf("status")[1]) + self.assertIn("[by slow session t-big batch r-9, message uds:/s/o]", self.box.wf("status")[1]) + self.box.wf("note", "--mem", "1G", "--for", "5m", "--by", "pilot", "edit") + self.assertEqual(self.box.ledger().get("r-2").by["name"], "pilot") + + def test_queued_run_starts_later_with_its_env(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + self.box.caller["FOO"] = "1" + self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "q", "--queue", "--", "y") + self.box.caller.pop("FOO") + self.box.wf("release", "r-1", "--stop") + self.box.wf("status") + starts = self.box.fake.ran("systemd-run") + self.assertEqual([[a for a in s if a.startswith("--setenv=")] for s in starts], [["--setenv=WF_RES_ID=r-1"], ["--setenv=FOO=1", "--setenv=WF_RES_ID=r-2"]]) + + def test_busy_exit_3(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + code, out, _ = self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--", "y") + # 20 − 6 − 2 − 10 unused = 2.0 free + self.assertEqual(code, 3) + self.assertEqual(out, 'busy: 10.0 GB held by proj "big" (r-1) until ~14:40; 2.0 GB free for agents; ' + 'retry after ~14:40 or work on something else; 12.0 GB really free beyond the reserve: ' + '--force starts it past the ledger (only if the holders will not use what they reserved)\n') + self.assertEqual([e.id for e in self.box.ledger().entries], ["r-1"]) + + def test_force_runs_past_unused_claim(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + code, out, _ = self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--force", "--", "y") + # really free: 20 − 6 reserve − 2 headroom = 12 ≥ 8 (r-1's unused 10 GB ignored) + self.assertEqual(code, 0) + self.assertTrue(out.startswith("r-2 started"), out) + self.assertEqual([(e.id, e.state) for e in self.box.ledger().entries], [("r-1", "running"), ("r-2", "running")]) + + def test_force_refused_when_memory_really_used(self): + self.box.available_gb = 12 # really free beyond reserve: 12 − 6 − 2 = 4 + self.box.wf("run", "--mem", "3G", "--for", "40m", "--title", "a", "--", "x") + code, out, _ = self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--", "y") + self.assertEqual(code, 3) + self.assertNotIn("--force", out) # hint only when --force would fit + code, out, _ = self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--force", "--", "y") + self.assertEqual((code, out), (3, "busy even with --force: only 4.0 GB really free beyond the reserve; " + "retry later or work on something else\n")) + self.assertEqual([e.id for e in self.box.ledger().entries], ["r-1"]) + + def test_never_fits(self): + code, _, err = self.box.wf("run", "--mem", "25G", "--for", "1h", "--title", "x", "--", "x") + self.assertEqual((code, err), (1, "wf: 25.0 GB can never fit (max 22.0 GB for agents)\n")) + + def test_systemd_run_failure(self): + self.box.fake.fail["systemd-run"] = "Failed to start transient service unit: boom" + code, _, err = self.box.wf("run", "--mem", "1G", "--for", "1m", "--title", "x", "--", "x") + self.assertEqual((code, err), (1, "wf: systemd-run: Failed to start transient service unit: boom\n")) + self.assertEqual(self.box.ledger().entries, []) + + def test_no_command(self): + code, _, err = self.box.wf("run", "--mem", "1G", "--for", "1m", "--title", "x") + self.assertEqual((code, err), (1, "wf: no command after --\n")) + + def test_usage_error_exit_2(self): + with redirect_stderr(io.StringIO()): + self.assertEqual(self.box.wf("run", "--for", "1m")[0], 2) + + +class Status(IOBase): + def test_status_and_exit_pruned(self): + self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build", "--", "make") + self.box.fake.units["wf-r-1.service"] = ("active", 2 * 1024 ** 3) + code, out, _ = self.box.wf("status") + self.assertEqual(out.splitlines()[0], 'r-1 proj "build" running 4.0 GB used 2.0 GB 1 cpu since 14:00 ETA ~14:40') + logs = self.box.env.state / "logs" + (logs / "r-1.rc").write_text("0\n") + (logs / "r-1.peak").write_text(str(3 * 1024 ** 3) + "\n") + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + code, out, _ = self.box.wf("status") + self.assertEqual(out.splitlines()[0], 'r-1 proj "build" done (exited) rc=0 peak 3.0 GB at 14:00') + + def test_status_warns_throttled(self): + self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build", "--", "make") + self.box.fake.units["wf-r-1.service"] = ("active", int(3.8 * 1024 ** 3)) + self.box.fake.cgroups["wf-r-1.service"] = "/x/wf-r-1.service" + self.box.psi["/x/wf-r-1.service"] = "some avg10=50.00 avg60=40.00 avg300=10.00 total=1\n" + out = self.box.wf("status")[1] + self.assertIn("r-1 throttled at its memory limit (3.8 of 4.0 GB, stalled 40% of the last minute): " + "likely too small; `wf res release r-1 --stop` and re-run with a bigger --mem\n", out) + + def test_status_warns_bare_units(self): + self.box.fake.bare["wh-t27.service"] = 2 * 1024 ** 3 + out = self.box.wf("status")[1] + self.assertIn("warning: 1 jobs outside wf res (bare systemd-run): wh-t27.service 2.0 GB; " + "start jobs with wf res run, also from project scripts\n", out) + self.assertIn(["systemctl", "--user", "list-units", "--type=service", "--state=running", "--no-legend", + "--plain"], self.box.fake.calls) + + def test_status_splits_slice_cache(self): + self.box.fake.units["agents.slice"] = ("active", 5 * 1024 ** 3) + self.box.fake.cgroups["agents.slice"] = "/a.slice" + self.box.stat["/a.slice"] = f"anon {3 * 1024 ** 3}\nfile {2 * 1024 ** 3}\nshmem {1024 ** 3 // 2}\n" + self.assertIn("unreserved agent memory 5.0 GB (1.5 GB of it file cache, reclaimable)\n", self.box.wf("status")[1]) + + def test_status_id_and_json(self): + self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build", "--", "make", "a b") + out = self.box.wf("status", "r-1")[1] + self.assertIn("cmd: make 'a b'", out) + self.assertEqual(json.loads(self.box.wf("status", "--json")[1])["entries"][0]["id"], "r-1") + + def test_show_failure_keeps_ledger(self): + self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build", "--", "make") + self.box.fake.fail["show"] = "Failed to connect to bus" + code, _, err = self.box.wf("status") + self.assertEqual((code, err), (1, "wf: systemctl show: Failed to connect to bus\n")) + self.assertEqual(self.box.ledger().get("r-1").state, "running") + + def test_corrupt_ledger_moved_aside(self): + self.box.env.state.mkdir(parents=True) + (self.box.env.state / "resources.json").write_text("{oops") + code, out, err = self.box.wf("status") + self.assertEqual(code, 0) + self.assertRegex(err, r"^wf: corrupt ledger: .*; moved to resources\.json\.bad-20261001-140000, starting empty\n$") + self.assertTrue((self.box.env.state / "resources.json.bad-20261001-140000").exists()) + self.assertEqual(out.splitlines()[0], "no reservations") + + +class LedgerCompat(IOBase): + """Readers of another code version (a long `wf res wait` started before a release) must keep entries.""" + + def _write(self, extra: dict, drop=()): + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "t", "--", "x") + path = self.box.env.state / "resources.json" + d = json.loads(path.read_text()) + for e in d["entries"]: + for k in drop: + e.pop(k, None) + e.update(extra) + path.write_text(json.dumps(d)) + return path + + def _kept(self, path): + for argv in (("status",), ("tick",), ("wait", "r-1", "--timeout", "1m")): + self.box.wf(*argv) + self.assertEqual(self.box.ledger().get("r-1").state, "running", argv) + self.assertEqual(sorted(p.name for p in path.parent.glob("resources.json*")), ["resources.json"]) + self.assertIn("r-1", self.box.wf("status")[1]) + + def test_pre_owner_ledger_kept(self): # written before e84d7ca/9533211: no by field + self._kept(self._write({}, drop=("by", "env"))) + + def test_future_fields_kept(self): # written by newer code: unknown keys + self._kept(self._write({"by": {"name": "x"}, "added_later": 1})) + + def test_corrupt_ledger_keeps_id_counter(self): + self.box.env.state.mkdir(parents=True) + (self.box.env.state / "resources.json").write_text('{"next": 652, "entries": [{"id": "r-651"}]}') + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "t", "--", "x") + self.assertEqual([e.id for e in self.box.ledger().entries], ["r-652"]) + + def test_start_clears_stale_results(self): + logs = self.box.env.state / "logs" + logs.mkdir(parents=True) + (logs / "r-1.rc").write_text("1\n") + (logs / "r-1.peak").write_text("5\n") + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "t", "--", "x") + self.assertFalse((logs / "r-1.rc").exists() or (logs / "r-1.peak").exists()) + + +class HookIO(IOBase): + def hook(self, text): + self.box.env.stdin = lambda: text + return self.box.wf("hook") + + def test_game_on_rewrites_bash(self): + self.box.wf("game", "on", "--for", "1h") + code, out, err = self.hook('{"tool_name": "Bash", "tool_input": {"command": "./app"}}') + self.assertEqual((code, err), (0, "")) + self.assertEqual(json.loads(out)["hookSpecificOutput"]["updatedInput"], + {"command": "unset DISPLAY WAYLAND_DISPLAY; ./app"}) + calls = len(self.box.fake.calls) + self.hook('{"tool_name": "Bash", "tool_input": {"command": "./app"}}') + self.assertEqual(len(self.box.fake.calls), calls) # read-only: no systemctl, no prune + + def test_silent_otherwise(self): + self.assertEqual(self.hook('{"tool_name": "Bash", "tool_input": {"command": "./app"}}'), (0, "", "")) + self.box.wf("game", "on") + self.assertEqual(self.hook("not json"), (0, "", "")) + (self.box.env.state / "resources.json").write_text("{oops") + self.assertEqual(self.hook('{"tool_name": "Bash", "tool_input": {"command": "x"}}'), (0, "", "")) + self.assertTrue((self.box.env.state / "resources.json").exists()) # hook never moves a bad ledger + + +class Release(IOBase): + def test_release_running_needs_stop(self): + self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build", "--", "make") + code, _, err = self.box.wf("release", "r-1") + self.assertEqual((code, err), (1, "wf: r-1 is running; --stop to kill it\n")) + code, out, _ = self.box.wf("release", "r-1", "--stop") + self.assertEqual((code, out), (0, "r-1 stopped\n")) + self.assertEqual(self.box.fake.ran("stop")[0], ["systemctl", "--user", "stop", "wf-r-1.service"]) + e = self.box.ledger().get("r-1") + self.assertEqual((e.state, e.why), ("done", "released")) + + def test_release_unknown(self): + self.assertEqual(self.box.wf("release", "r-7")[2], "wf: no entry 'r-7'\n") + + +class Dispatch(unittest.TestCase): + def test_wf_forwards_res(self): + r = subprocess.run([sys.executable, str(HERE / "wf.py"), "res", "-h"], capture_output=True, text=True) + self.assertEqual(r.returncode, 0) + self.assertIn("usage: wf res", r.stdout) + +class Wait(IOBase): + def test_wait_until_done(self): + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "t", "--", "x") + logs = self.box.env.state / "logs" + polls = [] + + def sleep(seconds): + polls.append(seconds) + self.box.tick(seconds) + if len(polls) == 2: + (logs / "r-1.rc").write_text("0\n") + (logs / "r-1.peak").write_text(str(1024 ** 3 // 2) + "\n") + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + self.box.env.sleep = sleep + code, out, _ = self.box.wf("wait", "r-1") + self.assertEqual((code, out, polls), (0, "r-1 done rc=0 peak 0.5 GB in 0 min\n", [15, 15])) + + def test_wait_warns_throttled_once(self): + self.box.wf("run", "--mem", "4G", "--for", "5m", "--title", "t", "--", "x") + self.box.fake.units["wf-r-1.service"] = ("active", int(3.8 * 1024 ** 3)) + self.box.fake.cgroups["wf-r-1.service"] = "/x/wf-r-1.service" + self.box.psi["/x/wf-r-1.service"] = "some avg10=50.00 avg60=40.00 avg300=10.00 total=1\n" + code, _, err = self.box.wf("wait", "r-1", "--timeout", "1m") + self.assertEqual(err, "wf: r-1 throttled at its memory limit (3.8 of 4.0 GB, stalled 40% of the last minute): " + "likely too small; `wf res release r-1 --stop` and re-run with a bigger --mem\n" + "wf: r-1 not done after 1m (running)\n") + + def test_wait_timeout(self): + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "t", "--", "x") + code, _, err = self.box.wf("wait", "r-1", "--timeout", "1m") + self.assertEqual((code, err), (1, "wf: r-1 not done after 1m (running)\n")) + + def test_wait_killed_job(self): + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "t", "--", "x") + self.box.fake.units["wf-r-1.service"] = ("failed", None) + self.assertEqual(self.box.wf("wait", "r-1")[1], + "r-1 done rc=? peak ? in 0 min (no exit code; see journalctl --user -u wf-r-1.service)\n") + + def test_wait_oom_killed_job(self): + self.box.wf("run", "--mem", "4G", "--for", "5m", "--title", "t", "--", "x") + self.box.fake.units["wf-r-1.service"] = ("failed", None) + self.box.fake.journal["wf-r-1.service"] = ( + "wf-r-1.service: systemd-oomd killed some process(es) in this unit.\n" + "wf-r-1.service: Failed with result 'oom-kill'.\n" + "wf-r-1.service: Consumed 11min 2.837s CPU time, 3.9G memory peak.\n") + self.assertEqual(self.box.wf("wait", "r-1")[1], "r-1 done rc=? peak 3.9 GB in 0 min " + "(killed: oom-kill by systemd-oomd, limit 4.0 GB; raise --mem)\n") + j = self.box.fake.ran("journalctl")[0] + self.assertEqual(j, ["journalctl", "--user", "-u", "wf-r-1.service", "-o", "cat", "--no-pager", + "--since", "@" + str(int(dt.datetime(2026, 10, 1, 14, 0, tzinfo=UTC).timestamp()))]) + out = self.box.wf("status")[1] + self.assertEqual(out.splitlines()[0], 'r-1 proj "t" done (killed: oom-kill by systemd-oomd, limit 4.0 GB; ' + 'raise --mem) rc=? peak 3.9 GB at 14:00') + + def test_rc_file_skips_journal(self): + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "t", "--", "x") + (self.box.env.state / "logs" / "r-1.rc").write_text("2\n") + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + self.assertEqual(self.box.wf("wait", "r-1")[1], "r-1 done rc=2 peak ? in 0 min\n") + self.assertEqual(self.box.fake.ran("journalctl"), []) + +class QueueIO(IOBase): + def test_queue_then_start_when_free(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + code, out, _ = self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--queue", "--", "y") + self.assertEqual((code, out), (0, "r-2 queued, position 1; est. start ~14:40; cancel: wf res release r-2\n")) + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + self.box.wf("status") + e = self.box.ledger().get("r-2") + self.assertEqual((e.state, e.unit), ("running", "wf-r-2.service")) + self.assertEqual(self.box.fake.ran("systemd-run")[-1][-2:], ["sh", "y"]) + + def test_direct_run_does_not_jump_queue(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--queue", "--", "y") + code, out, _ = self.box.wf("run", "--mem", "1G", "--for", "10m", "--title", "small", "--", "z") + self.assertEqual(code, 3) + + def test_queued_start_failure_moves_on(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--queue", "--", "y") + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + self.box.fake.fail["systemd-run"] = "boom" + code, out, err = self.box.wf("status") + self.assertEqual(code, 0) + e = self.box.ledger().get("r-2") + self.assertEqual((e.state, e.why), ("done", "start failed: systemd-run: boom")) + self.box.fake.fail.clear() + self.assertEqual(self.box.wf("status")[0], 0) + +class Note(IOBase): + def test_note_and_owner_exit(self): + code, out, _ = self.box.wf("note", "--mem", "3G", "--for", "20m", "dotnet test") + self.assertEqual((code, out), (0, "r-1 noted 3.0 GB until ~14:20 (then freed; nothing is killed)\n")) + self.assertEqual(self.box.ledger().get("r-1").owner, 4242) + self.box.alive.clear() + self.box.wf("status") + self.assertEqual(self.box.ledger().get("r-1").why, "owner gone") + + def test_note_busy(self): + self.box.wf("note", "--mem", "10G", "--for", "20m", "a") + code, out, _ = self.box.wf("note", "--mem", "3G", "--for", "20m", "b") + self.assertEqual((code, out), (3, 'busy: 10.0 GB held by proj "a" (r-1) until ~14:20; 2.0 GB free for agents; ' + 'retry after ~14:20 or work on something else; 12.0 GB really free ' + 'beyond the reserve: --force starts it past the ledger (only if the ' + 'holders will not use what they reserved)\n')) + code, out, _ = self.box.wf("note", "--mem", "3G", "--for", "20m", "--force", "b") + self.assertEqual(code, 0) + self.assertTrue(out.startswith("r-2 noted 3.0 GB"), out) + + +def _race_child(tmp, q): + import time as _t + box = Box(Path(tmp), available_gb=14) # budget 14 − 6 − 2 = 6 GB: one 5G note fits, not two + slow = box.env.meminfo + box.env.meminfo = lambda: (_t.sleep(0.3), slow())[1] + with redirect_stdout(io.StringIO()), redirect_stderr(io.StringIO()): + q.put(wf_res.main(["note", "--mem", "5G", "--for", "10m", "x"], box.env)) + + +class LockRace(unittest.TestCase): + def test_two_processes_never_overbook(self): + import multiprocessing as mp + ctx = mp.get_context("fork") + with tempfile.TemporaryDirectory() as tmp: + q = ctx.Queue() + ps = [ctx.Process(target=_race_child, args=(tmp, q)) for _ in range(2)] + for p in ps: + p.start() + for p in ps: + p.join(10) + self.assertEqual([p.exitcode for p in ps], [0, 0]) # a crashed child must fail, not hang q.get() + self.assertEqual(sorted([q.get(timeout=5), q.get(timeout=5)]), [0, 3]) + +class GameIO(IOBase): + def props(self): + return [c[5:] for c in self.box.fake.ran("set-property")] + + def test_on_off(self): + code, out, _ = self.box.wf("game", "on") + self.assertEqual((code, out), (0, "game on until 18:00 (4h); CPU/IO now yours\n")) + self.assertEqual(self.props()[-1], ["CPUWeight=5", "IOWeight=5", f"MemoryHigh={18 * 1024 ** 3}"]) + self.assertEqual(self.box.fake.ran("set-property")[-1][:5], + ["systemctl", "--user", "set-property", "--runtime", "agents.slice"]) + self.assertEqual(self.box.wf("game", "off")[1], "game off\n") + self.assertEqual(self.props()[-1], ["CPUWeight=20", "IOWeight=20", f"MemoryHigh={24 * 1024 ** 3}"]) + + def test_expiry_restores(self): + self.box.wf("game", "on", "--for", "1h") + self.box.tick(3601) + self.box.wf("status") + self.assertIsNone(self.box.ledger().game_until) + self.assertEqual(self.props()[-1][0], "CPUWeight=20") + + def test_game_refuses_new_job(self): + self.box.wf("game", "on") + # 20 − 12 − 2 = 6 GB budget + self.assertEqual(self.box.wf("run", "--mem", "7G", "--for", "5m", "--title", "x", "--", "x")[0], 3) + + def test_on_again_extends(self): + self.box.wf("game", "on", "--for", "2h") + self.box.wf("game", "on", "--for", "1h") + self.assertEqual(self.box.ledger().game_until.hour, 16) + + +class CleanIO(IOBase): + def make(self, rel, age_h, size=10): + p = self.box.env.scratch / rel + p.parent.mkdir(parents=True, exist_ok=True) + p.write_bytes(b"x" * size) + t = self.box.clock[0].timestamp() - age_h * 3600 + os.utime(p, (t, t)) + for d in [p.parent, *p.parent.parents]: + if d == self.box.env.scratch.parent: + break + os.utime(d, (t, t)) + return p + + def test_auto_clean_sweeps_tmp_litter(self): + t = self.box.tmp / "tmpdir" + (t / "old-empty").mkdir(parents=True) + (t / "full").mkdir() + (t / "full" / "x").write_text("x") + (t / "clr-debug-pipe-999999-1-in").write_text("") + (t / "keep.txt").write_text("x") + old = self.box.clock[0].timestamp() - 3 * 3600 + for n in ("old-empty", "full", "keep.txt"): + os.utime(t / n, (old, old)) + self.box.env.tmp_dirs = [t] + self.box.env.pid_alive = lambda p: p != 999999 and p in self.box.alive + self.box.wf("status") + self.assertEqual(sorted(os.listdir(t)), ["full", "keep.txt"]) + + def test_auto_clean_keeps_live_session(self): + live = self.make("-projects-a/live-1/scratchpad/notes.txt", 9) + dead = self.make("-projects-a/dead-2/scratchpad/x.txt", 9) + self.box.env.live_sessions = lambda: {"live-1"} + self.box.wf("status") + self.assertEqual((live.exists(), dead.exists()), (True, False)) + + def test_auto_clean_on_status_and_throttle(self): + old = self.make("-projects-a/s1/scratchpad/big.bin", 3) + new = self.make("-projects-b/s2/scratchpad/x.txt", 0.5) + self.box.wf("status") + self.assertFalse(old.exists()) + self.assertTrue(new.exists()) + self.assertTrue(any(c[2] == "reset-failed" for c in self.box.fake.calls if c[0] == "systemctl")) + again = self.make("-projects-c/s3/a.txt", 3) + self.box.wf("status") + self.assertTrue(again.exists()) # < 10 min since last clean + self.box.tick(601) + self.box.wf("status") + self.assertFalse(again.exists()) + + def test_clean_prints_and_lists_project_patterns(self): + self.make("-projects-a/s1/f.bin", 3, size=2048) + proj = self.box.tmp / "proj" + (proj / "workflow.toml").write_text('cleanup = ["out/prof", "out/logs/*.log:30d"]\n') + (proj / "out" / "prof").mkdir(parents=True) + (proj / "out" / "prof" / "p.dat").write_bytes(b"x" * 1024) + logs = proj / "out" / "logs" + logs.mkdir() + (logs / "new.log").write_text("n") + old = logs / "old.log" + old.write_text("o") + t = self.box.clock[0].timestamp() - 31 * 86400 + os.utime(old, (t, t)) + (logs / "link.log").symlink_to("/etc/hostname") + code, out, _ = self.box.wf("clean") + self.assertEqual(out.splitlines(), [ + f"freed 0 MB: {self.box.env.scratch / '-projects-a'}", + "would delete 0 MB: out/prof (wf res clean --yes)", + "would delete 0 MB: out/logs/old.log (wf res clean --yes)"]) + self.assertTrue((proj / "out" / "prof").exists()) + self.box.wf("clean", "--yes") + self.assertFalse((proj / "out" / "prof").exists()) + self.assertFalse(old.exists()) + self.assertTrue((logs / "new.log").exists()) + self.assertTrue((logs / "link.log").is_symlink()) + + def test_status_warnings(self): + self.box.env.tmp_used_gb = lambda: 8.0 + self.box.env.procs = lambda: [(10, 1, "claude", "/u/app.slice/tab.scope"), + (20, 1, "claude", "/u/agents.slice/wf-claude-20.scope")] + out = self.box.wf("status")[1].splitlines() + self.assertEqual(out[-2:], ["warning: /tmp (RAM) holds 8.0 GB; wf res clean", + "warning: 1 claude sessions outside agents.slice (wf res adopt)"]) + +class TimerIO(IOBase): + def test_timer_on_off(self): + code, out, _ = self.box.wf("timer", "on") + self.assertEqual((code, out), (0, "timer on: wf-res.timer every 1 min\n")) + units = self.box.env.units + self.assertIn("res tick", (units / "wf-res.service").read_text()) + self.assertTrue((units / "wf-res.timer").exists()) + self.assertIn(["systemctl", "--user", "enable", "--now", "wf-res.timer"], self.box.fake.calls) + self.assertEqual(self.box.wf("timer", "off")[1], "timer off\n") + self.assertFalse((units / "wf-res.timer").exists()) + self.assertIn(["systemctl", "--user", "disable", "--now", "wf-res.timer"], self.box.fake.calls) + + def test_adopt(self): + self.box.env.procs = lambda: [(101, 1, "claude", "/u/app.slice/t.scope"), (102, 101, "bash", "/u/app.slice/t.scope"), + (201, 1, "claude", "/u/app.slice/u.scope")] + code, out, _ = self.box.wf("adopt") + self.assertEqual(out, "adopted claude 101 (2 processes)\nadopted claude 201 (1 processes)\n") + self.assertEqual(self.box.fake.ran("StartTransientUnit")[0][8], "wf-claude-101.scope") + + def test_adopt_failure_reported(self): + self.box.env.procs = lambda: [(101, 1, "claude", "/u/app.slice/t.scope")] + self.box.fake.fail["busctl"] = "Call failed: No such process" + self.assertEqual(self.box.wf("adopt")[1], "claude 101: Call failed: No such process\n") + + def test_adopt_none(self): + self.assertEqual(self.box.wf("adopt")[1], "all claude sessions already in agents.slice\n") + + def test_tick_adopts_silently(self): + self.box.env.procs = lambda: [(101, 1, "claude", "/u/app.slice/t.scope")] + self.assertEqual(self.box.wf("tick"), (0, "", "")) + self.assertEqual(len(self.box.fake.ran("StartTransientUnit")), 1) + + def test_tick_caps_sessions(self): + self.box.fake.scopes = {"wf-claude-5.scope": {"Slice": "agents.slice", "MemoryHigh": "infinity", + "MemoryCurrent": str(1024 ** 3), "ControlGroup": "/a/5"}, + "init.scope": {"Slice": "-.slice", "MemoryHigh": "infinity", "MemoryCurrent": "1"}} + self.assertEqual(self.box.wf("tick"), (0, "", "")) + self.assertEqual(self.box.fake.ran("set-property"), + [["systemctl", "--user", "set-property", "--runtime", "wf-claude-5.scope", + f"MemoryHigh={6 * 1024 ** 3}"]]) + self.box.wf("tick") + self.assertEqual(len(self.box.fake.ran("set-property")), 1) # already capped → no call + + def test_status_warns_session_at_cap(self): + self.box.fake.scopes = {"wf-claude-5.scope": {"Slice": "agents.slice", "MemoryHigh": str(6 * 1024 ** 3), + "MemoryCurrent": str(int(5.7 * 1024 ** 3)), "ControlGroup": "/a/5"}} + self.box.psi["/a/5"] = "some avg10=50.00 avg60=40.00 avg300=10.00 total=1\n" + self.assertIn("warning: claude session 5 at its memory cap (5.7 of 6.0 GB, stalled 40% of the last minute): " + "run big work with wf res run\n", self.box.wf("status")[1]) + + def test_tick_silent_and_starts_queue(self): + self.box.wf("run", "--mem", "10G", "--for", "40m", "--title", "big", "--", "x") + self.box.wf("run", "--mem", "8G", "--for", "10m", "--title", "two", "--queue", "--", "y") + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + self.assertEqual(self.box.wf("tick"), (0, "", "")) + self.assertEqual(self.box.ledger().get("r-2").state, "running") + + def test_shell_init(self): + self.assertEqual(self.box.wf("shell-init")[1], + "alias claude='systemd-run --user --scope --quiet --slice=agents.slice claude'\n") + + +class Robust(IOBase): + def test_vanished_paths_are_skipped(self): + gone = self.box.tmp / "gone" + self.assertEqual((wf_res._newest(gone), wf_res._size(gone)), (0.0, 0)) + + def test_os_error_is_one_line(self): + def boom(): + raise PermissionError(13, "Permission denied", "/proc/meminfo") + self.box.env.meminfo = boom + code, _, err = self.box.wf("status") + self.assertEqual((code, err), (1, "wf: [Errno 13] Permission denied: '/proc/meminfo'\n")) + + +if __name__ == "__main__": + unittest.main() + + +class LockIO(IOBase): + """--lock / 'gate …' titles: one job per lock key and main tree at a time (shared gate checkout).""" + + def gate(self, title, *extra): + return self.box.wf("run", "--mem", "2G", "--for", "30m", "--title", title, *extra, "--", "x") + + def test_second_gate_refused_even_with_force(self): + self.assertEqual(self.gate("gate c61907f0")[0], 0) + for extra in ((), ("--force",)): + code, out, err = self.gate("gate 2c91cac7 (fix)", *extra) + self.assertEqual(code, 3, err) + self.assertIn("lock 'gate' held by r-1", out + err) + self.assertIn("--queue", out + err) + self.assertEqual(len(self.box.fake.ran("systemd-run")), 1) + + def test_second_gate_queues_and_starts_after_first(self): + self.gate("gate c61907f0") + code, out, _ = self.gate("gate 2c91cac7", "--queue") + self.assertEqual(code, 0) + self.assertTrue(out.startswith("r-2 queued, position 1 (lock held by r-1); est. start ~14:30"), out) + self.box.wf("status") # memory is free, lock is not + self.assertEqual(self.box.ledger().get("r-2").state, "queued") + other = self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "build", "--", "y") + self.assertEqual(other[0], 0) # unlocked work is not held by the locked queue + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + self.box.wf("status") + self.assertEqual(self.box.ledger().get("r-2").state, "running") + + def test_two_queued_gates_start_one_at_a_time(self): + self.gate("gate a") + self.gate("gate b", "--queue") + self.gate("gate c", "--queue") + self.box.fake.units["wf-r-1.service"] = ("inactive", None) + self.box.wf("status") + led = self.box.ledger() + self.assertEqual([led.get(i).state for i in ("r-2", "r-3")], ["running", "queued"]) + + def test_explicit_lock_and_unlocked_titles(self): + self.box.wf("run", "--mem", "2G", "--for", "30m", "--title", "e2e", "--lock", "e2e", "--", "x") + self.assertEqual(self.box.wf("run", "--mem", "2G", "--for", "30m", "--title", "e2e 2", "--lock", "e2e", + "--", "x")[0], 3) + self.assertEqual(self.gate("gate a")[0], 0) # other key + self.assertEqual(self.box.wf("run", "--mem", "2G", "--for", "30m", "--title", "gatekeeper", "--", "x")[0], 0) + self.assertEqual(self.box.ledger().get("r-1").lock, f"e2e@{self.box.tmp / 'proj'}") + + def test_lock_scoped_to_main_tree(self): + self.gate("gate a") + self.box.fake.git["rev-parse"] = str(self.box.tmp / "other" / ".git") # another project's tree + self.assertEqual(self.gate("gate b")[0], 0) + self.box.fake.git["rev-parse"] = str(self.box.tmp / "proj" / ".git") # lane worktree of proj + self.assertEqual(self.gate("gate c")[0], 3) + + def test_project_is_main_tree_name_from_lane_worktree(self): + self.box.fake.git["rev-parse"] = str(self.box.tmp / "home" / ".git") # cwd = lane worktree 'proj' of home + self.box.wf("run", "--mem", "1G", "--for", "1m", "--title", "build 1", "--", "x") + self.box.fake.git["rev-parse"] = "" # no git → root name + self.box.wf("run", "--mem", "1G", "--for", "1m", "--title", "build 2", "--", "x") + self.box.fake.git["rev-parse"] = str(self.box.tmp / "home" / ".git" / "modules" / "m") # submodule → root + self.box.wf("run", "--mem", "1G", "--for", "1m", "--title", "build 3", "--", "x") + self.assertEqual([self.box.ledger().get(f"r-{i}").project for i in (1, 2, 3)], ["home", "proj", "proj"]) + + def test_percent_args_reach_systemd_verbatim(self): + # r-671 'fatal: ambiguous argument %s"': caller quoting, not wf; argv passes through untouched + self.box.wf("run", "--mem", "1G", "--for", "5m", "--title", "log", "--", "git", "log", "--format=%h %s", "HEAD") + self.assertEqual(self.box.fake.ran("systemd-run")[0][-4:], ["git", "log", "--format=%h %s", "HEAD"]) + + +def hist_rec(id, title="build 1234567", peak=1.0, minutes=10.0, project="proj", mem=4.0, est=40): + return {"id": id, "project": project, "title": title, "mem_gb": mem, "est_min": est, "rc": 0, + "peak_gb": peak, "min": minutes} + + +def seed_history(box, recs): + box.env.state.mkdir(parents=True, exist_ok=True) + (box.env.state / "resources-history.jsonl").write_text("".join(json.dumps(r) + "\n" for r in recs)) + + +class HistoryIO(IOBase): + def finish(self, id, peak_gb): + logs = self.box.env.state / "logs" + (logs / f"{id}.rc").write_text("0\n") + (logs / f"{id}.peak").write_text(str(int(peak_gb * 1024 ** 3)) + "\n") + self.box.fake.units[f"wf-{id}.service"] = ("inactive", None) + + def test_pruned_run_kept_in_history(self): + self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build abc1234", "--", "make") + self.box.tick(12 * 60) + self.finish("r-1", 3.0) + self.box.wf("status") + hist = self.box.env.state / "resources-history.jsonl" + self.assertFalse(hist.exists()) # still in the ledger: not copied yet + self.box.tick(25 * 3600) + self.box.wf("status") + self.assertEqual(self.box.ledger().entries, []) + self.assertEqual([json.loads(x) for x in hist.read_text().splitlines()], [ + {"id": "r-1", "project": "proj", "title": "build abc1234", "mem_gb": 4.0, "est_min": 40, "rc": 0, + "peak_gb": 3.0, "min": 12.0}]) + + def test_history_trimmed(self): + seed_history(self.box, [hist_rec(f"r-{i}") for i in range(100, 100 + R.HIST_KEEP)]) + self.box.wf("run", "--mem", "4G", "--for", "40m", "--title", "build", "--", "make") + self.finish("r-1", 1.0) + self.box.wf("status") + self.box.tick(25 * 3600) + self.box.wf("status") + lines = (self.box.env.state / "resources-history.jsonl").read_text().splitlines() + self.assertEqual(len(lines), R.HIST_KEEP) + self.assertEqual((json.loads(lines[0])["id"], json.loads(lines[-1])["id"]), ("r-101", "r-1")) + + def test_run_hint_over_history(self): + # peaks 1,2,3 → p95 3 ×1.15 = 3.45 → 3.5 GB; durations 10,20,30 → p90 30 ×1.5 = 45 min + seed_history(self.box, [hist_rec("r-90", peak=1.0, minutes=10.0), hist_rec("r-91", peak=2.0, minutes=20.0), + hist_rec("r-92", title="build 89abcde (retry)", peak=3.0, minutes=30.0)]) + code, out, err = self.box.wf("run", "--mem", "8G", "--for", "40m", "--title", "build 7654321", "--", "make") + self.assertEqual((code, err), (0, "hint: history says ~3.5 GB / 45 min (3 runs)\n")) + code, out, err = self.box.wf("run", "--mem", "7G", "--for", "90m", "--title", "build", "--", "make") + self.assertEqual(err, "") + code, out, err = self.box.wf("run", "--mem", "1G", "--for", "91m", "--title", "build", "--", "make") + self.assertEqual(err, "hint: history says ~3.5 GB / 45 min (3 runs)\n") + code, out, err = self.box.wf("note", "--mem", "8G", "--for", "10m", "build") + self.assertEqual(err, ("hint: history says ~3.5 GB / 45 min (3 runs)\n")) + code, out, err = self.box.wf("run", "--mem", "8G", "--for", "40m", "--title", "other", "--", "make") + self.assertEqual(err, "") + + def test_hist_command(self): + seed_history(self.box, [hist_rec("r-90", peak=1.0, minutes=10.0), hist_rec("r-91", peak=2.0, minutes=20.0), + hist_rec("r-92", peak=3.0, minutes=30.0), hist_rec("r-93", project="zzz")]) + code, out, err = self.box.wf("hist") + self.assertEqual((code, err), (0, "")) + self.assertEqual(out, "proj · build · n 3 · req 4.0 GB · peak 2.0/3.0 GB · est 40m · dur 20m/30m · suggest 3.5 GB 45m\n" + "zzz · build · n 1 · req 4.0 GB · peak 1.0/1.0 GB · est 40m · dur 10m/10m · suggest - (< 3 runs)\n") + self.assertEqual(self.box.wf("hist", "--project", "zzz")[1].count("\n"), 1) diff --git a/tests/test_runner.py b/tests/test_runner.py new file mode 100644 index 0000000..a095f2b --- /dev/null +++ b/tests/test_runner.py @@ -0,0 +1,81 @@ +import sys +import unittest +from pathlib import Path + +HERE = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(HERE)) +sys.path.insert(0, str(HERE / "tests")) + +import test_cli # noqa: E402 + +TASKS = """\ +# Tasks — demo + +## Awaiting your decision + +## Pending + +- **t-ok** [P1] (1h): Good. Headless. + Done: tests green. + +- **t-nodone** [P1] (1h): No done line. + +- **t-owner** [P1] (1h): Owner bound. + Done: tests green. + Sessions: owner + +- **t-confirm** [P1] (1h): Confirm. + Done: owner confirms it works. + +- **t-report** [P1] (1h): Report. + - Done: report to the owner. + +- **t-parent** [P1] (5h): Parent. + Done: all slices. + Slices: [[t-parent-1]] + +- **t-parent-1** [P1] (1h): Slice. + Done: ok. + Model: sonnet + +- **t-doneparent** [P1] (10h): Slices all archived. + Done: all slices. + - Slices: [[t-gone]] + +- **t-blk** [P1] (1h) (blocked: [[a-x]]): Blocked. + Done: ok. +""" + + +class Runner(test_cli.Cli): + tasks_text = TASKS + + def col(self, *args): + return [l.split()[0] for l in self.ok("list", *args).splitlines()[:-1]] + + def test_runner_filter(self): + self.assertEqual(self.col("--runner"), ["t-ok", "t-parent-1"]) + self.assertEqual(self.col("--runner", "--model", "sonnet"), ["t-parent-1"]) + + def test_add_hint_without_done(self): + code, out, err = self.wf("add", "-p", "2", "-e", "1h", "Plain thing", hints=True) + self.assertEqual((code, err), (0, "hint: no Done line; add one (--done) so runners can pick it\n")) + + def test_add_no_hint_with_done_or_awaiting(self): + code, out, err = self.wf("add", "-p", "2", "-e", "1h", "--body", "Body thing", stdin="Done: x\n", hints=True) + self.assertEqual((code, err), (0, "")) + code, out, err = self.wf("add", "-s", "awaiting", "Which one?", hints=True) + self.assertEqual((code, err), (0, "")) + + def test_add_done_flag_writes_done_line(self): + code, out, err = self.wf("add", "-p", "2", "-e", "1h", "--done", "gate ALL GREEN", "Flagged thing", hints=True) + self.assertEqual((code, err), (0, "")) + self.assertIn("\n Done: gate ALL GREEN\n", self.ok("show", "t-flagged-thing") + "\n") + + def test_add_p0_without_done_warns(self): + code, out, err = self.wf("add", "-p", "0", "-e", "1h", "Urgent thing", hints=True) + self.assertEqual((code, err), (0, 'wf: warning: P0 without Done line is not runner-pickable; use --done "<text>"\n')) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_search.py b/tests/test_search.py new file mode 100644 index 0000000..4c02520 --- /dev/null +++ b/tests/test_search.py @@ -0,0 +1,78 @@ +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import search as S +from wflib import tasks as T + +TASKS = """\ +## Awaiting your decision + +## Pending + +- **t-guards** [P2] (1h): Guards notice lock picking. Port the notify function. + - Steps: reaction drop + Ref: docs/ai.md + +- **t-throw** [P2] (5h): Thrown weapons and grenades. NPCs throw items. + - Steps: guard against throwing into allies + +- **t-upgrade** [P3] (1h): Shop upgrenade typo. Fix spelling. + +## Needs human + +## Deferred +""" +ARCHIVE = "# Archive\n\n- 2026-09-01 **t-wear** Weapon wear — guards bust weapons\n- old entry about grenades\n" +DOCS = { + "docs/ai.md": "# AI notes\n\nIntro.\n\n## Guards and thieves\n\nText about catching.\n\n## Combat\n\nA guard attacks. Grenade use.\n", +} + + +def run(*words, **kw): + return S.search(list(words), T.parse(TASKS), ARCHIVE, DOCS, **kw) + + +class SearchTest(unittest.TestCase): + def test_title_outranks_body(self): + hits = run("guard", kinds={"task"}) + self.assertEqual([h.where for h in hits], ["t-guards", "t-throw"]) + self.assertEqual(hits[0].line, "Guards notice lock picking. Port the notify function.") + self.assertEqual(hits[1].line, "- Steps: guard against throwing into allies") + + def test_prefix_matches_word_start_only(self): + self.assertEqual([h.where for h in run("grenad", kinds={"task"})], ["t-throw"]) + + def test_all_words_outrank_one_strong_word(self): + hits = run("throwing", "allies", "guards", kinds={"task"}) + self.assertEqual([h.where for h in hits], ["t-throw", "t-guards"]) + + def test_kinds_and_tie_order(self): + hits = run("guard") + self.assertEqual([(h.kind, h.where) for h in hits], + [("task", "t-guards"), ("doc", "docs/ai.md:5"), ("archive", "archive:3"), + ("task", "t-throw"), ("doc", "docs/ai.md:9")]) + self.assertEqual(hits[1].label, "Guards and thieves") + self.assertEqual(hits[4].line, "A guard attacks. Grenade use.") + + def test_archive_filter(self): + hits = run("grenades", kinds={"archive"}) + self.assertEqual([(h.where, h.line) for h in hits], [("archive:4", "old entry about grenades")]) + + def test_limit(self): + self.assertEqual(len(run("guard", limit=2)), 2) + + def test_case_and_regex_chars(self): + self.assertEqual([h.where for h in run("GUARD", kinds={"task"})], ["t-guards", "t-throw"]) + self.assertEqual(run("a.i", "(x"), []) + + def test_id_matches(self): + self.assertEqual([h.where for h in run("t-throw", kinds={"task"})], ["t-throw"]) + + def test_no_words(self): + self.assertEqual(run(), []) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_sessions_field.py b/tests/test_sessions_field.py new file mode 100644 index 0000000..af1cb62 --- /dev/null +++ b/tests/test_sessions_field.py @@ -0,0 +1,197 @@ +import json +import os +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import tasks as T +from wflib import lanes as L +from test_cli import Cli +from test_check import Base, CLEAN + +SESS = """\ +# Tasks — demo + +## Awaiting your decision + +## Pending + +- **t-solo** [P1] (1h): Solo. + - Sessions: solo — lib bump + Model: sonnet + +- **t-own** [P1] (1h, interactive): Own. + +- **t-owner** [P2] (1h): Owner. + Sessions: owner + +- **t-par** [P2] (<1h): Par. + Sessions: parallel + +- **t-plain** [P3] (<1h): Plain. + +## Needs human + +## Deferred +""" + + +class SessionsLineTest(unittest.TestCase): + def setUp(self): + self.doc = T.parse(SESS) + + def test_values(self): + self.assertEqual([(i.id, i.sessions) for i in self.doc.section("pending").items], + [("t-solo", "solo"), ("t-own", "owner"), ("t-owner", "owner"), + ("t-par", "parallel"), ("t-plain", "parallel")]) + + def test_set_replaces_in_place_and_keeps_tail_order(self): + T.set_fields(self.doc, "t-solo", sessions="owner") + self.assertEqual(self.doc.item("t-solo").body, [" Sessions: owner", " Model: sonnet"]) + + def test_set_adds_before_model(self): + T.set_fields(self.doc, "t-plain", sessions="solo") + self.assertEqual(self.doc.item("t-plain").body, [" Sessions: solo"]) + + def test_set_empty_removes(self): + T.set_fields(self.doc, "t-owner", sessions="") + self.assertEqual(self.doc.item("t-owner").body, []) + + def test_set_drops_interactive_flag(self): + T.set_fields(self.doc, "t-own", sessions="owner") + item = self.doc.item("t-own") + self.assertEqual(item.lines(), ["- **t-own** [P1] (1h): Own.", " Sessions: owner"]) + + def test_set_bad_value(self): + with self.assertRaises(T.TaskError): + T.set_fields(self.doc, "t-plain", sessions="many") + + def test_note_goes_before_sessions_line(self): + T.add_note(self.doc, "t-owner", "x") + self.assertEqual(self.doc.item("t-owner").body, [" - x", " Sessions: owner"]) + + +class SessionsPickTest(unittest.TestCase): + def setUp(self): + self.doc = T.parse(SESS) + + def pick(self, lane, model, **kw): + return L.pick(self.doc, set(), L.DEFAULT_LANES, "1h", lane, model, **kw) + + def test_alone_solo_picked_owner_skipped(self): + item, skipped = self.pick(None, "sonnet") + self.assertEqual(item.id, "t-solo") + item, skipped = self.pick("slow", "opus", others=1) + self.assertEqual((item.id, [(i.id, why) for i, why in skipped]), + ("t-par", [("t-solo", "solo: 1 other live session"), + ("t-own", "owner: needs the owner (wf next --owner)"), + ("t-owner", "owner: needs the owner (wf next --owner)")])) + + def test_owner_present(self): + item, _ = self.pick("slow", "opus", others=1, owner=True) + self.assertEqual(item.id, "t-own") + + def test_solo_skipped_with_other_live_sessions(self): + item, skipped = self.pick(None, "sonnet", others=1) + self.assertEqual((item, [(i.id, why) for i, why in skipped]), + (None, [("t-solo", "solo: 1 other live session")])) + _, skipped = self.pick(None, "sonnet", others=2) + self.assertEqual(skipped[0][1], "solo: 2 other live sessions") + + def test_solo_running(self): + self.assertEqual(T.solo_running(self.doc, {"t-solo": "sonnet session uds:/a"}), + ("t-solo", "sonnet session uds:/a")) + self.assertIsNone(T.solo_running(self.doc, {"t-par": "x"})) + self.assertIsNone(T.solo_running(self.doc, {})) + + def test_solo_done_block(self): + sessions = {"opus": {"socket": "/o", "alive": True, "pid": 1}, + "haiku": {"socket": "/h", "alive": False, "pid": 2}, + "sonnet": {"socket": "/s", "alive": True, "pid": 3}} + self.assertEqual(L.solo_done_block(["t-solo"], sessions, "3"), + ["notify opus uds:/o: solo t-solo done, run wf next"]) + self.assertEqual(L.solo_done_block([], sessions, "3"), []) + + +class SessionsCheckTest(Base): + def test_bad_word_error_interactive_warning(self): + text = CLEAN.replace("- **t-one** [P1] (1h)", "- **t-one** [P1] (1h, interactive)").replace( + "## Needs human", "- **t-x** [P3] (1h): X.\n Sessions: lots\n\n## Needs human") + self.tasks(text) + errors, warnings = self.run_check() + self.assertTrue(any(e.endswith("t-x: Sessions 'lots' (want parallel, solo, owner)") for e in errors), errors) + self.assertTrue(any(w.endswith("t-one: 'interactive' flag: write 'Sessions: owner' " + "(wf set t-one --sessions owner)") for w in warnings), warnings) + + +class SessionsCliTest(Cli): + tasks_text = SESS + + def env(self, name, pid): + sock = self.root / f"{name}.sock" + sock.write_text("") + return {"CLAUDE_CODE_MESSAGING_SOCKET": str(sock), "CLAUDE_PID": str(pid), "CLAUDE_CODE_SESSION_ID": name} + + def register(self, lane, model, sock, pid): + d = self.root / ".wf" / "sessions" + d.mkdir(parents=True, exist_ok=True) + (d / f"{lane}.json").write_text(json.dumps({"lane": lane, "model": model, "socket": str(sock), "pid": pid, + "session": "s", "at": "2026-10-04T10:00"})) + + def test_list_marks(self): + self.assertEqual(self.ok("list"), + "t-solo P1 1h - sonnet slow [solo] Solo\n" + "t-own P1 1h - opus slow [owner] Own\n" + "t-owner P2 1h - opus slow [owner] Owner\n" + "t-par P2 <1h - opus fast Par\n" + "t-plain P3 <1h - opus fast Plain\n" + "pending 5 · human 0 · awaiting 0 · deferred 0\n") + + def test_add_and_set(self): + self.ok("add", "Four.", "-p", "3", "-e", "1h", "--sessions", "solo", "--model", "haiku") + self.assertEqual(self.item("t-four"), "- **t-four** [P3] (1h): Four.\n Sessions: solo\n Model: haiku\n") + self.ok("set", "t-four", "--sessions", "") + self.assertEqual(self.item("t-four"), "- **t-four** [P3] (1h): Four.\n Model: haiku\n") + self.fails("set", "t-four", "--sessions", "many", code=2) + self.fails("add", "Five.", "-p", "3", "-e", "1h", "--sessions", "many", code=2) + + def test_interactive_option_is_old_spelling(self): + code, out, err = self.wf("add", "Four.", "-p", "3", "-e", "1h", "--interactive") + self.assertEqual((code, err), (0, "wf: --interactive is now --sessions owner\n")) + self.assertEqual(self.item("t-four"), "- **t-four** [P3] (1h): Four.\n Sessions: owner\n") + code, out, err = self.wf("set", "t-plain", "--interactive", "yes") + self.assertEqual((code, err), (0, "wf: --interactive is now --sessions owner\n")) + self.assertEqual(self.item("t-plain"), "- **t-plain** [P3] (<1h): Plain.\n Sessions: owner\n") + + def test_next_owner(self): + self.ok("done", "t-solo", "-m", "ok") + out = self.ok("next", "--lane", "slow", "--as", "opus", env={"CLAUDE_PID": "", "CLAUDE_CODE_MESSAGING_SOCKET": ""}) + self.assertIn("- t-own: owner: needs the owner (wf next --owner)\n", out) + self.assertIn("===== Next task =====\n- **t-par**", out) # slow empty → fallback fast + self.assertTrue(self.ok("next", "--lane", "slow", "--as", "opus", "--owner", "--brief").startswith("- **t-own**")) + + def test_next_solo_with_other_live_session(self): + sock = self.root / "o.sock" + sock.write_text("") + self.register("fast", "opus", sock, os.getppid()) + err = self.fails("next", "--as", "sonnet", env=self.env("s", os.getpid())) + self.assertEqual(err, "wf: nothing pickable for all lanes (sonnet) in Pending\n") + out = self.wf("next", "--as", "sonnet", env=self.env("s", os.getpid()))[1] + self.assertIn("- t-solo: solo: 1 other live session\n", out) + + def test_solo_in_progress_blocks_others_and_done_notifies(self): + s, o = self.env("s", os.getpid()), self.env("o", os.getppid()) + self.ok("next", "--lane", "slow", "--as", "sonnet", "--brief", env=s) + self.ok("status", "t-solo", "progress", "x", env=s) + code, out, err = self.wf("next", "--lane", "fast", "--as", "opus", env=o) + self.assertEqual((code, err), (1, f"wf: solo t-solo in progress by sonnet session uds:{self.root / 's.sock'}: " + "wait (its wf done notifies you)\n")) + self.assertNotIn("Next task", out) + out = self.ok("done", "t-solo", "-m", "ok", env=s) + self.assertIn(f"notify fast uds:{self.root / 'o.sock'}: solo t-solo done, run wf next\n", out) + self.assertTrue(self.ok("next", "--lane", "fast", "--as", "opus", "--brief", env=o).startswith("- **t-par**")) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_setup.py b/tests/test_setup.py new file mode 100644 index 0000000..ea44c33 --- /dev/null +++ b/tests/test_setup.py @@ -0,0 +1,164 @@ +import os +import sys +import tempfile +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import config as C +from test_claims import FOUR, git +from test_cli import TOML, Cli + +SETUP = TOML + 'worktree_setup = ["mkdir -p out && ln -sfn \\"$WF_MAIN/out/data\\" out/data", "echo ran >> log.txt"]\n' + + +class SetupConfigTest(unittest.TestCase): + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.root = Path(self.tmp.name).resolve() + + def tearDown(self): + self.tmp.cleanup() + + def load(self, extra): + (self.root / "workflow.toml").write_text('format = 1\ntasks = "T.md"\narchive = "a.md"\n' + extra) + return C.load(self.root) + + def test_default_empty(self): + self.assertEqual(self.load("").worktree_setup, []) + + def test_list_of_commands(self): + self.assertEqual(self.load('worktree_setup = ["a", "b c"]\n').worktree_setup, ["a", "b c"]) + + def test_wrong_type(self): + with self.assertRaisesRegex(C.ConfigError, "'worktree_setup' must be a list of strings"): + self.load('worktree_setup = "a"\n') + + def test_unknown_key_still_errors(self): + with self.assertRaisesRegex(C.ConfigError, "unknown key 'worktree_setp'"): + self.load('worktree_setp = []\n') + + +class GitCli(Cli): + tasks_text = FOUR + toml = SETUP + + def setUp(self): + super().setUp() + git(self.root, "init", "-q", "-b", "master") + (self.root / ".gitignore").write_text(".worktrees/\n.wf/\nout/\n") + git(self.root, "add", "-A") + git(self.root, "commit", "-qm", "init") + + +class WorktreeCli(GitCli): + def setUp(self): + super().setUp() + self.wt = self.root / ".worktrees" / "opus" + git(self.root, "worktree", "add", "-q", str(self.wt), "-b", "opus/t-three") + (self.root / "out" / "data").mkdir(parents=True) + (self.root / "out" / "data" / "x.bin").write_text("main data\n") + + def setup_(self, *args, cwd=None): + return self.wf("setup", *args, project=False, cwd=cwd or self.wt) + + +class SetupCliTest(WorktreeCli): + def test_runs_commands_in_worktree_with_wf_main(self): + code, out, err = self.setup_() + self.assertEqual((code, err), (0, ""), out) + self.assertEqual((self.wt / "out" / "data" / "x.bin").read_text(), "main data\n") + self.assertEqual((self.wt / "log.txt").read_text(), "ran\n") + self.assertFalse((self.root / "log.txt").exists()) + self.assertIn("$ echo ran >> log.txt\n", out) + self.assertTrue(out.endswith("worktree_setup: 2 commands ok in .worktrees/opus\n"), out) + + def test_idempotent_rerun(self): + self.assertEqual(self.setup_()[0], 0) + self.assertEqual(self.setup_()[0], 0) + self.assertEqual((self.wt / "out" / "data" / "x.bin").read_text(), "main data\n") + + def test_from_subfolder_runs_at_worktree_project_root(self): + code, out, err = self.setup_(cwd=self.wt / "docs") + self.assertEqual((code, err), (0, ""), out) + self.assertTrue((self.wt / "log.txt").is_file()) + + def test_nonzero_exit_reported_and_rest_skipped(self): + (self.wt / "workflow.toml").write_text(TOML + 'worktree_setup = ["echo a > a.txt", "exit 3", "echo c > c.txt"]\n') + code, out, err = self.setup_() + self.assertEqual(code, 1, out + err) + self.assertEqual(err, "wf: worktree_setup 'exit 3' failed (exit 3): later commands skipped\n") + self.assertTrue((self.wt / "a.txt").is_file()) + self.assertFalse((self.wt / "c.txt").exists()) + + def test_outside_worktree_fails(self): + code, out, err = self.wf("setup", project=False, cwd=self.root) + self.assertEqual((code, err), (1, "wf: setup runs inside a linked git worktree (lane worktree)\n")) + + def test_none_configured(self): + (self.wt / "workflow.toml").write_text(TOML) + code, out, err = self.setup_() + self.assertEqual((code, out, err), (0, "no worktree_setup in workflow.toml: nothing to do\n", "")) + + +class WorktreeConfigTest(WorktreeCli): + """A lane worktree's own workflow.toml / area file: setup, gate, areas read it; --mark writes it.""" + + def edit_wt_toml(self, extra): + (self.wt / "workflow.toml").write_text(TOML + extra) + + def test_setup_and_gate_read_worktree_toml(self): + self.edit_wt_toml('worktree_setup = ["echo branch > b.txt"]\nquick_gate = ["echo g > g.txt"]\n') + code, out, err = self.setup_() + self.assertEqual((code, err), (0, ""), out) + self.assertEqual((self.wt / "b.txt").read_text(), "branch\n") + self.assertFalse((self.wt / "log.txt").exists()) + code, out, err = self.wf("gate", project=False, cwd=self.wt) + self.assertEqual((code, err), (0, ""), out) + self.assertTrue((self.wt / "g.txt").is_file()) + + def test_books_stay_main_tree(self): + self.edit_wt_toml("") + (self.wt / "TASKS.md").write_text("# other\n") + out = self.ok("show", "t-three", project=False, cwd=self.wt) + self.assertIn("t-three", out) + cfg = C.load_at(self.wt) + self.assertEqual((cfg.root, cfg.local, cfg.tasks), (self.root, self.wt, self.root / "TASKS.md")) + self.assertEqual(cfg.areas_file, self.wt / "CLAUDE.md") + + def test_areas_read_and_mark_worktree_copy(self): + notes = "## Areas\n### Core\n- Code map: `nothing_here`\n" + (self.root / "CLAUDE.md").write_text(notes.replace("Core", "Main")) + (self.wt / "CLAUDE.md").write_text(notes) + self.assertIn("Core: ", self.ok("areas", project=False, cwd=self.wt)) + self.ok("areas", "--mark", "Core", project=False, cwd=self.wt) + self.assertIn("- Checked: ", (self.wt / "CLAUDE.md").read_text()) + self.assertEqual((self.root / "CLAUDE.md").read_text(), notes.replace("Core", "Main")) + self.assertIn("Main: ", self.ok("areas", project=False, cwd=self.root)) + + def test_main_tree_ignores_worktree(self): + self.edit_wt_toml('quick_gate = ["exit 9"]\n') + self.assertIsNone(C.find_local(self.root)) + code, out, err = self.wf("gate", project=False, cwd=self.root) + self.assertEqual((code, out), (0, "no quick_gate in workflow.toml: nothing to do\n")) + + +class SetupHintTest(GitCli): + def env(self, name, pid): + sock = self.root / f"{name}.sock" + sock.write_text("") + return {"CLAUDE_CODE_MESSAGING_SOCKET": str(sock), "CLAUDE_PID": str(pid), "CLAUDE_CODE_SESSION_ID": name} + + def test_multi_session_hint_adds_setup(self): + son, me = self.env("s", os.getppid()), self.env("o", os.getpid()) + self.ok("next", "--lane", "slow", "--as", "sonnet", "--brief", env=son) + out = self.ok("next", "--lane", "fast", "--as", "opus", env=me) + self.assertIn(" git worktree add .worktrees/fast -b fast/<task> master && cd .worktrees/fast && wf setup" + " (wf there writes this TASKS.md)\n", out) + (self.root / ".worktrees" / "fast").mkdir(parents=True) + out = self.ok("next", "--lane", "fast", "--as", "opus", env=me) + self.assertIn(" cd .worktrees/fast && git switch -c fast/<task> master && wf setup" + " (wf there writes this TASKS.md)\n", out) + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_start.py b/tests/test_start.py new file mode 100644 index 0000000..e1a2d6e --- /dev/null +++ b/tests/test_start.py @@ -0,0 +1,109 @@ +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from test_claims import git +from test_cli import WF +from test_setup import GitCli + +NONE = {"CLAUDE_CODE_MESSAGING_SOCKET": ""} + + +class StartTest(GitCli): + def setUp(self): + super().setUp() + self.wt = self.root / ".worktrees" / "slow" + + def start(self, *extra, code=0): + got, out, err = self.wf("start", "t-three", "--worktree", ".worktrees/slow", "--branch", "slow/t-three", + *extra, project=False, env=NONE) + self.assertEqual(got, code, out + err) + self.assertNotIn("Traceback", err) + return out, err + + def branch(self): + return (self.wt / ".git").is_file() and \ + subprocess_out(self.wt, "branch", "--show-current") + + def test_new_worktree_setup_progress_ctx_finish_line(self): + out, _ = self.start() + self.assertEqual(self.branch(), "slow/t-three") + self.assertEqual((self.wt / "log.txt").read_text(), "ran\n") # worktree_setup ran there + self.assertIn("**t-three** [P2] (<1h) (in progress: slow/t-three): Three.", self.tasks()) + self.assertIn("worktree: .worktrees/slow (new, branch slow/t-three from master)\n", out) + self.assertEqual(out.count("- **t-three**"), 1, out) + self.assertIn("worktree_setup: 2 commands ok in .worktrees/slow\n", out) + self.assertIn("- **t-three** [P2] (<1h) (in progress: slow/t-three): Three.\n\nSection: Pending\n", out) + self.assertTrue(out.endswith( + f"Finish (after the work; fill in the quoted parts and the paths):\n" + f" cd {self.wt} && (make test) && python3 {WF} finish t-three -m \"<entry>\" --commit \"<msg + footer>\" <paths>\n"), + out) + + def test_existing_branch_new_worktree_reports_wip(self): + git(self.root, "branch", "slow/t-three") + out, _ = self.start() + self.assertEqual(self.branch(), "slow/t-three") + self.assertIn("worktree: .worktrees/slow (new, existing branch slow/t-three: earlier WIP, read the notes)\n", + out) + + def test_existing_clean_worktree_switches(self): + git(self.root, "worktree", "add", "-q", "--detach", str(self.wt), "master") + out, _ = self.start() + self.assertEqual(self.branch(), "slow/t-three") + self.assertIn("worktree: .worktrees/slow (switched to new branch slow/t-three from master)\n", out) + (self.wt / "log.txt").unlink() # setup output, not ignored here + out, _ = self.start() # rerun: already there + self.assertIn("worktree: .worktrees/slow (on slow/t-three)\n", out) + + def test_dirty_worktree_refused_without_recovery(self): + git(self.root, "worktree", "add", "-q", "--detach", str(self.wt), "master") + (self.wt / "DESIGN.md").write_text("changed\n") + _, err = self.start(code=1) + self.assertEqual(err, "wf: worktree .worktrees/slow has uncommitted changes: M DESIGN.md " + "(hand back, or --recovery when a dead worker left them)\n") + self.assertNotIn("in progress", self.tasks()) + self.assertFalse((self.wt / "log.txt").exists()) + + def test_recovery_keeps_dirty_wip_and_shows_it(self): + git(self.root, "worktree", "add", "-q", str(self.wt), "-b", "slow/t-three") + (self.wt / "DESIGN.md").write_text("changed\n") + out, _ = self.start("--recovery") + self.assertEqual((self.wt / "DESIGN.md").read_text(), "changed\n") + self.assertIn("recovery: uncommitted:\n M DESIGN.md\n", out) + self.assertIn("recovery: commits master..HEAD: none\n", out) + self.assertIn("(in progress: slow/t-three)", self.tasks()) + + def test_recovery_dirty_on_other_branch_refused(self): + git(self.root, "worktree", "add", "-q", "--detach", str(self.wt), "master") + (self.wt / "DESIGN.md").write_text("changed\n") + _, err = self.start("--recovery", code=1) + self.assertIn("uncommitted changes on another branch", err) + + def test_setup_failure_stops_before_progress(self): + (self.root / "workflow.toml").write_text(self.toml.replace('"echo ran >> log.txt"', '"false"')) + git(self.root, "commit", "-qam", "red setup") + _, err = self.start(code=1) + self.assertIn("worktree_setup 'false' failed", err) + self.assertNotIn("in progress", self.tasks()) + + def test_unknown_id_creates_nothing(self): + got, out, err = self.wf("start", "t-nope", "--worktree", ".worktrees/slow", "--branch", "b", + project=False, env=NONE) + self.assertEqual(got, 1, out + err) + self.assertFalse(self.wt.exists()) + + def test_inside_worktree_refused(self): + git(self.root, "worktree", "add", "-q", "--detach", str(self.wt), "master") + got, out, err = self.wf("start", "t-three", "--worktree", str(self.wt), "--branch", "b", + project=False, cwd=self.wt, env=NONE) + self.assertEqual((got, err), (1, "wf: start runs in the main tree (it creates the worktree)\n")) + + +def subprocess_out(cwd, *args): + import subprocess + return subprocess.run(["git", "-C", str(cwd), *args], capture_output=True, text=True).stdout.strip() + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_tasks.py b/tests/test_tasks.py new file mode 100644 index 0000000..64f219c --- /dev/null +++ b/tests/test_tasks.py @@ -0,0 +1,523 @@ +import sys +import unittest +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) +from wflib import tasks as T +from wflib import lanes as L + +SAMPLE = """\ +# Tasks — demo + +Commands: `make test`. + +## Awaiting your decision + +- **a-smoke**: Smoke screenshots. Need you present. + +## Pending + +Intro prose. + +- **t-first** [P1] (<1h): First thing. Do it well. + - Steps: one + - After: [[t-zero]], [[t-other]] + Ref: DESIGN.md#terrain, docs/x.md (skill rows) + +- **t-second** [P2] (5h, interactive) (in progress: master): Second: with colon. Goal here. + +- **t-third** [P3] (1h) (blocked: [[a-smoke]]): Third. + +Trailing prose after items. + +## Needs human + +- **t-play** [P1] (<1h): Play-test feel. + - [ ] keyboard works + - [ ] motion smooth + +## Deferred + +## Known caveats + +- plain bullet, mentions [[t-first]] +""" + + +class HeaderTest(unittest.TestCase): + def setUp(self): + self.doc = T.parse(SAMPLE) + + def item(self, id): + section, i = self.doc.find(id) + return section.items[i] + + def test_task_with_priority_and_effort(self): + it = self.item("t-first") + self.assertEqual((it.id, it.prio, it.effort, it.interactive, it.status, it.text), + ("t-first", 1, "<1h", False, None, "First thing. Do it well.")) + + def test_interactive_and_in_progress(self): + it = self.item("t-second") + self.assertEqual((it.prio, it.effort, it.interactive, it.status), + (2, "5h", True, "in progress: master")) + + def test_blocked(self): + it = self.item("t-third") + self.assertEqual(it.status, "blocked: [[a-smoke]]") + self.assertEqual(it.blocked_on, "a-smoke") + self.assertIsNone(self.item("t-first").blocked_on) + + def test_awaiting_item_has_no_priority(self): + it = self.item("a-smoke") + self.assertEqual((it.prio, it.effort, it.text), (None, None, "Smoke screenshots. Need you present.")) + + def test_header_rebuilds_the_line(self): + self.assertEqual(self.item("t-first").header(), "- **t-first** [P1] (<1h): First thing. Do it well.") + self.assertEqual(self.item("t-second").header(), + "- **t-second** [P2] (5h, interactive) (in progress: master): Second: with colon. Goal here.") + self.assertEqual(self.item("t-third").header(), "- **t-third** [P3] (1h) (blocked: [[a-smoke]]): Third.") + self.assertEqual(self.item("a-smoke").header(), "- **a-smoke**: Smoke screenshots. Need you present.") + + def test_title_and_goal(self): + it = self.item("t-second") + self.assertEqual(it.title, "Second: with colon") + self.assertEqual(it.goal, "Goal here.") + self.assertEqual(self.item("t-third").title, "Third") + self.assertEqual(self.item("t-third").goal, "") + + def test_after(self): + self.assertEqual(self.item("t-first").after, ["t-zero", "t-other"]) + self.assertEqual(self.item("t-second").after, []) + + def test_refs(self): + self.assertEqual(self.item("t-first").refs, [("DESIGN.md", "terrain"), ("docs/x.md", None)]) + self.assertEqual(self.item("t-second").refs, []) + + def test_body_lines(self): + self.assertEqual(self.item("t-play").body, [" - [ ] keyboard works", " - [ ] motion smooth"]) + + +class StructureTest(unittest.TestCase): + def test_round_trip(self): + self.assertEqual(T.render(T.parse(SAMPLE)), SAMPLE) + + def test_sections_and_keys(self): + doc = T.parse(SAMPLE) + self.assertEqual([(s.heading, s.key) for s in doc.sections], + [("Awaiting your decision", "awaiting"), ("Pending", "pending"), + ("Needs human", "human"), ("Deferred", "deferred"), ("Known caveats", None)]) + self.assertEqual([i.id for i in doc.section("pending").items], ["t-first", "t-second", "t-third"]) + self.assertEqual(doc.section("pending").prefix, ["", "Intro prose.", ""]) + self.assertEqual(doc.section("pending").suffix, ["Trailing prose after items.", ""]) + self.assertEqual(doc.ids(), {"a-smoke", "t-first", "t-second", "t-third", "t-play"}) + + def test_prose_section_bullets_are_not_items(self): + doc = T.parse(SAMPLE) + self.assertEqual(doc.sections[-1].items, []) + self.assertEqual(doc.sections[-1].prefix, ["", "- plain bullet, mentions [[t-first]]"]) + + def test_missing_section_raises(self): + doc = T.parse("# x\n\n## Pending\n") + with self.assertRaisesRegex(T.TaskError, "no '## Deferred' section"): + doc.section("deferred") + + def test_find_unknown_id_names_nearest(self): + with self.assertRaisesRegex(T.TaskError, r"unknown id 't-frist' \(nearest: t-first"): + T.parse(SAMPLE).find("t-frist") + + def test_items_are_separated_by_one_blank_line(self): + text = "## Pending\n- **t-a** [P1] (1h): A.\n\n\n\n- **t-b** [P1] (1h): B.\n body\n- **t-c** [P1] (1h): C.\n" + self.assertEqual(T.render(T.parse(text)), + "## Pending\n- **t-a** [P1] (1h): A.\n\n- **t-b** [P1] (1h): B.\n body\n\n- **t-c** [P1] (1h): C.\n") + + def test_malformed_header_is_kept_and_flagged(self): + text = "## Pending\n\n- **t-x** [P9] (1h) no colon\n body\n\n- **t-ok** [P1] (1h): Fine.\n" + doc = T.parse(text) + bad = doc.section("pending").items[0] + self.assertEqual(bad.raw, "- **t-x** [P9] (1h) no colon") + self.assertIn("header", bad.error) + self.assertEqual(bad.id, "t-x") + self.assertEqual(T.render(doc), text) + + def test_flush_left_line_ends_the_items_and_survives(self): + text = ("## Pending\n\n- **t-a** [P1] (1h): A.\n body\nstray continuation\n" + "- **t-b** [P1] (1h): B.\n\n## Deferred\n") + doc = T.parse(text) + pending = doc.section("pending") + self.assertEqual([i.id for i in pending.items], ["t-a"]) + self.assertEqual(pending.suffix, ["stray continuation", "- **t-b** [P1] (1h): B.", ""]) + self.assertEqual(T.render(doc), + "## Pending\n\n- **t-a** [P1] (1h): A.\n body\n\nstray continuation\n" + "- **t-b** [P1] (1h): B.\n\n## Deferred\n") + + def test_crlf_is_kept(self): + text = "## Pending\r\n\r\n- **t-a** [P1] (1h): A.\r\n body\r\n" + doc = T.parse(text) + self.assertEqual(doc.newline, "\r\n") + self.assertEqual(doc.section("pending").items[0].body, [" body"]) + self.assertEqual(T.render(doc), text) + + def test_missing_final_newline_is_added(self): + self.assertEqual(T.render(T.parse("## Pending\n\n- **t-a** [P1] (1h): A.")), + "## Pending\n\n- **t-a** [P1] (1h): A.\n") + + +if __name__ == "__main__": + unittest.main() + + +EDIT = """\ +## Awaiting your decision + +- **a-key**: Key needed. Which one? + +## Pending + +- **t-p0** [P0] (1h): Zero. + +- **t-p1** [P1] (1h): One. + - Steps: a + - After: [[t-p0]] + Ref: docs/a.md + +- **t-p2** [P2] (1h) (blocked: [[a-key]]): Two. + +- **t-p3** [P3] (1h): Three. + - After: [[t-gone]] + +## Needs human + +- **t-play** [P1] (<1h): Play. + - [ ] keyboard works + - [x] motion smooth + - [ ] gait feels right + +## Deferred +""" + + +def ids(doc, key): + return [i.id for i in doc.section(key).items] + + +class MakeIdTest(unittest.TestCase): + def test_slug_of_title(self): + self.assertEqual(T.make_id("Guards notice lock picking and busting", set()), + "t-guards-notice-lock-picking-and-busting") + + def test_long_title_is_cut_at_a_word(self): + self.assertEqual(T.make_id("Exe function map plus port re-verification of everything", set()), + "t-exe-function-map-plus-port-re") + self.assertEqual(T.make_id("Thrown weapons and grenades for everyone here", set()), + "t-thrown-weapons-and-grenades-for-everyone") + + def test_taken_id_gets_a_number(self): + self.assertEqual(T.make_id("Cave seams", {"t-cave-seams"}), "t-cave-seams-2") + self.assertEqual(T.make_id("Cave seams", {"t-cave-seams", "t-cave-seams-2"}), "t-cave-seams-3") + + def test_empty_slug_is_refused(self): + with self.assertRaisesRegex(T.TaskError, "--id"): + T.make_id("???", set()) + + def test_non_ascii_is_dropped(self): + self.assertEqual(T.make_id("Åäö test", set()), "t-test") + + def test_awaiting_prefix(self): + self.assertEqual(T.make_id("Smoke screenshots", set(), "a-"), "a-smoke-screenshots") + + def test_slice_id(self): + self.assertEqual(T.slice_id("t-map", {"t-map"}), "t-map-1") + self.assertEqual(T.slice_id("t-map", {"t-map", "t-map-1", "t-map-2"}), "t-map-3") + + +class EditTest(unittest.TestCase): + def setUp(self): + self.doc = T.parse(EDIT) + + def new(self, id, prio, body=()): + return T.Item(id=id, prio=prio, effort="1h", text="New.", body=list(body)) + + def test_insert_by_priority(self): + T.insert(self.doc, self.new("t-n", 1), "pending") + self.assertEqual(ids(self.doc, "pending"), ["t-p0", "t-p1", "t-n", "t-p2", "t-p3"]) + + def test_insert_into_empty_section(self): + T.insert(self.doc, self.new("t-n", 0), "deferred") + self.assertEqual(ids(self.doc, "deferred"), ["t-n"]) + + def test_insert_goes_after_its_dependency(self): + T.insert(self.doc, self.new("t-n", 1, [" - After: [[t-p3]]"]), "pending") + self.assertEqual(ids(self.doc, "pending"), ["t-p0", "t-p1", "t-p2", "t-p3", "t-n"]) + + def test_insert_refuses_taken_id(self): + with self.assertRaisesRegex(T.TaskError, "'t-p1' already exists"): + T.insert(self.doc, self.new("t-p1", 1), "pending") + + def test_insert_refuses_wrong_kind_for_section(self): + with self.assertRaisesRegex(T.TaskError, "a- items belong in Awaiting"): + T.insert(self.doc, T.Item(id="a-x", text="X."), "pending") + with self.assertRaisesRegex(T.TaskError, "needs a priority"): + T.insert(self.doc, T.Item(id="t-x", effort="1h", text="X."), "pending") + + def test_remove(self): + item = T.remove(self.doc, "t-p1") + self.assertEqual(item.body, [" - Steps: a", " - After: [[t-p0]]", " Ref: docs/a.md"]) + self.assertEqual(ids(self.doc, "pending"), ["t-p0", "t-p2", "t-p3"]) + + def test_remove_unknown_names_nearest(self): + with self.assertRaisesRegex(T.TaskError, r"unknown id 't-p9' \(nearest: t-p"): + T.remove(self.doc, "t-p9") + + def test_set_prio_moves_the_item_and_keeps_its_body(self): + T.set_prio(self.doc, "t-p1", 3) + self.assertEqual(ids(self.doc, "pending"), ["t-p0", "t-p2", "t-p3", "t-p1"]) + item = self.doc.item("t-p1") + self.assertEqual((item.prio, item.body[0]), (3, " - Steps: a")) + + def test_set_prio_range(self): + with self.assertRaisesRegex(T.TaskError, "priority 0-3"): + T.set_prio(self.doc, "t-p1", 4) + + def test_move_to_section(self): + T.move_to(self.doc, "t-p3", "deferred") + self.assertEqual(ids(self.doc, "pending"), ["t-p0", "t-p1", "t-p2"]) + self.assertEqual(ids(self.doc, "deferred"), ["t-p3"]) + + def test_move_to_refuses_kind_change(self): + with self.assertRaisesRegex(T.TaskError, "a- items belong in Awaiting"): + T.move_to(self.doc, "a-key", "pending") + + def test_move_rel_refuses_priority_break(self): + with self.assertRaisesRegex(T.TaskError, "breaks priority order.*--force"): + T.move_rel(self.doc, "t-p3", "t-p0", before=True) + self.assertEqual(ids(self.doc, "pending"), ["t-p0", "t-p1", "t-p2", "t-p3"]) + + def test_move_rel_with_force(self): + T.move_rel(self.doc, "t-p3", "t-p0", before=True, force=True) + self.assertEqual(ids(self.doc, "pending"), ["t-p3", "t-p0", "t-p1", "t-p2"]) + + def test_move_rel_never_puts_an_item_before_its_dependency(self): + with self.assertRaisesRegex(T.TaskError, "'t-p1' is After: \\[\\[t-p0\\]\\]"): + T.move_rel(self.doc, "t-p1", "t-p0", before=True, force=True) + + def test_move_rel_after_into_other_section(self): + T.move_rel(self.doc, "t-p1", "t-play", before=False) + self.assertEqual(ids(self.doc, "human"), ["t-play", "t-p1"]) + + def test_status(self): + T.set_status(self.doc, "t-p1", "in progress: feature/x") + self.assertEqual(self.doc.item("t-p1").header(), "- **t-p1** [P1] (1h) (in progress: feature/x): One.") + T.set_status(self.doc, "t-p1", "blocked: [[a-key]]") + self.assertEqual(self.doc.item("t-p1").blocked_on, "a-key") + T.set_status(self.doc, "t-p1", None) + self.assertEqual(self.doc.item("t-p1").header(), "- **t-p1** [P1] (1h): One.") + + def test_status_blocked_on_unknown_awaiting_item(self): + with self.assertRaisesRegex(T.TaskError, "'a-none' is not an open Awaiting item"): + T.set_status(self.doc, "t-p1", "blocked: [[a-none]]") + + def test_set_fields_header(self): + T.set_fields(self.doc, "t-p1", title="Uno", effort="5h", interactive=True) + self.assertEqual(self.doc.item("t-p1").header(), "- **t-p1** [P1] (5h, interactive): Uno.") + T.set_fields(self.doc, "t-p0", title="Nil") + self.assertEqual(self.doc.item("t-p0").text, "Nil.") + + def test_set_fields_title_keeps_goal(self): + doc = T.parse("## Pending\n\n- **t-a** [P1] (1h): Old name. The goal. More.\n") + T.set_fields(doc, "t-a", title="New name") + self.assertEqual(doc.item("t-a").text, "New name. The goal. More.") + + def test_set_fields_full_text_replaces_goal(self): + doc = T.parse("## Awaiting your decision\n\n- **a-q** : Colour. OK?\n\n## Pending\n\n" + "- **t-a** [P1] (1h): Old name. Old goal.\n".replace(" :", ":")) + T.set_fields(doc, "a-q", title="Colour red. OK?") + self.assertEqual(doc.item("a-q").text, "Colour red. OK?") + T.set_fields(doc, "t-a", title="New name. New goal") + self.assertEqual(doc.item("t-a").text, "New name. New goal.") + T.set_fields(doc, "t-a", title="Why not?") + self.assertEqual(doc.item("t-a").text, "Why not?") + + def test_set_fields_bad_effort(self): + with self.assertRaisesRegex(T.TaskError, "effort '2h'"): + T.set_fields(self.doc, "t-p1", effort="2h") + + def test_set_after(self): + T.set_fields(self.doc, "t-p1", after=[]) + self.assertEqual(self.doc.item("t-p1").body, [" - Steps: a", " Ref: docs/a.md"]) + T.set_fields(self.doc, "t-p1", after=["t-p0", "t-p2"]) + self.assertEqual(self.doc.item("t-p1").body, + [" - Steps: a", " - After: [[t-p0]], [[t-p2]]", " Ref: docs/a.md"]) + + def test_set_done(self): + doc = T.parse("## Pending\n\n- **t-a** [P1] (1h): A.\n - Steps: a\n - note\n Model: sonnet\n Ref: x.md\n") + T.set_fields(doc, "t-a", done="a works") + self.assertEqual(doc.item("t-a").body, + [" - Steps: a", " - note", " Done: a works", " Model: sonnet", " Ref: x.md"]) + T.set_fields(doc, "t-a", done=" b works ") + self.assertEqual(doc.item("t-a").body[2], " Done: b works") + T.set_fields(doc, "t-a", done="") + self.assertEqual(doc.item("t-a").body, [" - Steps: a", " - note", " Model: sonnet", " Ref: x.md"]) + + def test_set_refs(self): + T.set_fields(self.doc, "t-p1", refs=["docs/b.md#x (rows)", "DESIGN.md"]) + self.assertEqual(self.doc.item("t-p1").body[-1], " Ref: docs/b.md#x (rows), DESIGN.md") + T.set_fields(self.doc, "t-p0", refs=["docs/a.md"]) + self.assertEqual(self.doc.item("t-p0").body, [" Ref: docs/a.md"]) + T.set_fields(self.doc, "t-p0", refs=[]) + self.assertEqual(self.doc.item("t-p0").body, []) + + def test_add_note_goes_before_after_and_ref(self): + T.add_note(self.doc, "t-p1", "found the cause") + self.assertEqual(self.doc.item("t-p1").body, + [" - Steps: a", " - found the cause", " - After: [[t-p0]]", " Ref: docs/a.md"]) + + def test_set_body_keeps_after_and_ref(self): + T.set_body(self.doc, "t-p1", ["- Steps: b", " - sub", "", "- Done: c"]) + self.assertEqual(self.doc.item("t-p1").body, + [" - Steps: b", " - sub", "", " - Done: c", " - After: [[t-p0]]", " Ref: docs/a.md"]) + + def test_set_body_bare_after_ids_become_links(self): + T.set_body(self.doc, "t-p1", ["- Done: c", "- After: t-a, [[t-b]]"]) + self.assertEqual(self.doc.item("t-p1").body[:2], [" - Done: c", " - After: [[t-a]], [[t-b]]"]) + + def test_tick_by_number_and_text(self): + self.assertEqual(T.tick(self.doc, "t-play", "1"), "keyboard works") + self.assertEqual(T.tick(self.doc, "t-play", "gait"), "gait feels right") + self.assertEqual(self.doc.item("t-play").body, + [" - [x] keyboard works", " - [x] motion smooth", " - [x] gait feels right"]) + + def test_tick_errors(self): + with self.assertRaisesRegex(T.TaskError, "already ticked"): + T.tick(self.doc, "t-play", "2") + with self.assertRaisesRegex(T.TaskError, "no box matches 'jump'"): + T.tick(self.doc, "t-play", "jump") + with self.assertRaisesRegex(T.TaskError, "no box 7"): + T.tick(self.doc, "t-play", "7") + + def test_unblock(self): + self.assertEqual(T.unblock(self.doc, "a-key"), ["t-p2"]) + self.assertIsNone(self.doc.item("t-p2").status) + + +class PickTest(unittest.TestCase): + def test_skips_blocked_and_open_dependencies(self): + doc = T.parse(EDIT) + T.remove(doc, "t-p0") + T.remove(doc, "t-p1") + item, skipped = L.pick(doc, set(), L.DEFAULT_LANES, "1h") + self.assertIsNone(item) + self.assertEqual([(i.id, why) for i, why in skipped], + [("t-p2", "blocked: a-key"), ("t-p3", "after: t-gone")]) + + def test_archived_dependency_counts_as_done(self): + doc = T.parse(EDIT) + T.remove(doc, "t-p0") + T.remove(doc, "t-p1") + item, skipped = L.pick(doc, {"t-gone"}, L.DEFAULT_LANES, "1h") + self.assertEqual(item.id, "t-p3") + self.assertEqual([i.id for i, _ in skipped], ["t-p2"]) + + def test_first_item_when_free(self): + item, skipped = L.pick(T.parse(EDIT), set(), L.DEFAULT_LANES, "1h") + self.assertEqual((item.id, skipped), ("t-p0", [])) + + def test_open_dependency_in_file(self): + doc = T.parse(EDIT) + T.move_to(doc, "t-p0", "deferred") + item, skipped = L.pick(doc, set(), L.DEFAULT_LANES, "1h") + self.assertIsNone(item) + self.assertEqual(skipped[0][1], "after: t-p0") + + +class SliceTest(unittest.TestCase): + def test_open_slices(self): + doc = T.parse("## Pending\n\n- **t-p** [P1] (10h): Parent.\n\n- **t-p-2** [P1] (1h): Two.\n\n" + "- **t-px-1** [P1] (1h): Other.\n") + self.assertEqual(T.open_slices(doc, "t-p"), ["t-p-2"]) + + def test_open_slices_by_name(self): + doc = T.parse(SLICED) + self.assertEqual(T.open_slices(doc, "t-p"), ["t-map", "t-p-3"]) + self.assertEqual(T.parent_of(doc, "t-map"), "t-p") + self.assertEqual(T.parent_of(doc, "t-p-3"), "t-p") + self.assertIsNone(T.parent_of(doc, "t-q")) + + def test_pick_skips_parent_with_open_slices(self): + item, skipped = L.pick(T.parse(SLICED), {"t-done"}, L.DEFAULT_LANES, "1h") + self.assertEqual(item.id, "t-q") + self.assertEqual([(i.id, why) for i, why in skipped], + [("t-map", "after: t-x"), ("t-p-3", "after: t-map"), ("t-p", "open slices: t-map, t-p-3")]) + + +class RenameTest(unittest.TestCase): + def test_rename_changes_id_and_links(self): + doc = T.parse(SLICED) + n = T.rename(doc, "t-map", "t-map-scaffold", taken={"t-old"}) + self.assertEqual(n, 2) + self.assertEqual(T.render(doc), SLICED.replace("t-map", "t-map-scaffold")) + + def test_rename_awaiting_updates_blocked_status(self): + doc = T.parse("## Awaiting your decision\n\n- **a-a-foo**: Foo?\n\n## Pending\n\n" + "- **t-a** [P1] (1h) (blocked: [[a-a-foo]]): A.\n") + self.assertEqual(T.rename(doc, "a-a-foo", "a-foo", taken=set()), 1) + self.assertEqual(doc.item("t-a").blocked_on, "a-foo") + + def test_rename_refusals(self): + doc = T.parse(SLICED) + for new, err in [("t-q", "id 't-q' already exists"), ("t-old", "id 't-old' already used in the archive"), + ("a-map", "a rename keeps the kind"), ("t-Bad", "bad id 't-Bad'")]: + with self.assertRaisesRegex(T.TaskError, err): + T.rename(doc, "t-map", new, taken={"t-old"}) + with self.assertRaisesRegex(T.TaskError, "unknown id 't-none'"): + T.rename(doc, "t-none", "t-x", taken=set()) + self.assertEqual(T.render(doc), SLICED) + + +SLICED = """\ +## Pending + +- **t-p** [P1] (10h): Parent. + - Slices: [[t-done]], [[t-map]], [[t-p-3]] + +- **t-map** [P1] (1h): Map. + - After: [[t-x]] + +- **t-p-3** [P1] (1h): Three. + - After: [[t-map]] + +- **t-q** [P2] (1h): Free. +""" + + +class ArchiveTest(unittest.TestCase): + def test_line(self): + item = T.Item(id="t-x", prio=1, effort="1h", text="Cave seams. Close the slit.") + self.assertEqual(T.archive_line("2026-09-29", item, "closed with side walls"), + "- 2026-09-29 **t-x** Cave seams — closed with side walls") + + def test_prepend(self): + self.assertEqual(T.archive_prepend("# Archive (newest first)\n\n- old one\n- older\n", "- new"), + "# Archive (newest first)\n\n- new\n- old one\n- older\n") + + def test_prepend_to_empty_list(self): + self.assertEqual(T.archive_prepend("# Archive (newest first)\n", "- new"), + "# Archive (newest first)\n\n- new\n") + + def test_ids(self): + text = "# A\n\n- 2026-09-29 **t-x** X — done\n- old line without id\n- 2026-09-01 **t-y-2** Y — z\n" + self.assertEqual(T.archive_ids(text), {"t-x", "t-y-2"}) + + +class ParseBlockTest(unittest.TestCase): + def test_block_is_reindented(self): + item = T.parse_block("- **t-n** [P2] (1h): New. Goal.\n - Steps: a\n - sub\n Ref: docs/a.md\n") + self.assertEqual(item.lines(), ["- **t-n** [P2] (1h): New. Goal.", " - Steps: a", " - sub", " Ref: docs/a.md"]) + + def test_old_numbered_format_is_refused(self): + with self.assertRaisesRegex(T.TaskError, "old numbered format"): + T.parse_block("3. **[P2] Old** (Effort: 1h) — goal.") + + def test_bad_header_is_refused(self): + with self.assertRaisesRegex(T.TaskError, "bad header"): + T.parse_block("- **t-n** no colon") diff --git a/tests/test_usage.py b/tests/test_usage.py new file mode 100644 index 0000000..f2d55e3 --- /dev/null +++ b/tests/test_usage.py @@ -0,0 +1,337 @@ +import json +import os +import subprocess +import sys +import tempfile +import unittest +from pathlib import Path + +HERE = Path(__file__).resolve().parent.parent +WF = HERE / "wf.py" +sys.path.insert(0, str(HERE)) + +from wflib import usage # noqa: E402 + + +def entry(mid, req, model, ts, inp=0, cw5=0, cw1h=0, cr=0, out=0, type="assistant", stop=None, content=None, + uuid=None): + u = {"input_tokens": inp, "cache_creation_input_tokens": cw5 + cw1h, "cache_read_input_tokens": cr, + "output_tokens": out} + if cw1h: + u["cache_creation"] = {"ephemeral_5m_input_tokens": cw5, "ephemeral_1h_input_tokens": cw1h} + d = {"type": type, "requestId": req, "timestamp": ts, + "message": {"id": mid, "model": model, "role": "assistant", "usage": u}} + if stop: + d["message"]["stop_reason"] = stop + if content is not None: + d["message"]["content"] = content + if uuid: + d["uuid"] = uuid + return json.dumps(d) + + +# Subagent transcripts: stop_reason always null, output_tokens = message_start placeholder. +# e1: thinking (signature 1800 chars) + text 1000 chars + tool_use (input json 100 chars), placeholder 8. +# by hand: 0.3*(1800-800) + 0.3*1000 + (30 + 0.44*100) = 300 + 300 + 74 = 674 +# e2: same tool_use again but stop_reason set (final): reported 50 is trusted. +# e3: no content, placeholder 9 kept (nothing to estimate). +TOOL = {"type": "tool_use", "id": "t1", "name": "Bash", "input": {"a": "y" * 91}} # json.dumps → 100 chars +SUB = "\n".join([ + entry("e1", "q1", "claude-sonnet-5-5", "2026-10-04T10:00:00Z", out=8, uuid="u1", + content=[{"type": "thinking", "thinking": "", "signature": "z" * 1800}]), + entry("e1", "q1", "claude-sonnet-5-5", "2026-10-04T10:00:01Z", out=8, uuid="u2", + content=[{"type": "text", "text": "x" * 1000}]), + entry("e1", "q1", "claude-sonnet-5-5", "2026-10-04T10:00:01Z", out=8, uuid="u2", + content=[{"type": "text", "text": "x" * 1000}]), # duplicated line: counted once + entry("e1", "q1", "claude-sonnet-5-5", "2026-10-04T10:00:02Z", out=8, uuid="u3", content=[TOOL]), + entry("e2", "q2", "claude-sonnet-5-5", "2026-10-04T10:01:00Z", out=50, stop="tool_use", content=[TOOL]), + entry("e3", "q3", "claude-sonnet-5-5", "2026-10-04T10:02:00Z", out=9), +]) + "\n" + + +# main session, opus 5.5: m1 streamed as 2 entries (same usage), m2 as 2 entries (out 4000 then 10000). +# by hand: m1 = 1000*4 + 100000*5 + 2000*20 = 544000 µ$; m2 = 50000*8 + 1e6*0.2 + 10000*20 = 800000 µ$ +MAIN = "\n".join([ + json.dumps({"type": "user", "timestamp": "2026-10-04T09:00:00Z", "message": {"role": "user", "content": "hi"}}), + entry("m1", "r1", "claude-opus-5-5", "2026-10-04T09:00:01Z", inp=1000, cw5=100000, out=2000), + entry("m1", "r1", "claude-opus-5-5", "2026-10-04T09:00:02Z", inp=1000, cw5=100000, out=2000), + "not json", + json.dumps({"type": "attachment", "timestamp": "2026-10-04T09:00:03Z"}), + entry("m2", "r2", "claude-opus-5-5", "2026-10-04T09:01:00Z", cw1h=50000, cr=1000000, out=4000), + entry("m2", "r2", "claude-opus-5-5", "2026-10-04T09:01:05Z", cw1h=50000, cr=1000000, out=10000), + entry("m9", "r9", "<synthetic>", "2026-10-04T09:02:00Z", out=5), +]) + "\n" + +# subagent, sonnet 5.5: s1 = 10*2 + 20000*2.5 + 30000*0.2 + 500*10 = 61020 µ$; s2 = 50000*0.2 + 1500*10 = 25000 µ$ +AGENT = "\n".join([ + entry("s1", "q1", "claude-sonnet-5-5", "2026-10-04T10:00:00Z", inp=10, cw5=20000, cr=30000, out=500), + entry("s2", "q2", "claude-sonnet-5-5", "2026-10-04T12:00:00Z", cr=50000, out=1500), + entry("s3", "q3", "claude-mystery-1", "2026-10-04T12:30:00Z", inp=7, out=3), +]) + "\n" + +SID = "11111111-2222-3333-4444-555555555555" + + +class Parse(unittest.TestCase): + def test_dedupes_streamed_entries_and_takes_last_output(self): + got = usage.parse(MAIN) + self.assertEqual(list(got), ["claude-opus-5-5"]) + u = got["claude-opus-5-5"] + self.assertEqual((u.turns, u.inp, u.cw5, u.cw1h, u.cr, u.out), (2, 1000, 100000, 50000, 1000000, 12000)) + self.assertEqual(u.cw, 150000) + self.assertAlmostEqual(usage.cost("claude-opus-5-5", u), 1.344) + + def test_models_kept_apart_unknown_price_none(self): + got = usage.parse(AGENT) + s = got["claude-sonnet-5-5"] + self.assertEqual((s.turns, s.cr, s.out), (2, 80000, 2000)) + self.assertAlmostEqual(usage.cost("claude-sonnet-5-5", s), 0.08602) + self.assertIsNone(usage.cost("claude-mystery-1", got["claude-mystery-1"])) + + def test_since_until_filter_on_utc_timestamp(self): + s = usage.parse(AGENT, since="2026-10-04T11:00", until="2026-10-04T12:10")["claude-sonnet-5-5"] + self.assertEqual((s.turns, s.cr), (1, 50000)) + self.assertAlmostEqual(usage.cost("claude-sonnet-5-5", s), 0.025) + + def test_dated_model_id_priced_by_prefix(self): + u = usage.Usage(turns=1, inp=1000000, out=1000000) + self.assertAlmostEqual(usage.cost("claude-haiku-4-5-20251001", u), 6.0) + self.assertAlmostEqual(usage.cost("claude-opus-5", u), 30.0) + + def test_subagent_output_estimated_from_content(self): + u = usage.parse(SUB)["claude-sonnet-5-5"] + self.assertEqual((u.turns, u.out, u.est), (3, 674 + 50 + 9, 1)) + + def test_log_line_marks_estimate(self): + line = usage.log_line("T", "p", "t-x", "1h", "done", "a1", usage.parse(SUB)) + self.assertIn(" out=733 est=1 usd=", line) + self.assertNotIn(" est=", usage.log_line("T", "p", "t-x", "1h", "done", "a1", usage.parse(AGENT))) + + def test_log_line_lane_and_model(self): + by = usage.parse(SUB) + self.assertIn(" lane=fast model=sonnet effort=", usage.log_line("T", "p", "t-x", "1h", "done", "a1", by, lane="fast")) + self.assertIn(" lane=sonnet model=sonnet effort=", usage.log_line("T", "p", "t-x", "1h", "done", "a1", by)) + + def test_report_lane_model_key(self): + es = [{"lane": "fast", "model": "sonnet", "outcome": "done", "usd": 1.0, "turns": 10, "effort": "1h"}, + {"lane": "fast", "model": "opus", "outcome": "done", "usd": 2.0, "turns": 20, "effort": "1h"}, + {"lane": "opus", "outcome": "done", "usd": 3.0, "turns": 30, "effort": "1h"}] + self.assertEqual(usage.report(es, ("1h",)), [ + ("fast/opus", "all", 1, 1, 0, 2.0, 2.0, 2.0, 20, None), ("fast/opus", "1h", 1, 1, 0, 2.0, 2.0, 2.0, 20, None), + ("fast/sonnet", "all", 1, 1, 0, 1.0, 1.0, 1.0, 10, None), ("fast/sonnet", "1h", 1, 1, 0, 1.0, 1.0, 1.0, 10, None), + ("opus", "all", 1, 1, 0, 3.0, 3.0, 3.0, 30, None), ("opus", "1h", 1, 1, 0, 3.0, 3.0, 3.0, 30, None)]) + + def test_fmt(self): + self.assertEqual([usage.fmt(n) for n in (999, 1500, 8948413)], ["999", "1.5k", "8.95M"]) + + +def tcall(mid, ctx, name, inp, tid): + return json.dumps({"type": "assistant", "timestamp": "2026-10-04T10:00:00Z", "message": { + "id": mid, "model": "claude-sonnet-5-5", "usage": {"input_tokens": 10, "cache_read_input_tokens": ctx - 10}, + "content": [{"type": "tool_use", "id": tid, "name": name, "input": inp}]}}) + + +def tres(tid, text): + return json.dumps({"type": "user", "message": {"role": "user", "content": [ + {"type": "tool_result", "tool_use_id": tid, "content": text}]}}) + + +# agent A: ctx 1000 Read(400 chars=100 tok), 2000 grep(800=200), 3000 Edit, 4000 git commit. +# exp 3000 / mut 3000 / book 4000 of 10000; first edit after 2 calls, pre 3000 = 30%, context +2000 +EXA = "\n".join([ + tcall("a1", 1000, "Read", {"file_path": "/w/proj/.worktrees/x/src/a.py"}, "t1"), tres("t1", "x" * 400), + tcall("a2", 2000, "Bash", {"command": "grep -n foo src/a.py 2>/dev/null"}, "t2"), tres("t2", "y" * 800), + tcall("a3", 3000, "Edit", {"file_path": "/w/proj/src/a.py"}, "t3"), tres("t3", "ok"), + tcall("a4", 4000, "Bash", {"command": "git commit -m x"}, "t4"), tres("t4", "done"), +]) + "\n" +# agent B: 4 exploring calls of 1000 each, reads src/a.py (100 tok) + sed -n of b.py (400 chars) +EXB = "\n".join([ + tcall("b1", 1000, "Read", {"file_path": "/w/proj/src/a.py"}, "u1"), tres("u1", "x" * 400), + tcall("b2", 1000, "Bash", {"command": "sed -n 1,9p src/b.py"}, "u2"), tres("u2", "z" * 400), + tcall("b3", 1000, "Bash", {"command": "ls"}, "u3"), tres("u3", "q"), + tcall("b4", 1000, "Glob", {}, "u4"), tres("u4", "q"), +]) + "\n" + + +class Explore(unittest.TestCase): + def test_one_agent(self): + e = usage.explore(EXA, root="/w/proj") + self.assertEqual((e.calls, e.first_edit, e.pre, e.ctx_growth), (4, 2, 3000, 2000)) + self.assertEqual(e.cost, {"exp": 3000, "mut": 3000, "book": 4000}) + self.assertEqual(e.results, {"Read": 100, "grep": 200, "Edit": 0, "other": 1}) + self.assertEqual(e.files, {"src/a.py": 300}) + + def test_call_kind_book_is_writes_only(self): + kind = lambda c: usage.call_kind({"name": "Bash", "input": {"command": c}}) + for c in ("git worktree list", "git branch --list 'fast/*'", "git branch --show-current", "git status --short"): + self.assertEqual(kind(c), "exp", c) + for c in ("git worktree add .w/x -b x master", "git worktree remove .w/x", "git branch -d x", + "python3 /projects/public/workflow/wf.py finish t-x -m ok", "git switch -c x"): + self.assertEqual(kind(c), "book", c) + + def test_since_drops_early_calls(self): + e = usage.explore(EXA.replace("10:00:00Z", "09:00:00Z", 2), since="2026-10-04T10") + self.assertEqual(e.calls, 2) + + def test_report(self): + a, b = usage.explore(EXA, root="/w/proj"), usage.explore(EXB, root="/w/proj") + r = usage.explore_report([("A", a), ("B", b)], min_pre=4) + self.assertEqual(r["rows"], [("A", 4, 2, 0.3, 0.3), ("B", 4, 4, 1.0, 0.0)]) + self.assertEqual((r["total"], r["cost"]["exp"]), (14000, 7000)) + self.assertEqual((r["pre_n"], r["pre_share"], r["pre_calls"], r["pre_ctx"]), (1, 0.3, 2, 2000)) + self.assertEqual(r["results"]["sed-cat"], 100) + self.assertEqual(r["files"], [("src/a.py", 2, 400)]) + self.assertEqual(usage.explore_report([("A", a)], min_calls=5)["rows"], []) + + +class Cli(unittest.TestCase): + def setUp(self): + self.tmp = tempfile.TemporaryDirectory() + self.home = Path(self.tmp.name) + proj = self.home / "projects" / "-work-demo" + sub = proj / SID / "subagents" + sub.mkdir(parents=True) + (proj / f"{SID}.jsonl").write_text(MAIN) + (sub / "agent-abc123.jsonl").write_text(AGENT) + (sub / "agent-abc123.meta.json").write_text(json.dumps({"description": "wf-worker sonnet t-x"})) + + def tearDown(self): + self.tmp.cleanup() + + def wf(self, *args, sid="", cwd=None): + env = {**os.environ, "CLAUDE_CONFIG_DIR": str(self.home), "CLAUDE_CODE_SESSION_ID": sid} + r = subprocess.run([sys.executable, str(WF), "usage", *args], capture_output=True, text=True, + cwd=cwd or self.tmp.name, env=env, timeout=30) + return r.returncode, r.stdout, r.stderr + + def test_explore_skips_main_and_short_agents(self): + code, out, err = self.wf("--session", SID[:8], "--explore") + self.assertEqual((code, err), (0, "")) + self.assertEqual(out.strip(), "no subagent with >= 4 calls") + + def test_session_rows_and_total(self): + code, out, err = self.wf("--session", SID[:8]) + self.assertEqual((code, err), (0, "")) + rows = [l.split() for l in out.strip().split("\n")[1:]] + self.assertEqual(rows[0], ["main", "opus-5-5", "2", "1.0k", "150.0k", "1.00M", "12.0k", "1.34"]) + self.assertEqual(rows[1], ["abc123", "wf-worker", "sonnet", "t-x", + "sonnet-5-5", "2", "10", "20.0k", "80.0k", "2.0k", "0.09"]) + self.assertEqual(rows[2][-7:], ["mystery-1", "1", "7", "0", "0", "3", "?"]) + self.assertEqual(rows[3], ["total", "5", "1.0k", "170.0k", "1.08M", "14.0k", "1.43+"]) + + def test_session_from_env(self): + code, out, _ = self.wf(sid=SID) + self.assertEqual(code, 0) + self.assertIn("abc123", out) + + def test_agent_only(self): + code, out, err = self.wf("--agent", "agent-abc123", "--since", "2026-10-04T11:00") + self.assertEqual((code, err), (0, "")) + lines = out.strip().split("\n") + self.assertNotIn("main", out) + self.assertEqual(lines[1].split()[-7:], ["sonnet-5-5", "1", "0", "0", "50.0k", "1.5k", "0.03"]) + + def test_log_appends_one_key_value_line(self): + root = self.home / "demo" + root.mkdir() + (root / "workflow.toml").write_text("format = 1\n") + code, out, err = self.wf("--agent", "abc123", "--log", "t-x", "done", "--effort", "1h", cwd=root) + self.assertEqual((code, err), (0, "")) + code, out, err = self.wf("--agent", "abc123", "--log", "t-x", "handback", cwd=root) + log = (root / "out" / "wf-cost.log").read_text().split("\n") + self.assertEqual(len(log), 3) # two lines + trailing newline + # lane = model with most turns (sonnet 2 vs mystery 1); usd = known models only + self.assertRegex(log[0], r"^\d{4}-\d\d-\d\dT\d\d:\d\d:\d\dZ project=demo task=t-x lane=sonnet model=sonnet effort=1h " + r"outcome=done turns=3 in=17 cw=20000 cr=80000 out=2003 usd=0\.0860 agent=abc123$") + self.assertIn(" effort=- outcome=handback ", log[1]) + self.assertEqual(out, log[1] + "\n") + + def test_estimated_out_marked_with_tilde(self): + other = self.home / "projects" / "-work-demo" / "other-session" / "subagents" + other.mkdir(parents=True) + (other / "agent-est1.jsonl").write_text(SUB) + code, out, err = self.wf("--agent", "est1") + self.assertEqual((code, err), (0, "")) + self.assertEqual(out.strip().split("\n")[1].split()[-3:], ["0", "~733", "0.01"]) + + def test_log_needs_agent(self): + code, out, err = self.wf("--log", "t-x", "done") + self.assertEqual((code, err), (2, "wf: --log needs --agent\n")) + + def test_log_effort_must_be_an_estimate(self): + code, out, err = self.wf("--agent", "abc123", "--log", "t-x", "done", "--effort", "medium") + self.assertEqual(code, 2) + self.assertIn("invalid choice: 'medium'", err) + + def test_missing_agent_one_line_error(self): + code, out, err = self.wf("--agent", "nope") + self.assertEqual((code, out), (1, "")) + self.assertEqual(err, "wf: no transcript for agent nope\n") + + +LOG_A = """\ +2026-10-01T10:00:00Z project=a task=t-1 lane=sonnet effort=1h outcome=done turns=10 usd=0.50 +2026-10-01T11:00:00Z project=a task=t-2 lane=sonnet effort=1h outcome=handback turns=30 usd=1.50 +garbage line +""" +LOG_B = """\ +2026-10-02T10:00:00Z project=b task=t-3 lane=sonnet effort=<1h outcome=done turns=20 usd=0.40 +2026-10-02T11:00:00Z project=b task=t-4 lane=opus effort=5h outcome=done turns=80 usd=3.00 +2026-10-03T11:00:00Z project=b task=t-5 lane=opus effort=1h outcome=done turns=40 usd=2.00 +""" + + +class Report(unittest.TestCase): + def test_rows_by_lane_then_effort(self): + entries = usage.parse_log(LOG_A + LOG_B) + self.assertEqual(len(entries), 5) + rows = usage.report(entries, ("<1h", "1h", "5h")) + # lane, effort, n, done, other, median $, total $, $/done, median turns — all by hand + self.assertEqual(rows, [ + ("opus", "all", 2, 2, 0, 2.5, 5.0, 2.5, 60, None), + ("opus", "1h", 1, 1, 0, 2.0, 2.0, 2.0, 40, None), + ("opus", "5h", 1, 1, 0, 3.0, 3.0, 3.0, 80, None), + ("sonnet", "all", 3, 2, 1, 0.5, 2.4, 1.2, 20, None), + ("sonnet", "<1h", 1, 1, 0, 0.4, 0.4, 0.4, 20, None), + ("sonnet", "1h", 2, 1, 1, 1.0, 2.0, 2.0, 20, None), + ]) + + def test_duration(self): + line = usage.log_line("T", "p", "t-x", "1h", "done", "a1", usage.parse(SUB), dur=125) + self.assertIn(" dur=125 agent=a1", line) + self.assertNotIn("dur=", usage.log_line("T", "p", "t-x", "1h", "done", "a1", usage.parse(SUB))) + log = ("2026-10-04T10:00:00Z project=a task=t-6 lane=fast effort=<1h outcome=done turns=1 usd=1 dur=100\n" + "2026-10-04T10:00:00Z project=a task=t-7 lane=fast effort=<1h outcome=done turns=1 usd=1 dur=300\n" + "2026-10-04T10:00:00Z project=a task=t-8 lane=fast effort=<1h outcome=done turns=1 usd=1\n") + self.assertEqual(usage.report(usage.parse_log(log), ("<1h",))[0][9], 200) + + def test_done_gate_red_counts_as_done(self): + log = ("2026-10-04T10:00:00Z project=a task=t-6 lane=fast effort=<1h outcome=done+gate-red turns=12 usd=1.00\n" + "2026-10-04T11:00:00Z project=a task=t-7 lane=fast effort=<1h outcome=donex turns=8 usd=0.50\n") + rows = usage.report(usage.parse_log(log), ("<1h",)) + self.assertEqual(rows[0], ("fast", "all", 2, 1, 1, 0.75, 1.5, 1.5, 10, None)) + + def test_since_and_no_done(self): + entries = usage.parse_log(LOG_A + LOG_B, since="2026-10-01T10:30") + rows = usage.report(entries, ("1h",)) + self.assertIn(("sonnet", "all", 2, 1, 1, 0.95, 1.9, 1.9, 25, None), rows) + rows = usage.report(usage.parse_log(LOG_A, since="2026-10-01T10:30"), ("1h",)) + self.assertEqual(rows[0], ("sonnet", "all", 1, 0, 1, 1.5, 1.5, None, 30, None)) + + def test_cli_reads_every_project(self): + with tempfile.TemporaryDirectory() as tmp: + for name, text in (("a", LOG_A), ("b", LOG_B)): + (Path(tmp) / name / "out").mkdir(parents=True) + (Path(tmp) / name / "workflow.toml").write_text("format = 1\n") + (Path(tmp) / name / "out" / "wf-cost.log").write_text(text) + r = subprocess.run([sys.executable, str(WF), "usage", "--report"], capture_output=True, text=True, + cwd=tmp, env={**os.environ, "WF_ROOT": tmp}, timeout=30) + self.assertEqual((r.returncode, r.stderr), (0, "")) + lines = [l.split() for l in r.stdout.strip().split("\n")] + self.assertEqual(lines[0], ["lane", "effort", "n", "done", "other", "med$", "total$", "$/done", "med_turns", "med_dur"]) + self.assertEqual(lines[1], ["opus", "all", "2", "2", "0", "2.50", "5.00", "2.50", "60", "-"]) + self.assertEqual(lines[6], ["sonnet", "1h", "2", "1", "1", "1.00", "2.00", "2.00", "20", "-"]) + + +if __name__ == "__main__": + unittest.main() @@ -0,0 +1,2239 @@ +#!/usr/bin/env python3 +"""wf — shared task workflow tool. Run inside a project (folder with workflow.toml). + +Look: next [--brief] · list [filters] · show ID… · ctx ID|PATH#ANCHOR + search WORDS… · log [-n N] [WORDS] · projects · check +Change: add "Title. Goal." -p N -e EFFORT · done ID… -m "entry" · prio ID N + move ID SECTION | --before/--after ID · status ID progress NOTE|blocked A-ID|clear + set ID --title/--effort/--after/--ref/--model/--sessions/--cloud · note ID "line" + body ID <stdin (Model:/Sessions:/After:/Ref: kept) · tick ID N|TEXT · setup · merge (in a lane worktree) +Orch: orch pick LANE [--id ID] [--recovery WHY] · orch post ID LANE [--result LINE] [--agent A] [--duration S] [--no-next] +Other: report "what happened" · init · migrate [--write] · res … (shared memory/CPU ledger; wf res -h) + batch N [--lanes L,…] [--prep] | --status (unattended batch orchestrator; wf batch -h) + prep [N|all] [--lanes L,…] (= batch --prep: sonnet workers write Done for tasks without one; all = every such task) + +Every change command takes --dry-run. `wf CMD -h` for details. +Rules: /projects/CLAUDE.md. Design: docs/design.md. +""" +from __future__ import annotations + +import argparse +import contextlib +import datetime +import difflib +import glob +import io +import fcntl +import os +import re +import subprocess +import sys +import tempfile +from pathlib import Path + +HERE = Path(__file__).resolve().parent +sys.path.insert(0, str(HERE)) + +from wflib import check as checks # noqa: E402 +from wflib import areas as areas_mod # noqa: E402 +from wflib import config, lanes, ledgers, refs, search, tasks, usage # noqa: E402 +import json # noqa: E402 + +WIDTH = 100 +SKIP_DIRS = {".worktrees", "node_modules", "lib", "__pycache__"} + + +class Failure(Exception): + code = 1 + + +class Usage(Failure): + code = 2 + + +def old_spelling(old: str, new: str) -> None: + """One-line notice for an option kept one more release under its old name.""" + print(f"wf: {old} is now {new}", file=sys.stderr) + + +# ------------------------------------------------------------------ files + +def stamp(path: Path) -> tuple[int, int]: + s = path.stat() + return (s.st_mtime_ns, s.st_size) + + +def write_if_unchanged(path: Path, text: str, read_stamp: tuple[int, int]) -> None: + """Replace `path` by temp file + rename, unless it changed since `read_stamp`.""" + if stamp(path) != read_stamp: + raise Failure(f"{path.name} changed on disk since it was read: nothing written, run again") + mode = path.stat().st_mode & 0o777 + fd, tmp = tempfile.mkstemp(dir=path.parent, prefix=f".{path.name}.", suffix=".tmp") + try: + with os.fdopen(fd, "wb") as f: + f.write(text.encode("utf-8")) + os.chmod(tmp, mode) + if stamp(path) != read_stamp: + raise Failure(f"{path.name} changed on disk since it was read: nothing written, run again") + os.replace(tmp, path) + finally: + if os.path.exists(tmp): + os.unlink(tmp) + + +def read(path: Path) -> str: + return path.read_bytes().decode("utf-8", errors="replace") + + +class Project: + def __init__(self, cfg: config.Config): + self.cfg = cfg + for name, path in (("tasks", cfg.tasks), ("archive", cfg.archive)): + if not path.is_file(): + raise Failure(f"workflow.toml: {name} '{cfg.rel(path)}' does not exist") + self.tasks_stamp, self.archive_stamp = stamp(cfg.tasks), stamp(cfg.archive) + self.tasks_text, self.archive_text = read(cfg.tasks), read(cfg.archive) + self.doc = tasks.parse(self.tasks_text) + self.new_archive: str | None = None + + @property + def archived(self) -> set[str]: + return tasks.archive_ids(self.new_archive or self.archive_text) + + def taken(self) -> set[str]: + return self.doc.ids() | self.archived + + def resolve(self, given: str) -> str: + """Exact id, or a unique prefix of an open id (read commands).""" + known = self.doc.ids() + if given in known: + return given + hits = sorted(i for i in known if i.startswith(given)) + if len(hits) == 1: + return hits[0] + if hits: + raise Failure(f"'{given}' matches {', '.join(hits)}") + raise Failure(tasks.unknown_id(given, known)) + + def save(self, dry_run: bool) -> None: + cfg = self.cfg + new_tasks = tasks.render(self.doc) + before = {p.key for p in checks.check(cfg, slow=False)[0]} + after = checks.check(cfg, tasks_text=new_tasks, archive_text=self.new_archive, slow=False)[0] + added = [p for p in after if p.key not in before] + if added: + raise Failure("refused: nothing written, the change adds problems:\n" + "\n".join(f" {p}" for p in added)) + changes = [(cfg.tasks, self.tasks_text, new_tasks, self.tasks_stamp)] + if self.new_archive is not None: + changes.insert(0, (cfg.archive, self.archive_text, self.new_archive, self.archive_stamp)) + if dry_run: + for path, old, new, _ in changes: + name = cfg.rel(path) + sys.stdout.writelines(difflib.unified_diff( + old.replace("\r\n", "\n").splitlines(keepends=True), + new.replace("\r\n", "\n").splitlines(keepends=True), name, f"{name} (new)")) + return + written = [] + for path, old, new, read_stamp in changes: + if new == old: + continue + try: + write_if_unchanged(path, new, read_stamp) + except Failure as e: + if written: + raise Failure(f"{e}; {written[0]} was already written: remove its new top line(s)") from None + raise + written.append(cfg.rel(path)) + + +def load_project(args, write: bool = False) -> Project: + start = Path(args.project) if args.project else Path.cwd() + if not start.is_dir(): + raise Failure(f"--project {start}: no such folder") + cfg = config.load_at(start) + if cfg.format > config.FORMAT: + raise Failure(f"project format {cfg.format} is newer than this wf ({config.FORMAT}): update /projects/public/workflow") + if write and cfg.format < config.FORMAT: + raise Failure(f"TASKS format {cfg.format}, wf needs {config.FORMAT}: " + "run wf migrate --write (idle project, one commit)") + return Project(cfg) + + +# ------------------------------------------------------------------ output + +def show(item: tasks.Item) -> str: + return "\n".join(item.lines()) + + +def status_word(item: tasks.Item) -> str: + if not item.status: + return "-" + return "blkd" if item.status.startswith("blocked") else "prog" + + +def row(item: tasks.Item, width: int, cfg: config.Config) -> str: + prio = "--" if item.prio is None else f"P{item.prio}" + model = item.model if item.prio is not None else "-" + lane = lanes.lane_of(item, cfg.lanes, cfg.slice_above) if item.prio is not None else "-" + lane_w = max(len(l.name) for l in cfg.lanes) + mark = f"[{item.sessions}] " if item.prio is not None and item.sessions != "parallel" else "" + line = (f"{item.id:<{width}} {prio} {item.effort or '-':<4} {status_word(item):<4} {model:<6} " + f"{lane:<{lane_w}} {mark}{item.title}") + return line if len(line) <= WIDTH else line[:WIDTH - 1] + "…" + + +def counts(doc: tasks.Doc) -> str: + def n(key): + return sum(len(s.items) for s in doc.sections if s.key == key) + return " · ".join(f"{key} {n(key)}" for key in ("pending", "human", "awaiting", "deferred")) + + +def resolve_refs(cfg: config.Config, item: tasks.Item) -> list[str]: + out = [] + for path, anchor in item.refs: + out.append(refs.resolve(cfg.root, path, anchor)) + if anchor and cfg.anchors_index and cfg.anchors_specs and cfg.anchors_specs.is_dir() \ + and (cfg.root / path).resolve() == cfg.anchors_index.resolve(): + for spec in sorted(cfg.anchors_specs.rglob("*.md")): + text = read(spec) + if any(anchor in refs.ANCHOR_RE.findall(l) for l in text.split("\n")): + out.append(refs.resolve(cfg.root, cfg.rel(spec), anchor)) + return out + + +def inbox_path() -> Path: + return Path(os.environ.get("WF_INBOX") or HERE / "inbox.md") + + +def inbox_count() -> int: + path = inbox_path() + if not path.is_file(): + return 0 + return sum(1 for l in read(path).split("\n") if l.startswith("- ")) + + +# ------------------------------------------------------------------ sessions (lanes) + +def sessions_dir(cfg: config.Config) -> Path: + return cfg.root / ".wf" / "sessions" + + +def alive(pid, socket) -> bool: + try: + os.kill(int(pid), 0) + except (OSError, ValueError, TypeError): + return False + return bool(socket) and Path(socket).exists() + + +def lane_names(cfg: config.Config) -> list[str]: + return [l.name for l in cfg.lanes] + + +def check_lane(cfg: config.Config, lane: str | None) -> None: + if lane and lane not in lane_names(cfg): + raise Usage(f"unknown lane '{lane}' ({', '.join(lane_names(cfg))})") + + +def sessions(cfg: config.Config) -> dict[str, dict]: + """Registered sessions keyed by lane name or "all" (no lane: every lane).""" + out = {} + folder = sessions_dir(cfg) + for name in [*lane_names(cfg), "all"]: + try: + s = json.loads((folder / f"{name}.json").read_text()) + except (OSError, ValueError): + continue + if isinstance(s, dict): + s["alive"] = alive(s.get("pid"), s.get("socket")) + out[name] = s + return out + + +def my_session() -> tuple[str, str] | None: + """(socket, pid) of this agent session from Claude Code's env; None outside one.""" + socket, pid = os.environ.get("CLAUDE_CODE_MESSAGING_SOCKET"), os.environ.get("CLAUDE_PID") + return (socket, pid) if socket and pid else None + + +def wf_folder(cfg: config.Config, name: str) -> Path: + """.wf/<name> in the project root, created; .wf is git-ignored.""" + folder = cfg.root / ".wf" / name + folder.mkdir(parents=True, exist_ok=True) + ignore = folder.parent / ".gitignore" + if not ignore.exists(): + ignore.write_text("*\n") + return folder + + +@contextlib.contextmanager +def project_lock(root: Path): + """Exclusive project lock (.wf/lock): every wf write and merge-back runs under it. Not reentrant.""" + folder = root / ".wf" + folder.mkdir(parents=True, exist_ok=True) + ignore = folder / ".gitignore" + if not ignore.exists(): + ignore.write_text("*\n") + with open(folder / "lock", "w") as f: + fcntl.flock(f, fcntl.LOCK_EX) + yield + + +def register(cfg: config.Config, lane: str, model: str) -> str | None: + """Record this agent session as the lane's session (lane "all" = every lane; env from Claude Code; + skipped without it). Another live session already holding the lane keeps it; returns a warning line then.""" + me = my_session() + if not me: + return None + socket, pid = me + other = sessions(cfg).get(lane) + if other and other["alive"] and str(other.get("pid")) != pid: + return (f"another live {lane} session holds this lane: uds:{other['socket']} " + "(claims keep tasks apart; tell the owner if unintended)") + try: + record = {"lane": lane, "model": model, "socket": socket, "pid": int(pid) if pid.isdigit() else pid, + "session": os.environ.get("CLAUDE_CODE_SESSION_ID", ""), + "at": datetime.datetime.now().isoformat(timespec="seconds")} + (wf_folder(cfg, "sessions") / f"{lane}.json").write_text(json.dumps(record) + "\n") + except OSError as e: + print(f"wf: session not registered: {e}", file=sys.stderr) + return None + + +def unregister(cfg: config.Config) -> None: + """Drop session records whose pid is this session (orchestrator holds no lane).""" + me = my_session() + if not me: + return + for name in [*lane_names(cfg), "all"]: + f = sessions_dir(cfg) / f"{name}.json" + try: + if str(json.loads(f.read_text()).get("pid")) == me[1]: + f.unlink() + except (OSError, ValueError, AttributeError): + pass + + +# ------------------------------------------------------------------ claims (wf status progress) + +def claim(cfg: config.Config, id: str, lane: str = "") -> None: + """Mark task id as held by this session (skipped outside one). lane: the task's lane.""" + me = my_session() + if not me: + return + socket, pid = me + model = next((s.get("model", name) for name, s in sessions(cfg).items() + if str(s.get("pid")) == pid and s.get("socket") == socket), "") + try: + record = {"id": id, "lane": lane, "model": model, "socket": socket, "pid": int(pid) if pid.isdigit() else pid, + "session": os.environ.get("CLAUDE_CODE_SESSION_ID", ""), + "at": datetime.datetime.now().isoformat(timespec="seconds")} + (wf_folder(cfg, "claims") / f"{id}.json").write_text(json.dumps(record) + "\n") + except OSError as e: + print(f"wf: claim not recorded: {e}", file=sys.stderr) + + +def unclaim(cfg: config.Config, ids) -> None: + for id in ids: + try: + (cfg.root / ".wf" / "claims" / f"{id}.json").unlink(missing_ok=True) + except OSError: + pass + + +def held(p: "Project") -> dict[str, str]: + """id → holder, for tasks in progress claimed by another live session.""" + me = my_session() + pending = {i.id: i for i in p.doc.section("pending").items} + out = {} + for path in sorted((p.cfg.root / ".wf" / "claims").glob("*.json")): + try: + c = json.loads(path.read_text()) + except (OSError, ValueError): + continue + if not isinstance(c, dict) or (me and str(c.get("pid")) == me[1]) or not alive(c.get("pid"), c.get("socket")): + continue + item = pending.get(path.stem) + if item is None or not (item.status or "").startswith("in progress"): + continue + out[item.id] = f"{c.get('model') + ' ' if c.get('model') else ''}session uds:{c.get('socket')}" + return out + + +def stale(p: "Project") -> list: + """Items in progress whose claim is missing or held by a dead session.""" + out = [] + for item in p.doc.section("pending").items: + if not (item.status or "").startswith("in progress"): + continue + try: + c = json.loads((p.cfg.root / ".wf" / "claims" / f"{item.id}.json").read_text()) + except (OSError, ValueError): + c = None + if not isinstance(c, dict) or not alive(c.get("pid"), c.get("socket")): + out.append(item) + return out + + +def other_live(cfg: config.Config) -> int: + """Live sessions in the project other than this one.""" + me = my_session() + return sum(1 for s in sessions(cfg).values() if s["alive"] and not (me and str(s.get("pid")) == me[1])) + + +def multi_lines(p: Project, args, lane: str, extra: int) -> list[str]: + """Worktree mode advice when more than one session is live in the project (git only).""" + n = sum(1 for s in sessions(p.cfg).values() if s["alive"]) + extra + main_top = config.git_top(p.cfg.root) + if n < 2 or not main_top or not (main_top / ".git").is_dir(): + return [] + head = f"{n} live sessions here: " + wt = config.linked_worktree(Path(args.project or Path.cwd()).resolve()) + if wt: + return [head + f"you are in worktree {p.cfg.rel(wt[0])} (branch {config.git_branch(wt[1])}); " + "wf writes the main tree's TASKS.md"] + folder, master = p.cfg.root / ".worktrees" / lane, config.git_branch(main_top / ".git") or "master" + cmd = (f"cd {p.cfg.rel(folder)} && git switch -c {lane}/<task> {master}" if folder.is_dir() + else f"git worktree add {p.cfg.rel(folder)} -b {lane}/<task> {master}") + if p.cfg.worktree_setup: + cmd += " && wf setup" if folder.is_dir() else f" && cd {p.cfg.rel(folder)} && wf setup" + return [head + f"work in your lane's worktree, never on {master}:", f" {cmd} (wf there writes this TASKS.md)"] + + +def lane_lines(p: Project, me: str | None, runner: bool = False) -> list[str]: + a = (p.doc, p.archived, p.cfg.lanes, p.cfg.slice_above, held(p)) + return lanes.lanes_block(lanes.counts(*a, runner=runner), sessions(p.cfg), me, + lanes.not_ready(*a) if runner else None) + + +# ------------------------------------------------------------------ read commands + +def cmd_next(args) -> int: + p = load_project(args) + lane = args.lane + check_lane(p.cfg, lane) + model = args.as_ or "haiku" + out: list[str] = [] + warning = None + if not args.as_: + out.append("no --as: treated as haiku; pass --as " + "|".join(tasks.MODELS)) + elif warning := register(p.cfg, lane or "all", model): + out.append(warning) + claims = held(p) + if solo := tasks.solo_running(p.doc, claims): + if out: + print("\n".join(out)) + raise Failure(f"solo {solo[0]} in progress by {solo[1]}: wait (its wf done notifies you)") + item, skipped = lanes.pick(p.doc, p.archived, p.cfg.lanes, p.cfg.slice_above, lane, model, claims, + other_live(p.cfg), args.owner) + multi = multi_lines(p, args, lane or "all", 1 if warning else 0) + if not args.brief: + awaiting = p.doc.section("awaiting").items + if awaiting: + out += ["===== Awaiting your decision (mention, don't block) =====", + *(f"- {i.id}: {i.text}" for i in awaiting), ""] + human = p.doc.section("human").items + if human: + out += ["===== Needs human (not picked) =====", *(f"- {i.id}: {i.title}" for i in human), ""] + flight = ledgers.in_flight(p.cfg.root, p.cfg.ledgers) if p.cfg.ledgers and p.cfg.ledgers.is_dir() else [] + if flight: + out += ["===== In-flight plans =====", *flight, ""] + if skipped: + out += ["===== Skipped =====", *(f"- {i.id}: {why}" for i, why in skipped), ""] + if multi: + out += ["===== Multi-session =====", *multi, ""] + lines = lane_lines(p, lane) + if len(lines) > 1: + out += ["===== Lanes =====", *lines, ""] + waits = lanes.waiting_block(lanes.cross_waits(p.doc, p.archived, lane, p.cfg.lanes, p.cfg.slice_above), + sessions(p.cfg), p.cfg.lanes, p.cfg.slice_above) if lane else [] + if waits: + out += ["===== Waiting on other lanes =====", *waits, ""] + if item: + if not args.brief: + out.append("===== Next task =====") + out.append(show(item)) + if lanes.is_slice_job(item, p.cfg.slice_above): + out += ["", "===== Slice job (no code) =====", + f"effort {item.effort} > slice_above {p.cfg.slice_above}: split it, don't implement.", + f' wf add "<title>. <goal>" -e <1h|1h> --parent {item.id} --model haiku|sonnet|opus' + " (then wf body: Steps/Done/Ref)", + f' wf note {item.id} "sliced into …" · wf status {item.id} clear · not wf done' + " (the parent waits on its slices)"] + if not args.brief: + for block in resolve_refs(p.cfg, item): + out += ["", block] + n = 0 if args.brief else inbox_count() + if n: + out += ["", f"workflow inbox: {n} reports (triage: workflow session)"] + if hint := ctx_hint_line(p.cfg): + out += ["", hint] + while out and not out[-1]: + out.pop() + if out: + print("\n".join(out)) + if not item: + raise Failure(f"nothing pickable for {lane or 'all lanes'} ({model}) in Pending") + return 0 + + +def cmd_areas(args) -> int: + p = load_project(args, write=bool(args.mark)) + file = p.cfg.areas_file + rel = p.cfg.areas_rel + text = file.read_text(encoding="utf-8", errors="replace") if file.is_file() else "" + found = areas_mod.parse(text) + if not found: + print(f"no areas in {rel}") + return 0 + names = ", ".join(a.name for a in found) + wanted = args.name or args.mark + if wanted and wanted not in [a.name for a in found]: + hits = [a.name for a in found if a.name.startswith(wanted)] + if len(hits) == 1: + if args.name: + args.name = hits[0] + else: + args.mark = hits[0] + wanted = hits[0] + elif hits: + raise Usage(f"ambiguous area '{wanted}' in {rel}: " + ", ".join(hits)) + else: + raise Usage(f"no area '{wanted}' in {rel} ({names})") + if args.mark: + r = git_run(p.cfg.code, "rev-parse", "--short", "HEAD") + if r.returncode != 0: + raise Failure("git rev-parse HEAD failed: " + r.stderr.strip()) + text = areas_mod.mark(text, args.mark, r.stdout.strip()) + file.write_text(text, encoding="utf-8") + found = areas_mod.parse(text) + for a in found: + if wanted and a.name != wanted: + continue + stale, miss, commits = areas_mod.status(p.cfg.code, a, p.cfg.area_stale_commits) + line = " · ".join([f"{a.name}: " + ("stale" if stale else "ok")] + + ([f"missing: {', '.join(miss)}"] if miss else []) + + ([f"{commits} commits since {a.checked}"] if commits is not None else ["no Checked"] if not a.checked else [])) + print(line) + if not wanted: + for folder, n in uncovered_areas(p, args): + print(f"uncovered: {folder} ({n} file{'s' * (n != 1)}): no area's Paths covers it") + return 0 + + +def cmd_lanes(args) -> int: + p = load_project(args) + check_lane(p.cfg, args.lane) + if args.unregister: + unregister(p.cfg) + elif args.lane or args.as_: + register(p.cfg, args.lane or "all", args.as_ or "haiku") + if args.wait is not None: + import time + poll = float(os.environ.get("WF_LANES_POLL", "30")) + end = time.monotonic() + args.wait + while True: + p = load_project(args) + if (p.cfg.root / "out" / "wf-batch.stop").exists(): + print("stop requested: out/wf-batch.stop") + return 2 + if any(c[0] for c in lanes.counts(p.doc, p.archived, p.cfg.lanes, p.cfg.slice_above, held(p), runner=True).values()): + break + left = end - time.monotonic() + if left <= 0: + print("\n".join(lane_lines(p, args.lane, True))) + return 1 + time.sleep(min(poll, left)) + print("\n".join(lane_lines(p, args.lane, True))) + return 0 + + +def runner_ids(p: Project) -> set[str]: + """Pending ids a headless worker can take: not blocked, After done, runner-ready, no open slices, not sliced out.""" + return {i.id for i in p.doc.section("pending").items + if not i.error and not (i.status or "").startswith("blocked") and all(a in p.archived for a in i.after) + and i.runner_ready and not tasks.open_slices(p.doc, i.id) and not lanes.sliced_out(i, p.cfg.slice_above)} + + +def cmd_list(args) -> int: + p = load_project(args) + check_lane(p.cfg, args.lane) + keys = list(tasks.SECTIONS) if args.section == "all" else [args.section] + ready = None + if args.ready: + ready = {i.id for i in p.doc.section("pending").items + if not i.error and not (i.status or "").startswith("blocked") + and all(a in p.archived for a in i.after)} + if args.runner: + ready = runner_ids(p) + stale_ids = {i.id for i in stale(p)} if args.stale else None + want_ref = tuple(args.ref.split("#", 1)) if args.ref else None + + def keep(i: tasks.Item) -> bool: + if args.prio is not None and (i.prio is None or i.prio > args.prio): + return False + if ready is not None and i.id not in ready: + return False + if args.blocked and status_word(i) != "blkd": + return False + if args.progress and status_word(i) != "prog": + return False + if stale_ids is not None and i.id not in stale_ids: + return False + if args.model and (i.prio is None or i.model != args.model): + return False + if args.lane and (i.prio is None or lanes.lane_of(i, p.cfg.lanes, p.cfg.slice_above) != args.lane): + return False + if want_ref and not any(path == want_ref[0] and (len(want_ref) == 1 or anchor == want_ref[1]) + for path, anchor in i.refs): + return False + return True + + groups = [(s, [i for i in s.items if keep(i)]) for k in keys for s in p.doc.sections if s.key == k] + if args.runner and args.lane: # pick order of the lane, then its fallback lane's + order = lanes.ranked(p.doc, p.archived, p.cfg.lanes, p.cfg.slice_above, args.lane) + keys, groups = ["pending"], [(None, [i for i in order if i.id in ready + and (args.prio is None or i.prio <= args.prio)])] + if args.n is not None: + left = args.n + for n, (s, items) in enumerate(groups): + groups[n] = (s, items[:left]) + left -= len(groups[n][1]) + width = max((len(i.id) for _, items in groups for i in items), default=0) + for s, items in groups: + if len(keys) > 1: + print(f"== {s.heading}") + for i in items: + print(row(i, width, p.cfg)) + print(counts(p.doc)) + return 0 + + +def cmd_show(args) -> int: + p = load_project(args) + print("\n\n".join(show(p.doc.item(p.resolve(i))) for i in args.ids)) + return 0 + + +def cmd_ctx(args) -> int: + p = load_project(args) + target = args.target + if "#" in target or "/" in target or (p.cfg.root / target).exists(): + path, _, anchor = target.partition("#") + users = [i.id for i in p.doc.all_items() + if any(r == path and (not anchor or a == anchor) for r, a in i.refs)] + print(refs.resolve(p.cfg.root, path, anchor or None)) + print("\nTasks: " + (", ".join(users) if users else "none")) + return 0 + if target not in p.doc.ids() and target in p.archived: + line = next(l for l in p.archive_text.replace("\r\n", "\n").split("\n") + if (m := tasks.ARCHIVE_ID_RE.match(l)) and m.group(1) == target) + print(f"done: {line}") + return 0 + id = p.resolve(target) + section, _ = p.doc.find(id) + item = p.doc.item(id) + where = {i.id: s.heading for s in p.doc.sections for i in s.items} + out = [show(item), "", f"Section: {section.heading}"] + if item.after: + out.append("After: " + ", ".join( + f"{a} ({'open, ' + where[a] if a in where else 'done' if a in p.archived else 'unknown'})" + for a in item.after)) + needed = [i.id for i in p.doc.all_items() if id in i.after] + blocks = [i.id for i in p.doc.all_items() if i.blocked_on == id] + linked = [i.id for i in p.doc.all_items() + if i.id != id and i.id not in needed and i.id not in blocks + and id in refs.links("\n".join(i.lines()))] + for label, ids in (("Needed by", needed), ("Blocks", blocks), ("Linked from", linked)): + if ids: + out.append(f"{label}: {', '.join(ids)}") + if p.cfg.verify and not id.startswith("a-"): + out.append("Verify (run before wf finish): " + " · ".join(p.cfg.verify)) + for block in resolve_refs(p.cfg, item): + out += ["", block] + out += area_blocks(p, "\n".join(item.lines())) + print("\n".join(out)) + return 0 + + +def area_blocks(p: Project, text: str) -> list[str]: + """Notes of the areas the task text names (area name, anchor or a path no other area lists) → ctx lines.""" + file = p.cfg.areas_file + if not file.is_file(): + return [] + notes = file.read_text(encoding="utf-8", errors="replace") + out = [] + found = areas_mod.parse(notes) + skip = areas_mod.shared_paths(found) + for a in found: + if areas_mod.matches(a, text, skip): + out += ["", f"Area {a.name} ({p.cfg.areas_rel}):", areas_mod.block(notes, a.name)] + return out + + +def cmd_search(args) -> int: + p = load_project(args) + kinds = {k for k, on in (("task", args.tasks), ("archive", args.archive), ("doc", args.docs)) if on} or None + docs = {} + if kinds is None or "doc" in kinds: + docs = {p.cfg.rel(f): read(f) for f in checks.doc_files(p.cfg)} + hits = search.search(args.words, p.doc, p.archive_text, docs, kinds=kinds, limit=args.n, + archive_name=p.cfg.rel(p.cfg.archive)) + if not hits: + raise Failure("no hits") + for h in hits: + if h.kind == "task": + header = p.doc.item(h.where).text if h.where in p.doc.ids() else "" + text = h.line if h.line == header else f"{h.label} — {h.line}" + elif h.kind == "archive": + text = re.sub(r"^\d{4}-\d\d-\d\d \*\*([^*]+)\*\*", r"\1", h.line) + else: + text = f"{h.label} — {h.line}" if h.line and h.line != h.label else h.label + line = f"{h.where} {text}" + print(line if len(line) <= 2 * WIDTH else line[:2 * WIDTH - 1] + "…") + return 0 + + +def cmd_log(args) -> int: + p = load_project(args) + lines = [l for l in p.archive_text.replace("\r\n", "\n").split("\n") if l.startswith("- ")] + words = [w.lower() for w in args.words] + lines = [l for l in lines if all(w in l.lower() for w in words)] + print("\n".join(lines[:args.n])) + return 0 + + +def cmd_check(args) -> int: + p = load_project(args) + errors, warnings = checks.check(p.cfg) + warnings += checks.merged_tool_worktrees(Path(os.environ.get("WF_TOOL_ROOT") or HERE)) + for w in warnings: + print(f"warn: {w}") + for e in errors: + print(f"ERROR: {e}") + print(("OK: " if not errors else "") + f"{len(errors)} errors · {len(warnings)} warnings") + return 1 if errors else 0 + + +def find_projects(root: Path, depth: int = 3) -> list[Path]: + out = [] + + def walk(folder: Path, left: int) -> None: + if (folder / "wf.py").is_file() and (folder / "wflib").is_dir(): + return # a wf checkout (its templates/ is no project) + if (folder / config.NAME).is_file(): + out.append(folder) + return + if left == 0: + return + try: + children = sorted(c for c in folder.iterdir() if c.is_dir() and not c.is_symlink()) + except OSError: + return + for c in children: + if not c.name.startswith(".") and c.name not in SKIP_DIRS: + walk(c, left - 1) + + walk(root, depth) + return out + + +def default_root(tool: Path) -> Path: + """Folder `wf projects` scans: the tool's parent, else its grandparent (tool under e.g. public/), if it holds projects.""" + return next((d for d in (tool.parent, tool.parent.parent) if find_projects(d)), tool.parent) + + +def projects_root() -> Path: + return Path(os.environ.get("WF_ROOT") or default_root(HERE)) + + +def cmd_projects(args) -> int: + root = projects_root() + rows = [] + for folder in find_projects(root): + name = str(folder.relative_to(root)) + try: + cfg = config.load(folder) + doc = tasks.parse(read(cfg.tasks)) + archived = tasks.archive_ids(read(cfg.archive)) if cfg.archive.is_file() else set() + errors = len(checks.check(cfg, slow=False)[0]) + item = lanes.pick(doc, archived, cfg.lanes, cfg.slice_above, owner=True)[0] + n = {k: sum(len(s.items) for s in doc.sections if s.key == k) for k in tasks.SECTIONS} + rows.append((name, f"pending {n['pending']} human {n['human']} awaiting {n['awaiting']} " + f"errors {errors} next: " + (f"{item.id} {item.title}" if item else "-"))) + except (config.ConfigError, tasks.TaskError, OSError) as e: + rows.append((name, f"broken: {e}")) + width = max((len(n) for n, _ in rows), default=0) + for name, text in rows: + line = f"{name:<{width}} {text}" + print(line if len(line) <= WIDTH else line[:WIDTH - 1] + "…") + if not rows: + raise Failure(f"no projects under {root}") + return 0 + + +def claude_projects() -> Path: + return Path(os.environ.get("CLAUDE_CONFIG_DIR") or Path.home() / ".claude") / "projects" + + +def since_time(value: str | None) -> str | None: + """ISO (UTC) as is; 90m / 2h / 1d = that long ago.""" + m = re.fullmatch(r"(\d+)([mhd])", value or "") + if not m: + return value + delta = datetime.timedelta(**{{"m": "minutes", "h": "hours", "d": "days"}[m[2]]: int(m[1])}) + return (datetime.datetime.now(datetime.timezone.utc) - delta).strftime("%Y-%m-%dT%H:%M:%S") + + +def agent_label(path: Path) -> str: + id = path.stem.removeprefix("agent-") + try: + desc = json.loads(path.with_name(path.stem + ".meta.json").read_text()).get("description") or "" + except (OSError, ValueError, AttributeError): + desc = "" + return f"{id} {desc}".strip() + + +def usage_transcripts(args) -> list[tuple[str, Path]]: + """(label, jsonl): one agent, or a session's main transcript + its subagents.""" + base = claude_projects() + if args.agent: + id = args.agent.removeprefix("agent-") + found = sorted(base.glob(f"*/*/subagents/agent-{id}.jsonl")) + if not found: + raise Failure(f"no transcript for agent {args.agent}") + return [(agent_label(found[0]), found[0])] + sid = args.session or os.environ.get("CLAUDE_CODE_SESSION_ID") + if sid: + found = sorted(base.glob(f"*/{sid}*.jsonl")) + if len(found) != 1: + raise Failure(f"session {sid}: {len(found)} transcripts match") + main = found[0] + else: + folder = base / re.sub(r"[^A-Za-z0-9]", "-", str(Path.cwd())) + found = sorted(folder.glob("*.jsonl"), key=lambda p: p.stat().st_mtime) + if not found: + raise Failure(f"no transcripts in {folder} (--session ID)") + main = found[-1] + agents = sorted((main.parent / main.stem / "subagents").glob("agent-*.jsonl")) + return [("main", main)] + [(agent_label(p), p) for p in agents] + + +def cmd_usage(args) -> int: + since = since_time(args.since) + if args.report: + return usage_report(since) + if args.explore: + return usage_explore(args, since) + if args.log and not args.agent: + raise Usage("--log needs --agent") + if args.log: + root = config.find_root(Path(args.project) if args.project else Path.cwd()) + path = usage_transcripts(args)[0][1] + now = datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + line = usage.log_line(now, root.name, args.log[0], args.effort or "-", args.log[1], + path.stem.removeprefix("agent-"), usage.parse(read(path), since=since), lane=args.lane, dur=args.duration) + (root / "out").mkdir(exist_ok=True) + with open(root / "out" / "wf-cost.log", "a") as f: + f.write(line + "\n") + print(line) + return 0 + rows, total, dollars, unknown = [], usage.Usage(), 0.0, False + for label, path in usage_transcripts(args): + for model, u in usage.parse(read(path), since=since).items(): + c = usage.cost(model, u) + total.add(u) + dollars += c or 0 + unknown |= c is None + rows.append((label, usage.short(model), u, "?" if c is None else f"{c:.2f}")) + width = min(max([len(r[0]) for r in rows] + [5]), 40) + + def nums(u): + out = ("~" if u.est else "") + usage.fmt(u.out) # ~ = estimated from content (subagent) + return " ".join(f"{usage.fmt(n):>7}" for n in (u.inp, u.cw, u.cr)) + f" {out:>7}" + + print(f"{'agent':<{width}} {'model':<12} {'turns':>6} {'in':>7} {'cw':>7} {'cr':>7} {'out':>7} {'$':>7}") + for label, model, u, c in rows: + label = label if len(label) <= width else label[:width - 1] + "…" + print(f"{label:<{width}} {model:<12} {u.turns:>6} {nums(u)} {c:>7}") + print(f"{'total':<{width}} {'':<12} {total.turns:>6} {nums(total)} {dollars:>7.2f}{'+' if unknown else ''}") + return 0 + + +def usage_explore(args, since: str | None) -> int: + root = str(Path.cwd()) + agents = [(label, usage.explore(read(path), since=since, root=root)) + for label, path in usage_transcripts(args) if label != "main"] + r = usage.explore_report(agents) + if not r["rows"]: + print("no subagent with >= 4 calls") + return 0 + width = min(max([len(x[0]) for x in r["rows"]] + [5]), 40) + print(f"{'agent':<{width}} {'calls':>5} {'1stEdit':>7} {'explore':>7} {'pre-edit':>8}") + for label, calls, fe, exp, pre in r["rows"]: + label = label if len(label) <= width else label[:width - 1] + "…" + print(f"{label:<{width}} {calls:>5} {fe:>7} {exp:>7.0%} {pre:>8.0%}") + tot = r["total"] or 1 + print(f"ALL {r['agents']} agents, {usage.fmt(r['total'])} billed input: " + + ", ".join(f"{k} {100 * v // tot}%" for k, v in sorted(r["cost"].items(), key=lambda kv: -kv[1]))) + if r["pre_n"]: + print(f"before 1st edit (median, {r['pre_n']} agents): {r['pre_share']:.0%} of input, " + f"{r['pre_calls']:g} calls, +{usage.fmt(int(r['pre_ctx']))} context") + rt = sum(r["results"].values()) or 1 + print("result tokens: " + ", ".join(f"{k} {usage.fmt(v)} {100 * v // rt}%" + for k, v in sorted(r["results"].items(), key=lambda kv: -kv[1]))) + if r["files"]: + print("files read by >= 2 agents (agents, tokens):") + for f, a, n in r["files"]: + print(f"{a:>4} {usage.fmt(n):>7} {f}") + return 0 + + +def fmt_dur(s) -> str: + return "-" if s is None else f"{int(s) // 3600}h{int(s) % 3600 // 60:02d}m" if s >= 3600 else f"{int(s) // 60}m{int(s) % 60:02d}s" + + +def usage_report(since: str | None) -> int: + root = projects_root() + entries = [] + for folder in find_projects(root): + log = folder / "out" / "wf-cost.log" + if log.is_file(): + entries += usage.parse_log(read(log), since=since) + if not entries: + raise Failure(f"no out/wf-cost.log lines in the projects under {root}") + print(f"{'lane':<8} {'effort':<6} {'n':>4} {'done':>4} {'other':>5} {'med$':>7} {'total$':>8} {'$/done':>7} med_turns med_dur") + for lane, effort, n, done, other, med, total, per, turns, dur in usage.report(entries, tasks.EFFORTS): + per = "-" if per is None else f"{per:.2f}" + print(f"{lane:<8} {effort:<6} {n:>4} {done:>4} {other:>5} {med:>7.2f} {total:>8.2f} {per:>7} {turns:>9} {fmt_dur(dur):>8}") + return 0 + + +# ------------------------------------------------------------------ write commands + +def split_list(value: str) -> list[str]: + return tasks.split_refs(value) + + +def stdin_lines() -> list[str]: + return sys.stdin.read().replace("\r\n", "\n").split("\n") + + +def set_slices_line(parent: tasks.Item, id: str) -> None: + at = parent._line(tasks.SLICES_RE) + if at is None: + parent.body.insert(tasks._tail_start(parent), f"{tasks.INDENT}- Slices: [[{id}]]") + else: + parent.body[at] = parent.body[at].rstrip() + f", [[{id}]]" + + +def cmd_add(args) -> int: + p = load_project(args, write=True) + doc = p.doc + if args.title == "-": + item = tasks.parse_block(sys.stdin.read()) + key = args.section or ("awaiting" if item.id.startswith("a-") else "pending") + else: + key = args.section or "pending" + prio, after = args.prio, split_list(args.after or "") + text = " ".join(args.title.split()) + leading = tasks.LEADING_ID_RE.match(text) + if leading and not args.id: + args.id, text = leading.groups() + if args.parent: + parent_section, _ = doc.find(args.parent) + parent = doc.item(args.parent) + key = args.section or parent_section.key + prio = parent.prio if prio is None else prio + id = args.id or tasks.slice_id(args.parent, p.taken()) + m = re.match(re.escape(args.parent) + r"-(\d+)$", id) + prev = parent.slices or ([f"{args.parent}-{int(m.group(1)) - 1}"] if m and int(m.group(1)) > 1 else []) + if not after: + # a deferred previous slice would block a live one forever: chain to the last live slice + live = [s for s in prev if s in p.archived or key == "deferred" + or (s in doc.ids() and doc.find(s)[0].key != "deferred")] + after = live[-1:] + task = key != "awaiting" + if task and prio is None: + raise Usage("add: a task needs -p 0-3") + if task and args.effort is None: + raise Usage("add: a task needs -e " + "|".join(tasks.EFFORTS)) + if task and args.effort not in tasks.EFFORTS: + raise Failure(f"effort '{args.effort}' (want {', '.join(tasks.EFFORTS)})") + if not text: + raise Usage("add: empty title") + if not text.endswith((".", "?", "!")): + text += "." + if not args.parent: + title = text.split(". ", 1)[0] + id = args.id or tasks.make_id(title, p.taken(), "t-" if task else "a-") + if not tasks.ID_RE.match(id): + raise Failure(f"bad id '{id}' (want t-… or a-…, lowercase a-z 0-9 -)") + if id in p.archived: + raise Failure(f"id '{id}' already used in the archive (ids are never reused)") + item = tasks.Item(id=id, prio=prio if task else None, effort=args.effort if task else None, text=text) + if args.body: + item.body = tasks.indent_body(stdin_lines()) + if args.interactive: + old_spelling("--interactive", "--sessions owner") + args.sessions = args.sessions or "owner" + if args.done and task: + item.body.append(f"{tasks.INDENT}Done: {args.done}") + if args.sessions and args.sessions != "parallel" and task: + item.body.append(f"{tasks.INDENT}Sessions: {args.sessions}") + if args.model and task: + item.body.append(f"{tasks.INDENT}Model: {args.model}") + if args.cloud and task: + item.body.append(f"{tasks.INDENT}Cloud: {args.cloud}") + if after: + item.body.append(f"{tasks.INDENT}- After: " + ", ".join(f"[[{a}]]" for a in after)) + if args.ref: + item.body.append(f"{tasks.INDENT}Ref: " + ", ".join(split_list(args.ref))) + if args.parent: + set_slices_line(parent, id) + tasks.insert(doc, item, key) + p.save(args.dry_run) + if not args.dry_run: + print(item.header()) + if item.prio is not None and item.done_text is None: + if item.prio == 0: + print("wf: warning: P0 without Done line is not runner-pickable; use --done \"<text>\"", file=sys.stderr) + else: + print("hint: no Done line; add one (--done) so runners can pick it", file=sys.stderr) + return 0 + + +RELEARN_LINE = "re-learned anything (>3 greps to find)? one anchor line → that area's code map" + + +def area_tasks(p: Project) -> list[str]: + """Add t-map-<area> refresh tasks for stale areas without an open one; returns output lines.""" + file = p.cfg.areas_file + if not file.is_file(): + return [] + out = [] + for a in areas_mod.parse(file.read_text(encoding="utf-8", errors="replace")): + if not areas_mod.status(p.cfg.code, a, p.cfg.area_stale_commits)[0]: + continue + prefix = f"t-map-{a.slug}" + if any(i == prefix or i.startswith(prefix + "-") for i in p.doc.ids()): + continue + id, n = prefix, 1 + while id in p.archived: + n += 1 + id = f"{prefix}-{n}" + item = tasks.Item(id, prio=1, effort="<1h", text=f"Refresh {a.name} area map.") + item.body = tasks.indent_body([ + f"Steps: wf areas {a.name}; fix missing anchors + test recipe from git log --stat " + f"<Checked>..HEAD -- <paths>; wf areas --mark {a.name}.", + f"Done: wf areas {a.name} shows ok.", + "Model: sonnet"]) + tasks.insert(p.doc, item, "pending") + out.append(f"added {id} (area map stale)") + return out + + +def uncovered_areas(p: Project, args) -> list[tuple[str, int]]: + """Folders this task's diff (branch since its base + uncommitted) touches outside every area's Paths.""" + file = p.cfg.areas_file + if not file.is_file() or p.cfg.code_root: # code in another repo: this diff is not its code + return [] + start = Path(getattr(args, "project", None) or Path.cwd()).resolve() + here = _here(start, p.cfg) + found = [areas_mod.with_anchor_paths(here, a) + for a in areas_mod.parse(file.read_text(encoding="utf-8", errors="replace"))] + if not any(a.paths for a in found): + return [] + wt = config.linked_worktree(start) + if wt: + base = config.git_branch(wt[2] / ".git") or "master" + else: + base = next((b for b in ("master", "main") + if git_run(here, "rev-parse", "-q", "--verify", f"refs/heads/{b}").returncode == 0), None) + books = {p.cfg.rel(f) for f in (p.cfg.tasks, p.cfg.archive, p.cfg.root / config.NAME)} | {p.cfg.areas_rel} + files = [f for f in areas_mod.changed_files(here, base) if f not in books] + return areas_mod.uncovered(files, found, p.cfg.area_ignore) + + +UNCOVERED_MAX = 3 # map tasks one wf done adds at most + + +def uncovered_tasks(p: Project, args) -> list[str]: + """Add t-map-<folder> tasks for folders the diff touches outside every area's Paths (once per id).""" + out = [] + for folder, n in uncovered_areas(p, args)[:UNCOVERED_MAX]: + id = f"t-map-{areas_mod.slug(folder)}" + if any(i == id or i.startswith(id + "-") for i in [*p.doc.ids(), *p.archived]): + continue + rel = p.cfg.areas_rel + item = tasks.Item(id, prio=1, effort="<1h", text=f"Map {folder} area. No area's Paths covers it.") + item.body = tasks.indent_body([ + f"Steps: git log --stat -- {folder}; new ### section under ## Areas in {rel} (Code map anchors, " + f"Test recipe, Paths: {folder}/) or add {folder}/ to an existing area's Paths; wf areas --mark <area>.", + f"Done: wf areas lists the area ok and no longer names {folder} uncovered.", + "Model: sonnet"]) + tasks.insert(p.doc, item, "pending") + out.append(f"added {id} (uncovered area: {folder}, {n} file{'s' * (n != 1)})") + return out + + +def cmd_done(args) -> int: + p = load_project(args, write=True) + doc = p.doc + task_ids = [i for i in args.ids if not i.startswith("a-")] + if len(task_ids) == 1 and not args.m: + raise Usage('done: -m "<entry>" is required for a single task') + if len(task_ids) > 1 and args.m: + raise Usage("done: -m goes with one task; several ids use each task's goal") + out = [] + today = datetime.date.today().isoformat() + for id in args.ids: + doc.find(id) + if not config.linked_worktree(Path(args.project or Path.cwd()).resolve()): + for id in task_ids: + if wt := claimed_worktree(p.cfg.root, doc.item(id)): + raise Failure(f"{id} is in progress in worktree {wt[0]} (branch {wt[1]}): run wf finish/done there " + f"(its bookkeeping goes in with wf merge), or wf status {id} clear first") + before = lanes.pickable_ids(doc, p.archived) + lane = lambda i: lanes.lane_of(i, p.cfg.lanes, p.cfg.slice_above) + done_lanes = {lane(doc.item(i)) for i in task_ids} + solo = [i for i in task_ids if doc.item(i).sessions == "solo"] + for id in args.ids: + if id.startswith("a-"): + tasks.remove(doc, id) + out.append(f"removed: {id}") + freed = tasks.unblock(doc, id) + if freed: + out.append("unblocked: " + ", ".join(freed)) + continue + open_ = [s for s in tasks.open_slices(doc, id) if s not in args.ids] + if open_: + raise Failure(f"'{id}' has open slices: {', '.join(open_)}") + item = tasks.remove(doc, id) + line = tasks.archive_line(today, item, args.m or item.goal) + p.new_archive = tasks.archive_prepend(p.new_archive or p.archive_text, line) + out.append(f"done: {id} → {p.cfg.rel(p.cfg.archive)}") + parent = tasks.parent_of(doc, id) + if parent and not tasks.open_slices(doc, parent): + out.append(f"last slice of {parent}: finish it with wf done {parent} -m \"…\"") + if task_ids and not args.dry_run: + out += area_tasks(p) + out += uncovered_tasks(p, args) + p.save(args.dry_run) + if args.dry_run: + return 0 + unclaim(p.cfg, task_ids) + after = lanes.pickable_ids(doc, p.archived | set(task_ids)) + freed = [i for s in doc.sections if s.key == "pending" for i in s.items + if i.id in after - before and lane(i) not in done_lanes] + if task_ids: + out += lanes.notify_block(freed, sessions(p.cfg), p.cfg.lanes, p.cfg.slice_above) + me = my_session() + out += lanes.solo_done_block(solo, sessions(p.cfg), me[1] if me else None) + if task_ids: + if p.cfg.verify: + out += ["verify:", *(f" {v}" for v in p.cfg.verify)] + out += ["checklist:", f" - {RELEARN_LINE}", *(f" - {d}" for d in p.cfg.done)] + if not getattr(args, "finishing", False): + out += worktree_done_lines(p, args) + if hint := ctx_hint_line(p.cfg): + out.append(hint) + print("\n".join(out)) + return 0 + + +CTX_TAIL = 1 << 20 # bytes of the transcript read for the last request + + +def ctx_hint_line(cfg) -> str | None: + """This session's prompt size over cfg.ctx_hint → a /clear hint; no session id / transcript → None.""" + sid = os.environ.get("CLAUDE_CODE_SESSION_ID") + if not cfg.ctx_hint or not sid: + return None + found = sorted(claude_projects().glob(f"*/{glob.escape(sid)}.jsonl")) + if len(found) != 1: + return None + try: + with open(found[0], "rb") as f: + f.seek(max(0, f.seek(0, 2) - CTX_TAIL)) + text = f.read().decode("utf-8", "replace") + except OSError: + return None + return usage.ctx_hint(usage.context_tokens(text), cfg.ctx_hint) + + +def claimed_worktree(root: Path, item: tasks.Item) -> tuple[str, str] | None: + """(worktree path, branch) of a linked worktree on the branch named by item's 'in progress: <branch>'.""" + status = item.status or "" + if not status.startswith("in progress: "): + return None + words = status.removeprefix("in progress: ").split() + branch = words[0] if words else "" + path = None + for line in git_run(root, "worktree", "list", "--porcelain").stdout.splitlines(): + if line.startswith("worktree "): + path = line.removeprefix("worktree ") + elif branch and line == f"branch refs/heads/{branch}" and path and Path(path).resolve() != root.resolve(): + if config.linked_worktree(Path(path)): + return path, branch + return None + + +def worktree_done_lines(p: Project, args) -> list[str]: + """Merge-back steps when wf done runs in a linked worktree (multi-session mode).""" + wt = config.linked_worktree(Path(args.project or Path.cwd()).resolve()) + if not wt: + return [] + branch, main = config.git_branch(wt[1]), wt[2] + master = config.git_branch(main / ".git") or "master" + files = " ".join(str(f.relative_to(main)) for f in (p.cfg.tasks, p.cfg.archive)) + warn = [] + if not branch and (n := unmerged_count(wt[0], master)): + warn = [f"detached HEAD has {n} commit{'s' * (n != 1)} not in {master}: wf merge merges " + f"{'them' if n != 1 else 'it'} (never leave them unmerged)"] + return warn + [f"worktree mode (branch {branch or 'none: detached HEAD'}), after verify: commit your code here (explicit paths), then:", + f" wf merge (rebase, ff-merge into {master}, commit {files}, push home; " + f"conflict → git rebase {master}, resolve, verify, wf merge again)"] + + +def git_run(cwd: Path, *args: str) -> subprocess.CompletedProcess: + return subprocess.run(["git", *args], cwd=cwd, capture_output=True, text=True) + + +def unmerged_count(top: Path, master: str) -> int: + """Commits on this worktree's HEAD not in master (0 when git fails).""" + r = git_run(top, "rev-list", "--count", f"{master}..HEAD") + return int(r.stdout.strip()) if r.returncode == 0 and r.stdout.strip().isdigit() else 0 + + +def cmd_merge(args) -> int: + """Merge-back of a lane worktree's branch, under the project lock (main() takes it).""" + start = Path(args.project or Path.cwd()).resolve() + wt = config.linked_worktree(start) + if not wt: + raise Failure("merge runs inside a linked git worktree (lane worktree)") + top, gitdir, main = wt + branch = config.git_branch(gitdir) + master = config.git_branch(main / ".git") or "master" + cfg = config.load_at(start) + files = [str(f.relative_to(main)) for f in (cfg.tasks, cfg.archive)] + if git_run(top, "status", "--porcelain").stdout.strip(): + raise Failure("worktree has uncommitted changes: commit them first") + if not branch and not unmerged_count(top, master): + # after a merge: notes / follow-ups written since → bookkeeping commit only + if not git_run(main, "status", "--porcelain", "--", *files).stdout.strip(): + raise Failure("worktree is on a detached HEAD: nothing to merge") + r = git_run(main, "commit", "-q", "-m", args.m or "bookkeeping", "--", *files) + if r.returncode: + raise Failure(f"bookkeeping commit failed: {(r.stderr.strip() or r.stdout.strip() or 'git error').splitlines()[-1]}") + print(f"committed {' '.join(files)}") + return 0 + out = [] # a branch, or a detached HEAD with commits not in master (never left behind) + if git_run(top, "rebase", master).returncode: + git_run(top, "rebase", "--abort") + raise Failure(f"rebase onto {master} conflicts: git rebase {master}, resolve, verify, then wf merge again") + out.append(f"rebased onto {master}") + args.merged_sha = git_run(top, "rev-parse", "--short", "HEAD").stdout.strip() + r = git_run(main, "merge", "--ff-only", branch or git_run(top, "rev-parse", "HEAD").stdout.strip()) + if r.returncode: + raise Failure(f"ff-merge into {master} failed: {(r.stderr.strip() or 'git error').splitlines()[-1]}") + out.append(f"fast-forwarded {master}") + if git_run(main, "status", "--porcelain", "--", *files).stdout.strip(): + msg = args.m or (f"{branch.rsplit('/', 1)[-1]} done" if branch else "bookkeeping") + r = git_run(main, "commit", "-q", "-m", msg, "--", *files) + if r.returncode: + raise Failure(f"bookkeeping commit failed: {(r.stderr.strip() or r.stdout.strip() or 'git error').splitlines()[-1]}") + out.append(f"committed {' '.join(files)}") + git_run(top, "switch", "-q", "--detach", master) + if branch: + git_run(top, "branch", "-q", "-d", branch) + if not args.no_push and _push_home(main): + out.append("pushed home") + out.append(f"merged {branch or 'detached HEAD'} into {master}") + print("\n".join(out)) + return 0 + + +def _push_home(main: Path) -> bool: + """git push home --all/--tags when remote home exists; True if pushed.""" + if "home" not in git_run(main, "remote").stdout.split(): + return False + for extra in ("--all", "--tags"): + r = git_run(main, "push", "-q", "home", extra) + if r.returncode: + raise Failure(f"push home failed: {(r.stderr.strip() or 'git error').splitlines()[-1]}") + return True + + +def _dirty_outside(top: Path, keep: list[Path]) -> list[str]: + """Uncommitted paths of the repo at `top` not under any of `keep` (absolute paths).""" + stray = [] + for line in git_run(top, "status", "--porcelain").stdout.splitlines(): + rel = line[3:].split(" -> ")[-1].strip('"').rstrip("/") + path = (top / rel).resolve() + if not any(path == k or k in path.parents for k in keep): + stray.append(rel) + return stray + + +GATE_RED_HELP = ( + "\ngate red procedure (also for a red gate run after wf finish, on the merged sha):" + "\n- culprit = your diff (git show --stat <sha>) explains the failure -> fix now (new commit, wf merge) or hand back" + "\n- not your diff (earlier task broke it; see git log of the failing area) -> wf add -p 0 --model <m> -e <effort>" + ' --done "gate ALL GREEN" -b "Fix gate red <test>. <goal>" with body line "Steps: <failing test, culprit sha/id>"' + "\n- then report result: done+gate-red <culprit sha or id> <fix id> (your task is done, the lane goes on with the fix)") + + +def cmd_finish(args) -> int: + """gate → done → commit the given paths → merge (lane worktree), in one call; checks first, so a refusal + leaves the task open.""" + if args.paths and not args.commit: + raise Usage("finish: paths need --commit MSG") + start = Path(args.project or Path.cwd()).resolve() + root = config.find_root(start) + cfg = config.load(root) + wt = config.linked_worktree(start) + top = Path(git_run(start, "rev-parse", "--show-toplevel").stdout.strip() or start) + paths = [(start / f).resolve() for f in args.paths] + books = [] if wt else [cfg.tasks.resolve(), cfg.archive.resolve()] + for f, path in zip(args.paths, paths): + if path != top and top not in path.parents: + raise Failure(f"path {f} is outside this repo ({top}): commit it in its own repo first, " + "then wf finish without it") + if wt and path in (top / cfg.tasks.relative_to(wt[2]), top / cfg.archive.relative_to(wt[2])): + raise Failure(f"path {f} = this worktree's copy of the books: wf finish writes the main tree's " + "and wf merge commits them; drop it") + if not path.exists() and git_run(top, "ls-files", "--error-unmatch", "--", str(path)).returncode: + raise Failure(f"path {f} does not exist and is not tracked: nothing to commit") + if wt and paths and not git_run(top, "status", "--porcelain", "--", *map(str, paths)).stdout.strip(): + raise Failure("nothing to commit in the --commit paths: drop them (wf finish without paths) or fix them") + if stray := _dirty_outside(top, paths + books): + raise Failure("uncommitted changes outside the --commit paths: " + " ".join(stray)) + if cfg.quick_gate: + _run_lines(cfg.quick_gate, _here(start, cfg), cfg, + "quick_gate '{line}' red (exit {rc}): fix it, then wf finish again; later commands skipped" + GATE_RED_HELP) + print(f"quick_gate: {len(cfg.quick_gate)} green", flush=True) + with project_lock(root): + p = load_project(args) + if args.id not in p.doc.ids() and args.id in p.archived: + print(f"{args.id} already done ({cfg.rel(cfg.archive)}): resuming with commit + merge", flush=True) + else: + done = argparse.Namespace(project=args.project, ids=[args.id], m=args.m, dry_run=False, finishing=True) + cmd_done(done) + sys.stdout.flush() + resume = f" ({args.id} is done already: fix it, then rerun the same wf finish, it resumes)" + commit = list(dict.fromkeys(map(str, [*paths, *(books if args.commit or not wt else [])]))) + if commit: + r = git_run(top, "add", "--", *map(str, paths)) if paths else None + if r and r.returncode: + raise Failure(f"git add failed: {(r.stderr.strip() or 'git error').splitlines()[-1]}" + resume) + msg = args.commit or f"{args.id} done" + r = git_run(top, "commit", "-q", "-m", msg, "--", *commit) + if r.returncode: + raise Failure(f"commit failed: {(r.stderr.strip() or r.stdout.strip() or 'git error').splitlines()[-1]}" + + resume) + print("committed " + " ".join(os.path.relpath(c, top) for c in commit), flush=True) + if not wt: + sha = git_run(top, "rev-parse", "--short", "HEAD").stdout.strip() + tool = getattr(args, "tool_commit", None) + print(f"report: commit {sha}" + (f" tool {tool}" if tool else ""), flush=True) + if not wt and not args.no_push and _push_home(top): + print("pushed home", flush=True) + if wt: + ns = argparse.Namespace(project=args.project, m=None, no_push=args.no_push) + rc = cmd_merge(ns) + if not rc and getattr(ns, "merged_sha", ""): + tool = getattr(args, "tool_commit", None) + print(f"report: commit {ns.merged_sha}" + (f" tool {tool}" if tool else ""), flush=True) + return rc + return 0 + + +def cmd_wip(args) -> int: + """Wrap-up in one call (lane worktree): commit the given paths on the branch, note the state, status clear.""" + start = Path(args.project or Path.cwd()).resolve() + if not config.linked_worktree(start): + raise Failure("wip runs inside a lane worktree") + top = Path(git_run(start, "rev-parse", "--show-toplevel").stdout.strip() or start) + paths = [str((start / f).resolve()) for f in args.paths] + if paths: + r = git_run(top, "add", "--", *paths) + if r.returncode: + raise Failure(f"git add failed: {(r.stderr.strip() or 'git error').splitlines()[-1]}") + r = git_run(top, "commit", "-q", "-m", args.commit or f"{args.id} WIP", "--", *paths) + if r.returncode: + raise Failure(f"commit failed: {(r.stderr.strip() or r.stdout.strip() or 'git error').splitlines()[-1]}") + print("committed " + " ".join(os.path.relpath(c, top) for c in paths), flush=True) + p = load_project(args) + id = p.resolve(args.id) + with project_lock(p.cfg.root): + cmd_note(argparse.Namespace(project=args.project, id=id, line=args.m, dry_run=False)) + cmd_status(argparse.Namespace(project=args.project, id=id, kind="clear", value=[], dry_run=False, + clear_stale=False)) + return 0 + + +def _here(start: Path, cfg) -> Path: + """Project folder to run config commands in: the worktree's copy of cfg.root, else cfg.root.""" + wt = config.linked_worktree(start) + if not wt: + return cfg.root + top, _, main = wt + try: + return top / cfg.root.relative_to(main) + except ValueError: + return top + + +def _run_lines(lines: list[str], here: Path, cfg, fail: str) -> None: + """Run shell lines in `here` with env WF_MAIN = main tree project folder; first nonzero → Failure.""" + env = {**os.environ, "WF_MAIN": str(cfg.root)} + for line in lines: + print(f"$ {line}", flush=True) + rc = subprocess.run(line, shell=True, cwd=here, env=env).returncode + if rc: + raise Failure(fail.format(line=line, rc=rc)) + + +def cmd_setup(args) -> int: + """Run workflow.toml worktree_setup in this lane worktree (cwd = its project folder, env WF_MAIN).""" + start = Path(args.project or Path.cwd()).resolve() + wt = config.linked_worktree(start) + if not wt: + raise Failure("setup runs inside a linked git worktree (lane worktree)") + cfg = config.load_at(start) + if not cfg.worktree_setup: + print(f"no worktree_setup in {config.NAME}: nothing to do") + return 0 + _run_lines(cfg.worktree_setup, _here(start, cfg), cfg, + "worktree_setup '{line}' failed (exit {rc}): later commands skipped") + n = len(cfg.worktree_setup) + print(f"worktree_setup: {n} command{'s' * (n != 1)} ok in {os.path.relpath(wt[0], cfg.root)}") + return 0 + + +def cmd_start(args) -> int: + """Worker setup in one call (main tree): worktree + branch (clean check), worktree_setup, status progress, + then wf ctx and a ready verify && wf finish line.""" + start = Path(args.project or Path.cwd()).resolve() + if config.linked_worktree(start): + raise Failure("start runs in the main tree (it creates the worktree)") + p = load_project(args) + id = p.resolve(args.id) + main = Path(git_run(start, "rev-parse", "--show-toplevel").stdout.strip() or p.cfg.root) + master = config.git_branch(main / ".git") or "master" + wt = (Path.cwd() / args.worktree).resolve() + shown = p.cfg.rel(wt) + b = args.branch + + def git_ok(cwd, *a): + r = git_run(cwd, *a) + if r.returncode: + raise Failure(f"git {a[0]} failed: {(r.stderr.strip() or r.stdout.strip() or 'git error').splitlines()[-1]}") + return r.stdout + + has_branch = not git_run(main, "rev-parse", "--verify", "-q", f"refs/heads/{b}").returncode + extra = [] + if not wt.exists(): + if has_branch: + git_ok(main, "worktree", "add", "-q", str(wt), b) + how = f"new, existing branch {b}: earlier WIP, read the notes" + else: + git_ok(main, "worktree", "add", "-q", str(wt), "-b", b, master) + how = f"new, branch {b} from {master}" + else: + dirty = git_run(wt, "status", "--porcelain").stdout.rstrip("\n") + cur = config.git_branch(Path(git_run(wt, "rev-parse", "--absolute-git-dir").stdout.strip())) + if dirty and not args.recovery: + raise Failure(f"worktree {shown} has uncommitted changes: " + + " ".join(l.strip() for l in dirty.splitlines()) + + " (hand back, or --recovery when a dead worker left them)") + if dirty and cur != b: + raise Failure(f"worktree {shown} has uncommitted changes on another branch ({cur or 'detached'}), " + f"not {b}: hand back") + if cur == b: + how = f"on {b}" + elif has_branch: + git_ok(wt, "switch", "-q", b) + how = f"switched to existing branch {b}: earlier WIP, read the notes" + else: + git_ok(wt, "switch", "-q", "-c", b, master) + how = f"switched to new branch {b} from {master}" + if args.recovery: + log = git_run(wt, "log", "--oneline", f"{master}..HEAD").stdout.strip() + extra = [f"recovery: uncommitted:\n{dirty}" if dirty else "recovery: uncommitted: none", + f"recovery: commits {master}..HEAD:" + (f"\n{log}" if log else " none")] + print(f"worktree: {shown} ({how})", *extra, sep="\n", flush=True) + try: + here = wt / p.cfg.root.relative_to(main) + except ValueError: + here = wt + cmd_setup(argparse.Namespace(project=str(here))) + sys.stdout.flush() + with project_lock(p.cfg.root), contextlib.redirect_stdout(io.StringIO()): + cmd_status(argparse.Namespace(project=args.project, id=id, kind="progress", value=[b], dry_run=False, + clear_stale=False)) + print(flush=True) + cmd_ctx(argparse.Namespace(project=args.project, target=id)) + steps = [f"cd {here}", *(f"({v})" for v in p.cfg.verify), + f'python3 {HERE / "wf.py"} finish {id} -m "<entry>" --commit "<msg + footer>" <paths>'] + print("\nFinish (after the work; fill in the quoted parts and the paths):\n " + " && ".join(steps)) + return 0 + + +def cmd_gate(args) -> int: + """Run workflow.toml quick_gate (fast regression check) in this tree's project folder, env WF_MAIN.""" + start = Path(args.project or Path.cwd()).resolve() + cfg = config.load_at(start) + if not cfg.quick_gate: + print(f"no quick_gate in {config.NAME}: nothing to do") + return 0 + _run_lines(cfg.quick_gate, _here(start, cfg), cfg, + "quick_gate '{line}' red (exit {rc}): fix it before wf done; later commands skipped" + GATE_RED_HELP) + n = len(cfg.quick_gate) + print(f"quick_gate: {n} command{'s' * (n != 1)} green") + return 0 + + +# ------------------------------------------------------------------ orchestrator + +GO_ON = ("done", "done+gate-red", "sliced") # outcomes after which the lane picks its next task + + +def orch_main(args) -> tuple[Path, Project]: + """(main tree, project) for wf orch; refuses a linked worktree.""" + start = Path(args.project or Path.cwd()).resolve() + if config.linked_worktree(start): + raise Failure("orch runs in the main tree (the orchestrator's)") + p = load_project(args) + main = Path(git_run(start, "rev-parse", "--show-toplevel").stdout.strip() or p.cfg.root) + return main, p + + +def orch_record(cfg: config.Config, id: str) -> dict: + try: + r = json.loads((cfg.root / ".wf" / "orch" / f"{id}.json").read_text()) + except (OSError, ValueError): + return {} + return r if isinstance(r, dict) else {} + + +def live_session_cwds() -> list[Path]: + """cwd of every live Claude Code session (~/.claude/sessions/*.json).""" + out = [] + for f in (claude_projects().parent / "sessions").glob("*.json"): + try: + s = json.loads(f.read_text()) + os.kill(int(s["pid"]), 0) + out.append(Path(s["cwd"]).resolve()) + except (OSError, ValueError, KeyError, TypeError): + continue + return out + + +def worktree_free(path: Path, branch: str, busy: set[Path], cwds: list[Path]) -> bool: + """No live session in it, no other picked worker on it, and clean on a detached HEAD (or on branch = recovery).""" + if path in busy: + return False + if not path.exists(): + return True + if any(c == path or path in c.parents for c in cwds): + return False + cur = config.git_branch(Path(git_run(path, "rev-parse", "--absolute-git-dir").stdout.strip() or path / ".git")) + if cur == branch: + return True + return not cur and not git_run(path, "status", "--porcelain").stdout.strip() + + +def lane_worktree(main: Path, p: Project, lane: str, id: str) -> Path: + """First free main/.worktrees/<lane>[-n] for branch <lane>/<id> (worktree_free; other picked workers' trees busy).""" + cfg, busy = p.cfg, set() + for f in (cfg.root / ".wf" / "orch").glob("*.json"): + r = orch_record(cfg, f.stem) + other = p.doc.item(f.stem) if f.stem in p.doc.ids() else None + if f.stem != id and r.get("worktree") and other and (other.status or "").startswith("in progress"): + busy.add(Path(r["worktree"])) + cwds = live_session_cwds() + n = 1 + while not worktree_free(wt := main / ".worktrees" / (lane if n == 1 else f"{lane}-{n}"), f"{lane}/{id}", busy, cwds): + n += 1 + return wt + + +CLOUD = "cloud" # virtual lane: wf cloud send instead of a local worker (spec cloud-lane §4.6) + + +def cloud_pick(main: Path, p: Project, id: str | None) -> int: + """wf orch pick cloud: ledger check, first fitting task across the lanes (cloud.pick_key) -> wf cloud send.""" + import wf_cloud + from wflib import cloud + cfg = p.cfg + if not cfg.cloud: + print(f"stop lane {CLOUD}: none fit (project not opted in: workflow.toml cloud = true)") + return 0 + with wf_cloud.locked(wf_cloud.state_dir()) as led: + run, cap, bal, res = cloud.running(led), led["max_parallel"], cloud.balance(led), led["reserve_per_task"] + if run >= cap: + print(f"stop lane {CLOUD}: max parallel ({run} running >= max_parallel {cap})") + return 0 + if bal < res: + print(f"stop lane {CLOUD}: ledger (balance ${bal:.2f} < reserve ${res:.2f})") + return 0 + if id: + item = p.doc.item(p.resolve(id)) + if why := cloud.fit(item, cfg): + print(f"stop lane {CLOUD}: none fit ({item.id}: {why})") + return 0 + else: + ready, seen = runner_ids(p), [] + for l in cfg.lanes: + seen += [i for i in lanes.ranked(p.doc, p.archived, cfg.lanes, cfg.slice_above, l.name) + if i not in seen and i.id in ready and not (i.status or "").startswith("in progress")] + fits = sorted((i for i in seen if cloud.fit(i, cfg) is None), key=cloud.pick_key) + if not fits: + print(f"stop lane {CLOUD}: none fit (wf cloud fit rules: wf set ID --cloud yes skips the regex list, opus only)") + return 0 + item = fits[0] + buf = io.StringIO() + with contextlib.redirect_stdout(buf): + code = wf_cloud.main(["send", item.id, "--project", str(cfg.root)]) + sent = buf.getvalue().strip() + if code == 3: + print(f"stop lane {CLOUD}: ledger (send refused, see above)") + return 0 + if code: + if sent: + print(sent) + raise Failure(f"wf cloud send {item.id} failed (exit {code}): lane {CLOUD} stops, tell the owner") + rec = json.loads(wf_cloud.record_path(cfg, item.id).read_text()) + (wf_folder(cfg, "orch") / f"{item.id}.json").write_text(json.dumps( + {"id": item.id, "lane": CLOUD, "task_lane": rec["lane"], "model": item.model, "effort": item.effort, + "sid": rec["sid"], "branch": f"{rec['lane']}/{item.id}", "at": int(datetime.datetime.now().timestamp())}) + "\n") + print(f"pick: {item.id} (lane {CLOUD}, model {item.model}, effort {item.effort or '-'}) · {sent.splitlines()[-1]}" + f" · claimed (in progress: cloud:{rec['sid']})") + print(f"agent: none (cloud session). On each wake (>= 10 min apart): wf cloud pull --all; per ended task " + f"'<id>: <state> …' → wf orch post <id> {CLOUD} --result <state> [--commit <sha>]") + return 0 + + +def orch_pick(main: Path, p: Project, lane: str, id: str | None = None, recovery: str | None = None) -> int: + cfg = p.cfg + stop = cfg.root / "out" / "wf-batch.stop" + if stop.exists() and not recovery: + print(f"stop: {cfg.rel(stop)} exists: spawn nothing (let running workers finish)") + return 0 + if lane == CLOUD: + return cloud_pick(main, p, id) + if id: + id = p.resolve(id) + else: + ready = runner_ids(p) + order = [i for i in lanes.ranked(p.doc, p.archived, cfg.lanes, cfg.slice_above, lane) + if i.id in ready and not (i.status or "").startswith("in progress")] + if not order: + print(f"none: lane {lane} has no runner-ready task (wf list --runner --lane {lane}; wf lanes)") + return 0 + id = order[0].id + item = p.doc.item(id) + branch = f"{lane}/{id}" + wt = lane_worktree(main, p, lane, id) + with project_lock(cfg.root), contextlib.redirect_stdout(io.StringIO()): + cmd_status(argparse.Namespace(project=str(cfg.root), id=id, kind="progress", value=["worker"], + dry_run=False, clear_stale=False)) + model = item.model + (wf_folder(cfg, "orch") / f"{id}.json").write_text(json.dumps( + {"id": id, "lane": lane, "model": model, "effort": item.effort, "worktree": str(wt), "branch": branch, + "at": int(datetime.datetime.now().timestamp())}) + "\n") + kind = ", slice job: the worker only slices" if lanes.is_slice_job(item, cfg.slice_above) else "" + print(f"pick: {id} (lane {lane}, model {model}, effort {item.effort or '-'}{kind}) · claimed (in progress: worker)") + print(f"agent: subagent_type wf-worker, model {model}, no isolation; prompt:") + print(f"Task: {id} Lane: {lane} Model: {model}\nMain tree: {main} Worktree: {wt} Branch: {branch}" + + (f"\nRecovery: {recovery}" if recovery else "") + "\nFinal message: the 4 report lines only.") + return 0 + + +def cost_line(root: Path, agent: str, task: str, outcome: str, effort: str | None, lane: str, dur: int | None) -> str: + """Append one usage line for a worker to <root>/out/wf-cost.log; returns it.""" + path = usage_transcripts(argparse.Namespace(agent=agent, session=None))[0][1] + now = datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + line = usage.log_line(now, root.name, task, effort or "-", outcome, path.stem.removeprefix("agent-"), + usage.parse(read(path)), lane=lane, dur=dur) + (root / "out").mkdir(exist_ok=True) + with open(root / "out" / "wf-cost.log", "a") as f: + f.write(line + "\n") + return line + + +def archived_meta(cfg: config.Config, id: str) -> tuple[str | None, str | None]: + """(model, effort) of a task already archived: its block in TASKS.md just before the commit that removed it.""" + rel = str(cfg.tasks.relative_to(cfg.root)) + r = git_run(cfg.root, "log", "-n1", "--format=%H", f"-S**{id}**", "--", rel) + if r.returncode or not r.stdout.strip(): + return None, None + old = git_run(cfg.root, "show", f"{r.stdout.strip()}^:{rel}") + try: + it = tasks.parse(old.stdout).item(id) if old.returncode == 0 and id in tasks.parse(old.stdout).ids() else None + except Exception: + it = None + return (it.model, it.effort) if it else (None, None) + + +def orch_post(main: Path, p: Project, args) -> int: + cfg, id, lane = p.cfg, args.id, args.lane + rec = orch_record(cfg, id) + wt = Path(rec.get("worktree") or main / ".worktrees" / lane) + item = p.doc.item(id) if id in p.doc.ids() else None + if item is None and id not in p.archived: + raise Failure(tasks.unknown_id(id, p.doc.ids())) + words = (args.result or "").split() + outcome = words[0] if words else ("done" if id in p.archived else "no-report") + if outcome == "done" and item is not None and id not in p.archived and tasks.open_slices(p.doc, id): + outcome = "sliced" # slice job: worker said done but the task is open with slices + master = config.git_branch(main / ".git") or "master" + files = [str(f.relative_to(main)) for f in (cfg.tasks, cfg.archive)] + problems, out = [], [] + with project_lock(cfg.root): + if outcome.startswith("done"): + if id not in p.archived: + problems.append(f"no archive line for {id}") + if (wt / ".git").exists(): + dirty = git_run(wt, "status", "--porcelain").stdout.strip() + if dirty: + problems.append(f"worktree {wt} dirty: " + " ".join(l.strip() for l in dirty.splitlines())) + elif git_run(wt, "merge-base", "--is-ancestor", "HEAD", master).returncode: + buf = io.StringIO() + try: + with contextlib.redirect_stdout(buf): + cmd_merge(argparse.Namespace(project=str(wt), m=None, no_push=args.no_push)) + out.append(f"merged worktree HEAD: {buf.getvalue().strip().splitlines()[-1]}") + except Failure as e: + problems.append(f"wf merge in {wt}: {e}") + branch = rec.get("branch") or f"{lane}/{id}" + if not git_run(main, "rev-parse", "--verify", "-q", f"refs/heads/{branch}").returncode: + problems.append(f"branch {branch} still there") + errors = checks.check(config.load(cfg.root))[0] + if errors: + problems.append(f"wf check: {len(errors)} errors: {errors[0]}") + elif git_run(main, "status", "--porcelain", "--", *files).stdout.strip(): + r = git_run(main, "commit", "-q", "-m", f"{id} {outcome} (orchestrator)", "--", *files) + out.append("committed leftover " + " ".join(files) if not r.returncode + else f"leftover commit failed: {(r.stderr.strip() or 'git error').splitlines()[-1]}") + raised = bool(item and rec.get("model") and not outcome.startswith("done") + and tasks.MODELS.index(item.model) > tasks.MODELS.index(rec["model"])) + final = "post-check-red" if problems else ("model-raised" if raised else outcome) + am, ae = (None, None) if item or (rec.get("model") and rec.get("effort")) else archived_meta(cfg, id) + model = rec.get("model") or (item.model if item else am) or "?" + dur = args.duration if args.duration else ( # None/0 (orchestrator lacked duration_ms) → since the pick + int(datetime.datetime.now().timestamp()) - rec["at"] if isinstance(rec.get("at"), int) else None) + commit = args.commit or (git_run(main, "rev-parse", "--short", master).stdout.strip() + if outcome.startswith("done") else "-") + now = datetime.datetime.now().isoformat(timespec="seconds") + (cfg.root / "out").mkdir(exist_ok=True) + with open(cfg.root / "out" / "wf-orch.log", "a") as f: + f.write(f"{now} {lane} {model} {id} {final} {commit} {fmt_dur(dur) if dur is not None else '-'}" + + (f" ({'; '.join(problems)})" if problems else "") + "\n") + if args.agent: + try: + out.append("cost: " + cost_line(cfg.root, args.agent, id, final, rec.get("effort") or (item and item.effort) or ae, + lane, dur)) + except Failure as e: + out.append(f"cost: not logged ({e})") + if final != "post-check-red": # kept: the re-post after the fix still knows the pick time + (cfg.root / ".wf" / "orch" / f"{id}.json").unlink(missing_ok=True) + print(f"post: {id} {final}" + "".join(f"\n {l}" for l in problems + out)) + if outcome == "done+gate-red" and len(words) >= 3: + q = load_project(args) + fix = q.doc.item(words[2]) if words[2] in q.doc.ids() else None + if fix and not fix.runner_ready: + print(f"fix {fix.id} not runner-ready: add its Done/Model (wf set {fix.id} --done … --model …), " + f"then wf orch pick {lane} --id {fix.id}") + return 0 + if final not in GO_ON and final != "model-raised": + print(f"stop lane {lane}: {final} → tell the owner" + + (f" (crash/no report: one fresh worker: wf orch pick {lane} --id {id} --recovery \"<why>\", then stop)" + if final == "no-report" else "")) + return 0 + if args.no_pick: + return 0 + print() + return orch_pick(main, load_project(args), lane) + + +def cmd_orch(args) -> int: + main, p = orch_main(args) + if args.lane != CLOUD or CLOUD in lane_names(p.cfg): + check_lane(p.cfg, args.lane) + if args.action == "pick": + return orch_pick(main, p, args.lane, args.id, args.recovery) + return orch_post(main, p, args) + + +def change(args, action) -> int: + """Run one edit on one item, save, print its new header.""" + p = load_project(args, write=True) + note = action(p) + p.save(args.dry_run) + if not args.dry_run: + print(note if note else p.doc.item(args.id).header()) + return 0 + + +def cmd_prio(args) -> int: + return change(args, lambda p: tasks.set_prio(p.doc, args.id, args.prio)) + + +def cmd_move(args) -> int: + targets = [t for t in (args.section, args.before, args.after) if t] + if len(targets) != 1: + raise Usage("move: give a section, or --before ID, or --after ID") + if args.section: + return change(args, lambda p: tasks.move_to(p.doc, args.id, args.section)) + return change(args, lambda p: tasks.move_rel(p.doc, args.id, args.before or args.after, + before=bool(args.before), force=args.force)) + + +def clear_stale(args) -> int: + p = load_project(args, write=True) + items = stale(p) + for i in items: + tasks.set_status(p.doc, i.id, None) + p.save(args.dry_run) + if not args.dry_run: + unclaim(p.cfg, [i.id for i in items]) + for i in items: + print(f"cleared {i.id}") + return 0 + + +def cmd_status(args) -> int: + if args.clear_stale: + if args.id or args.kind or args.value: + raise Usage("status: --clear-stale takes no id/kind/value") + return clear_stale(args) + if not args.id or not args.kind: + raise Usage("status: ID and progress NOTE|blocked A-ID|clear (or --clear-stale)") + if args.kind == "clear": + if args.value: + raise Usage("status: clear takes no value") + status = None + elif not args.value: + raise Usage("status: progress needs a note, blocked needs an a-id") + elif args.kind == "progress": + status = "in progress: " + " ".join(args.value) + else: + status = f"blocked: [[{args.value[0]}]]" + code = change(args, lambda p: tasks.set_status(p.doc, args.id, status)) + if not args.dry_run: + proj = load_project(args) + cfg = proj.cfg + if args.kind == "progress": + claim(cfg, args.id, lanes.lane_of(proj.doc.item(args.id), cfg.lanes, cfg.slice_above)) + else: + unclaim(cfg, [args.id]) + return code + + +def cmd_set(args) -> int: + if args.interactive is not None: + old_spelling("--interactive", "--sessions owner") + if args.sessions is None: + args.sessions = "owner" if args.interactive == "yes" else "" + fields = dict( + title=args.title, effort=args.effort, + after=None if args.after is None else split_list(args.after), + refs=None if args.ref is None else split_list(args.ref), model=args.model, sessions=args.sessions, + done=args.done, cloud=args.cloud) + if args.model not in (None, "", *tasks.MODELS): + raise Usage(f"set: --model {args.model} (want {', '.join(tasks.MODELS)}, or \"\" to remove)") + if args.sessions not in (None, "", *tasks.SESSIONS): + raise Usage(f"set: --sessions {args.sessions} (want {', '.join(tasks.SESSIONS)}, or \"\" to remove)") + if args.cloud not in (None, "", *tasks.CLOUDS): + raise Usage(f"set: --cloud {args.cloud} (want {', '.join(tasks.CLOUDS)}, or \"\" to remove)") + if all(v is None for v in fields.values()): + raise Usage("set: nothing to set (--title --effort --after --ref --model --sessions --cloud --done)") + return change(args, lambda p: tasks.set_fields(p.doc, args.id, **fields)) + + +def cmd_rename(args) -> int: + def action(p): + n = tasks.rename(p.doc, args.id, args.new, p.archived) + return f"renamed: {args.id} → {args.new} ({n} link{'' if n == 1 else 's'})" + return change(args, action) + + +def cmd_note(args) -> int: + return change(args, lambda p: tasks.add_note(p.doc, args.id, " ".join(args.line.split()))) + + +def cmd_body(args) -> int: + if args.text: + raise Failure("body text goes on stdin (wf body ID <<'EOF' … EOF), not as an argument") + lines = stdin_lines() + return change(args, lambda p: tasks.set_body(p.doc, args.id, lines)) + + +def cmd_tick(args) -> int: + return change(args, lambda p: "ticked: " + tasks.tick(p.doc, args.id, args.which)) + + +# ------------------------------------------------------------------ other commands + +def wf_version() -> str: + try: + out = subprocess.run(["git", "-C", str(HERE), "rev-parse", "--short=7", "HEAD"], + capture_output=True, text=True, timeout=10) + return out.stdout.strip() if out.returncode == 0 and out.stdout.strip() else "0000000" + except (OSError, subprocess.SubprocessError): + return "0000000" + + +def cmd_report(args) -> int: + text = " ".join(args.text.split()) + if not text: + raise Usage("report: say what happened") + start = Path(args.project) if args.project else Path.cwd() + try: + project = config.find_root(start).name + except config.ConfigError: + project = start.resolve().name + cmd = f" (cmd: {' '.join(args.cmd.split())})" if args.cmd else "" + line = f"- {datetime.date.today().isoformat()} {project} {args.kind}: {text}{cmd} @{wf_version()}\n" + fd = os.open(inbox_path(), os.O_WRONLY | os.O_APPEND | os.O_CREAT, 0o644) + try: + os.write(fd, line.encode("utf-8")) + finally: + os.close(fd) + print("reported (workflow inbox); carry on") + return 0 + + +def cmd_init(args) -> int: + root = (Path(args.project) if args.project else Path.cwd()).resolve() + if (root / config.NAME).exists(): + raise Failure(f"{root}/{config.NAME} exists already") + templates = HERE / "templates" + for target, source in ((config.NAME, "workflow.toml"), ("TASKS.md", "TASKS.md"), + ("tasks/archive.md", "archive.md"), ("CLAUDE.md", "CLAUDE.md")): + path = root / target + if path.exists(): + print(f"kept {target} (exists" + ("; old format → wf migrate)" if target == "TASKS.md" else ")")) + continue + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(read(templates / source).replace("{name}", root.name), encoding="utf-8") + print(f"wrote {target}") + return 0 + + +def cmd_migrate(args) -> int: + from wflib import migrate + p = load_project(args) + cfg = p.cfg + new, report = migrate.migrate(p.tasks_text, taken=p.archived) + bump = cfg.format < config.FORMAT + if new == p.tasks_text and not bump: + print("nothing to migrate") + return 0 + name = cfg.rel(cfg.tasks) + sys.stdout.writelines(difflib.unified_diff( + p.tasks_text.replace("\r\n", "\n").splitlines(keepends=True), + new.replace("\r\n", "\n").splitlines(keepends=True), name, f"{name} (new)")) + if report.ids: + print("\nids:") + for title, id in report.ids.items(): + print(f" {id} {title}") + if report.notes: + print("\nnotes:") + for note in report.notes: + print(f" {note}") + errors = checks.check(cfg, tasks_text=new, slow=False)[0] + for e in errors: + print(f"ERROR: {e}") + print(f"check after migrate: {len(errors)} errors") + if not args.write: + print("dry run: nothing written (wf migrate --write)") + return 0 + wrote = [] + if new != p.tasks_text: + write_if_unchanged(cfg.tasks, new, p.tasks_stamp) + wrote.append(name) + if bump: + toml = cfg.root / config.NAME + text = read(toml) + changed = re.sub(r"(?m)^format(\s*)=(\s*)\d+", rf"format\g<1>=\g<2>{config.FORMAT}", text, count=1) + write_if_unchanged(toml, changed, stamp(toml)) + wrote.append(f"{config.NAME} (format = {config.FORMAT})") + print("wrote " + ", ".join(wrote)) + return 0 + + +# ------------------------------------------------------------------ main + +def parser() -> argparse.ArgumentParser: + ap = argparse.ArgumentParser(prog="wf", description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--project", metavar="DIR", help="project folder (default: found from the working directory)") + sub = ap.add_subparsers(dest="command", metavar="COMMAND") + sections = list(tasks.SECTIONS) + + def cmd(name, func, help, write=False): + sp = sub.add_parser(name, help=help, description=help) + sp.set_defaults(func=func, locks=write) + if write: + sp.add_argument("--dry-run", action="store_true", help="print the diff, write nothing") + return sp + + sp = cmd("next", cmd_next, "cold start: awaiting, needs human, plans in flight, next task with its refs") + sp.add_argument("--brief", "-b", action="store_true", help="the task only") + sp.add_argument("--lane", help="lane name (wf lanes); none = all lanes") + sp.add_argument("--as", dest="as_", choices=tasks.MODELS, + help="your model: takes tasks with Model ≤ it (default haiku)") + sp.add_argument("--owner", action="store_true", help="the owner is present: Sessions: owner tasks pickable") + + sp = cmd("areas", cmd_areas, "area notes: anchors missing, commits since Checked (stale = missing or " + "≥ area_stale_commits); no NAME → also uncovered: folders the diff since master touches outside " + "every Paths (area_ignore skipped); --mark NAME stamps HEAD", write=True) + sp.add_argument("name", nargs="?", help="one area only") + sp.add_argument("--mark", metavar="NAME", help="set the area's Checked to HEAD (after a refresh)") + sp = cmd("lanes", cmd_lanes, "lanes: pickable/waiting per lane, the session of each " + "(--lane/--as registers yours; neither lane = all)") + sp.add_argument("--lane", help="your lane (wf lanes lists them)") + sp.add_argument("--as", dest="as_", choices=tasks.MODELS, help="your model (default haiku)") + sp.add_argument("--unregister", action="store_true", help="drop this session's lane record (orchestrators hold none)") + sp.add_argument("--wait", type=float, metavar="SECS", + help="block until any lane has a pickable task (exit 0, prints the lanes line) or SECS pass " + "(exit 1, prints the lanes line); out/wf-batch.stop exists (wf batch --stop) → exit 2 at once; " + "polls every 30 s (env WF_LANES_POLL)") + sp = cmd("list", cmd_list, "one line per item: id, priority, effort, status, model, lane, title") + sp.add_argument("-s", "--section", choices=[*sections, "all"], default="pending") + sp.add_argument("-p", "--prio", type=int, choices=range(4), help="this priority or higher") + sp.add_argument("--ready", action="store_true", help="pickable now") + sp.add_argument("--runner", action="store_true", + help="--ready + Done line, no Sessions: owner, Done not about owner/confirming, no open slices") + sp.add_argument("--blocked", action="store_true") + sp.add_argument("--progress", action="store_true") + sp.add_argument("--stale", action="store_true", help="in progress with no live claim (dead or no holder)") + sp.add_argument("--model", choices=tasks.MODELS, help="tasks with this Model line (none = opus)") + sp.add_argument("--lane", help="tasks of this lane (by effort); with --runner: in pick order") + sp.add_argument("--ref", metavar="PATH[#ANCHOR]") + sp.add_argument("-n", type=int, metavar="N") + + sp = cmd("show", cmd_show, "the raw item(s)") + sp.add_argument("ids", nargs="+", metavar="ID") + + sp = cmd("ctx", cmd_ctx, "item with refs resolved, its relations and the areas it names; or a doc section and its tasks") + sp.add_argument("target", metavar="ID|PATH#ANCHOR") + + sp = cmd("search", cmd_search, "ranked search over tasks, archive and docs") + sp.add_argument("words", nargs="+") + sp.add_argument("--tasks", action="store_true") + sp.add_argument("--archive", action="store_true") + sp.add_argument("--docs", action="store_true") + sp.add_argument("-n", type=int, default=15, metavar="N") + + sp = cmd("log", cmd_log, "newest archive lines") + sp.add_argument("words", nargs="*") + sp.add_argument("-n", type=int, default=10, metavar="N") + + sp = cmd("usage", cmd_usage, "tokens and API-price $ per agent of a session, from Claude Code transcripts") + sp.add_argument("--session", metavar="ID", help="session id or prefix (default: this session, else newest here)") + sp.add_argument("--agent", metavar="ID", help="one subagent only (agent id from its completion notice)") + sp.add_argument("--since", metavar="TIME", help="ISO UTC time, or 90m / 2h / 1d ago") + sp.add_argument("--log", nargs=2, metavar=("TASK", "OUTCOME"), + help="with --agent: append one line to <project>/out/wf-cost.log (orchestrator, per worker)") + sp.add_argument("--effort", choices=tasks.EFFORTS, metavar="|".join(tasks.EFFORTS), + help="the task's estimate (not reasoning effort), for --log") + sp.add_argument("--duration", type=int, metavar="S", help="agent wall time in seconds (duration_ms / 1000), for --log") + sp.add_argument("--lane", metavar="L", help="lane of the logged worker, for --log (default: its model)") + sp.add_argument("--explore", action="store_true", + help="exploring cost per subagent: pct billed input exploring / before the first edit, " + "result tokens per tool, files read by >= 2 agents (same --session / --since)") + sp.add_argument("--report", action="store_true", + help="per lane and effort: n, done, $ (median, total, per done), turns, duration, from every project's log") + + cmd("projects", cmd_projects, "every wf project: counts, errors, next task") + cmd("check", cmd_check, "validate ids, links, refs, order") + + sp = cmd("add", cmd_add, "new item; `add -` reads a whole item block from stdin", write=True) + sp.add_argument("title", metavar='"Title. Goal."|-') + sp.add_argument("-p", "--prio", type=int, choices=range(4)) + sp.add_argument("-e", "--effort", metavar="|".join(tasks.EFFORTS)) + sp.add_argument("-s", "--section", choices=sections) + sp.add_argument("--id", help="the id (default: from the title, or <parent>-N with --parent); " + "a leading 'ID: ' in the text works too") + sp.add_argument("--after", metavar="ID,…") + sp.add_argument("--ref", metavar="REF,…") + sp.add_argument("--parent", metavar="ID", help="add as the next slice of ID, After: the previous one") + sp.add_argument("--interactive", action="store_true", help=argparse.SUPPRESS) + sp.add_argument("--model", choices=tasks.MODELS, help="Model line (default: none = opus)") + sp.add_argument("--done", metavar="TEXT", help="Done line (makes the task runner-pickable)") + sp.add_argument("--cloud", choices=tasks.CLOUDS, help="Cloud line: yes = cloud lane may take it (opus only), no = never") + sp.add_argument("--sessions", choices=tasks.SESSIONS, + help="Sessions line: solo = no other live session, owner = owner present (default parallel)") + sp.add_argument("-b", "--body", action="store_true", help="read body lines from stdin") + + sp = cmd("done", cmd_done, "finish task(s): remove, archive line, print verify + checklist", write=True) + sp.add_argument("ids", nargs="+", metavar="ID") + sp.add_argument("-m", metavar="ENTRY", help="archive entry, ≤2 lines") + + sp = cmd("prio", cmd_prio, "set priority, reposition", write=True) + sp.add_argument("id") + sp.add_argument("prio", type=int, choices=range(4)) + + sp = cmd("move", cmd_move, "move to a section, or before/after another item", write=True) + sp.add_argument("id") + sp.add_argument("section", nargs="?", choices=sections) + sp.add_argument("--before", metavar="ID") + sp.add_argument("--after", metavar="ID") + sp.add_argument("--force", action="store_true", help="allow breaking priority order") + + sp = cmd("status", cmd_status, "progress NOTE | blocked A-ID | clear", write=True) + sp.add_argument("id", nargs="?") + sp.add_argument("kind", nargs="?", choices=["progress", "blocked", "clear"]) + sp.add_argument("value", nargs="*") + sp.add_argument("--clear-stale", action="store_true", help="clear in progress of tasks with no live claim") + + sp = cmd("set", cmd_set, "change header fields, Done:, Model:, After:, Ref:", write=True) + sp.add_argument("id") + sp.add_argument("--title", help="new title, goal kept; 'Title. Goal.' or a question replaces the text") + sp.add_argument("--effort") + sp.add_argument("--after", metavar="ID,…", help='"" removes the line') + sp.add_argument("--ref", metavar="REF,…", help='"" removes the line') + sp.add_argument("--interactive", choices=["yes", "no"], help=argparse.SUPPRESS) + sp.add_argument("--sessions", metavar="|".join(tasks.SESSIONS), help='Sessions line; "" removes it (= parallel)') + sp.add_argument("--model", metavar="|".join(tasks.MODELS), help='Model line; "" removes it (= opus)') + sp.add_argument("--cloud", metavar="|".join(tasks.CLOUDS), help='Cloud line (yes: skip the fit regex list, opus only; no: never cloud); "" removes it') + sp.add_argument("--done", metavar="TEXT", help='Done line (makes it runner-pickable); "" removes it') + + sp = cmd("rename", cmd_rename, "change an open item's id and every [[link]] to it", write=True) + sp.add_argument("id") + sp.add_argument("new", metavar="NEW") + + sp = cmd("note", cmd_note, "append one line to the body", write=True) + sp.add_argument("id") + sp.add_argument("line") + + sp = cmd("body", cmd_body, "replace the body with stdin; Model:, Sessions:, After:, Ref: lines kept", write=True) + sp.formatter_class = argparse.RawDescriptionHelpFormatter + sp.epilog = "example (body text goes on stdin, not as an argument):\n wf body ID <<'EOF'\n - Steps: …\n - Done: …\n EOF" + sp.add_argument("id") + sp.add_argument("text", nargs="*", help=argparse.SUPPRESS) + + sp = cmd("tick", cmd_tick, "check a `- [ ]` box by number or text", write=True) + sp.add_argument("id") + sp.add_argument("which", metavar="N|TEXT") + sp = cmd("merge", cmd_merge, "in a lane worktree: rebase, ff-merge into master, commit TASKS/archive, " + "detach, push home (under the project lock)") + sp.set_defaults(locks=True) + sp.add_argument("-m", help="bookkeeping commit message (default '<id> done', id from the branch)") + sp.add_argument("--no-push", action="store_true", help="skip git push home") + + sp = cmd("finish", cmd_finish, "worker's last step in one call: quick_gate → done → commit PATHS (explicit) → " + "wf merge (lane worktree; main tree: TASKS/archive go into the commit). Run verify first. " + "Refuses before done if the gate is red, files outside PATHS are uncommitted, or a PATH is outside " + "this repo / missing / unchanged; an id already archived resumes (commit + merge). " + "Prints the done output (notify / checklist) and the merge lines, then 'report: commit <sha> [tool <sha>]' " + "(lane worktree) = the sha for your report") + sp.add_argument("id") + sp.add_argument("-m", required=True, metavar="ENTRY", help="archive entry, ≤2 lines") + sp.add_argument("--commit", metavar="MSG", help="commit message for PATHS (incl. footer lines)") + sp.add_argument("paths", nargs="*", metavar="PATH", help="files/folders to commit (relative to cwd)") + sp.add_argument("--no-push", action="store_true", help="skip git push home") + sp.add_argument("--tool-commit", metavar="SHA", help="tool-repo sha merged by hand; added to the report line") + + sp = cmd("wip", cmd_wip, "wrap-up in one call (lane worktree): commit PATHS on the branch, note MSG " + "(state + next step), status clear") + sp.add_argument("id") + sp.add_argument("-m", required=True, metavar="NOTE", help="state + next step, one line") + sp.add_argument("--commit", metavar="MSG", help="commit message (default '<id> WIP')") + sp.add_argument("paths", nargs="*", metavar="PATH", help="files/folders to commit (relative to cwd)") + + cmd("setup", cmd_setup, "in a lane worktree: run workflow.toml worktree_setup (cwd = worktree, env WF_MAIN = " + "main tree project folder); provides git-ignored inputs, commands must be idempotent") + + sp = cmd("start", cmd_start, "worker setup in one call, in the main tree: worktree at PATH on BRANCH " + "(new: from master or the existing branch; existing: must be clean, switched to BRANCH), " + "worktree_setup, status progress BRANCH, then wf ctx and a ready verify && wf finish line") + sp.add_argument("id") + sp.add_argument("--worktree", required=True, metavar="PATH", help="worktree folder (relative to cwd)") + sp.add_argument("--branch", required=True, metavar="BRANCH") + sp.add_argument("--recovery", action="store_true", + help="a dead worker's WIP: uncommitted changes on BRANCH kept, printed with its commits") + + cmd("gate", cmd_gate, "run workflow.toml quick_gate (fast regression check a worker runs before wf done; " + "cwd = this tree's project folder, env WF_MAIN = main tree project folder); red = exit 1, nothing = exit 0") + + sp = cmd("report", cmd_report, "report a workflow problem or idea to the workflow inbox") + sp.add_argument("text") + sp.add_argument("--kind", choices=["bug", "idea", "friction"], default="bug") + sp.add_argument("--cmd", metavar="COMMAND") + + cmd("init", cmd_init, "make this folder a wf project") + + sp = cmd("migrate", cmd_migrate, "convert a numbered TASKS.md to the id format (dry run unless --write)") + sp.add_argument("--write", action="store_true") + + sp = cmd("orch", cmd_orch, "orchestrator, main tree: pick LANE (pick + claim + free worktree + agent prompt) · " + "post ID LANE (post-check, merge if needed, orch log + cost line, next pick)") + osub = sp.add_subparsers(dest="action", required=True) + op = osub.add_parser("pick", help="first runner-ready task of LANE: claim it, choose a free worktree " + "(.worktrees/LANE, -2, …), print the wf-worker prompt; 'none:'/'stop:' line = spawn nothing. " + "LANE cloud: ledger check, first cloud-fitting task of any lane (opus first) -> wf cloud send; " + "'stop lane cloud: ledger|none fit|max parallel'") + op.add_argument("lane") + op.add_argument("--id", help="this task instead of the pick (a P0 fix, a recovery)") + op.add_argument("--recovery", metavar="WHY", help="prompt gets 'Recovery: WHY' (crashed worker, one retry)") + op = osub.add_parser("post", help="after the worker's report: post-check (done: archive line, branch gone, " + "worktree clean and in master, else wf merge; wf check), leftover commit otherwise, " + "out/wf-orch.log line, out/wf-cost.log line (--agent), then the next pick or 'stop lane'") + op.add_argument("id") + op.add_argument("lane") + op.add_argument("--result", metavar="LINE", help="the report's result line (default: done if archived, else no-report)") + op.add_argument("--commit", metavar="SHA", help="the report's commit (default: master's sha when done)") + op.add_argument("--agent", metavar="ID", help="agent id from the completion notice: cost line (wf usage --log)") + op.add_argument("--duration", type=int, metavar="S", help="agent duration_ms / 1000 (default or 0: since the pick)") + op.add_argument("--no-pick", "--no-next", dest="no_pick", action="store_true", + help="no next pick, claims nothing (batch end, stop file, cross-project rollout)") + op.add_argument("--no-push", action="store_true", help="the wf merge it may run skips git push home") + + cmd("res", None, "shared memory/CPU ledger for agent jobs (wf res -h)") + cmd("cloud", None, "cloud-lane ledger (wf cloud -h)") + cmd("prep", None, "prep tasks without Done: wf prep [N|all] = wf batch --prep (wf batch -h)") + cmd("batch", None, "start an unattended batch orchestrator (claude -p) under wf res (wf batch -h)") + return ap + + +def main(argv: list[str]) -> int: + if argv[:1] == ["cloud"]: + import wf_cloud + return wf_cloud.main(argv[1:]) + if argv[:1] == ["res"]: + import wf_res + return wf_res.main(argv[1:]) + if argv[:1] == ["prep"]: + import wf_res + return wf_res.prep_main(argv[1:]) + if argv[:1] == ["batch"]: + import wf_res + return wf_res.batch_main(argv[1:]) + ap = parser() + args = ap.parse_args(argv) + if not args.command: + ap.print_usage(sys.stderr) + return 2 + try: + if getattr(args, "locks", False): + root = config.find_root(Path(args.project) if args.project else Path.cwd()) + with project_lock(root): + return args.func(args) + return args.func(args) + except Failure as e: + print(f"wf: {e}", file=sys.stderr) + return e.code + except (tasks.TaskError, config.ConfigError) as e: + print(f"wf: {e}", file=sys.stderr) + return 1 + except BrokenPipeError: + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/wf_cloud.py b/wf_cloud.py new file mode 100644 index 0000000..10f1b0a --- /dev/null +++ b/wf_cloud.py @@ -0,0 +1,720 @@ +"""wf cloud — cloud-lane ledger IO. Pure logic: wflib/cloud.py. + + ledger [--set-balance USD] [--budget USD] print budget/spent/balance/running; reconcile / change cap + send ID [--dry-run] [--project DIR] snapshot the project's master (+ cloud_include) into + out/cloud/ID, prompt from templates/cloud-prompt.md, reserve in the ledger (refused: exit 3), + `claude --cloud`, claim `in progress: cloud:<sid>`, record .wf/cloud/ID.json. + --dry-run: snapshot size + prompt, nothing sent / reserved / kept + archive SID | --ended archive ended sessions in the claude.ai app (POST + /v1/code/sessions/SID/archive with the CLI's login token, never printed); pull does it after each end + and prints `archive SID failed: …` on failure (pull's exit code unchanged); --ended retries the rest. + pull ID | --all [--export FILE] [--no-push] teleport + /export, parse the final WF-RESULT message: + done -> check sha256/bytes, gunzip, refuse paths outside the project / TASKS / archive / .wf / out / + cloud_include, `git am --3way` on <lane>/<id> from the recorded master in a free lane worktree, + quick_gate, wf done -m "<WF-REPORT> (cloud <sid>)" + wf merge; awaiting -> wf add -s awaiting + + status blocked; handback / red gate / bad patch -> note "Recovery: cloud attempt <sid> — <why>", + status clear, patch kept in out/cloud/ID.patch. No result: running (exit 4), > 24h: lost. + Every end charges the ledger, deletes the snapshot + record and archives the session. + Exit: 0 ended, 4 still running, 1 local error (worktree busy, branch exists: session kept, pull again). + +Library (for send/pull): send(folder, prompt) -> sid, export(folder, sid, out) -> path drive the real +CLI under a pty (env WF_CLAUDE = claude binary, WF_CLAUDE_JSON = trust file, default ~/.claude.json). +""" +from __future__ import annotations + +import argparse +import contextlib +import datetime as dt +import fcntl +import os +import pty +import re +import select +import shutil +import signal +import struct +import subprocess +import sys +import termios +import time +from pathlib import Path + +HERE = Path(__file__).resolve().parent +sys.path.insert(0, str(HERE)) + +from wflib import cloud # noqa: E402 + +LEDGER, LOCK = "cloud.json", "cloud.lock" + + +def state_dir() -> Path: + base = os.environ.get("XDG_STATE_HOME") or str(Path.home() / ".local/state") + return Path(base) / "wf" + + +@contextlib.contextmanager +def locked(state: Path): + """Exclusive flock; yields the ledger dict; atomic write-back when the block ends without error.""" + state.mkdir(parents=True, exist_ok=True) + with open(state / LOCK, "w") as lock: + fcntl.flock(lock, fcntl.LOCK_EX) + path = state / LEDGER + led = cloud.loads(path.read_text() if path.exists() else "") + yield led + tmp = path.with_name(path.name + ".tmp") + tmp.write_text(cloud.dumps(led)) + os.replace(tmp, path) + + +# --- pty driver (spec §3 F1-F3, F7, F8; §4.1, §4.3) ---------------------------------------------- + +CLAUDE = os.environ.get("WF_CLAUDE", "claude") +DOWN, ENTER = "\x1b[B", "\r" +CAP = cloud.CAP # bytes; tests lower it +SETTLE = 2.0 # s: TUI settle after resume / between Enter retries while /export runs + + +class Pty: + """A child process on a pseudo-terminal (claude --cloud refuses pipes, F2). `out` = raw output so far.""" + + def __init__(self, argv: list[str], cwd: Path, cols: int = 200, rows: int = 50): + self.master, slave = pty.openpty() + fcntl.ioctl(slave, termios.TIOCSWINSZ, struct.pack("HHHH", rows, cols, 0, 0)) + env = {**os.environ, "TERM": os.environ.get("TERM") or "xterm-256color"} + self.proc = subprocess.Popen(argv, cwd=cwd, stdin=slave, stdout=slave, stderr=slave, env=env, + start_new_session=True, close_fds=True) + os.close(slave) + self.out = "" + + def _read(self, wait: float) -> bool: + """Read what is there within `wait` s. False = child output closed.""" + r, _, _ = select.select([self.master], [], [], wait) + if not r: + return True + try: + data = os.read(self.master, 65536) + except OSError: # EIO: child side closed + return False + if not data: + return False + self.out += data.decode("utf-8", "replace") + return True + + def expect(self, pred, timeout: float) -> bool: + """Read until pred(plain text so far) or timeout / EOF. Returns pred's final truth.""" + end = time.monotonic() + timeout + while not pred(cloud.strip_ansi(self.out)): + left = end - time.monotonic() + if left <= 0 or not self._read(min(left, 0.5)): + return bool(pred(cloud.strip_ansi(self.out))) + return True + + def alive(self) -> bool: + return self.proc.poll() is None + + def write(self, s: str) -> None: + os.write(self.master, s.encode()) + + def type(self, s: str) -> None: + """Text then Enter, Enter as its own write (a TUI reads a pasted CR as text).""" + self.write(s) + time.sleep(0.3) + self.write(ENTER) + + def close(self, timeout: float = 10) -> int: + """Wait for exit (draining output), then TERM / KILL the process group. Returns the exit code.""" + end = time.monotonic() + timeout + while self.alive() and time.monotonic() < end: + self._read(0.2) + for sig in (signal.SIGTERM, signal.SIGKILL): + if not self.alive(): + break + with contextlib.suppress(ProcessLookupError): + os.killpg(self.proc.pid, sig) + with contextlib.suppress(subprocess.TimeoutExpired): + self.proc.wait(3) + with contextlib.suppress(OSError): + while self._read(0): + pass + os.close(self.master) + return self.proc.wait() + + +def claude_json() -> Path: + return Path(os.environ.get("WF_CLAUDE_JSON") or Path.home() / ".claude.json") + + +def pre_trust(folder: Path) -> None: + """Accept the trust dialog for folder in ~/.claude.json, atomic rewrite (F3).""" + path = claude_json() + text = path.read_text() if path.exists() else "" + new = cloud.trust(text, str(folder.resolve())) + tmp = path.with_name(f"{path.name}.wf-{os.getpid()}.tmp") + tmp.write_text(new) + if path.exists(): + os.chmod(tmp, path.stat().st_mode & 0o777) + os.replace(tmp, path) + + +def _answer_trust(p: Pty, plain: str, done: set) -> None: + """Fallback when the pre-accept did not take: Yes = Down, Enter (F3). Once per run.""" + if "trust" not in done and cloud.TRUST_PROMPT.search(plain): + done.add("trust") + p.write(DOWN) + time.sleep(0.3) + p.write(ENTER) + + +def send(folder: Path, prompt: str, timeout: float = 300) -> str: + """`claude --cloud <prompt>` in folder under a pty -> session id (F1). CloudError with the last 5 lines.""" + folder = Path(folder) + pre_trust(folder) + p, seen = Pty([CLAUDE, "--cloud", prompt], folder), set() + + def ready(plain: str) -> bool: + _answer_trust(p, plain, seen) + return cloud.parse_sid(plain) is not None and "Resume with:" in plain or not p.alive() + + p.expect(ready, timeout) + code = p.close(5) + sid = cloud.parse_sid(p.out) + if not sid: + tail = " | ".join(cloud.last_lines(p.out)) or "(no output)" + raise cloud.CloudError(f"claude --cloud gave no session id (exit {code}): {tail}") + return sid + + +def _teleport_export(folder: Path, sid: str, out: Path, timeout: float) -> tuple[bool, str]: + """One teleport -> /export -> /exit round. (exported?, raw output).""" + p, seen = Pty([CLAUDE, "--teleport", sid], folder), set() + end = time.monotonic() + timeout + try: + def resumed(plain: str) -> bool: + _answer_trust(p, plain, seen) + return bool(cloud.RESUMED.search(plain)) or not p.alive() + + if not p.expect(resumed, timeout) or not p.alive(): + return False, p.out + p.expect(lambda _: False, SETTLE) # let the TUI settle + p.type(f"/export {out}") + while not (out.exists() and out.stat().st_size) and p.alive() and time.monotonic() < end: + p.expect(lambda _: False, SETTLE) + if not (out.exists() and out.stat().st_size): + p.write(ENTER) # autocomplete eats the first Enter (F8) + ok = out.exists() and out.stat().st_size > 0 + if p.alive(): + p.type("/exit") + return ok, p.out + finally: + p.close(10) + + +def export(folder: Path, sid: str, out: Path, timeout: float = 300) -> Path: + """Teleport into sid, `/export out`, `/exit`; no model call (F7, F8). Teleport may exit at + 'Checking out branch' without resuming: retried once, then CloudError with the last 5 lines.""" + folder, out = Path(folder), Path(out).resolve() + pre_trust(folder) + out.parent.mkdir(parents=True, exist_ok=True) + raw = "" + for _ in range(2): + with contextlib.suppress(FileNotFoundError): + out.unlink() + ok, raw = _teleport_export(folder, sid, out, timeout) + if ok: + return out + tail = " | ".join(cloud.last_lines(raw)) or "(no output)" + raise cloud.CloudError(f"teleport {sid}: no export after 2 tries: {tail}") + + +# --- send (spec §4.1, §4.2) ---------------------------------------------------------------------- + +def git(cwd: Path, *args: str, check: bool = True) -> str: + r = subprocess.run(["git", "-C", str(cwd), *args], capture_output=True, text=True) + if check and r.returncode: + raise cloud.CloudError(f"git {args[0]}: {(r.stderr.strip() or r.stdout.strip() or 'failed').splitlines()[-1]}") + return r.stdout.strip() + + +def snapshot(cfg, folder: Path) -> tuple[str, str, int]: + """folder = git archive of the project's master minus TASKS/archive/.wf/.worktrees/out, plus + cloud_include, as a fresh repo with one commit 'base <sha>' on main. -> (master sha, base sha, packed bytes).""" + from wflib import config + top = Path(git(cfg.root, "rev-parse", "--show-toplevel")) + master = config.git_branch(top / ".git") or "master" + sha = git(top, "rev-parse", master) + for inc in cfg.cloud_include: + if why := cloud.include_problem(inc): + raise cloud.CloudError(why) + if not (cfg.root / inc).exists(): + raise cloud.CloudError(f"cloud_include '{inc}': not found in {cfg.root}") + if folder.exists(): + shutil.rmtree(folder) + folder.mkdir(parents=True) + archive = subprocess.run(["git", "-C", str(top), "archive", "--format=tar", sha], capture_output=True) + if archive.returncode: + raise cloud.CloudError(f"git archive: {archive.stderr.decode(errors='replace').strip()}") + subprocess.run(["tar", "-x", "-C", str(folder)], input=archive.stdout, check=True) + rel = cfg.root.resolve().relative_to(top.resolve()) + drop = [cfg.tasks, cfg.archive] + [cfg.root / n for n in cloud.NEVER] + for path in drop: + target = folder / Path(path).resolve().relative_to(top.resolve()) + if target.is_dir() and not target.is_symlink(): + shutil.rmtree(target) + elif target.exists() or target.is_symlink(): + target.unlink() + for inc in cfg.cloud_include: + src, dst = cfg.root / inc, folder / rel / inc + dst.parent.mkdir(parents=True, exist_ok=True) + if src.is_dir(): + shutil.copytree(src, dst, symlinks=True, dirs_exist_ok=True) + else: + shutil.copy2(src, dst) + who = ["-c", "user.name=wf", "-c", "user.email=wf@localhost", "-c", "commit.gpgsign=false"] + git(folder, "init", "-q", "-b", "main") + git(folder, "add", "-A", "-f") + git(folder, *who, "commit", "-q", "--allow-empty", "--no-verify", "-m", f"base {sha}") + git(folder, "repack", "-adq") + size = sum(p.stat().st_size for p in (folder / ".git" / "objects" / "pack").glob("*.pack")) + return sha, git(folder, "rev-parse", "HEAD"), size + + +def credentials_path() -> Path: + """The CLI's login file (env WF_CLAUDE_CREDENTIALS, else $CLAUDE_CONFIG_DIR or ~/.claude).""" + if os.environ.get("WF_CLAUDE_CREDENTIALS"): + return Path(os.environ["WF_CLAUDE_CREDENTIALS"]) + return Path(os.environ.get("CLAUDE_CONFIG_DIR") or Path.home() / ".claude") / ".credentials.json" + + +def cli_version() -> str: + try: + out = subprocess.run([CLAUDE, "--version"], capture_output=True, text=True, timeout=20).stdout + except (OSError, subprocess.SubprocessError): + out = "" + m = re.match(r"\s*(\d+\.\d+\.\d+)", out) + return m[1] if m else "unknown" + + +def post(url: str, headers: dict, timeout: float = 10) -> tuple[int, str]: + """POST {} -> (status, body); network failure -> CloudError.""" + import urllib.error + import urllib.request + req = urllib.request.Request(url, data=b"{}", headers=headers, method="POST") + try: + with urllib.request.urlopen(req, timeout=timeout) as r: + return r.status, r.read(500).decode("utf-8", "replace") + except urllib.error.HTTPError as e: + return e.code, e.read(500).decode("utf-8", "replace") + except (urllib.error.URLError, OSError) as e: + raise cloud.CloudError(str(getattr(e, "reason", e))) + + +def archive(state: Path, sid: str) -> str | None: + """Archive one session in the app (undocumented endpoint, spec §4.3): None ok (ledger row marked), else why.""" + try: + try: + creds = credentials_path().read_text() + except OSError: + creds = "" + url, headers = cloud.archive_request(sid, creds, cli_version()) + why = cloud.archive_problem(*post(url, headers)) + except cloud.CloudError as e: + why = str(e) + if why is None: + with locked(state) as led: + with contextlib.suppress(cloud.CloudError): + cloud.find(led, sid)["archived"] = True + return why + + +def cmd_archive(a, state: Path) -> int: + if a.ended: + with locked(state) as led: + sids = cloud.unarchived(led) + if not sids: + print("no ended cloud session left to archive") + return 0 + else: + sids = [a.sid] + code = 0 + for sid in sids: + if why := archive(state, sid): + print(f"wf: archive {sid} failed: {why}", file=sys.stderr) + code = 1 + else: + print(f"archived {sid}") + return code + + +def drop(folder: Path) -> None: + """Delete a snapshot folder and its parents out/cloud, out when that leaves them empty.""" + shutil.rmtree(folder, ignore_errors=True) + for parent in (folder.parent, folder.parent.parent): + with contextlib.suppress(OSError): + parent.rmdir() + + +def record_path(cfg, id: str) -> Path: + return cfg.root / ".wf" / "cloud" / f"{id}.json" + + +def cmd_send(a, state: Path, now: dt.datetime) -> int: + import json + import wf + from wflib import lanes + p = wf.load_project(a) + cfg, id = p.cfg, a.id + item = p.doc.item(p.resolve(id)) + id = item.id + if not cfg.cloud: + raise cloud.CloudError("project not opted in to the cloud lane (workflow.toml: cloud = true)") + rec = record_path(cfg, id) + if rec.exists() and not a.dry_run: + raise cloud.CloudError(f"{id} already sent ({json.loads(rec.read_text()).get('sid')}): wf cloud pull it first") + folder = cfg.root / "out" / "cloud" / (id + ".dry-run" if a.dry_run else id) # never the live snapshot + try: + sha, base, size = snapshot(cfg, folder) + if why := cloud.size_refusal(size, CAP): + raise cloud.CloudError(why) + text = "\n".join(item.lines()) + recipe = "\n".join(wf.area_blocks(p, text)).strip() + prompt = cloud.fill((HERE / "templates" / "cloud-prompt.md").read_text(), id, base, text, recipe, + cfg.cloud_note) + if a.dry_run: + print(f"snapshot {size / 1e6:.1f} MB (master {sha[:12]}, base {base[:12]})\n\n{prompt}") + return 0 + pending = f"pending:{id}:{os.getpid()}" + with locked(state) as led: + if why := cloud.refusal(led): + print(f"wf: cloud refused: {why}", file=sys.stderr) + drop(folder) + return 3 + cloud.add(led, id, str(cfg.root), pending, cloud.MODEL, now) + try: + sid = send(folder, prompt) + except BaseException: + with locked(state) as led: + led["entries"] = [e for e in led["entries"] if e["sid"] != pending] + raise + with locked(state) as led: + cloud.find(led, pending)["sid"] = sid + except BaseException: + drop(folder) + raise + finally: + if a.dry_run: + drop(folder) + lane = lanes.lane_of(item, cfg.lanes, cfg.slice_above) + rec.parent.mkdir(parents=True, exist_ok=True) + if not (rec.parent.parent / ".gitignore").exists(): + (rec.parent.parent / ".gitignore").write_text("*\n") + rec.write_text(json.dumps({"id": id, "sid": sid, "project": str(cfg.root), "lane": lane, "master": sha, + "base": base, "folder": str(folder), "sent": now.isoformat(timespec="seconds"), + "model": cloud.MODEL, "bytes": size}, indent=2) + "\n") + code = wf.main(["--project", str(cfg.root), "status", id, "progress", f"cloud:{sid}"]) + print(f"sent {id}: {sid} ({size / 1e6:.1f} MB){'' if code == 0 else '; claim failed, see above'}") + return 0 if code == 0 else 1 + + +# --- pull (spec §4.3) ---------------------------------------------------------------------------- + +class Busy(cloud.CloudError): + """Local obstacle (lane worktree / branch / setup): the session stays out, pull again later.""" + + +def age_text(td: dt.timedelta) -> str: + m = int(td.total_seconds() // 60) + return f"{m // 60}h{m % 60:02d}m" if m >= 60 else f"{m}m" + + +def apply_patch(p, rec: dict, lane: str, patch: Path) -> tuple[Path, Path, str]: + """Lane worktree on a new branch <lane>/<id> at the recorded master, `git am --3way` the patch, worktree_setup. + -> (worktree, its project folder, branch). Busy: worktree/branch unusable; CloudError: am failed (undone).""" + import argparse + import wf + cfg, id = p.cfg, rec["id"] + main = Path(git(cfg.root, "rev-parse", "--show-toplevel")) + branch = f"{lane}/{id}" + if not subprocess.run(["git", "-C", str(main), "rev-parse", "--verify", "-q", f"refs/heads/{branch}"], + capture_output=True).returncode: + raise Busy(f"branch {branch} exists (local WIP?): merge or delete it, then pull again") + wt = wf.lane_worktree(main, p, lane, id) + if not wt.exists(): + git(main, "worktree", "add", "-q", "--detach", str(wt), rec["master"]) + git(wt, "switch", "-q", "-c", branch, rec["master"]) + here = wt / cfg.root.resolve().relative_to(main.resolve()) + + def undo(): + git(wt, "am", "--abort", check=False) + git(wt, "reset", "-q", "--hard", check=False) + git(wt, "switch", "-q", "--detach", rec["master"], check=False) + git(main, "branch", "-q", "-D", branch, check=False) + + r = subprocess.run(["git", "-C", str(wt), "am", "-q", "--3way", str(patch)], capture_output=True, text=True) + if r.returncode: + undo() + tail = (r.stderr.strip() or r.stdout.strip() or "failed").splitlines()[-1] + raise cloud.CloudError(f"git am --3way: {tail}") + try: + wf.cmd_setup(argparse.Namespace(project=str(here))) + except wf.Failure as e: + undo() + raise Busy(str(e)) + return wt, here, branch + + +def pull_one(a, p, rec_path: Path, state: Path, now: dt.datetime) -> int: + """One sent task: export, parse, apply/gate/done/merge or awaiting/handback/lost; ends the session + (ledger charge, snapshot + record deleted). 0 ended, 4 running, 1 local error (session kept).""" + import argparse + import contextlib as cl + import io + import json + import wf + from wflib import lanes + cfg = p.cfg + rec = json.loads(rec_path.read_text()) + id, sid = rec["id"], rec["sid"] + age = now - dt.datetime.fromisoformat(rec["sent"]) + lane = rec.get("lane") or lanes.lane_of(p.doc.item(id), cfg.lanes, cfg.slice_above) + out_dir = cfg.root / "out" / "cloud" + kept = out_dir / f"{id}.patch" + + def quiet(*argv): + buf = io.StringIO() + with cl.redirect_stdout(buf): + code = wf.main(["--project", str(cfg.root), *argv]) + if code: + raise cloud.CloudError(f"wf {argv[0]} {id} failed (exit {code})") + return buf.getvalue() + + def finish(state_: str, res=None, line=""): + with locked(state) as led: + try: + e = cloud.end(led, sid, state_, res.usage if res else None) + charge = f"${e['usd']:.2f} {e['usd_source']}" + except cloud.CloudError as err: + charge = f"no ledger charge: {err}" + drop(Path(rec["folder"])) + rec_path.unlink(missing_ok=True) + print(f"{id}: {state_} ({sid}, {charge}){': ' + line if line else ''}") + if why := archive(state, sid): + print(f"archive {sid} failed: {why}; archive by hand (wf cloud archive --ended)") + return 0 + + def keep(raw: bytes | None) -> str: + if not raw or not raw.strip(): + return "" + out_dir.mkdir(parents=True, exist_ok=True) + kept.write_bytes(raw) + return f"; patch {cfg.rel(kept)}" + + def handback(why: str, res=None, raw=None, state_="handback"): + extra = keep(raw) + quiet("note", id, f"Recovery: cloud attempt {sid} — {why}{extra}") + quiet("status", id, "clear") + return finish(state_, res, why + extra) + + # 1. export (or a given file), F12 parse + exp = out_dir / f"{id}.export.txt" + try: + if a.export: + text = Path(a.export).read_text() + else: + text = export(Path(rec["folder"]), sid, exp).read_text() + except (cloud.CloudError, OSError) as e: + if age > cloud.LOST_AFTER: + return handback(f"lost: no export after {age_text(age)} ({e})", state_="lost") + raise + finally: + exp.unlink(missing_ok=True) + lines = cloud.final_message(text) + if lines is None and (kl := cloud.keyless_message(text)): + res, raw = cloud.keyless_result(kl), None + with cl.suppress(cloud.CloudError): + raw = cloud.decode_patch(res) + return handback("bad result: no WF-RESULT key", res, raw) + if lines is None: + if age > cloud.LOST_AFTER: + return handback(f"lost: no WF-RESULT after {age_text(age)}", state_="lost") + print(f"{id}: running ({sid}, sent {age_text(age)} ago)") + return 4 + try: + res = cloud.parse_result(lines) + except cloud.CloudError as e: + return handback(f"bad result: {e}") + if res.usage and res.model and cloud.U.price(res.model): + with locked(state) as led: + with cl.suppress(cloud.CloudError): + cloud.find(led, sid)["model"] = res.model + raw, why = None, None + with cl.suppress(cloud.CloudError): + raw = cloud.decode_patch(res) + + # 2. awaiting / handback from the session + if res.state == "awaiting": + extra = keep(raw) + header = quiet("add", "-s", "awaiting", f"{res.report} [[{id}]]") + m = re.search(r"\*\*(a-[^*]+)\*\*", header) + if not m: + raise cloud.CloudError(f"wf add printed no a-id: {header.strip()}") + quiet("status", id, "blocked", m[1]) + if extra: + quiet("note", id, f"cloud attempt {sid}: partial work{extra}") + return finish("awaiting", res, m[1] + extra) + if res.state == "handback": + return handback(f"cloud handback: {res.report}", res, raw) + + # 3. done: patch checks, apply, gate, done + merge + try: + raw = cloud.decode_patch(res) + except cloud.CloudError as e: + return handback(str(e), res) + patch = raw.decode("utf-8", errors="replace") + files = cloud.patch_files(patch) + main = Path(git(cfg.root, "rev-parse", "--show-toplevel")).resolve() + proj = str(cfg.root.resolve().relative_to(main)) + never = [cfg.rel(cfg.tasks), cfg.rel(cfg.archive), *cloud.NEVER, *cfg.cloud_include] + if not files: + return handback("done with an empty patch", res, raw) + if why := cloud.patch_problem(files, "" if proj == "." else proj, never): + return handback(why, res, raw) + out_dir.mkdir(parents=True, exist_ok=True) + kept.write_bytes(raw) + try: + wt, here, branch = apply_patch(p, rec, lane, kept) + except Busy: + kept.unlink(missing_ok=True) + raise + except cloud.CloudError as e: + return handback(str(e), res, raw) + try: + if cfg.quick_gate: + wf._run_lines(cfg.quick_gate, here, cfg, "quick_gate '{line}' red (exit {rc})") + except wf.Failure as e: + git(wt, "reset", "-q", "--hard", check=False) + git(wt, "switch", "-q", "--detach", rec["master"], check=False) + git(main, "branch", "-q", "-D", branch, check=False) + return handback(str(e), res, raw) + kept.unlink(missing_ok=True) + sys.stdout.flush() + merged, err = "", None + with wf.project_lock(cfg.root): + wf.cmd_done(argparse.Namespace(project=str(cfg.root), ids=[id], m=f"{res.report} (cloud {sid})", + dry_run=False, finishing=True)) + sys.stdout.flush() + ns = argparse.Namespace(project=str(here), m=None, no_push=a.no_push) + try: + wf.cmd_merge(ns) + merged = getattr(ns, "merged_sha", "") + except wf.Failure as e: + err = e + finish("done", res, f"merged {merged}" if merged else f"merge failed: {err}") + if err: + print(f"wf: {id} done, not merged: {err} (in {wt})", file=sys.stderr) + return 1 + print(f"report: commit {merged}") + return 0 + + +@contextlib.contextmanager +def pull_claim(rec: Path): + """Exclusive non-blocking flock on the task's record: yields False when another pull holds it or already + ended it (record gone), so two pulls of one id never race into 'branch … exists (local WIP?)'.""" + try: + fd = os.open(rec, os.O_RDONLY) + except FileNotFoundError: + yield False + return + try: + try: + fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB) + except BlockingIOError: + yield False + return + yield rec.exists() + finally: + os.close(fd) + + +def cmd_pull(a, state: Path, now: dt.datetime) -> int: + import wf + p = wf.load_project(a) + folder = p.cfg.root / ".wf" / "cloud" + if a.all: + recs = sorted(folder.glob("*.json")) + if not recs: + print("no task out in the cloud") + return 0 + else: + id = p.resolve(a.id) + recs = [folder / f"{id}.json"] + if not recs[0].exists(): + raise cloud.CloudError(f"{id} is not out in the cloud (no {p.cfg.rel(recs[0])})") + codes = [] + for rec in recs: + try: + with pull_claim(rec) as mine: + if not mine: # a concurrent pull (batch sidecar + orchestrator) has it: not WIP, not an error + print(f"{rec.stem}: pulled by another wf cloud pull, skipped") + codes.append(0) + continue + codes.append(pull_one(a, wf.load_project(a), rec, state, now)) + except (cloud.CloudError, wf.Failure) as e: + print(f"wf: {rec.stem}: {e}", file=sys.stderr) + codes.append(1) + sys.stdout.flush() + return 1 if 1 in codes else 4 if 4 in codes else 0 + + +def main(argv: list[str], state: Path | None = None, now: dt.datetime | None = None) -> int: + ap = argparse.ArgumentParser(prog="wf cloud", description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + sub = ap.add_subparsers(dest="sub", required=True) + lp = sub.add_parser("ledger", help="print budget/spent/balance/running") + lp.add_argument("--set-balance", type=float, metavar="USD", help="balance shown on claude.ai: spent = budget - balance") + lp.add_argument("--budget", type=float, metavar="USD", help="change the cap") + sp = sub.add_parser("send", help="snapshot + prompt -> claude --cloud; ledger reserve, claim, record") + sp.add_argument("id") + sp.add_argument("--dry-run", action="store_true", help="print snapshot size + prompt; send nothing") + sp.add_argument("--project", metavar="DIR", help="project folder (default: from the working directory)") + pp = sub.add_parser("pull", help="export result -> apply patch, gate, done/merge | awaiting | handback; ledger charge; an id another pull holds → '<id>: pulled by another …, skipped' (rc 0)") + g = pp.add_mutually_exclusive_group(required=True) + g.add_argument("id", nargs="?") + g.add_argument("--all", action="store_true", help="every task of the project out in the cloud") + pp.add_argument("--project", metavar="DIR", help="project folder (default: from the working directory)") + pp.add_argument("--export", metavar="FILE", help="parse this /export file instead of teleporting (one id)") + pp.add_argument("--no-push", action="store_true", help="merge without git push home") + ar = sub.add_parser("archive", help="archive ended sessions in the app (pull does it; this retries)") + g = ar.add_mutually_exclusive_group(required=True) + g.add_argument("sid", nargs="?") + g.add_argument("--ended", action="store_true", help="every ended ledger session not archived yet") + a = ap.parse_args(argv) + now = now or dt.datetime.now().astimezone() + if a.sub == "archive": + return cmd_archive(a, state or state_dir()) + if a.sub in ("send", "pull"): + import wf + from wflib import config, tasks + try: + if a.sub == "pull": + if a.export and a.all: + raise cloud.CloudError("--export goes with one id, not --all") + return cmd_pull(a, state or state_dir(), now) + return cmd_send(a, state or state_dir(), now) + except (cloud.CloudError, wf.Failure, tasks.TaskError, config.ConfigError) as e: + print(f"wf: {e}", file=sys.stderr) + return 1 + try: + with locked(state or state_dir()) as led: + if a.budget is not None: + led["budget"] = a.budget + if a.set_balance is not None: + cloud.set_balance(led, a.set_balance, now) + print(cloud.summary(led)) + except cloud.CloudError as e: + print(f"wf: {e}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/wf_res.py b/wf_res.py new file mode 100644 index 0000000..cc9a415 --- /dev/null +++ b/wf_res.py @@ -0,0 +1,1208 @@ +"""wf res — shared memory/CPU ledger for agent jobs (files, /proc, systemd, printing). +Pure logic: wflib/res.py. Spec: docs/resource-ledger.md. + + run --mem 10G [--cpus N] --for 40m --title T [--queue|--force] [--by NAME] -- CMD… start a job or get `busy` (exit 3) + status [ID] [--json] · wait ID · release ID [--stop] · note --mem 3G --for 20m [--force] [--by NAME] TITLE + game on [--for 4h] | off · clean [--yes] · timer on|off · tick · shell-init +""" +from __future__ import annotations + +import re +import argparse +import contextlib +import datetime as dt +import fcntl +import json +import os +import shlex +import shutil +import stat +import subprocess +import sys +import time +import tomllib +from dataclasses import dataclass +from pathlib import Path +from typing import Callable + +HERE = Path(__file__).resolve().parent +sys.path.insert(0, str(HERE)) + +from wflib import res # noqa: E402 + +LEDGER, LOCK = "resources.json", "resources.lock" +HISTORY = "resources-history.jsonl" # finished runs pruned from the ledger: reservation sizes (wf res hist) +ACTIVE = {"active", "activating", "deactivating", "reloading", "refreshing"} + + +@dataclass +class Env: + """Everything wf res touches outside its own state; tests swap the callables.""" + state: Path + config: Path + units: Path + scratch: Path + run: Callable[[list[str]], subprocess.CompletedProcess] + meminfo: Callable[[], str] + nproc: int + now: Callable[[], dt.datetime] + pid_alive: Callable[[int], bool] + owner: Callable[[], int] + root: Callable[[], Path] + cwd: Callable[[], str] + procs: Callable[[], list[tuple[int, int, str, str]]] # (pid, ppid, comm, cgroup) + sleep: Callable[[float], None] + tmp_used_gb: Callable[[], float] + stdin: Callable[[], str] = lambda: sys.stdin.read() + live_sessions: Callable[[], set[str]] = lambda: _live_sessions() + tmp_dirs: tuple = (Path("/tmp"), Path("/var/tmp")) # swept for litter (res.litter_victims) + cgread: Callable[[str, str], str] = lambda cg, name: _read(f"/sys/fs/cgroup{cg}/{name}") # ControlGroup, file + environ: Callable[[], dict] = lambda: dict(os.environ) + + +# ---------------------------------------------------------------- real machine + +def _read(path: str) -> str: + try: + return Path(path).read_text() + except OSError: + return "" + + +def _live_sessions() -> set[str]: + """Session ids of running Claude processes (Claude Code's ~/.claude/sessions/<pid>.json); unreadable → empty.""" + out = set() + for f in (Path.home() / ".claude" / "sessions").glob("*.json"): + try: + stat = Path(f"/proc/{f.stem}/stat").read_text() + sid = res.live_session_id(f.read_text(), stat.rsplit(")", 1)[1].split()[19]) + except (OSError, IndexError): + continue + if sid: + out.add(sid) + return out + + +def _run(argv): + return subprocess.run(argv, capture_output=True, text=True) + + +def _xdg(var: str, default: str) -> Path: + return Path(os.environ.get(var) or Path.home() / default) + + +def _argv0(d: Path) -> str: + try: + return (d / "cmdline").read_bytes().split(b"\0", 1)[0].decode(errors="replace") + except OSError: + return "" + + +def _comm(pid: int) -> str: + """Process name; `claude` for any Claude Code process (res.proc_name).""" + d = Path(f"/proc/{pid}") + try: + comm = (d / "comm").read_text().strip() + except OSError: + return "" + return res.proc_name(comm, _argv0(d)) + + +def _ppid(pid: int) -> int: + return int(Path(f"/proc/{pid}/stat").read_text().rsplit(")", 1)[1].split()[1]) + + +def _owner() -> int: + """Nearest ancestor named `claude` (the session outlives its shells), else wf's parent.""" + parent = p = os.getppid() + while p > 1: + if _comm(p) == "claude": + return p + try: + p = _ppid(p) + except (OSError, ValueError, IndexError): + break + return parent + + +def _root() -> Path: + r = _run(["git", "rev-parse", "--show-toplevel"]) + return Path(r.stdout.strip()) if r.returncode == 0 and r.stdout.strip() else Path.cwd() + + +def _procs() -> list[tuple[int, int, str, str]]: + out = [] + for d in Path("/proc").iterdir(): + if not d.name.isdigit(): + continue + try: + stat = (d / "stat").read_text() + comm = stat[stat.index("(") + 1:stat.rindex(")")] + ppid = int(stat.rsplit(")", 1)[1].split()[1]) + cgroup = (d / "cgroup").read_text().strip().split(":", 2)[2] + except (OSError, ValueError, IndexError): + continue + out.append((int(d.name), ppid, res.proc_name(comm, _argv0(d)), cgroup)) + return out + + +def _tmp_used_gb() -> float: + st = os.statvfs("/tmp") + return (st.f_blocks - st.f_bfree) * st.f_frsize / res.GB + + +def real_env() -> Env: + runtime = Path(os.environ.get("XDG_RUNTIME_DIR") or f"/run/user/{os.getuid()}") + if not (runtime / "systemd").is_dir(): + raise res.ResError("wf res needs a systemd user session") + return Env(state=_xdg("XDG_STATE_HOME", ".local/state") / "wf", + config=_xdg("XDG_CONFIG_HOME", ".config") / "wf" / "resources.toml", + units=_xdg("XDG_CONFIG_HOME", ".config") / "systemd" / "user", + scratch=Path(f"/tmp/claude-{os.getuid()}"), run=_run, + meminfo=lambda: Path("/proc/meminfo").read_text(), nproc=os.cpu_count() or 1, + now=lambda: dt.datetime.now().astimezone(), pid_alive=lambda p: Path(f"/proc/{p}").exists(), + owner=_owner, root=_root, cwd=os.getcwd, procs=_procs, sleep=time.sleep, + tmp_used_gb=_tmp_used_gb) + + +# ---------------------------------------------------------------- ledger file + +def warn(msg: str) -> None: + print(f"wf: {msg}", file=sys.stderr) + + +def write_atomic(path: Path, text: str) -> None: + tmp = path.with_name(path.name + ".tmp") + tmp.write_text(text) + os.replace(tmp, path) + + +@contextlib.contextmanager +def locked(env: Env): + """Exclusive flock; yields the ledger; writes it back atomically when the block ends without error.""" + env.state.mkdir(parents=True, exist_ok=True) + with open(env.state / LOCK, "w") as lock: + fcntl.flock(lock, fcntl.LOCK_EX) + path = env.state / LEDGER + text = path.read_text() if path.exists() else "" + try: + led = res.loads(text) + except res.ResError as e: + bad = path.with_name(f"{LEDGER}.bad-{env.now():%Y%m%d-%H%M%S}") + path.rename(bad) + warn(f"{e}; moved to {bad.name}, starting empty") + led = res.Ledger(next=res.salvage_next(text)) + yield led + write_atomic(path, res.dumps(led)) + + +# ---------------------------------------------------------------- systemd + +def systemctl(env: Env, args: list[str]) -> str: + r = env.run(["systemctl", "--user", *args]) + if r.returncode != 0: + first = (r.stderr or r.stdout).strip().splitlines() + raise res.ResError(f"systemctl {args[0]}: " + (first[0] if first else f"exit {r.returncode}")) + return r.stdout + + +def ensure_slice(env: Env) -> None: + f = env.units / res.SLICE + if not f.exists(): + env.units.mkdir(parents=True, exist_ok=True) + f.write_text(res.slice_unit_text()) + systemctl(env, ["daemon-reload"]) + + +def _result(env: Env, e: res.Entry) -> tuple: + """(rc, peak_gb, why); no .rc (wrapper killed with the job) → why from the journal.""" + def read(name): + try: + return (env.state / "logs" / name).read_text().strip() + except OSError: + return "" + rc, peak = read(f"{e.id}.rc"), read(f"{e.id}.peak") + rc, peak = int(rc) if rc.lstrip("-").isdigit() else None, int(peak) / res.GB if peak.isdigit() else None + if rc is not None: + return rc, peak, "" + argv = ["journalctl", "--user", "-u", e.unit, "-o", "cat", "--no-pager"] + if e.started: + argv += ["--since", f"@{int(e.started.timestamp())}"] + why, jpeak = res.journal_reason(env.run(argv).stdout or "", e.mem_gb) + return rc, peak if peak is not None else jpeak, why or f"no exit code; see journalctl --user -u {e.unit}" + + +def bare_units(env: Env) -> list: + names = [l.split()[0] for l in systemctl(env, ["list-units", "--type=service", "--state=running", "--no-legend", + "--plain"]).splitlines() if l.strip()] + if not names: + return [] + return res.bare_units(res.show_units(systemctl(env, ["show", "-p", "Id,Transient,Slice,MemoryCurrent", *names]))) + + +def session_scopes(env: Env) -> dict[str, dict[str, str]]: + """Running scopes in agents.slice (Claude sessions) → show props.""" + names = [l.split()[0] for l in systemctl(env, ["list-units", "--type=scope", "--state=running", "--no-legend", + "--plain"]).splitlines() if l.strip()] + if not names: + return {} + blocks = res.show_units(systemctl(env, ["show", "-p", "Id,Slice,MemoryHigh,MemoryCurrent,ControlGroup", *names])) + return {n: b for n, b in blocks.items() if b.get("Slice") == res.SLICE} + + +def cap_sessions(env: Env, cfg: res.Config) -> None: + for unit, value in res.session_caps(session_scopes(env), cfg.session_mem_gb): + systemctl(env, ["set-property", "--runtime", unit, f"MemoryHigh={value}"]) + + +def gather(env: Env, led: res.Ledger) -> res.Facts: + total, available = res.meminfo(env.meminfo()) + names = sorted({e.unit for e in led.entries if e.state == "running"}) + blocks = res.show_units(systemctl(env, ["show", "-p", "Id,ActiveState,MemoryCurrent,ControlGroup", + res.SLICE, *names])) + units = {} + for n in names: + b = blocks.get(n, {}) + active, cg = b.get("ActiveState") in ACTIVE, b.get("ControlGroup", "") + units[n] = res.Unit(active, res.gb_or_none(b.get("MemoryCurrent", "")), + res.psi_some_avg60(env.cgread(cg, "memory.pressure")) if active and cg else None) + scg = blocks.get(res.SLICE, {}).get("ControlGroup", "") + return res.Facts( + now=env.now(), total_gb=total, available_gb=available, nproc=env.nproc, units=units, + live={e.owner for e in led.entries if e.state == "note" and env.pid_alive(e.owner)}, + results={e.id: _result(env, e) for e in led.entries if e.state == "running" and not units[e.unit].active}, + slice_gb=res.gb_or_none(blocks.get(res.SLICE, {}).get("MemoryCurrent", "")) or 0.0, + slice_cache_gb=res.cache_gb(env.cgread(scg, "memory.stat")) if scg else 0.0) + + +def start(env: Env, e: res.Entry, now: dt.datetime) -> str | None: + """Launch e as a unit; on success e is running. Returns an error line or None.""" + logs = env.state / "logs" + logs.mkdir(parents=True, exist_ok=True) + e.unit, e.log = f"wf-{e.id}.service", str(logs / f"{e.id}.log") + for stale in (f"{e.id}.rc", f"{e.id}.peak"): # a reused id must not inherit an old exit code + (logs / stale).unlink(missing_ok=True) + ensure_slice(env) + r = env.run(res.run_argv(e, str(logs))) + if r.returncode != 0: + first = (r.stderr or r.stdout).strip().splitlines() + return "systemd-run: " + (first[0] if first else f"exit {r.returncode}") + e.state, e.started = "running", now + e.expires = now + dt.timedelta(minutes=2 * e.est_min) + return None + + +def housekeep(env: Env, cfg: res.Config, led: res.Ledger, facts: res.Facts) -> list[str]: + """Every command, under the lock: prune, start queued work.""" + done = [e for e in led.entries if e.state == "done"] + lines = res.prune(led, facts) + kept = {id(e) for e in led.entries} + save_history(env, [r for r in (res.history_record(e) for e in done if id(e) not in kept) if r]) + for e in res.to_start(cfg, led, facts): + err = start(env, e, facts.now) + if err: + e.state, e.ended, e.why = "done", facts.now, f"start failed: {err}" + lines.append(f"{e.id} {e.why}") + else: + lines.append(f"{e.id} started (queued)") + if res.gaming(led, facts.now) or "game off (expired)" in lines: + apply_slice(env, cfg, led, facts) + if not led.last_clean or facts.now - led.last_clean >= res.CLEAN_EVERY: + lines.extend(clean_auto(env, cfg, led, facts.now)) + return lines + + +def read_history(env: Env) -> list[dict]: + path = env.state / HISTORY + out = [] + for line in (path.read_text().splitlines() if path.exists() else []): + try: + r = json.loads(line) + except ValueError: + continue + if isinstance(r, dict): + out.append(r) + return out + + +def save_history(env: Env, recs: list[dict]) -> None: + """Append runs leaving the ledger; keep the newest res.HIST_KEEP (call under the lock).""" + if not recs: + return + path = env.state / HISTORY + old = path.read_text().splitlines() if path.exists() else [] + lines = old + [json.dumps(r) for r in recs] + if len(lines) > res.HIST_KEEP: + write_atomic(path, "".join(x + "\n" for x in lines[-res.HIST_KEEP:])) + else: + with open(path, "a") as f: + f.write("".join(json.dumps(r) + "\n" for r in recs)) + + +def suggestion(env: Env, led: res.Ledger, title: str) -> tuple[float, int, int] | None: + """History's (mem GB, minutes, n) for this project's kind of title (res.suggest), else None.""" + return res.suggest(res.all_runs(read_history(env), led), main_tree(env).name, res.title_kind(title)) + + +def hint(env: Env, led: res.Ledger, title: str, mem: float, est: int) -> None: + line = res.hint_line(mem, est, suggestion(env, led, title)) + if line: + print(line, file=sys.stderr) + + +def session(env: Env, cfg: res.Config, led: res.Ledger) -> res.Facts: + facts = gather(env, led) + housekeep(env, cfg, led, facts) + return facts + + +# ---------------------------------------------------------------- commands + +def _git(env: Env, *args) -> str: + r = env.run(["git", "-C", env.cwd(), *args]) + return r.stdout.strip() if r.returncode == 0 else "" + + +def main_tree(env: Env) -> Path: + """Project's main tree: git common dir's parent (same for its lane worktrees), else env.root().""" + common = _git(env, "rev-parse", "--path-format=absolute", "--git-common-dir") + return Path(common).parent if common and Path(common).name == ".git" else env.root() + + +def _session_records(env: Env) -> list[dict]: + """wf lane session records (.wf/sessions/*.json) of this project: main tree + root.""" + out = [] + for top in dict.fromkeys([main_tree(env), env.root()]): + for f in sorted((top / ".wf" / "sessions").glob("*.json")): + try: + out.append(json.loads(f.read_text())) + except (OSError, ValueError): + pass + return out + + +def owner_by(env: Env, name: str = "") -> dict: + return res.owner_by(env.environ(), _git(env, "branch", "--show-current"), _session_records(env), name) + + +def _new_entry(env, led, facts, args, mem, est, state) -> res.Entry: + return res.Entry(id=led.new_id(), project=main_tree(env).name, owner=env.owner(), title=args.title, mem_gb=mem, + cpus=args.cpus, est_min=est, state=state, cwd=env.cwd(), queued=facts.now, + by=owner_by(env, getattr(args, "by", "") or "")) + + +def force_busy(cfg, led, facts, mem, cpus, room) -> str: + """--force refused: not even the really free memory (beyond the user reserve) holds it.""" + free = f"{res.fmt_gb(max(0.0, room[0]))}" if mem > room[0] + res.EPS else f"{max(0, room[1])} cpus" + return f"busy even with --force: only {free} really free beyond the reserve; retry later or work on something else" + + +def cmd_run(env, cfg, args) -> int: + mem, est = res.parse_size(args.mem), res.parse_duration(args.for_) + cmd = args.cmd[1:] if args.cmd[:1] == ["--"] else args.cmd + if not cmd: + raise res.ResError("no command after --") + fail = busy = out = None + shown = env.run(["systemctl", "--user", "show-environment"]) + caller_env = res.env_diff(env.environ(), res.parse_show_environment(shown.stdout if shown.returncode == 0 else "")) + lock = res.lock_for(args.title, getattr(args, "lock", ""), str(main_tree(env))) + with locked(env) as led: + facts = session(env, cfg, led) + hint(env, led, args.title, mem, est) + fail = res.never_fits(cfg, facts, mem, args.cpus) + holder = res.lock_holder(led, lock) + if not fail and holder and not args.queue: + busy = res.lock_busy_line(holder, facts) + elif not fail and holder: + e = _new_entry(env, led, facts, args, mem, est, "queued") + e.cmd, e.env, e.lock = cmd, caller_env, lock + led.entries.append(e) + when = res.queue_estimate(cfg, led, facts, e) + out = (f"{e.id} queued, position {res.position(led, e)} (lock held by {holder.id}); est. start " + + (f"~{res.hhmm(when)}" if when else "unknown") + f"; cancel: wf res release {e.id}") + elif not fail: + q_gb, q_cpus = res.queued_total(led) + room = res.force_room(cfg, led, facts) if args.force else res.budget(cfg, led, facts) + if args.force: + q_gb, q_cpus = 0.0, 0 + if res.fits(mem + q_gb, args.cpus + q_cpus, room): + e = _new_entry(env, led, facts, args, mem, est, "queued") + e.cmd, e.env, e.lock = cmd, caller_env, lock + fail = start(env, e, facts.now) + if not fail: + led.entries.append(e) + out = f"{e.id} started; log {e.log}; ETA ~{res.hhmm(e.eta())} (estimate: not killed when over)" + elif args.queue: + e = _new_entry(env, led, facts, args, mem, est, "queued") + e.cmd, e.env, e.lock = cmd, caller_env, lock + led.entries.append(e) + when = res.queue_estimate(cfg, led, facts, e) + out = (f"{e.id} queued, position {res.position(led, e)}; est. start " + + (f"~{res.hhmm(when)}" if when else "unknown") + f"; cancel: wf res release {e.id}") + elif args.force: + busy = force_busy(cfg, led, facts, mem, args.cpus, room) + else: + busy = res.busy_line(cfg, led, facts, mem + q_gb, args.cpus + q_cpus, + res.force_hint(cfg, led, facts, mem, args.cpus)) + if fail: + raise res.ResError(fail) + if busy: + raise res.Busy(busy) + print(out) + return 0 + + +def cmd_status(env, cfg, args) -> int: + with locked(env) as led: + facts = session(env, cfg, led) + if args.json: + text = res.status_json(cfg, led, facts) + elif args.id: + text = "\n".join(res.detail_lines(led.get(args.id))) + else: + outside = len(res.adopt_groups(env.procs())) + scopes = session_scopes(env) + stall = {n: res.psi_some_avg60(env.cgread(b.get("ControlGroup", ""), "memory.pressure")) or 0.0 + for n, b in scopes.items() if b.get("ControlGroup")} + text = "\n".join(res.status_lines(cfg, led, facts) + res.throttle_lines(led, facts) + + res.session_cap_lines(scopes, stall, cfg.session_mem_gb) + + res.warning_lines(facts.total_gb, env.tmp_used_gb(), outside, bare_units(env))) + print(text) + return 0 + + +def cmd_hist(env, cfg, args) -> int: + with locked(env) as led: + runs = res.all_runs(read_history(env), led) + print("\n".join(res.hist_lines(runs, args.project))) + return 0 + + +def cmd_release(env, cfg, args) -> int: + fail = out = None + with locked(env) as led: + facts = session(env, cfg, led) + e = led.get(args.id) + if e.state == "running" and not args.stop: + fail = f"{e.id} is running; --stop to kill it" + elif e.state == "running": + systemctl(env, ["stop", e.unit]) + e.rc, e.peak_gb, _ = _result(env, e) + e.state, e.ended, e.why = "done", facts.now, "released" + out = f"{e.id} stopped" + else: + led.entries.remove(e) + out = f"{e.id} released" + if fail: + raise res.ResError(fail) + print(out) + return 0 + + +POLL_SECONDS = 15 + + +def cmd_wait(env, cfg, args) -> int: + limit = res.parse_duration(args.timeout) + deadline = env.now() + dt.timedelta(minutes=limit) + warned = False + while True: + with locked(env) as led: + facts = session(env, cfg, led) + e = led.get(args.id) + line, state = (res.done_line(e) if e.state == "done" else None), e.state + hot = res.throttle_lines(res.Ledger(entries=[e]), facts) + if hot and not warned: # once per wait + warn(hot[0]) + warned = True + if line: + print(line) + return 0 + if env.now() >= deadline: + raise res.ResError(f"{args.id} not done after {res.fmt_dur(limit)} ({state})") + env.sleep(POLL_SECONDS) +def cmd_note(env, cfg, args) -> int: + mem, est = res.parse_size(args.mem), res.parse_duration(args.for_) + fail = busy = out = None + shown = env.run(["systemctl", "--user", "show-environment"]) + caller_env = res.env_diff(env.environ(), res.parse_show_environment(shown.stdout if shown.returncode == 0 else "")) + with locked(env) as led: + facts = session(env, cfg, led) + hint(env, led, args.title, mem, est) + fail = res.never_fits(cfg, facts, mem, args.cpus) + q_gb, q_cpus = (0.0, 0) if args.force else res.queued_total(led) + room = res.force_room(cfg, led, facts) if args.force else res.budget(cfg, led, facts) + if not fail and res.fits(mem + q_gb, args.cpus + q_cpus, room): + e = _new_entry(env, led, facts, args, mem, est, "note") + e.started, e.expires = facts.now, facts.now + dt.timedelta(minutes=est) + led.entries.append(e) + out = f"{e.id} noted {res.fmt_gb(mem)} until ~{res.hhmm(e.expires)} (then freed; nothing is killed)" + elif not fail and args.force: + busy = force_busy(cfg, led, facts, mem, args.cpus, room) + elif not fail: + busy = res.busy_line(cfg, led, facts, mem + q_gb, args.cpus + q_cpus, + res.force_hint(cfg, led, facts, mem, args.cpus)) + if fail: + raise res.ResError(fail) + if busy: + raise res.Busy(busy) + print(out) + return 0 +def apply_slice(env: Env, cfg: res.Config, led: res.Ledger, facts: res.Facts) -> None: + ensure_slice(env) + props = res.slice_props(cfg, led, facts) + systemctl(env, ["set-property", "--runtime", res.SLICE, *(f"{k}={v}" for k, v in props.items())]) + + +def cmd_game(env, cfg, args) -> int: + with locked(env) as led: + facts = session(env, cfg, led) + if args.state == "on": + minutes = res.parse_duration(args.for_) if args.for_ else int(cfg.game_hours * 60) + led.game_until = max(led.game_until or facts.now, facts.now + dt.timedelta(minutes=minutes)) + lines = res.game_on_lines(cfg, led, facts, minutes) + else: + led.game_until = None + lines = ["game off"] + apply_slice(env, cfg, led, facts) + print("\n".join(lines)) + return 0 +def _lstat(path) -> os.stat_result | None: + """None when the path vanished meanwhile (scratch is shared with running sessions).""" + try: + return os.lstat(path) + except OSError: + return None + + +def _newest(p: Path) -> float: + st = _lstat(p) + newest = st.st_mtime if st else 0.0 + if p.is_dir() and not p.is_symlink(): + for dirpath, dirnames, filenames in os.walk(p, followlinks=False): + for name in dirnames + filenames: + if st := _lstat(os.path.join(dirpath, name)): + newest = max(newest, st.st_mtime) + return newest + + +def _size(p: Path) -> int: + if p.is_symlink() or not p.is_dir(): + st = _lstat(p) + return st.st_size if st else 0 + return sum(st.st_size for d, _, fs in os.walk(p, followlinks=False) for f in fs + if (st := _lstat(os.path.join(d, f)))) + + +def _remove(p: Path) -> None: + if p.is_dir() and not p.is_symlink(): + shutil.rmtree(p, ignore_errors=True) + else: + p.unlink(missing_ok=True) + + +def scan_scratch(env: Env) -> dict: + tree = {} + if not env.scratch.is_dir(): + return tree + for top in env.scratch.iterdir(): + kids = {} + if top.is_dir() and not top.is_symlink(): + kids = {str(c): _newest(c) for c in top.iterdir()} + tree[str(top)] = (_newest(top), kids) + return tree + + +def scan_litter(env: Env) -> list[tuple[str, str, float]]: + uid, out = os.getuid(), [] + for d in env.tmp_dirs: + try: + names = os.listdir(d) + except OSError: + continue + for n in names: + p = Path(d) / n + st = _lstat(p) + if st is None or st.st_uid != uid: + continue + if stat.S_ISDIR(st.st_mode): + try: + kind = "dir" if os.listdir(p) else "emptydir" + except OSError: + continue + else: + kind = "file" if stat.S_ISREG(st.st_mode) and st.st_size else "other" + out.append((str(p), kind, st.st_mtime)) + return out + + +def sweep_litter(env: Env, cfg: res.Config, now: dt.datetime) -> int: + n = 0 + for path in res.litter_victims(scan_litter(env), now.timestamp(), cfg.scratch_hours, env.pid_alive): + try: + (os.rmdir if os.path.isdir(path) and not os.path.islink(path) else os.unlink)(path) + n += 1 + except OSError: + pass # refilled or gone meanwhile + return n + + +def clean_auto(env: Env, cfg: res.Config, led: res.Ledger, now: dt.datetime) -> list[str]: + """Scratch untouched for scratch_hours (D4) and failed wf units; no confirmation needed.""" + lines = [] + for path in res.scratch_victims(scan_scratch(env), now.timestamp(), cfg.scratch_hours, + env.live_sessions()): + p = Path(path) + size = _size(p) + _remove(p) + lines.append(f"freed {res.fmt_bytes(size)}: {p}") + swept = sweep_litter(env, cfg, now) + if swept: + lines.append(f"swept {swept} temp-dir litter entries (empty dirs, dead .NET pipes)") + env.run(["systemctl", "--user", "reset-failed", "wf-r-*"]) + led.last_clean = now + return lines + + +def project_rules(env: Env) -> tuple[Path, list[tuple[str, int | None]]]: + root = env.root() + toml = root / "workflow.toml" + if not toml.exists(): + return root, [] + try: + patterns = tomllib.loads(toml.read_text()).get("cleanup", []) + except tomllib.TOMLDecodeError as e: + raise res.ResError(f"workflow.toml: {e}") from None + return root, [res.cleanup_rule(p) for p in patterns] + + +def project_matches(env: Env, now: dt.datetime) -> list[Path]: + root, rules = project_rules(env) + out = [] + for pattern, days in rules: + for p in sorted(root.glob(pattern)): + if p.is_symlink() or not p.resolve().is_relative_to(root.resolve()): + continue + if days is not None and now.timestamp() - _newest(p) < days * 86400: + continue + out.append(p) + return out + + +def cmd_clean(env, cfg, args) -> int: + with locked(env) as led: + facts = gather(env, led) + lines = [l for l in housekeep(env, cfg, led, facts) if l.startswith("freed ")] + lines += clean_auto(env, cfg, led, facts.now) + root = env.root() + for p in project_matches(env, facts.now): + size, rel = res.fmt_bytes(_size(p)), p.relative_to(root) + if args.yes: + _remove(p) + lines.append(f"deleted {size}: {rel}") + else: + lines.append(f"would delete {size}: {rel} (wf res clean --yes)") + print("\n".join(lines) if lines else "nothing to clean") + return 0 + + +def cmd_timer(env, cfg, args) -> int: + service, timer = env.units / "wf-res.service", env.units / "wf-res.timer" + if args.state == "on": + env.units.mkdir(parents=True, exist_ok=True) + service.write_text(res.service_text(sys.executable, str(HERE / "wf.py"))) + timer.write_text(res.timer_text()) + systemctl(env, ["daemon-reload"]) + systemctl(env, ["enable", "--now", "wf-res.timer"]) + print("timer on: wf-res.timer every 1 min") + else: + if timer.exists(): + systemctl(env, ["disable", "--now", "wf-res.timer"]) + service.unlink(missing_ok=True) + timer.unlink(missing_ok=True) + systemctl(env, ["daemon-reload"]) + print("timer off") + return 0 + + +def adopt(env: Env) -> list[str]: + """Move running Claude sessions (and their descendants) into agents.slice; one line per session.""" + lines = [] + groups = res.adopt_groups(env.procs()) + if groups: + ensure_slice(env) + for root, pids in groups.items(): + r = env.run(res.adopt_argv(root, pids)) + if r.returncode == 0: + lines.append(f"adopted claude {root} ({len(pids)} processes)") + else: + first = (r.stderr or r.stdout).strip().splitlines() + lines.append(f"claude {root}: " + (first[0] if first else f"exit {r.returncode}")) + return lines + + +def cmd_adopt(env, cfg, args) -> int: + print("\n".join(adopt(env)) or "all claude sessions already in agents.slice") + return 0 + + +def cmd_tick(env, cfg, args) -> int: + with locked(env) as led: + session(env, cfg, led) + adopt(env) # silent; a session that vanished mid-move is retried next minute + try: + cap_sessions(env, cfg) + except res.ResError: + pass # a session that ended between list and set-property; retried next minute + return 0 + + +def cmd_hook(env, cfg, args) -> int: + """PreToolUse hook; read-only (no lock, no systemctl) and silent on any error: it runs before every command.""" + try: + payload = json.loads(env.stdin()) + led = res.loads((env.state / LEDGER).read_text()) + out = res.hook_output(payload, led.game_until, env.now()) + except (OSError, ValueError, AttributeError, res.ResError): + return 0 + if out: + print(json.dumps(out)) + return 0 + + +def cmd_shell_init(env, cfg, args) -> int: + print(res.shell_init_line()) + return 0 +# ---------------------------------------------------------------- main + +def parser() -> argparse.ArgumentParser: + ap = argparse.ArgumentParser(prog="wf res", description=__doc__, + formatter_class=argparse.RawDescriptionHelpFormatter) + sub = ap.add_subparsers(dest="command", metavar="COMMAND", required=True) + + def cmd(name, func, help): + sp = sub.add_parser(name, help=help, description=help) + sp.set_defaults(func=func) + return sp + + sp = cmd("run", cmd_run, "reserve and start a job as a systemd user unit; busy → exit 3") + sp.add_argument("--mem", required=True, metavar="SIZE", help="e.g. 10G, 512M") + sp.add_argument("--cpus", type=int, default=1, choices=range(1, 1025), metavar="N") + sp.add_argument("--for", dest="for_", required=True, metavar="DURATION", help="estimate, e.g. 40m, 2h (ETA and queue; the job is never killed at it)") + sp.add_argument("--title", required=True) + sp.add_argument("--queue", action="store_true", help="wait in the FIFO instead of exit 3") + sp.add_argument("--force", action="store_true", help=FORCE_HELP) + sp.add_argument("--by", default="", metavar="NAME", help="your session name (ListAgents), shown as the owner; default: WF_SESSION_NAME, else the wf lane session record") + sp.add_argument("--lock", default="", metavar="KEY", help="one queued/running job per KEY in this project (main tree " + "and its lane worktrees) at a time, e.g. a shared checkout; a title starting with 'gate' locks " + "'gate' unasked; held → exit 3, --queue waits for it, --force does not override it") + sp.add_argument("cmd", nargs=argparse.REMAINDER, metavar="-- CMD …") + + sp = cmd("status", cmd_status, "reservations, budget, game mode, warnings") + sp.add_argument("id", nargs="?") + sp.add_argument("--json", action="store_true") + + sp = cmd("wait", cmd_wait, "poll every 15 s until the entry is done; prints rc and peak") + sp.add_argument("id") + sp.add_argument("--timeout", default="2h", metavar="DURATION") + sp = cmd("hist", cmd_hist, "reservation sizes from history, per project and title kind (title minus trailing " + "sha/hex words and (…)): runs n, request, peak p50/p95, estimate, duration p50/p90, suggestion " + "(≥ 3 runs: mem = p95 peak ×1.15 ≥ 0.2 GB, for = p90 duration ×1.5 ≥ 5 min); run/note print a hint " + "when asking > 2× it") + sp.add_argument("--project", metavar="P", help="only this project (dir name of its main tree)") + sp = cmd("note", cmd_note, "reserve for foreground work; freed when this Claude session exits or at --for") + sp.add_argument("--mem", required=True, metavar="SIZE") + sp.add_argument("--cpus", type=int, default=1, choices=range(1, 1025), metavar="N") + sp.add_argument("--for", dest="for_", required=True, metavar="DURATION", help="reservation freed after this (nothing is killed)") + sp.add_argument("--force", action="store_true", help=FORCE_HELP) + sp.add_argument("--by", default="", metavar="NAME", help="your session name (ListAgents), shown as the owner; default: WF_SESSION_NAME, else the wf lane session record") + sp.add_argument("title") + sp = cmd("game", cmd_game, "gaming reserve on (auto-off, default 4h) or off; jobs slow, never stop") + sp.add_argument("state", choices=["on", "off"]) + sp.add_argument("--for", dest="for_", metavar="DURATION") + sp = cmd("clean", cmd_clean, "delete stale /tmp/claude-* scratch (auto) and list project cleanup patterns") + sp.add_argument("--yes", action="store_true", help="also delete the project's workflow.toml cleanup matches") + + cmd("adopt", cmd_adopt, "move running Claude sessions into agents.slice (no restart); the timer does this every minute") + sp = cmd("timer", cmd_timer, "install/remove the 1-minute wf-res.timer (queue, game expiry, clean)") + sp.add_argument("state", choices=["on", "off"]) + cmd("tick", cmd_tick, "housekeeping once (run by the timer); silent") + cmd("hook", cmd_hook, "Claude Code PreToolUse hook (Bash): no display for agent commands while game is on") + cmd("shell-init", cmd_shell_init, "print the alias that starts Claude inside agents.slice (add to ~/.bashrc)") + sp = cmd("release", cmd_release, "drop an entry (cancels a queued job, frees a note); --stop kills a running job") + sp.add_argument("id") + sp.add_argument("--stop", action="store_true") + return ap + + +# ---------------------------------------------------------------- wf batch + +FORCE_HELP = ("memory is really free though the ledger says busy (unused reservations, stale notes): fit against " + "MemAvailable minus the user reserve and headroom only, ignoring claims and the queue; still recorded; " + "the busy line names --force when it would fit") +CEILING_VAR = "CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS" +BATCH_TITLE = "wf-batch" +BATCH_FOR = "6h" # default --for and ceiling (no history) +PREP_MEM = "1G" +STOP_FILE = "out/wf-batch.stop" + + +def cmd_batch(env, cfg, args) -> int: + if args.stop: + f = env.root() / STOP_FILE + f.parent.mkdir(exist_ok=True) + f.touch() + print(f"stop requested: {STOP_FILE} (a batch or loop stops before its next spawn; running workers finish)") + return 0 + if args.status: + return batch_status(env, cfg) + if args.time_left: + return batch_time_left(env, args.time_left) + if args.n is None: + batch_parser().print_usage(sys.stderr) + print("wf batch: error: N is required (or --status)", file=sys.stderr) + return 2 + if args.left: + if args.prep: + raise res.ResError("--left is for work batches, not --prep (use --for)") + left = res.parse_duration(args.left) + p90, runs = _task_p90(env) + k = res.batch_fit(args.n, left, p90) + line = f"fit: {k} of {args.n} tasks (left {res.fmt_dur(left)}, task p90 {_p90_text(p90, runs)})" + if not k: + print(f"{line}; nothing started") + return 0 + print(line) + args.n = k + args.for_ = args.for_ or args.left + sug = None + if not args.prep and (args.mem is None or args.for_ is None): # prep jobs share the title, not the size + with locked(env) as led: + sug = suggestion(env, led, BATCH_TITLE) + minutes = res.parse_duration(args.for_ or BATCH_FOR) # the ceiling: never shortened by history + if args.for_ is None: + args.for_ = f"{sug[1]}m" if sug else BATCH_FOR + if args.mem is None: # prep: one claude -p, subagents in-process, no builds → small + args.mem = PREP_MEM if args.prep else f"{sug[0]:g}G" if sug else f"{min(8.0, max(4.0, 1.5 * args.n)):g}G" + claude = shutil.which("claude", path=env.environ().get("PATH")) + if not claude: + raise res.ResError("claude not found on PATH") + out = f"out/wf-batch-{env.now().strftime('%Y-%m-%d-%H%M')}.md" + names = [x.strip() for x in (args.lanes or "").split(",") if x.strip()] + stamp = env.now().strftime('%Y-%m-%d %H:%M') + if args.prep: + ids, tasks_rel = prep_ids(env.root(), set(names) if names else None, args.n) + if not ids: + print(f"prep: no pending task without Done (lanes: {', '.join(names) or 'all'}); nothing started") + return 0 + prompt = (HERE / "templates" / "prep-prompt.md").read_text().strip().format( + here=HERE, root=env.root(), ids=" ".join(ids), out=out, tasks=tasks_rel) + header = f"prep K={args.n}: {' '.join(ids)}" + else: + lanes = ", ".join(names) if names else \ + "every lane with ready tasks, see wf lanes" + prompt = (HERE / "templates" / "batch-prompt.md").read_text().strip().format( + here=HERE, n=args.n, lanes=lanes, out=out, stop=env.root() / STOP_FILE, + deadline=(env.now() + dt.timedelta(minutes=minutes)).isoformat(timespec="minutes")) + header = f"N={args.n}, lanes: {lanes}" + summary = env.root() / out + if not args.dry_run: # dry-run touches nothing; an existing summary (same minute) is appended to, never clobbered + summary.parent.mkdir(exist_ok=True) + with summary.open("a") as f: + f.write(f"# wf-batch {env.root().name} {stamp} ({header})\n") + cmd = [claude, "-p", "--model", args.model, "--permission-mode", "auto", "--permission-prompts", "none", prompt] + if not args.prep and _cloud_project(env.root()): + cmd = [sys.executable, str(HERE / "wf_res.py"), SIDECAR, "--root", str(env.root()), "--summary", out, + "--every", str(SIDECAR_EVERY), "--", *cmd] + ceiling = str(minutes * 60_000) + if args.dry_run: + print(f"{CEILING_VAR}={ceiling} wf res run --mem {args.mem} --for {args.for_} --title {BATCH_TITLE} -- " + + " ".join(shlex.quote(a) for a in cmd)) + return 0 + stop = env.root() / STOP_FILE + if stop.exists(): + stop.unlink() + print(f"warning: stale {STOP_FILE} cleared; starting batch", file=sys.stderr) + base = env.environ + env.environ = lambda: {**base(), CEILING_VAR: ceiling} + try: + code = cmd_run(env, cfg, argparse.Namespace(mem=args.mem, cpus=1, for_=args.for_, title=BATCH_TITLE, + queue=False, force=args.force, cmd=cmd)) + finally: + env.environ = base + if stop.exists(): + stop.unlink() + print(f"summary: {out} · status: wf batch --status") + return code + + +SIDECAR = "batch-sidecar" +SIDECAR_EVERY = 600 # s between the sidecar's wf cloud pull --all runs +PULLED = re.compile(r"^(\S+): (\w+) \(") +TAIL_CAP = 24 * 3600 # s: the tail after the orchestrator exits pulls at most this long +TAIL_MEM_GB, TAIL_CPUS = 0.2, 1 # the job's reservation during the tail + + +def _cloud_project(root: Path) -> bool: + from wflib import config + try: + return config.load(config.find_root(root)).cloud + except (config.ConfigError, OSError): + return False + + +def sidecar(child: list[str], root: Path, summary: Path, every: float, wf_cmd: list[str], + cap: float = TAIL_CAP, shrink: Callable[[], str] | None = None) -> int: + """Run the batch orchestrator `child`; while it lives, every `every` s (until the stop file appears): + wf cloud pull --all, each ended '<id>: <state> (…' line → wf orch post <id> cloud --result <state> + [--commit <sha>] --no-pick + one summary line. After it exits with .wf/cloud records left: the tail — + shrink() the job's reservation, keep pulling every `every` s (no picks, no sends) until no record is left, + the stop file appears or `cap` s passed; one summary line says why it ended. Returns the child's exit code.""" + proc = subprocess.Popen(child) + while True: + try: + rc = proc.wait(timeout=every) + break + except subprocess.TimeoutExpired: + pass + if (root / STOP_FILE).exists(): + print("sidecar: stop file, no more pulls", flush=True) + return proc.wait() + _pull_post(root, summary, wf_cmd) + if not _cloud_records(root): + return rc + print(f"sidecar: orchestrator exited ({rc}), cloud records left: {' '.join(_cloud_records(root))}; " + f"pulling every {every:g}s (no picks) until none, stop file or {cap / 3600:g}h", flush=True) + if shrink: + print(shrink(), flush=True) + t0, pulls = time.monotonic(), 0 + while True: + time.sleep(every) + if (root / STOP_FILE).exists(): + why = "stop file" + elif time.monotonic() - t0 >= cap: + why = f"{cap / 3600:g}h cap" + else: + _pull_post(root, summary, wf_cmd) + pulls += 1 + if _cloud_records(root): + continue + why = "no cloud records left" + break + left = " ".join(_cloud_records(root)) or "none" + line = f"cloud tail ended: {why}; {pulls} pulls; left: {left}" + print(f"sidecar: {line}", flush=True) + with summary.open("a") as f: + f.write(f"{time.strftime('%H:%M')} {line} (sidecar)\n") + return rc + + +def _cloud_records(root: Path) -> list[str]: + return sorted(p.stem for p in (root / ".wf" / "cloud").glob("*.json")) + + +def _pull_post(root: Path, summary: Path, wf_cmd: list[str]) -> None: + pull = subprocess.run([*wf_cmd, "cloud", "pull", "--all", "--project", str(root)], cwd=root, + capture_output=True, text=True) + print(pull.stdout + pull.stderr, end="", flush=True) + for line in pull.stdout.splitlines(): + m = PULLED.match(line) + if not m or m[2] == "running" or line.startswith("archive "): + continue + sha = re.search(r"merged ([0-9a-f]{7,40})\b", line) + post = [*wf_cmd, "orch", "post", m[1], "cloud", "--result", m[2], "--no-pick"] + post += ["--commit", sha[1]] if sha else [] + r = subprocess.run(post, cwd=root, capture_output=True, text=True) + print(r.stdout + r.stderr, end="", flush=True) + with summary.open("a") as f: + f.write(f"{time.strftime('%H:%M')} cloud {m[1]} {m[2]} {sha[1] if sha else '-'} (sidecar pull)\n") + + +def shrink_reservation(env: Env, res_id: str, mem_gb: float = TAIL_MEM_GB, cpus: int = TAIL_CPUS) -> str: + """The batch job's ledger entry down to the tail's size (ledger only: the unit's MemoryMax stays, a pull + must never be OOM-killed).""" + with locked(env) as led: + e = next((x for x in led.entries if x.id == res_id), None) + if not e or e.state != "running": + return f"sidecar: {res_id} not running in the ledger; reservation kept" + old = e.mem_gb + e.mem_gb, e.cpus = min(e.mem_gb, mem_gb), min(e.cpus, cpus) + return (f"sidecar: {res_id} reservation {res.fmt_gb(old)} -> {res.fmt_gb(e.mem_gb)}, {e.cpus} cpu " + "for the cloud tail") + + +def sidecar_main(argv: list[str]) -> int: + ap = argparse.ArgumentParser(prog=f"wf_res.py {SIDECAR}", description="wf batch job wrapper of cloud projects " + "(started by wf batch, not by hand): runs the orchestrator and pulls cloud tasks") + ap.add_argument("--root", type=Path, required=True) + ap.add_argument("--summary", required=True, help="summary file, relative to --root") + ap.add_argument("--every", type=float, default=SIDECAR_EVERY) + ap.add_argument("--wf", default=f"{sys.executable} {HERE / 'wf.py'}", help="wf command (tests: a fake)") + ap.add_argument("--cap", type=float, default=TAIL_CAP, help="s the tail (cloud records left after the " + "orchestrator exits) pulls at most") + ap.add_argument("cmd", nargs=argparse.REMAINDER) + a = ap.parse_args(argv) + cmd = a.cmd[1:] if a.cmd[:1] == ["--"] else a.cmd + if not cmd: + ap.error("no command after --") + rid = os.environ.get(res.RES_ID_VAR, "") + + def shrink() -> str: + try: + return shrink_reservation(real_env(), rid) + except (res.ResError, OSError) as e: + return f"sidecar: reservation of {rid} kept: {e}" + return sidecar(cmd, a.root, a.root / a.summary, a.every, shlex.split(a.wf), a.cap, shrink if rid else None) + + +def prep_ids(root: Path, only: set[str] | None, k: int) -> tuple[list[str], str]: + """Up to k ids a prep batch works (wflib.lanes.prep_targets) + the task file path relative to root.""" + from wflib import config, lanes, tasks + try: + cfg = config.load(config.find_root(root)) + doc = tasks.parse(cfg.tasks.read_text(encoding="utf-8")) + archived = tasks.archive_ids(cfg.archive.read_text(encoding="utf-8")) + except (config.ConfigError, OSError) as e: + raise res.ResError(f"batch --prep: {e}") from None + items = lanes.prep_targets(doc, archived, cfg.lanes, cfg.slice_above, only) + return [i.id for i in items[:k]], cfg.rel(cfg.tasks) + + +def _orch_progress(root: Path, stem: str) -> list[str]: + """out/wf-orch.log lines (start 'YYYY-MM-DD HH:MM' or 'T') at/after the batch start (from the summary name).""" + m = re.fullmatch(r"wf-batch-(\d{4}-\d\d-\d\d)-(\d\d)(\d\d)", stem) + log = root / "out" / "wf-orch.log" + if not m or not log.is_file(): + return [] + start = f"{m[1]} {m[2]}:{m[3]}" + done = [ln for ln in log.read_text().splitlines() + if re.match(r"\d{4}-\d\d-\d\d[ T]\d\d:\d\d", ln) and ln[:16].replace("T", " ") >= start] + return ["finished since batch start (out/wf-orch.log):"] + done[-30:] if done else [] + + +def _task_p90(env) -> tuple[int, int]: + log = env.root() / "out" / "wf-orch.log" + return res.task_p90(log.read_text(errors="replace") if log.exists() else "") + + +def _p90_text(p90: int, runs: int) -> str: + return f"{res.fmt_dur(p90)} from {runs} runs" if runs >= res.TASK_P90_MIN_RUNS else \ + f"{res.fmt_dur(p90)} default ({runs} runs)" + + +def batch_time_left(env, deadline: str) -> int: + """Per-round check of a running batch: spawn while the time to the deadline is ≥ the task p90.""" + try: + end = dt.datetime.fromisoformat(deadline) + except ValueError: + raise res.ResError(f"--time-left '{deadline}' (want an ISO time, e.g. 2026-10-01T20:00+02:00)") + now = env.now() + if end.tzinfo is None: + end = end.replace(tzinfo=now.tzinfo) + left = max(0, int((end - now).total_seconds() // 60)) + p90, runs = _task_p90(env) + word = "spawn" if left >= p90 else "stop" + print(f"time left {res.fmt_dur(left)}, task p90 {res.fmt_dur(p90)} ({runs} runs): {word}") + return 0 + + +def batch_status(env, cfg) -> int: + root = env.root() + with locked(env) as led: + facts = session(env, cfg, led) + mine = [e for e in led.entries if e.title == BATCH_TITLE and e.project == main_tree(env).name] + lines = [res.entry_line(mine[-1], led, facts), f" log {mine[-1].log}"] if mine else [] + summaries = sorted((root / "out").glob("wf-batch-*.md")) + if summaries: + rel = summaries[-1].relative_to(root) + lines += [f"{rel}:"] + summaries[-1].read_text().rstrip("\n").splitlines()[-30:] + lines += _orch_progress(root, summaries[-1].stem) + if (root / STOP_FILE).exists(): + lines.append(f"stop requested: {STOP_FILE}") + print("\n".join(lines) or "no batch in this project") + return 0 + + +def _count(text: str) -> int: + if not text.isdigit() or int(text) < 1: + raise argparse.ArgumentTypeError("N must be ≥ 1") + return int(text) + + +def batch_parser() -> argparse.ArgumentParser: + ap = argparse.ArgumentParser( + prog="wf batch", description="Start an unattended batch orchestrator: claude -p (auto mode, no prompts) " + "as a wf res job, prompt templates/batch-prompt.md; it works up to N tasks and writes out/wf-batch-*.md. " + "Owner allow rule (once): Bash(python3 /projects/public/workflow/wf.py batch:*).") + ap.add_argument("n", nargs="?", type=_count, metavar="N", help="at most N tasks") + ap.add_argument("--lanes", metavar="L,…", help="lanes to work (default: every lane with ready tasks)") + ap.add_argument("--model", default="opus", help="the batch orchestrator's model (default opus)") + ap.add_argument("--for", dest="for_", default=None, metavar="DURATION", + help=f"estimate and {CEILING_VAR} (default: estimate from wf res hist for wf-batch in this " + f"project when ≥ 3 runs (not --prep), else {BATCH_FOR}; the ceiling stays {BATCH_FOR} unless given)") + ap.add_argument("--mem", default=None, metavar="SIZE", + help="reservation (default: wf res hist suggestion for wf-batch in this project when ≥ 3 runs, " + "else max(4G, 1.5G x N), cap 8G; --prep: 1G; explicit wins; a request > 2x history prints a hint)") + ap.add_argument("--dry-run", action="store_true", help="print the command, start nothing") + ap.add_argument("--force", action="store_true", help=FORCE_HELP) + ap.add_argument("--prep", action="store_true", + help="prep batch: for up to N pending tasks without Done (pickable first, then priority) sonnet " + "workers write Done (+Model) from the task text, unclear ones → awaiting; prompt " + "templates/prep-prompt.md; summary lines 'id → Done: …' | 'id → awaiting a-…'. " + "None to prep → prints so, starts nothing") + ap.add_argument("--left", metavar="DURATION", + help="time left to the pilot's deadline: N becomes min(N, floor(left / task p90)) (p90 of done/" + f"handback durations in out/wf-orch.log, < {res.TASK_P90_MIN_RUNS} runs → {res.TASK_P90_DEFAULT}m), " + "prints 'fit: K of N tasks (…)'; K = 0 → nothing started (rc 0); --for defaults to DURATION") + ap.add_argument("--time-left", metavar="DEADLINE", + help="(run by the batch orchestrator before each round) print 'time left X, task p90 Y (n runs): " + "spawn|stop' for the ISO DEADLINE its prompt names (= start + ceiling); stop = left < p90") + ap.add_argument("--status", action="store_true", help="newest batch job of this project + its summary file + tasks finished since start (out/wf-orch.log)") + ap.add_argument("--stop", action="store_true", + help=f"graceful stop: create {STOP_FILE}; batch and loop check it before every spawn (latency ≤ the running " + "tasks), stop with 'stopped: stop file' and removes it") + ap.set_defaults(func=cmd_batch) + return ap + + +def batch_main(argv: list[str], env: Env | None = None) -> int: + return main(argv, env, batch_parser()) + + +PREP_ALL = 1000 # "wf prep all": more than any task list + + +def prep_main(argv: list[str], env: Env | None = None) -> int: + """wf prep [N|all] [batch opts] = wf batch --prep N (all = every not-ready task).""" + if argv and argv[0] == "all": + argv = [str(PREP_ALL), *argv[1:]] + return batch_main([*argv, "--prep"], env) + + +def load_cfg(env: Env) -> res.Config: + return res.load_config(env.config.read_text() if env.config.exists() else "") + + +def main(argv: list[str], env: Env | None = None, ap: argparse.ArgumentParser | None = None) -> int: + try: + args = (ap or parser()).parse_args(argv) + except SystemExit as e: # usage error → 2, -h → 0; callers (and tests) get a code, not an exception + return e.code if isinstance(e.code, int) else 2 + try: + env = env or real_env() + return args.func(env, load_cfg(env), args) + except res.Busy as e: + print(e) + return 3 + except res.ResError as e: + warn(str(e)) + return 1 + except BrokenPipeError: + return 0 + except OSError as e: + warn(str(e)) + return 1 + + +if __name__ == "__main__": + sys.exit(sidecar_main(sys.argv[2:]) if sys.argv[1:2] == [SIDECAR] else main(sys.argv[1:])) diff --git a/wflib/__init__.py b/wflib/__init__.py new file mode 100644 index 0000000..e69de29 --- /dev/null +++ b/wflib/__init__.py diff --git a/wflib/areas.py b/wflib/areas.py new file mode 100644 index 0000000..ce23d12 --- /dev/null +++ b/wflib/areas.py @@ -0,0 +1,199 @@ +"""Area notes (`## Areas` in the project's CLAUDE.md): code-map anchors, staleness, Checked marker.""" +from __future__ import annotations + +import fnmatch +import re +import subprocess +from dataclasses import dataclass, replace +from pathlib import Path + + +@dataclass(frozen=True) +class Area: + name: str + slug: str + anchors: list[str] + paths: list[str] + checked: str | None + line: int + + +def slug(name: str) -> str: + return re.sub(r"[^a-z0-9]+", "-", name.lower()).strip("-") + + +def parse(text: str) -> list[Area]: + out: list[Area] = [] + in_areas = False + cur: dict | None = None + + def flush(): + if cur: + out.append(Area(cur["name"], slug(cur["name"]), cur["anchors"], cur["paths"], cur["checked"], cur["line"])) + + for n, line in enumerate(text.replace("\r\n", "\n").split("\n"), 1): + if line.startswith("## "): + flush() + cur = None + in_areas = line[3:].strip().lower() == "areas" + elif in_areas and line.startswith("### "): + flush() + cur = {"name": line[4:].strip(), "anchors": [], "paths": [], "checked": None, "line": n} + elif cur is not None: + if line.startswith("- Code map:"): + cur["anchors"] = re.findall(r"`([^`]+)`", line) + elif line.startswith("- Paths:"): + cur["paths"] = line[len("- Paths:"):].split() + elif line.startswith("- Checked:"): + words = line[len("- Checked:"):].split() + cur["checked"] = words[0] if words and re.fullmatch(r"[0-9a-fA-F]{4,40}", words[0]) else None + flush() + return out + + +def matches(area: Area, text: str, skip: set[str] = frozenset()) -> bool: + """text (a task) names the area, one of its anchors or Paths (whole word / path, case-insensitive name); + Paths in `skip` (shared by several areas) never match.""" + def word(tok: str, flags=0) -> bool: + return re.search(r"(?<![\w.])" + re.escape(tok) + r"(?![\w])", text, flags) is not None + paths = [p.removeprefix("./").rstrip("/") for p in area.paths if not any(c in p for c in "*?[")] + paths = [p for p in paths if p not in skip] + return (word(area.name, re.I) or any(word(t) for t in area.anchors) + or any(word(p) for p in paths if p)) + + +def shared_paths(found: list[Area]) -> set[str]: + """Paths (normalised) listed by more than one area.""" + seen: dict[str, int] = {} + for a in found: + for p in {p.removeprefix("./").rstrip("/") for p in a.paths}: + seen[p] = seen.get(p, 0) + 1 + return {p for p, n in seen.items() if n > 1} + + +def block(text: str, name: str) -> str: + """The area's notes: its ### heading up to the next heading, trailing blank lines dropped; unknown → "".""" + lines = text.replace("\r\n", "\n").split("\n") + in_areas, out = False, None + for line in lines: + if line.startswith("#"): + if out is not None: + break + if line.startswith("## "): + in_areas = line[3:].strip().lower() == "areas" + elif in_areas and line.startswith("### ") and line[4:].strip() == name: + out = [line] + elif out is not None: + out.append(line) + while out and not out[-1].strip(): + out.pop() + return "\n".join(out or []) + + +def covers(paths: list[str], file: str) -> bool: + """file (project-relative) matches one of an area's Paths: dir prefix, exact file or glob.""" + for p in paths: + p = p.removeprefix("./").rstrip("/") + if file == p or file.startswith(p + "/") or fnmatch.fnmatchcase(file, p): + return True + return False + + +def uncovered(files: list[str], found: list[Area], ignore: list[str]) -> list[tuple[str, int]]: + """Folders (≤ 2 levels) of files no area's Paths covers → [(folder, file count)], most files first. + Root files, dot folders and `ignore` prefixes skipped; no area with Paths → [].""" + if not any(a.paths for a in found): + return [] + count: dict[str, int] = {} + for f in files: + parts = f.split("/")[:-1] + if not parts or any(x.startswith(".") for x in parts) or covers(ignore, f): + continue + if any(covers(a.paths, f) for a in found): + continue + d = "/".join(parts[:2]) + count[d] = count.get(d, 0) + 1 + return sorted(count.items(), key=lambda kv: (-kv[1], kv[0])) + + +def changed_files(here: Path, base: str | None) -> list[str]: + """Files (relative to here) changed since merge-base(HEAD, base) incl. uncommitted + untracked; + base None → HEAD (uncommitted only).""" + ref = "HEAD" + if base: + r = _git(here, "merge-base", "HEAD", base) + if r.returncode == 0 and r.stdout.strip(): + ref = r.stdout.strip() + out = [] + for args in (("diff", "--name-only", "--relative", ref), ("ls-files", "--others", "--exclude-standard")): + r = _git(here, *args) + if r.returncode == 0: + out += [l for l in r.stdout.splitlines() if l] + return sorted(set(out)) + + +def with_anchor_paths(root: Path, area: Area) -> Area: + """No Paths → the folders of the files holding its anchors (git grep) stand in for them.""" + if area.paths or not area.anchors: + return area + args = [x for tok in area.anchors for x in ("-e", tok)] + r = _git(root, "grep", "-l", "-F", *args) + files = r.stdout.splitlines() if r.returncode == 0 else [] + return replace(area, paths=sorted({f.rsplit("/", 1)[0] if "/" in f else f for f in files})) + + +def _git(root: Path, *args: str) -> subprocess.CompletedProcess: + return subprocess.run(["git", "-C", str(root), *args], capture_output=True, text=True) + + +def missing(root: Path, area: Area) -> list[str]: + out = [] + for tok in area.anchors: + r = _git(root, "grep", "-F", "-q", "-e", tok, *(["--", *area.paths] if area.paths else [])) + if r.returncode != 0: + out.append(tok) + return out + + +def commits_since(root: Path, area: Area) -> int | None: + if not area.checked or _git(root, "cat-file", "-e", f"{area.checked}^{{commit}}").returncode != 0: + return None + r = _git(root, "rev-list", "--count", f"{area.checked}..HEAD", *(["--", *area.paths] if area.paths else [])) + return int(r.stdout.strip()) if r.returncode == 0 and r.stdout.strip().isdigit() else None + + +def status(root: Path, area: Area, limit: int) -> tuple[bool, list[str], int | None]: + miss = missing(root, area) + commits = commits_since(root, area) + return (bool(miss) or (commits is not None and commits >= limit)), miss, commits + + +def mark(text: str, name: str, sha: str) -> str: + lines = text.split("\n") + in_areas = False + start = None + for i, line in enumerate(lines): + if line.startswith("## "): + if start is not None: + break + in_areas = line[3:].strip().lower() == "areas" + elif in_areas and line.startswith("### "): + if start is not None: + break + if line[4:].strip() == name: + start = i + if start is None: + return text + end = start + 1 + while end < len(lines) and not lines[end].startswith("#"): + end += 1 + block_end = end + while block_end > start + 1 and not lines[block_end - 1].strip(): + block_end -= 1 + new = f"- Checked: {sha}" + for i in range(start + 1, block_end): + if lines[i].startswith("- Checked:"): + lines[i] = new + return "\n".join(lines) + lines.insert(block_end, new) + return "\n".join(lines) diff --git a/wflib/check.py b/wflib/check.py new file mode 100644 index 0000000..f86b2fe --- /dev/null +++ b/wflib/check.py @@ -0,0 +1,378 @@ +"""Validation: everything `wf check` reports.""" +from __future__ import annotations + +import re +import subprocess +import time +from dataclasses import dataclass +from pathlib import Path + +from . import areas, refs, tasks +from .config import NAME, Config + +NUMBERED_RE = re.compile(r"^\d+\. ") +NUMBER_REF_RE = re.compile(r"(?<![\w/&#])#\d+\b") +MD_LINK_RE = re.compile(r"\[[^\]]*\]\(([^)\s]+)\)") +TASK_LINK_RE = re.compile(r"^[ta]-") +EFFORTS = ", ".join(tasks.EFFORTS) +KINDS = "a- items belong in Awaiting, t- elsewhere" +STALE_DAYS = 30 + + +@dataclass(frozen=True) +class Problem: + path: str + line: int | None + id: str + message: str + + @property + def key(self) -> tuple[str, str, str]: + return (self.path, self.id, self.message) + + def __str__(self) -> str: + where = self.path + (f":{self.line}" if self.line else "") + return f"{where}: " + (f"{self.id}: " if self.id else "") + self.message + + +def _read(path: Path) -> str: + return path.read_text(encoding="utf-8", errors="replace").replace("\r\n", "\n") + + +def _prose_lines(text: str): + """(1-based line number, line) outside code fences.""" + fenced = False + for n, line in enumerate(text.split("\n"), 1): + if refs.FENCE_RE.match(line): + fenced = not fenced + elif not fenced: + yield n, line + + +def _link_problems(path: str, text: str, known: set[str]) -> list[Problem]: + return [Problem(path, n, f"[[{id}]]", "no such id in TASKS or archive") + for n, line in _prose_lines(text) + for id in dict.fromkeys(refs.links(line)) + if TASK_LINK_RE.match(id) and id not in known] + + +def _cycle(start: str, edges: dict[str, list[str]]) -> list[str] | None: + def walk(node: str, path: list[str]) -> list[str] | None: + for nxt in edges.get(node, []): + if nxt == start: + return path + [nxt] + if nxt not in path and (found := walk(nxt, path + [nxt])): + return found + return None + return walk(start, [start]) + + +def _index_sections(cfg: Config) -> list[refs.DocSection] | None: + """Headings under the anchors index section; None when no [anchors].""" + if not cfg.anchors_index or not cfg.anchors_index.is_file(): + return None + found = refs.sections(_read(cfg.anchors_index)) + if not cfg.anchors_section: + return [s for s in found if s.level] + top = next((s for s in found if s.level and s.heading == cfg.anchors_section), None) + if top is None: + return [] + return [s for s in found if s.level > top.level and top.start < s.start < top.end] + + +def check_tasks(text: str, archive_text: str, cfg: Config) -> tuple[list[Problem], list[Problem]]: + path = cfg.rel(cfg.tasks) + errors: list[Problem] = [] + warnings: list[Problem] = [] + doc = tasks.parse(text) + archived = tasks.archive_ids(archive_text) + + def err(line, id, message): + errors.append(Problem(path, line, id, message)) + + for key, heading in tasks.SECTIONS.items(): + if not any(s.key == key for s in doc.sections): + err(None, "", f"no '## {heading}' section") + + errors += _link_problems(path, text, doc.ids() | archived) + + index = _index_sections(cfg) + index_anchors = {a for s in index for a in s.anchors} if index is not None else None + open_awaiting = {i.id for s in doc.sections if s.key == "awaiting" for i in s.items} + deferred = {i.id for s in doc.sections if s.key == "deferred" for i in s.items} + edges = {i.id: [a for a in i.after if a in doc.ids()] for i in doc.all_items()} + seen: dict[str, int] = {} + in_cycle: set[str] = set() + + for section in doc.sections: + if section.key is None: + continue + for n, line in enumerate(section.prefix, section.line + 1): + if NUMBERED_RE.match(line): + err(n, "", "old numbered item (wf migrate)") + for n, line in enumerate(section.suffix, section.suffix_line or 0): + if NUMBERED_RE.match(line): + err(n, "", "old numbered item (wf migrate)") + elif line.startswith(tasks.ITEM_START): + err(n, tasks.parse_header(line).id, + f"item after a flush-left prose line (line {section.suffix_line}): indent or move the prose") + for at, item in enumerate(section.items): + if item.error: + err(item.line, item.id, item.error) + continue + if not tasks.ID_RE.match(item.id): + err(item.line, item.id, "bad id (want t-… or a-…, lowercase a-z 0-9 -)") + continue + if item.id in seen: + err(item.line, item.id, f"duplicate id (also line {seen[item.id]})") + seen.setdefault(item.id, item.line) + if item.id in archived: + err(item.line, item.id, "id already used in the archive (ids are never reused)") + if item.id.startswith("a-") != (section.key == "awaiting"): + err(item.line, item.id, f"{item.id[:2]} item in {section.heading} ({KINDS})") + continue + if section.key == "awaiting": + continue + if item.prio is None: + err(item.line, item.id, "no priority [P0]-[P3]") + elif item.prio > 3: + err(item.line, item.id, f"priority P{item.prio} (want P0-P3)") + if item.effort is None: + err(item.line, item.id, f"no effort (want {EFFORTS})") + elif item.effort not in tasks.EFFORTS: + err(item.line, item.id, f"effort '{item.effort}' (want {EFFORTS})") + if item.status and item.status.startswith("blocked:") and item.blocked_on not in open_awaiting: + err(item.line, item.id, f"blocked on '{item.blocked_on}', which is not an open Awaiting item") + if item.status and re.fullmatch(r"in progress:\s*", item.status): + warnings.append(Problem(path, item.line, item.id, "in progress without a branch or note")) + if item.id not in in_cycle and (cycle := _cycle(item.id, edges)): + in_cycle.update(cycle) + err(item.line, item.id, "After: cycle " + " → ".join(cycle)) + later = {o.id for o in section.items[at + 1:]} + for dep in item.after: + if dep in later: + err(item.line, item.id, f"placed before '{dep}', which it is After:") + for label, pattern in (("After", tasks.AFTER_RE), ("Slices", tasks.SLICES_RE)): + if (i := item._line(pattern)) is None: + continue + rest = tasks.LINK_RE.sub(" ", pattern.match(item.body[i]).group(1)) + for word in re.split(r"[\s,;]+", rest): + if tasks.ID_RE.match(word): + err(item.line + 1 + i, item.id, f"{label}: '{word}' is not a link (want [[{word}]])") + if (i := item._line(tasks.AFTER_RE)) is not None: + rest = tasks.LINK_RE.sub(" ", tasks.AFTER_RE.match(item.body[i]).group(1)) + if re.search(r"[A-Za-z0-9]", rest): + warnings.append(Problem(path, item.line + 1 + i, item.id, + "After: line has prose; every [[id]] in it is a dependency")) + for dep in item.after: + if dep in deferred and section.key != "deferred": + warnings.append(Problem(path, item.line + 1 + i, item.id, + f"After: '{dep}' is deferred (never runs; blocks this task)")) + models = [(n, m.group(1).split()) for n, l in enumerate(item.body) if (m := tasks.MODEL_RE.match(l))] + for n, words in models[:1]: + if not words or words[0] not in tasks.MODELS: + err(item.line + 1 + n, item.id, + f"Model '{' '.join(words)}' (want {', '.join(tasks.MODELS)})") + elif len(words) > 1: + warnings.append(Problem(path, item.line + 1 + n, item.id, f"Model line '{' '.join(words)}': " + f"write 'Model: {words[0]}' (wf set --model)")) + if len(models) > 1: + err(item.line + 1 + models[1][0], item.id, "two Model lines") + sess = [(n, m.group(1).split()) for n, l in enumerate(item.body) if (m := tasks.SESSIONS_RE.match(l))] + for n, words in sess[:1]: + if not words or words[0] not in tasks.SESSIONS: + err(item.line + 1 + n, item.id, + f"Sessions '{' '.join(words)}' (want {', '.join(tasks.SESSIONS)})") + if len(sess) > 1: + err(item.line + 1 + sess[1][0], item.id, "two Sessions lines") + clouds = [(n, m.group(1).split()) for n, l in enumerate(item.body) if (m := tasks.CLOUD_RE.match(l))] + for n, words in clouds[:1]: + if not words or words[0] not in tasks.CLOUDS: + err(item.line + 1 + n, item.id, f"Cloud '{' '.join(words)}' (want {', '.join(tasks.CLOUDS)})") + if len(clouds) > 1: + err(item.line + 1 + clouds[1][0], item.id, "two Cloud lines") + if item.interactive: + warnings.append(Problem(path, item.line, item.id, "'interactive' flag: write 'Sessions: owner' " + f"(wf set {item.id} --sessions owner)")) + if section.key == "pending" and item.human_done_match and item.sessions != "owner": + warnings.append(Problem(path, item.line, item.id, f"not runner-ready: Done reads as human action " + f"('{item.human_done_match}'); reword or set Sessions: owner")) + ref_at = item._line(tasks.REF_RE) + ref_line = item.line + 1 + ref_at if ref_at is not None else item.line + for target, anchor in item.refs: + shown = target + (f"#{anchor}" if anchor else "") + file = cfg.root / target + if not file.exists(): + err(ref_line, item.id, f"Ref '{shown}' does not exist") + elif anchor and file.is_file(): + if anchor not in refs.anchors(_read(file)): + err(ref_line, item.id, f"Ref '{shown}': no such anchor") + elif index_anchors is not None and file.resolve() == cfg.anchors_index.resolve() \ + and anchor not in index_anchors: + err(ref_line, item.id, f"Ref '{shown}' is not a heading under '{cfg.anchors_section}'") + + for n, line in _prose_lines(text): + for m in NUMBER_REF_RE.finditer(refs.strip_code(line)): + warnings.append(Problem(path, n, "", f"'{m.group(0)}': number ref (tasks have ids: [[t-…]])")) + return errors, warnings + + +def _anchor_problems(cfg: Config) -> list[Problem]: + index = _index_sections(cfg) + if index is None: + return [] + out: list[Problem] = [] + path = cfg.rel(cfg.anchors_index) + lines = _read(cfg.anchors_index).split("\n") + where = f"'{cfg.anchors_section}'" if cfg.anchors_section else "the index" + seen: dict[str, int] = {} + for s in index: + slug = refs.slugify(s.heading) + if slug in seen: + out.append(Problem(path, s.start + 1, "", f"two {where} headings give anchor '{slug}' (also line {seen[slug]})")) + seen.setdefault(slug, s.start + 1) + for n in range(s.start, s.end): + for target in MD_LINK_RE.findall(refs.strip_code(lines[n])): + if re.match(r"^[a-z][a-z0-9+.-]*:", target): + continue + file_part, _, frag = target.partition("#") + file = (cfg.anchors_index.parent / file_part) if file_part else cfg.anchors_index + if not file.exists(): + out.append(Problem(path, n + 1, "", f"link '{target}': file does not exist")) + elif frag and file.suffix == ".md" and frag not in refs.anchors(_read(file)): + out.append(Problem(path, n + 1, "", f"link '{target}': no such anchor")) + if cfg.anchors_specs and cfg.anchors_specs.is_dir(): + known = {a for s in index for a in s.anchors} + for spec in sorted(cfg.anchors_specs.rglob("*.md")): + for n, line in _prose_lines(_read(spec)): + for anchor in refs.ANCHOR_RE.findall(line): + if anchor not in known: + out.append(Problem(cfg.rel(spec), n, "", + f"explicit anchor '{anchor}' has no heading under {where} in {path}")) + return out + + +def doc_files(cfg: Config) -> list[Path]: + out: list[Path] = [] + for d in cfg.docs: + if d.is_file(): + out.append(d) + elif d.is_dir(): + out += sorted(p for p in d.rglob("*.md") if p.is_file()) + skip = {cfg.tasks.resolve(), cfg.archive.resolve()} + return [p for p in dict.fromkeys(out) if p.resolve() not in skip] + + +def _age_days(root: Path, path: Path, line: int) -> float | None: + try: + out = subprocess.run(["git", "-C", str(root), "blame", "-L", f"{line},{line}", "--porcelain", "--", str(path)], + capture_output=True, text=True, timeout=10) + except (OSError, subprocess.SubprocessError): + return None + m = re.search(r"^author-time (\d+)$", out.stdout, re.M) if out.returncode == 0 else None + if not m or re.match(r"^0{40}", out.stdout): + return None + return (time.time() - int(m.group(1))) / 86400 + + +def _stale_awaiting(cfg: Config, text: str) -> list[Problem]: + doc = tasks.parse(text) + waited = {i.blocked_on for i in doc.all_items()} | {l for i in doc.all_items() if i.id.startswith("t-") + for l in refs.links("\n".join(i.lines()))} + out = [] + for s in doc.sections: + if s.key != "awaiting": + continue + for item in s.items: + if item.id in waited or item.error: + continue + age = _age_days(cfg.root, cfg.tasks, item.line) + if age is not None and age > STALE_DAYS: + out.append(Problem(cfg.rel(cfg.tasks), item.line, item.id, + f"waiting {int(age)} days, no task references it")) + return out + + +def _git(root: Path, *args: str) -> str | None: + try: + out = subprocess.run(["git", "-C", str(root), *args], capture_output=True, text=True, timeout=10) + except (OSError, subprocess.SubprocessError): + return None + return out.stdout.strip() if out.returncode == 0 else None + + +def merged_tool_worktrees(tool_root: Path) -> list[Problem]: + """Worktrees of the wf tool repo whose branch is behind master (merged, leftover after release). + A branch equal to master, or a worktree created < 2h ago, is not flagged.""" + listing = _git(tool_root, "worktree", "list", "--porcelain") + master = _git(tool_root, "rev-parse", "master") + if not listing or not master: + return [] + out = [] + for block in listing.split("\n\n"): + path = branch = head = None + for l in block.splitlines(): + if l.startswith("worktree "): + path = l[9:] + elif l.startswith("HEAD "): + head = l[5:] + elif l.startswith("branch refs/heads/"): + branch = l[18:] + if not path or not branch or branch == "master" or head == master: + continue + try: # created < 2h ago: a running worker's fresh worktree, not a leftover + if time.time() - (Path(path) / ".git").stat().st_mtime < 7200: + continue + except OSError: + pass + if _git(tool_root, "merge-base", "--is-ancestor", branch, "master") is not None: + out.append(Problem("wf-tool", None, "", f"tool worktree '{path}' (branch {branch}) is merged into " + f"master: git worktree remove it, git branch -d {branch}")) + return out + + +def _area_problems(cfg: Config) -> list[Problem]: + file = cfg.areas_file + if not file.is_file(): + return [] + out = [] + for a in areas.parse(_read(file)): + for tok in areas.missing(cfg.code, a): + out.append(Problem(cfg.rel(file), a.line, "", f"area {a.name}: anchor {tok} not found")) + if a.checked and areas.commits_since(cfg.code, a) is None: + out.append(Problem(cfg.rel(file), a.line, "", f"area {a.name}: Checked {a.checked} unknown")) + return out + + +def check(cfg: Config, tasks_text: str | None = None, archive_text: str | None = None, + slow: bool = True) -> tuple[list[Problem], list[Problem]]: + """All errors and warnings of the project. `tasks_text` / `archive_text` + replace what is on disk (to judge a change before it is written); + `slow=False` leaves out the git-based warnings.""" + errors: list[Problem] = [] + for name, path in (("tasks", cfg.tasks), ("archive", cfg.archive)): + given = tasks_text if name == "tasks" else archive_text + if given is None and not path.is_file(): + errors.append(Problem(NAME, None, "", f"{name} '{cfg.rel(path)}' does not exist")) + for name, path in [("docs", d) for d in cfg.docs] + [("anchors.index", cfg.anchors_index), + ("anchors.specs", cfg.anchors_specs)]: + if path is not None and not path.exists(): + errors.append(Problem(NAME, None, "", f"{name} '{cfg.rel(path)}' does not exist")) + if tasks_text is None and not cfg.tasks.is_file(): + return errors, [] + text = (_read(cfg.tasks) if tasks_text is None else tasks_text).replace("\r\n", "\n") + archive = archive_text if archive_text is not None else (_read(cfg.archive) if cfg.archive.is_file() else "") + found, warnings = check_tasks(text, archive, cfg) + errors += found + known = tasks.parse(text).ids() | tasks.archive_ids(archive) + for file in doc_files(cfg): + errors += _link_problems(cfg.rel(file), _read(file), known) + errors += _anchor_problems(cfg) + if slow: + warnings += _stale_awaiting(cfg, text) + warnings += _area_problems(cfg) + order: dict[str, int] = {} + for p in errors + warnings: + order.setdefault(p.path, len(order)) + by_place = lambda p: (order[p.path], p.line or 0) + return sorted(errors, key=by_place), sorted(warnings, key=by_place) diff --git a/wflib/cloud.py b/wflib/cloud.py new file mode 100644 index 0000000..5878b87 --- /dev/null +++ b/wflib/cloud.py @@ -0,0 +1,442 @@ +"""Cloud-lane ledger: pure logic (spec docs/specs cloud-lane-design §4.4). IO: wf_cloud.py.""" +from __future__ import annotations + +import base64 +import binascii +import dataclasses +import datetime as dt +import gzip +import hashlib +import json +import re + +from . import tasks as TK, usage as U + +OVERHEAD = 1.15 # title generation, setup: what the session can't see +LOST_AFTER = dt.timedelta(hours=24) +ENDED = ("done", "awaiting", "handback", "lost") + + +class CloudError(Exception): + pass + + +def new() -> dict: + return {"budget": 240.0, "spent": 0.0, "reserve_per_task": 4.0, "max_parallel": 3, "entries": []} + + +def loads(text: str) -> dict: + if not text.strip(): + return new() + try: + led = json.loads(text) + led["budget"], led["spent"], led["entries"] + except (ValueError, KeyError, TypeError) as e: + raise CloudError(f"cloud ledger unreadable: {e}") + return {**new(), **led} + + +def dumps(led: dict) -> str: + return json.dumps(led, indent=1) + "\n" + + +def running(led: dict) -> int: + return sum(e["state"] == "running" for e in led["entries"]) + + +def balance(led: dict) -> float: + return led["budget"] - led["spent"] - led["reserve_per_task"] * running(led) + + +def refusal(led: dict) -> str | None: + """One line why `send` must refuse (exit 3), or None.""" + bal, res = balance(led), led["reserve_per_task"] + if bal < res: + return f"cloud busy: balance ${bal:.2f} < reserve ${res:.2f}" + if running(led) >= led["max_parallel"]: + return f"cloud busy: {running(led)} running >= max_parallel {led['max_parallel']}" + return None + + +def add(led: dict, id: str, project: str, sid: str, model: str, now: dt.datetime) -> dict: + e = {"id": id, "project": project, "sid": sid, "model": model, "sent": now.isoformat(), + "state": "running", "usd": None, "usd_source": None} + led["entries"].append(e) + return e + + +def find(led: dict, sid: str) -> dict: + for e in led["entries"]: + if e["sid"] == sid: + return e + raise CloudError(f"no cloud entry with sid {sid}") + + +def charge_usd(model: str, u: U.Usage | None, reserve: float) -> tuple[float, str]: + """(usd, source): WF-USAGE x PRICES + 15%; missing/unknown model -> reserve, 'est'.""" + c = U.cost(model, u) if u is not None else None + return (c * OVERHEAD, "self") if c is not None else (reserve, "est") + + +def end(led: dict, sid: str, state: str, u: U.Usage | None = None) -> dict: + """Running entry ends: state, charge added to spent.""" + if state not in ENDED: + raise CloudError(f"bad end state {state}") + e = find(led, sid) + if e["state"] != "running": + raise CloudError(f"entry {sid} is {e['state']}, not running") + if state == "lost": + usd, src = led["reserve_per_task"], "est" + else: + usd, src = charge_usd(e["model"], u, led["reserve_per_task"]) + e.update(state=state, usd=usd, usd_source=src) + led["spent"] += usd + return e + + +def expire(led: dict, now: dt.datetime) -> list[dict]: + """Running > 24h (no WF-RESULT seen) -> lost, charged reserve. Returns the lost entries.""" + out = [] + for e in led["entries"]: + if e["state"] == "running" and now - dt.datetime.fromisoformat(e["sent"]) > LOST_AFTER: + out.append(end(led, e["sid"], "lost")) + return out + + +def set_balance(led: dict, bal: float, now: dt.datetime) -> dict: + """Owner reconcile (balance read from claude.ai): spent = budget - balance.""" + old = led["spent"] + led["spent"] = led["budget"] - bal + row = {"id": "-", "project": "-", "sid": "-", "model": "-", "sent": now.isoformat(), "state": "done", + "usd": led["spent"] - old, "usd_source": "owner"} + led["entries"].append(row) + return row + + +def summary(led: dict) -> str: + return (f"budget ${led['budget']:.2f} spent ${led['spent']:.2f} balance ${balance(led):.2f} " + f"running {running(led)}/{led['max_parallel']} (reserve ${led['reserve_per_task']:.2f}/task)") + + + +# --- archive ended sessions (spec §4.3 Archive): the CLI's internal archiveRemoteSession, CLI 2.1.291 --- +API = "https://api.anthropic.com" +_ARCH_SID = re.compile(r"session_[A-Za-z0-9]+") + + +def archive_request(sid: str, creds: str, version: str) -> tuple[str, dict]: + """(url, headers) for POST .../archive; creds = the CLI's .credentials.json text. Never logged.""" + if not _ARCH_SID.fullmatch(sid): + raise CloudError(f"{sid!r} is not a cloud session id") + try: + d = json.loads(creds) + tok = d["claudeAiOauth"]["accessToken"] + except (ValueError, KeyError, TypeError): + tok = None + if not tok: + raise CloudError("no claude.ai login token (run claude once)") + h = {"Authorization": f"Bearer {tok}", "Content-Type": "application/json", + "anthropic-version": "2023-06-01", "User-Agent": f"claude-code/{version}"} + if d.get("trustedDeviceToken"): + h["X-Trusted-Device-Token"] = d["trustedDeviceToken"] + return f"{API}/v1/code/sessions/{sid}/archive", h + + +def archive_problem(status: int, body: str) -> str | None: + """None = archived (200, or 409 = already); else one short why.""" + if status in (200, 409): + return None + if status == 401: + return "HTTP 401 (login expired? run claude once)" + return f"HTTP {status}: {' '.join(body.split())}"[:120] + + +def unarchived(led: dict) -> list[str]: + """sids of ended sessions not archived yet (owner reconcile rows and pending sids skipped).""" + return [e["sid"] for e in led["entries"] + if e["state"] in ENDED and _ARCH_SID.fullmatch(e["sid"]) and not e.get("archived")] + +# --- pty driver helpers (spec §3 F1-F3, F7, F8) --------------------------------------------------- +import re # noqa: E402 + +_ANSI = re.compile(r"\x1b(\[[0-?]*[ -/]*[@-~]|\][^\x07\x1b]*(\x07|\x1b\\)|[PX^_][^\x1b]*\x1b\\|[@-Z\\-_])") +_SID = re.compile(r"(session_[A-Za-z0-9]{8,})") + + +def strip_ansi(text: str) -> str: + """Terminal output -> plain text: cursor-forward (CSI n C) -> n spaces, cursor-to-column (CSI n G) + -> one space (the TUI draws word gaps that way), other escape sequences dropped, CR -> LF.""" + text = re.sub(r"\x1b\[(\d*)C", lambda m: " " * int(m.group(1) or 1), text) + text = re.sub(r"\x1b\[\d*G", " ", text) + return _ANSI.sub("", text).replace("\r\n", "\n").replace("\r", "\n") + + +def parse_sid(text: str) -> str | None: + """Session id from `claude --cloud` output: 'Resume with: claude --teleport <sid>', else the View URL.""" + plain = strip_ansi(text) + for pat in (r"Resume with:\s*claude --teleport\s+" + _SID.pattern, r"View:\s*\S*?/code/" + _SID.pattern): + m = re.findall(pat, plain) + if m: + return m[-1] + return None + + +def last_lines(text: str, n: int = 5) -> list[str]: + return [ln.strip() for ln in strip_ansi(text).splitlines() if ln.strip()][-n:] + + +def trust(claude_json: str, path: str) -> str: + """~/.claude.json text with projects[path].hasTrustDialogAccepted = true (F3). Other keys kept.""" + try: + data = json.loads(claude_json) if claude_json.strip() else {} + projects = data.setdefault("projects", {}) + if not isinstance(projects, dict): + raise TypeError("projects is not an object") + except (ValueError, TypeError, AttributeError) as e: + raise CloudError(f"~/.claude.json unreadable: {e}") + projects.setdefault(path, {})["hasTrustDialogAccepted"] = True + return json.dumps(data, indent=2) + "\n" + + +TRUST_PROMPT = re.compile(r"trust\s*(the\s*files\s*in\s*)?this\s*folder|Do\s*you\s*trust", re.I) +RESUMED = re.compile(r"Session\s*resumed", re.I) # teleport reached the prompt (F7) + + +# --- send: snapshot + prompt (spec §4.1, §4.2) ----------------------------------------------------- + +MODEL = "claude-opus-5-5" # what cloud sessions run (F4): ledger rows price with it +CAP = 90_000_000 # bytes: packed snapshot limit (upload ≤ 100 MB, F1) +NEVER = (".wf", ".worktrees", "out") # + TASKS / archive: never in a snapshot (cloud_include re-adds) + + +def include_problem(path: str) -> str | None: + """Why a cloud_include entry is unusable (absolute, leaves the project, empty), None = fine.""" + parts = path.replace("\\", "/").strip("/").split("/") + if not path.strip() or path.startswith("/") or ".." in parts or parts in (["."], [".git"]) or parts[0] == ".git": + return f"cloud_include '{path}': must be a path inside the project" + return None + + +def size_refusal(size: int, cap: int = CAP) -> str | None: + if size > cap: + return f"snapshot {size / 1e6:.0f} MB > {cap / 1e6:.0f} MB; trim cloud_include" + return None + + +def fill(template: str, id: str, base: str, task: str, recipe: str, note: str | None) -> str: + """templates/cloud-prompt.md with its {{…}} fields filled.""" + values = {"id": id, "base": base, "task": task.strip(), "recipe": recipe.strip() or "(none: see CLAUDE.md)", + "note": f"\n{note.strip()}\n" if note and note.strip() else ""} + unknown = sorted(set(re.findall(r"\{\{(\w+)\}\}", template)) - set(values)) + if unknown: + raise CloudError(f"cloud prompt template: unknown field {{{{{unknown[0]}}}}}") + out = re.sub(r"\{\{(\w+)\}\}", lambda m: values[m.group(1)], template) # one pass: task text kept as is + return out + + +# The export renders the prompt as '❯ first line' + ' ' continuations and every assistant message as +# '● first line' + ' ' continuations (F12): only an assistant message can start with '● WF-RESULT'. +RESULT = re.compile(r"^● WF-RESULT\b") + + +def final_message(export: str) -> list[str] | None: + """Lines of the last assistant message that starts with WF-RESULT (bullet / indent and trailing blanks + removed), None = no result yet. The prompt never matches, whatever it contains.""" + lines = export.replace("\r\n", "\n").split("\n") + starts = [i for i, ln in enumerate(lines) if RESULT.match(ln)] + if not starts: + return None + out = [] + for ln in lines[starts[-1]:]: + if out and ln[:2] not in (" ", ""): + break # next message / prompt + out.append(ln[2:].rstrip()) + while out and not out[-1]: + out.pop() + return out + + +def keyless_message(export: str) -> list[str] | None: + """No WF-RESULT anywhere, but the last assistant message ends in a bare WF-PATCH-END line (the session + dropped the key words): its lines (as final_message), else None. A non-empty prompt after it (a redo + sent) -> None: still running.""" + lines = export.replace("\r\n", "\n").split("\n") + blocks: list[tuple[str, list[str]]] = [] + for ln in lines: + if ln[:2] in (" ", "") and blocks: + blocks[-1][1].append(ln[2:].rstrip()) + elif ln.strip(): + blocks.append((ln[:1], [ln[2:].rstrip()])) + found = None + for kind, body in blocks: + if kind == "❯" and any(b.strip() for b in body): + found = None # a later prompt: the message before it is answered + elif kind == "●": + text = [b for b in body if b.strip()] + if text and text[-1].strip() == "WF-PATCH-END": + found = body + if found is None: + return None + out = list(found) + while out and not out[-1]: + out.pop() + return out + + +def keyless_result(lines: list[str]) -> Result: + """keyless_message() lines -> a handback Result: usage + patch (sha256=HEX bytes=N header anywhere before + WF-PATCH-END, base64 after it) when found, so the patch can be kept.""" + text = " ".join(ln.strip() for ln in lines[:-1]) + r = Result("handback", "(no WF-RESULT key)") + if m := _USAGE.search(text): + r.usage = U.Usage(turns=1, inp=int(m[1]), cw5=int(m[2]), cr=int(m[3]), out=int(m[4])) + r.model = m[5] + blob = "".join(text.split()) + if m := _HEADER.search(blob): + r.sha, r.nbytes, r.b64, r.has_patch = m[1].lower(), int(m[2]), blob[m.end():], True + return r + + +# --- pull: result + patch (spec §4.3) --------------------------------------------------------------- + +KEYS = ("WF-RESULT", "WF-REPORT", "WF-USAGE", "WF-PATCH-BEGIN", "WF-PATCH-END") +STATES = ("done", "awaiting", "handback") +_USAGE = re.compile(r"in=(\d+)\s+cw=(\d+)\s+cr=(\d+)\s+out=(\d+)(?:\s+model=(\S+))?") +# whitespace removed: the export may wrap the header anywhere; gzip base64 starts 'H4sI', never a digit +_HEADER = re.compile(r"sha256=([0-9a-fA-F]{64})bytes=(\d+)") + + +@dataclasses.dataclass +class Result: + state: str + report: str + usage: U.Usage | None = None + model: str | None = None + sha: str | None = None + nbytes: int | None = None + b64: str = "" + has_patch: bool = False + + +def parse_result(lines: list[str]) -> Result: + """final_message() lines -> Result. A line starting with a key opens its section, other lines continue + the open one (the export wraps long lines). CloudError: bad state / missing patch markers.""" + sec: dict[str, list[str]] = {} + cur = None + for ln in lines: + s = ln.strip() + key = next((k for k in KEYS if s == k or s.startswith(k + " ")), None) + if key: + cur = key + sec.setdefault(key, []).append(s[len(key):].strip()) + elif cur: + sec[cur].append(s) + state = " ".join(sec.get("WF-RESULT", [""])).strip().lower() + if state not in STATES: + raise CloudError(f"bad WF-RESULT '{state}' (want {'/'.join(STATES)})") + r = Result(state, " ".join(" ".join(sec.get("WF-REPORT", [])).split()) or "(no report)") + if m := _USAGE.search(" ".join(sec.get("WF-USAGE", []))): + r.usage = U.Usage(turns=1, inp=int(m[1]), cw5=int(m[2]), cr=int(m[3]), out=int(m[4])) + r.model = m[5] + if "WF-PATCH-BEGIN" in sec: + if "WF-PATCH-END" not in sec: + raise CloudError("WF-PATCH-BEGIN without WF-PATCH-END (truncated message)") + blob = "".join("".join(sec["WF-PATCH-BEGIN"]).split()) + m = _HEADER.match(blob) + if not m: + raise CloudError("WF-PATCH-BEGIN needs sha256=HEX bytes=N") + r.sha, r.nbytes, r.b64, r.has_patch = m[1].lower(), int(m[2]), blob[m.end():], True + return r + + +def decode_patch(r: Result) -> bytes: + """base64 -> check bytes + sha256 -> gunzip. CloudError on any mismatch.""" + if not r.has_patch: + raise CloudError("no WF-PATCH-BEGIN/END in the result") + try: + gz = base64.b64decode(r.b64, validate=True) + except (binascii.Error, ValueError) as e: + raise CloudError(f"patch base64 broken: {e}") + if len(gz) != r.nbytes: + raise CloudError(f"patch bytes {len(gz)} != {r.nbytes} announced") + if hashlib.sha256(gz).hexdigest() != r.sha: + raise CloudError("patch sha256 mismatch") + try: + return gzip.decompress(gz) + except (OSError, EOFError) as e: + raise CloudError(f"patch gzip broken: {e}") + + +_DIFF = re.compile(r'^diff --git "?a/(.+?)"? "?b/(.+?)"?$') +_MOVE = re.compile(r"^(?:rename|copy) (?:from|to) (.+)$") + + +def patch_files(patch: str) -> list[str]: + """Repo-relative paths a format-patch touches (both sides of renames/copies), in order, unique.""" + out = [] + for ln in patch.split("\n"): + if m := _DIFF.match(ln): + out += [m[1], m[2]] + elif m := _MOVE.match(ln): + out.append(m[1].strip('"')) + return list(dict.fromkeys(out)) + + +def patch_problem(files: list[str], project: str, never: list[str]) -> str | None: + """Why a patch may not be applied, None = fine. files: repo-relative; project: the project folder + relative to the repo ('' = repo root); never: project-relative paths the cloud may not touch + (TASKS, archive, .wf, .worktrees, out, cloud_include).""" + pre = project.strip("/") + for f in files: + parts = f.split("/") + if not f or f.startswith("/") or ".." in parts or parts[0] == ".git": + return f"patch touches {f}: outside the tree" + if pre and not (f + "/").startswith(pre + "/"): + return f"patch touches {f}: outside the project folder {pre}" + rel = f[len(pre) + 1:] if pre else f + for n in never: + n = n.strip("/") + if rel == n or rel.startswith(n + "/"): + return f"patch touches {f}: {n} is never sent (wf bookkeeping / out / cloud_include)" + return None + + +# ------------------------------------------------------------------ fit (spec §4.5) + +# (reason, regex) over title + body; `Cloud: yes` skips them (Model opus still required). wf res before the generic wf-step rule. +UNFIT = ( + ("names `wf res` (local resource ledger)", re.compile(r"\bwf res\b")), + ("names `wf` commands as steps", re.compile(r"\bwf (?!res\b)[a-z][a-z-]*")), + ("needs a GUI", re.compile(r"\bGUI\b|\bscreenshot", re.I)), + ("needs a LAN host", re.compile(r"\bLAN\b|\b192\.168\.\d+\.\d+|\b10\.\d+\.\d+\.\d+|\b[\w-]+\.local\b|\bssh [\w@.-]+")), + ("needs a live server", re.compile(r"live server|running server|\blocalhost\b|127\.0\.0\.1", re.I)), +) + + +def fit(item: TK.Item, cfg) -> str | None: + """None when the task fits the cloud lane, else the reason it does not (spec §4.5).""" + if not cfg.cloud: + return "project not opted in (cloud = true)" + if item.cloud == "no": + return "Cloud: no" + if not item.runner_ready: + return "not runner-ready (needs Done, not owner-bound)" + if item.sessions in ("owner", "solo"): + return f"Sessions: {item.sessions}" + if item.effort not in TK.EFFORTS or TK.EFFORTS.index(item.effort) > TK.EFFORTS.index(cfg.slice_above): + return "slice job (effort above slice_above)" + if item.model != "opus": + return f"Model {item.model} stays local" + if item.cloud == "yes": + return None + text = "\n".join([item.text, *item.body]) + for reason, rx in UNFIT: + if rx.search(text): + return reason + return None + + +def pick_key(item: TK.Item) -> tuple: + """wf orch pick cloud order (spec §4.5): Cloud: yes first, then prio, then larger effort (all fits are opus).""" + eff = TK.EFFORTS.index(item.effort) if item.effort in TK.EFFORTS else -1 + return (item.cloud != "yes", item.prio if item.prio is not None else 9, -eff) diff --git a/wflib/config.py b/wflib/config.py new file mode 100644 index 0000000..8f75c89 --- /dev/null +++ b/wflib/config.py @@ -0,0 +1,225 @@ +"""Find the project (nearest ancestor with workflow.toml) and load its config.""" +from __future__ import annotations + +import tomllib +from dataclasses import dataclass, field +from pathlib import Path + +from . import lanes as lanes_mod + +NAME = "workflow.toml" +FORMAT = 1 +KEYS = {"format", "tasks", "archive", "docs", "verify", "done", "ledgers", "anchors", "ctx_hint", "worktree_setup", "quick_gate", + "slice_above", "lanes", "areas", "code_root", "area_stale_commits", "area_ignore", + "cloud", "cloud_include", "cloud_note"} +AREA_IGNORE = ("tests", "test", "docs", "doc") +ANCHOR_KEYS = {"index", "index_section", "specs"} + + +class ConfigError(Exception): + pass + + +@dataclass +class Config: + root: Path + format: int + tasks: Path + archive: Path + docs: list[Path] = field(default_factory=list) + verify: list[str] = field(default_factory=list) + done: list[str] = field(default_factory=list) + ledgers: Path | None = None + anchors_index: Path | None = None + anchors_section: str | None = None + anchors_specs: Path | None = None + ctx_hint: int = 100_000 # tokens; wf done/next suggest /clear above it (0 = off) + worktree_setup: list[str] = field(default_factory=list) # `wf setup`: cwd = worktree, env WF_MAIN + quick_gate: list[str] = field(default_factory=list) # `wf gate`: worker runs it before `wf done` + slice_above: str = "1h" + lanes: tuple = lanes_mod.DEFAULT_LANES + areas: Path | None = None # area notes file; None = CLAUDE.md + code_root: Path | None = None # repo area anchors / Checked / staleness refer to; None = root + area_stale_commits: int = 20 + area_ignore: list[str] = field(default_factory=lambda: list(AREA_IGNORE)) # never an uncovered area + local: Path | None = None # lane worktree's copy of root whose LOCAL_KEYS were used + cloud: bool = False # cloud lane opt-in (wf cloud send) + cloud_include: list[str] = field(default_factory=list) # git-ignored paths force-added to the snapshot + cloud_note: str | None = None # appended to the cloud prompt + + @property + def areas_file(self) -> Path: + return self.areas or (self.local or self.root) / "CLAUDE.md" + + @property + def code(self) -> Path: + return self.code_root or self.root + + def rel(self, path: Path) -> str: + try: + return str(path.relative_to(self.root)) + except ValueError: + return str(path) + + @property + def areas_rel(self) -> str: + """areas_file relative to its project folder (the lane worktree's copy or root).""" + try: + return str(self.areas_file.relative_to(self.local or self.root)) + except ValueError: + return str(self.areas_file) + + +def git_top(folder: Path) -> Path | None: + """The work tree top holding folder (.git dir or file), or None.""" + for f in (folder, *folder.parents): + if (f / ".git").exists(): + return f + return None + + +def git_branch(gitdir: Path) -> str | None: + try: + head = (gitdir / "HEAD").read_text().strip() + except OSError: + return None + return head.removeprefix("ref: refs/heads/") if head.startswith("ref: refs/heads/") else None + + +def linked_worktree(folder: Path) -> tuple[Path, Path, Path] | None: + """(worktree top, its gitdir, main tree top) when folder is in a linked git worktree.""" + top = git_top(folder) + if not top or not (top / ".git").is_file(): + return None + try: + line = (top / ".git").read_text().strip() + gitdir = (top / line.removeprefix("gitdir:").strip()).resolve() + common = (gitdir / (gitdir / "commondir").read_text().strip()).resolve() + except OSError: + return None + if common.name != ".git" or not line.startswith("gitdir:"): + return None + return top, gitdir, common.parent + + +def find_root(start: Path) -> Path: + """Folder with workflow.toml at or above start; in a linked git worktree the main tree's + same folder, when it is a project too (one TASKS.md for all worktrees).""" + start = start.resolve() + for folder in (start, *start.parents): + if (folder / NAME).is_file(): + wt = linked_worktree(folder) + if wt: + main = wt[2] / folder.relative_to(wt[0]) + if (main / NAME).is_file(): + return main + return folder + raise ConfigError(f"no {NAME} in {start} or above: not a wf project (wf init makes one)") + + +def find_local(start: Path) -> Path | None: + """In a linked worktree whose project root is the main tree's (find_root): the worktree's own copy + of that folder when it has a workflow.toml, else None.""" + start = start.resolve() + wt = linked_worktree(start) + if not wt: + return None + root = find_root(start) + try: + local = wt[0] / root.relative_to(wt[2]) + except ValueError: + return None + return local if local != root and (local / NAME).is_file() else None + + +def load_at(start: Path) -> "Config": + """Config for a command run in start: the main tree's books, branch-testable keys (LOCAL_KEYS: + worktree_setup, quick_gate, areas file, area settings) from a lane worktree's own workflow.toml.""" + return load(find_root(start), find_local(start)) + + +def _strings(data: dict, key: str) -> list[str]: + value = data.get(key, []) + if not isinstance(value, list) or not all(isinstance(v, str) for v in value): + raise ConfigError(f"{NAME}: '{key}' must be a list of strings") + return value + + +def _string(data: dict, key: str, shown: str | None = None) -> str | None: + value = data.get(key) + if value is not None and not isinstance(value, str): + raise ConfigError(f"{NAME}: '{shown or key}' must be a string") + return value + + +LOCAL_KEYS = ("worktree_setup", "quick_gate", "areas", "area_stale_commits", "area_ignore") + + +def _read(path: Path) -> dict: + try: + return tomllib.loads(path.read_text(encoding="utf-8")) + except (tomllib.TOMLDecodeError, OSError) as e: + raise ConfigError(f"{path.name}: {e}") from None + + +def load(root: Path, local: Path | None = None) -> Config: + """root's workflow.toml; with local (a lane worktree's copy of root), LOCAL_KEYS come from local's + workflow.toml and the areas file resolves there (a branch can test and --mark them before merge).""" + data = _read(root / NAME) + if local: + mine = _read(local / NAME) + for key in LOCAL_KEYS: + data.pop(key, None) + if key in mine: + data[key] = mine[key] + for key in data: + if key not in KEYS: + raise ConfigError(f"{NAME}: unknown key '{key}'") + anchors = data.get("anchors", {}) + if not isinstance(anchors, dict): + raise ConfigError(f"{NAME}: 'anchors' must be a table") + for key in anchors: + if key not in ANCHOR_KEYS: + raise ConfigError(f"{NAME}: unknown key 'anchors.{key}'") + for key in ("format", "tasks", "archive"): + if key not in data: + raise ConfigError(f"{NAME}: missing '{key}'") + if not isinstance(data["format"], int) or isinstance(data["format"], bool): + raise ConfigError(f"{NAME}: 'format' must be a number") + ctx_hint = data.get("ctx_hint", 100_000) + if not isinstance(ctx_hint, int) or isinstance(ctx_hint, bool) or ctx_hint < 0: + raise ConfigError(f"{NAME}: 'ctx_hint' must be a number of tokens (0 = off)") + index = _string(anchors, "index", "anchors.index") + specs = _string(anchors, "specs", "anchors.specs") + ledgers = _string(data, "ledgers") + stale = data.get("area_stale_commits", 20) + if not isinstance(stale, int) or isinstance(stale, bool) or stale < 1: + raise ConfigError(f"{NAME}: 'area_stale_commits' must be a number ≥ 1") + slice_above = _string(data, "slice_above") or "1h" + try: + lane_list = lanes_mod.parse_lanes(data.get("lanes"), slice_above) + except ValueError as e: + raise ConfigError(f"{NAME}: {e}") from None + areas = _string(data, "areas") + code_root = _string(data, "code_root") + if not isinstance(data.get("cloud", False), bool): + raise ConfigError(f"{NAME}: 'cloud' must be true or false") + return Config( + root=root, format=data["format"], + tasks=root / _string(data, "tasks"), archive=root / _string(data, "archive"), + docs=[root / d for d in _strings(data, "docs")], + verify=_strings(data, "verify"), done=_strings(data, "done"), + ledgers=root / ledgers if ledgers else None, + anchors_index=root / index if index else None, + anchors_section=_string(anchors, "index_section", "anchors.index_section"), + anchors_specs=root / specs if specs else None, + ctx_hint=ctx_hint, + worktree_setup=_strings(data, "worktree_setup"), quick_gate=_strings(data, "quick_gate"), + slice_above=slice_above, lanes=lane_list, areas=(local or root) / areas if areas else None, + code_root=(root / code_root).resolve() if code_root else None, + area_stale_commits=stale, + area_ignore=_strings(data, "area_ignore") if "area_ignore" in data else list(AREA_IGNORE), + local=local, + cloud=data.get("cloud", False), cloud_include=_strings(data, "cloud_include"), + cloud_note=_string(data, "cloud_note"), + ) diff --git a/wflib/lanes.py b/wflib/lanes.py new file mode 100644 index 0000000..d9b6fe6 --- /dev/null +++ b/wflib/lanes.py @@ -0,0 +1,297 @@ +"""Lanes (by task size): config records, membership, pick order, output blocks. Sessions are dicts +{"lane", "model", "socket", "alive", …} keyed by lane name or "all" (registry I/O lives in wf.py).""" +import re +from dataclasses import dataclass + +from . import tasks + +LANE_KEYS = {"efforts", "order", "fallback", "slices"} +ORDERS = ("unblock", "priority") +LANE_NAME_RE = re.compile(r"^[a-z][a-z0-9-]*$") + + +@dataclass(frozen=True) +class Lane: + name: str + efforts: tuple[str, ...] + order: str = "priority" + fallback: str | None = None + slices: bool = False + + +DEFAULT_LANES = (Lane("fast", ("<1h",), "unblock", None, True), Lane("slow", ("1h",), "priority", "fast", False)) + + +def parse_lanes(table: dict | None, slice_above: str) -> tuple[Lane, ...]: + """[lanes.*] of workflow.toml → lanes in file order; None → DEFAULT_LANES. Raises ValueError.""" + if slice_above not in tasks.EFFORTS: + raise ValueError(f"slice_above '{slice_above}' (want {', '.join(tasks.EFFORTS)})") + allowed = tasks.EFFORTS[:tasks.EFFORTS.index(slice_above) + 1] + if table is None: + out = DEFAULT_LANES + else: + if not isinstance(table, dict): + raise ValueError("'lanes' must be a table of tables") + out = [] + for name, t in table.items(): + if not LANE_NAME_RE.match(name): + raise ValueError(f"bad lane name '{name}' (want a-z 0-9 -)") + if name == "all": + raise ValueError("lane name 'all' is reserved") + if not isinstance(t, dict): + raise ValueError(f"'lanes.{name}' must be a table") + for k in t: + if k not in LANE_KEYS: + raise ValueError(f"unknown key 'lanes.{name}.{k}'") + efforts = t.get("efforts") + if not isinstance(efforts, list) or not efforts or not all(isinstance(e, str) for e in efforts): + raise ValueError(f"lane {name}: 'efforts' must be a non-empty list of strings") + order = t.get("order", "priority") + if order not in ORDERS: + raise ValueError(f"lane {name}: order '{order}' (want {', '.join(ORDERS)})") + fallback, slices = t.get("fallback"), t.get("slices", False) + if fallback is not None and not isinstance(fallback, str) or not isinstance(slices, bool): + raise ValueError(f"lane {name}: 'fallback' must be a string, 'slices' true/false") + out.append(Lane(name, tuple(efforts), order, fallback, slices)) + names = [l.name for l in out] + owner: dict[str, str] = {} + for l in out: + for e in l.efforts: + if e not in tasks.EFFORTS: + raise ValueError(f"lane {l.name}: effort '{e}' (want {', '.join(tasks.EFFORTS)})") + if e not in allowed: + raise ValueError(f"lane {l.name}: effort '{e}' above slice_above '{slice_above}'") + if e in owner: + raise ValueError(f"effort '{e}' in lanes {owner[e]} and {l.name}") + owner[e] = l.name + if l.fallback is not None and (l.fallback not in names or l.fallback == l.name): + raise ValueError(f"lane {l.name}: fallback '{l.fallback}' is not another lane") + for e in allowed: + if e not in owner: + raise ValueError(f"effort '{e}' (≤ slice_above) in no lane") + by = {l.name: l for l in out} + for l in out: + if l.fallback and by[l.fallback].fallback == l.name: + a, b = sorted((l.name, l.fallback)) + raise ValueError(f"fallback cycle: {a} ↔ {b}") + flagged = [l for l in out if l.slices] + if len(flagged) > 1: + raise ValueError("slices = true on more than one lane") + if not flagged: + out = [Lane(l.name, l.efforts, l.order, l.fallback, owner["<1h"] == l.name) for l in out] + return tuple(out) + + + +def _pending(doc: tasks.Doc) -> list[tasks.Item]: + return [i for s in doc.sections if s.key == "pending" for i in s.items if not i.error] + + +def counts(doc: tasks.Doc, archived: set[str], lanes, slice_above: str, + held: dict[str, str] | None = None, runner: bool = False) -> dict[str, tuple[int, int, int]]: + """lane → (pickable, waiting, pickable slice jobs) over Pending, in lane order. held: claimed ids count waiting. + runner: pickable = what `list --runner` shows (has a usable Done); the rest is not_ready().""" + out = {l.name: [0, 0, 0] for l in lanes} + for item in _pending(doc): + c = out[lane_of(item, lanes, slice_above)] + if _why(doc, item, archived, slice_above, held, 0, True) is None: + if runner and not item.runner_ready: + continue + c[0] += 1 + c[2] += is_slice_job(item, slice_above) + else: + c[1] += 1 + return {name: tuple(c) for name, c in out.items()} + + +def not_ready(doc: tasks.Doc, archived: set[str], lanes, slice_above: str, + held: dict[str, str] | None = None) -> dict[str, int]: + """lane → pickable items without a Done that a prep worker can fill (owner-bound ones not counted).""" + out = {l.name: 0 for l in lanes} + for item in _pending(doc): + if _why(doc, item, archived, slice_above, held, 0, True) is None \ + and not (item.done_text and item.done_text.strip()) and item.sessions != "owner": + out[lane_of(item, lanes, slice_above)] += 1 + return out + + +def prep_targets(doc: tasks.Doc, archived: set[str], lanes, slice_above: str, + only: set[str] | None = None) -> list[tasks.Item]: + """Pending items a prep worker can make runner-ready: no Done, not owner-bound, no status, no open + slices; only: lane names. Order: pickable first, then priority, then file order.""" + out = [] + for n, item in enumerate(_pending(doc)): + if item.done_text and item.done_text.strip() or item.sessions == "owner" or item.status \ + or tasks.open_slices(doc, item.id): + continue + if only is not None and lane_of(item, lanes, slice_above) not in only: + continue + waits = _why(doc, item, archived, slice_above, None, 0, True) is not None + out.append(((waits, 9 if item.prio is None else item.prio, n), item)) + return [item for _, item in sorted(out, key=lambda x: x[0])] + + +def cross_waits(doc: tasks.Doc, archived: set[str], lane: str, lanes, + slice_above: str) -> list[tuple[tasks.Item, tasks.Item]]: + """(my waiting item, pickable item of another lane it waits on, transitively).""" + items = _pending(doc) + wait = tasks.waiters(doc) + name = {i.id: lane_of(i, lanes, slice_above) for i in items} + out = [] + for mine in items: + if name[mine.id] != lane or tasks.why_not(doc, mine, archived) is None: + continue + for other in items: + if name[other.id] != lane and mine.id in wait.get(other.id, ()) \ + and tasks.why_not(doc, other, archived) is None: + out.append((mine, other)) + return out + + +def lanes_block(counts: dict[str, tuple[int, int, int]], sessions: dict[str, dict], me: str | None, + nready: dict[str, int] | None = None) -> list[str]: + """One line per lane (always all: lanes are few). nready: lane → not runner-ready count.""" + out = [] + for name, (p, w, sl) in counts.items(): + line = (f"{name}{' (you)' if name == me else ''}: {p} pickable" + + (f" ({sl} slice{'s' if sl > 1 else ''})" if sl else "") + f" · {w} waiting") + if (nready or {}).get(name): + line += f" · {nready[name]} not runner-ready (no Done)" + if name != me: + s = sessions.get(name) + if s and s.get("alive"): + line += f" · session uds:{s['socket']} (alive, {s.get('model') or '?'})" + else: + line += " · no session" + (f" → orchestrator or owner: start one (wf next --lane {name})" if p else "") + out.append(line) + return out + + +def waiting_block(waits: list[tuple[tasks.Item, tasks.Item]], sessions: dict[str, dict], lanes, + slice_above: str) -> list[str]: + out = [] + for mine, other in waits: + name = lane_of(other, lanes, slice_above) + s = sessions.get(name) + head = f"- {mine.id} waits on {other.id} ({name} lane) → " + if s and s.get("alive"): + out.append(head + f'message uds:{s["socket"]}: "{other.id} blocks my {mine.id}, please take it"') + else: + out.append(head + f"no {name} session: tell the owner") + return out + + +def pickable_ids(doc: tasks.Doc, archived: set[str]) -> set[str]: + return {i.id for i in _pending(doc) if tasks.why_not(doc, i, archived) is None} + + +def notify_block(freed: list[tasks.Item], sessions: dict[str, dict], lanes, slice_above: str) -> list[str]: + """One line per lane with newly pickable tasks, in lane order.""" + out = [] + for lane in lanes: + ids = ", ".join(i.id for i in freed if lane_of(i, lanes, slice_above) == lane.name) + if not ids: + continue + s = sessions.get(lane.name) + if s and s.get("alive"): + out.append(f"notify {lane.name} uds:{s['socket']}: now pickable {ids}") + else: + out.append(f"{lane.name} work now pickable: {ids} (no session: tell the owner)") + return out + + +def solo_done_block(solo: list[str], sessions: dict[str, dict], me_pid: str | None) -> list[str]: + """After solo task(s) are done: one line per other live session, so it can pick again.""" + if not solo: + return [] + ids = ", ".join(solo) + return [f"notify {name} uds:{s['socket']}: solo {ids} done, run wf next" + for name, s in sessions.items() if s.get("alive") and str(s.get("pid")) != me_pid] + + +MODEL_RANK = {m: n for n, m in enumerate(tasks.MODELS)} + + +def over(effort: str | None, slice_above: str) -> bool: + return effort in tasks.EFFORTS and tasks.EFFORTS.index(effort) > tasks.EFFORTS.index(slice_above) + + +def is_slice_job(item: tasks.Item, slice_above: str) -> bool: + """Too big to implement and not sliced yet: the slices lane splits it.""" + return over(item.effort, slice_above) and not item.slices + + +def sliced_out(item: tasks.Item, slice_above: str) -> bool: + """Too big and already sliced once: never picked again (wf done it or add slices by hand).""" + return over(item.effort, slice_above) and bool(item.slices) + + +def lane_of(item: tasks.Item, lanes: tuple[Lane, ...], slice_above: str) -> str: + if over(item.effort, slice_above): + return next(l.name for l in lanes if l.slices) + effort = item.effort if item.effort in tasks.EFFORTS else "<1h" + return next(l.name for l in lanes if effort in l.efforts) + + +def _why(doc, item, archived, slice_above, held, others, owner) -> str | None: + why = tasks.why_not(doc, item, archived, held, others, owner) + if why is None and not item.error and sliced_out(item, slice_above): + return f"sliced: all slices done → wf done {item.id} or add slices" + return why + + +def rank(doc: tasks.Doc, items: list[tasks.Item], order: str, lane: dict[str, str]) -> dict[str, tuple]: + """Sort key per item. lane: item id → lane name (for 'waited on by another lane').""" + wait = tasks.waiters(doc) + by_id = {i.id: i for s in doc.sections if s.key in ("pending", "human") for i in s.items if not i.error} + out = {} + for n, item in enumerate(items): + ws = [by_id[w] for w in wait.get(item.id, ())] + prio = min([p for p in [item.prio, *(w.prio for w in ws)] if p is not None], default=9) + cross = any(lane.get(w.id) != lane.get(item.id) for w in ws) + out[item.id] = ((not ws, -len(ws), prio, n) if order == "unblock" + else (prio, not cross, -len(ws), n)) + return out + + +def _candidates(doc, lanes, slice_above, lane: str | None, model: str | None) -> list[tasks.Item]: + items = doc.section("pending").items + return [i for i in items if i.error or ( + (lane is None or lane_of(i, lanes, slice_above) == lane) + and (model is None or MODEL_RANK[i.model] <= MODEL_RANK[model]))] + + +def _ordered(doc, lanes, slice_above, lane: str | None, model: str | None) -> list[tasks.Item]: + items = _candidates(doc, lanes, slice_above, lane, model) + order = next((l.order for l in lanes if l.name == lane), "priority") + names = {i.id: lane_of(i, lanes, slice_above) for i in doc.section("pending").items if not i.error} + keys = rank(doc, items, order, names) + return sorted(items, key=lambda i: keys.get(i.id, (9, True, 0, items.index(i)))) + + +def pick(doc, archived, lanes, slice_above, lane=None, model=None, held=None, others=0, owner=False): + """Best pickable Pending item of the lane (None = all lanes, priority order), Model ≤ model; + the lane empty → its fallback lane (one hop). Plus unpickable items ranked before it, with the reason.""" + skipped = [] + hops = [lane] + fb = next((l.fallback for l in lanes if l.name == lane), None) + if fb: + hops.append(fb) + for name in hops: + for item in _ordered(doc, lanes, slice_above, name, model): + why = _why(doc, item, archived, slice_above, held, others, owner) + if why is None: + return item, skipped + if (item, why) not in skipped: + skipped.append((item, why)) + return None, skipped + + +def ranked(doc, archived, lanes, slice_above, lane: str | None) -> list[tasks.Item]: + """Every pickable item of the lane in pick order, then its fallback lane's (runner lists).""" + hops = [lane] + [l.fallback for l in lanes if l.name == lane and l.fallback] + out = [] + for name in hops: + out += [i for i in _ordered(doc, lanes, slice_above, name, None) + if not i.error and _why(doc, i, archived, slice_above, None, 0, False) is None and i not in out] + return out diff --git a/wflib/ledgers.py b/wflib/ledgers.py new file mode 100644 index 0000000..e22caa5 --- /dev/null +++ b/wflib/ledgers.py @@ -0,0 +1,35 @@ +"""In-flight plan ledgers (superpowers executing-plans / subagent-driven-development): +<ledgers>/<plan>/progress.md, first line `# SDD ledger — plan: <path>`.""" +from __future__ import annotations + +import re +from pathlib import Path + +TASK_RE = re.compile(r"^### Task (\S+?):", re.M) +DONE_RE = re.compile(r"^Task (\S+?): complete", re.M) + + +def summary(plan_text: str, ledger_text: str) -> str: + """'done 1, 2; resume at Task 3 (…)' / 'done 1, 2, 3; all tasks done; Final review: …'.""" + planned = TASK_RE.findall(plan_text) + done = DONE_RE.findall(ledger_text) + left = [t for t in planned if t not in done] + finals = [l for l in ledger_text.split("\n") if l.startswith("Final review")] + head = f"done {', '.join(done) if done else 'none'}" + if left: + return f"{head}; resume at Task {left[0]} (task-start PLAN {left[0]})" + return f"{head}; all tasks done; {finals[-1] if finals else 'final review not dispatched yet'}" + + +def in_flight(root: Path, ledgers: Path) -> list[str]: + out = [] + for ledger in sorted(ledgers.glob("*/progress.md")): + text = ledger.read_text(encoding="utf-8", errors="replace") + first = text.split("\n", 1)[0] + plan = first.split("plan:", 1)[1].strip() if "plan:" in first else "" + plan_path = root / plan + plan_text = plan_path.read_text(encoding="utf-8", errors="replace") if plan and plan_path.is_file() else "" + rulings = sum(1 for l in text.split("\n") if "Ruling:" in l) + out.append(f"- {plan or ledger.parent.name}: {summary(plan_text, text)} " + f"({rulings} rulings; ledger {ledger.relative_to(root)})") + return out diff --git a/wflib/migrate.py b/wflib/migrate.py new file mode 100644 index 0000000..04f606c --- /dev/null +++ b/wflib/migrate.py @@ -0,0 +1,237 @@ +"""One-time conversion of the numbered TASKS.md format to the id format. + +Old: `N. **[P1] Title** (Effort: 1h) (in progress: x) — goal.` + indented body, + `After: <title>`, `Reference: …`, Awaiting bullets without ids. +New: docs/design.md, section 3. Input already in the new format comes back unchanged. +""" +from __future__ import annotations + +import re +from dataclasses import dataclass, field + +from . import tasks + +NUMBERED_RE = re.compile(r"^\d+\. ") +OLD_HEADER_RE = re.compile(r"^\d+\. (?:N\. )?\*\*\[P(\d)\] (.+?)\*\*(.*)$") +SLICE_RE = re.compile(r"^(.+?) (\d+)/\d+(?::|$)") +MD_LINK_RE = re.compile(r"\[[^\]]*\]\(([^)\s]+)\)") +AFTER_RE = re.compile(r"^(\s*)(?:-\s+)?After:\s*(.*)$") +REFERENCE_RE = re.compile(r"^\s*(?:-\s+)?(?:Reference|Ref):\s*(.*)$") +BULLET_RE = re.compile(r"^- (?:\*\*(.+?)\*\*:?\s*)?(.*)$") + + +@dataclass +class Report: + ids: dict[str, str] = field(default_factory=dict) # old title -> new id + notes: list[str] = field(default_factory=list) + + +@dataclass +class _Old: + title: str + item: tasks.Item + after: list[str] + refs: str | None + + +def _groups(rest: str) -> tuple[list[str], str]: + """Leading balanced '(…)' groups of `rest`, and what follows them.""" + out, i = [], 0 + while True: + j = i + while j < len(rest) and rest[j] == " ": + j += 1 + if j >= len(rest) or rest[j] != "(": + return out, rest[i:] + depth, k = 0, j + while k < len(rest): + depth += rest[k] == "(" + depth -= rest[k] == ")" + k += 1 + if depth == 0: + break + if depth: + return out, rest[i:] + out.append(rest[j + 1:k - 1]) + i = k + + +def _id_title(title: str) -> str: + return re.sub(r"\s*\([^)]*\)", "", title).strip() or title + + +def _body(lines: list[str]) -> list[str]: + """Re-indented body. A line one column deeper than the top level is a + child when the line above it opens a list (ends with ':'), else a typo.""" + out: list[str] = [] + child = False + for line in tasks.indent_body(lines): + if re.match(r"^ \S", line): + line = (2 * tasks.INDENT if child else tasks.INDENT) + line[3:] + elif re.match(r"^ \S", line): + child = line.rstrip().endswith(":") + out.append(line) + return out + + +def _is_path(part: str) -> bool: + first = part.split(" ", 1)[0] + return "/" in first or bool(re.search(r"\.\w+(#|$)", first)) + + +def _refs(text: str) -> tuple[str | None, str | None]: + """(Ref: value, leftover words that name no file).""" + parts: list[str] = [] + for part in tasks.split_refs(MD_LINK_RE.sub(r"\1", text).replace(" · ", ", ").strip().rstrip(".")): + if _is_path(part): + parts.append(part) + elif parts: + last = parts[-1] + parts[-1] = f"{last[:-1]}; {part})" if last.endswith(")") else f"{last} ({part})" + else: + return None, text.strip() + return (", ".join(parts) or None), None + + +def _sentence(title: str, goal: str) -> str: + title = title.strip().rstrip(".").replace(". ", ", ") + return f"{title}. {goal.strip()}" if goal.strip() else f"{title}." + + +def _split_body(lines: list[str]) -> tuple[list[str], list[str], str | None]: + """(body re-indented, After: titles, Ref: value).""" + body, after, ref, notes = [], [], None, [] + for line in lines: + if (m := AFTER_RE.match(line)) and not tasks.LINK_RE.search(line): + after.append(m.group(2).strip()) + elif (m := REFERENCE_RE.match(line)): + ref, words = _refs(m.group(1)) + if words: + notes.append(f"{tasks.INDENT}- Reference: {words}") + else: + body.append(line) + return _body(body) + notes, after, ref + + +def _task(lines: list[str], taken: set[str]) -> _Old | None: + m = OLD_HEADER_RE.match(lines[0]) + if not m: + return None + title = m.group(2).strip() + groups, rest = _groups(m.group(3)) + effort, interactive, status, extra = None, False, None, [] + for g in groups: + if g.startswith("Effort:"): + parts = [p.strip() for p in g[len("Effort:"):].split(",")] + effort, interactive = parts[0], "interactive" in parts[1:] + elif g.startswith(("in progress:", "blocked:")): + status = g.replace("(", "[").replace(")", "]") + else: + extra.append(f"({g})") + goal = re.sub(r"^\s*[—–-]\s*", "", rest).strip() + goal = " ".join(extra + ([goal] if goal else [])) + slice_ = SLICE_RE.match(title) + if slice_: + base = tasks.make_id(_id_title(slice_.group(1)), set()) + id = f"{base}-{int(slice_.group(2))}" + n = 2 + while id in taken: + id, n = f"{base}-{int(slice_.group(2))}-{n}", n + 1 + else: + id = tasks.make_id(_id_title(title), taken) + body, after, ref = _split_body(lines[1:]) + item = tasks.Item(id=id, prio=int(m.group(1)), effort=effort, interactive=interactive, + status=status, text=_sentence(title, goal), body=body) + return _Old(title, item, after, ref) + + +def _awaiting(lines: list[str], taken: set[str]) -> _Old: + m = BULLET_RE.match(lines[0]) + bold, rest = m.group(1), m.group(2).strip() + if bold: + title, text = bold.strip(), _sentence(bold, rest) + else: + title = rest.split(". ", 1)[0].rstrip(".") + text = rest + id = tasks.make_id(_id_title(title), taken, "a-") + body, _, _ = _split_body(lines[1:]) + return _Old(title, tasks.Item(id=id, text=text, body=body), [], None) + + +def _is_new_item(line: str) -> bool: + m = re.match(r"^- \*\*([^*\s]+)\*\*", line) + return bool(m and tasks.ID_RE.match(m.group(1))) + + +def migrate(text: str, taken: set[str] | None = None) -> tuple[str, Report]: + taken = set(taken or ()) + report = Report() + newline = "\r\n" if "\r\n" in text else "\n" + lines = text.replace("\r\n", "\n").split("\n") + while lines and not lines[-1].strip(): + lines.pop() + taken |= {m.group(1) for l in lines if (m := re.match(r"^- \*\*([^*\s]+)\*\*", l)) and _is_new_item(l)} + + out: list[object] = [] # str lines and _Old items, in order + key = None + i = 0 + while i < len(lines): + line = lines[i] + if line.startswith("## "): + heading = line[3:].strip() + key = next((k for k, h in tasks.SECTIONS.items() if heading == h or heading.startswith(h + " (")), None) + if key and heading != tasks.SECTIONS[key]: + out += [f"## {tasks.SECTIONS[key]}", "", heading[len(tasks.SECTIONS[key]):].strip()] + else: + out.append(line) + i += 1 + continue + old_task = key in tasks.TASK_KEYS and NUMBERED_RE.match(line) + old_bullet = key == "awaiting" and line.startswith("- ") and not _is_new_item(line) + if not (old_task or old_bullet): + out.append(line) + i += 1 + continue + j = i + 1 + while j < len(lines) and (not lines[j].strip() or lines[j][0].isspace()): + j += 1 + block = lines[i:j] + while block and not block[-1].strip(): + block.pop() + old = _task(block, taken) if old_task else _awaiting(block, taken) + if old is None: + report.notes.append(f"line {i + 1}: numbered item not understood, left as it is: {line}") + out += lines[i:j] + else: + taken.add(old.item.id) + report.ids[old.title] = old.item.id + out += [old, ""] + i = j + + by_title = {t.lower().rstrip("."): id for t, id in report.ids.items()} + result: list[str] = [] + for entry in out: + if isinstance(entry, str): + result.append(entry) + continue + item, ids = entry.item, [] + for wanted in entry.after: + whole = by_title.get(wanted.lower().rstrip(".")) + parts = [whole] if whole else [by_title.get(p.strip().lower().rstrip(".")) for p in wanted.split("; ")] + if not any(parts): + report.notes.append(f"{item.id}: dropped 'After: {wanted}' (no open task with that title: done)") + ids += [p for p in parts if p] + if ids: + item.body.append(f"{tasks.INDENT}- After: " + ", ".join(f"[[{a}]]" for a in ids)) + if entry.refs: + item.body.append(f"{tasks.INDENT}Ref: {entry.refs}") + result += item.lines() + + doc = tasks.parse("\n".join(result) + "\n") + for key, heading in tasks.SECTIONS.items(): + if not any(s.key == key for s in doc.sections): + if doc.sections and doc.sections[-1].lines()[-1].strip(): + doc.sections[-1].suffix.append("") if doc.sections[-1].key else doc.sections[-1].prefix.append("") + doc.sections.append(tasks.Section(heading, key, prefix=[""])) + doc.newline = newline + return tasks.render(doc), report diff --git a/wflib/refs.py b/wflib/refs.py new file mode 100644 index 0000000..d1e853d --- /dev/null +++ b/wflib/refs.py @@ -0,0 +1,107 @@ +"""Links, anchors and doc sections: what `[[id]]` and `Ref: path#anchor` point at.""" +from __future__ import annotations + +import re +from dataclasses import dataclass +from pathlib import Path + +LINK_RE = re.compile(r"\[\[([^\]|#]+)(?:#[^\]|]+)?(?:\|[^\]]+)?\]\]") +HEADING_RE = re.compile(r"^(#{1,6})\s+(.*?)\s*$") +ANCHOR_RE = re.compile(r'<a\s+id="([^"]+)"\s*>\s*</a>', re.I) +FENCE_RE = re.compile(r"^\s*(```|~~~)") + + +def strip_code(text: str) -> str: + text = re.sub(r"```.*?```", "", text, flags=re.S) + return re.sub(r"`[^`\n]*`", "", text) + + +def links(text: str) -> list[str]: + return [m.group(1).strip() for m in LINK_RE.finditer(strip_code(text))] + + +def slugify(heading: str) -> str: + text = ANCHOR_RE.sub("", heading).strip().lower() + text = re.sub(r"[^\w\s-]", "", text) + return re.sub(r"[\s_]+", "-", text).strip("-").replace("--", "-") + + +@dataclass +class DocSection: + level: int # 0 = an anchor standing alone, not a heading + heading: str + anchors: set[str] + start: int # 0-based line of the heading (or of the lone anchor) + end: int # exclusive + + +def _lines(text: str) -> list[str]: + lines = text.replace("\r\n", "\n").split("\n") + while lines and not lines[-1].strip(): + lines.pop() + return lines + + +def _is_lone_anchor(line: str) -> bool: + return bool(ANCHOR_RE.search(line)) and not ANCHOR_RE.sub("", line).strip() + + +def sections(text: str) -> list[DocSection]: + lines = _lines(text) + out: list[DocSection] = [] + fenced = False + for n, line in enumerate(lines): + if FENCE_RE.match(line): + fenced = not fenced + if fenced or FENCE_RE.match(line): + continue + m = HEADING_RE.match(line) + tags = set(ANCHOR_RE.findall(line)) + if m: + heading = ANCHOR_RE.sub("", m.group(2)).strip() + above = set(ANCHOR_RE.findall(lines[n - 1])) if n and _is_lone_anchor(lines[n - 1]) else set() + out.append(DocSection(len(m.group(1)), heading, {slugify(heading)} | tags | above, n, len(lines))) + elif tags and not (_is_lone_anchor(line) and n + 1 < len(lines) and HEADING_RE.match(lines[n + 1])): + out.append(DocSection(0, "", tags, n, len(lines))) + headings = [s for s in out if s.level] + for s in out: + nxt = next((h for h in headings if h.start > s.start and (s.level == 0 or h.level <= s.level)), None) + if nxt: + above = nxt.start - 1 + s.end = above if above > s.start and _is_lone_anchor(lines[above]) else nxt.start + return out + + +def anchors(text: str) -> set[str]: + return {a for s in sections(text) for a in s.anchors} + + +def find(text: str, anchor: str) -> DocSection | None: + found = sections(text) + return next((s for s in found if anchor in s.anchors and s.level), None) or \ + next((s for s in found if anchor in s.anchors), None) + + +def resolve(root: Path, path: str, anchor: str | None, limit: int = 80) -> str: + target = root / path + shown = path.rstrip("/") + if target.is_dir(): + return f"{shown}/ (directory)" + if not target.is_file(): + return f"(missing: {shown})" + text = target.read_text(encoding="utf-8", errors="replace") + lines = _lines(text) + if anchor is None: + heading = next((m.group(2) for l in lines if (m := HEADING_RE.match(l))), "") + goal = next((l[len("**Goal:**"):].strip() for l in lines if l.startswith("**Goal:**")), "") + return shown + (f": {heading}" if heading else "") + (f" — {goal}" if goal else "") + s = find(text, anchor) + if s is None: + return f"(no anchor '{anchor}' in {shown})" + body = lines[s.start:s.end] + while body and not body[-1].strip(): + body.pop() + out = [f"===== {shown}#{anchor} (line {s.start + 1}) =====", *body[:limit]] + if len(body) > limit: + out.append(f"… ({len(body) - limit} more lines: {shown}:{s.start + limit + 1})") + return "\n".join(out) diff --git a/wflib/res.py b/wflib/res.py new file mode 100644 index 0000000..fd1fba9 --- /dev/null +++ b/wflib/res.py @@ -0,0 +1,993 @@ +"""wf res — pure part: parsing, ledger, liveness, capacity, queue, game, scratch, output lines. +Text and data in, data and lines out; no files, no processes. +Spec: docs/resource-ledger.md.""" +from __future__ import annotations + +import datetime as dt +import json +import math +import re +import shlex +import tomllib +from dataclasses import dataclass, field, fields +from typing import Callable + +GB = 1024 ** 3 +SLICE = "agents.slice" +JOBS_SLICE = "agents-jobs.slice" +DONE_KEEP = dt.timedelta(hours=24) +CLEAN_EVERY = dt.timedelta(minutes=10) +EPS = 1e-9 + + +class ResError(Exception): + """Expected failure: `wf: <msg>`, exit 1.""" + + +class Busy(Exception): + """Request does not fit now: the busy line, exit 3.""" + + +# ---------------------------------------------------------------- parsing + +SIZE_RE = re.compile(r"^(\d+(?:\.\d+)?)([MG])B?$", re.I) +DUR_RE = re.compile(r"^(?:(\d+)h)?(?:(\d+)m)?$") + + +def parse_size(text: str) -> float: + """'10G' → 10.0, '512M' → 0.5 (GiB).""" + m = SIZE_RE.match(text.strip()) + if not m: + raise ResError(f"size '{text}' (want e.g. 10G or 512M)") + n = float(m.group(1)) + return n if m.group(2).upper() == "G" else n / 1024 + + +def parse_duration(text: str) -> int: + """'40m' → 40, '4h' → 240, '1h30m' → 90, '90' → 90 (minutes, > 0).""" + t = text.strip() + if t.isdigit(): + n = int(t) + else: + m = DUR_RE.match(t) + if not t or not m or not any(m.groups()): + raise ResError(f"duration '{text}' (want e.g. 40m, 4h, 1h30m)") + n = int(m.group(1) or 0) * 60 + int(m.group(2) or 0) + if n <= 0: + raise ResError(f"duration '{text}' must be > 0") + return n + + +def meminfo(text: str) -> tuple[float, float]: + """(MemTotal, MemAvailable) in GB from /proc/meminfo.""" + vals = {} + for line in text.splitlines(): + key, _, rest = line.partition(":") + if key in ("MemTotal", "MemAvailable"): + vals[key] = int(rest.split()[0]) * 1024 / GB + if len(vals) != 2: + raise ResError("no MemTotal/MemAvailable in /proc/meminfo") + return vals["MemTotal"], vals["MemAvailable"] + + +def show_units(text: str) -> dict[str, dict[str, str]]: + """`systemctl show -p Id,…` output (blank-line separated blocks) → {Id: {prop: value}}.""" + out = {} + for block in text.strip().split("\n\n"): + props = dict(l.split("=", 1) for l in block.splitlines() if "=" in l) + if "Id" in props: + out[props["Id"]] = props + return out + + +def psi_some_avg60(text: str) -> float | None: + """cgroup memory.pressure → `some avg60` (percent).""" + for line in text.splitlines(): + if line.startswith("some "): + vals = dict(kv.split("=", 1) for kv in line.split()[1:] if "=" in kv) + try: + return float(vals["avg60"]) + except (KeyError, ValueError): + return None + return None + + +def cache_gb(stat: str) -> float: + """cgroup memory.stat → reclaimable file cache GB: `file` − `shmem` (tmpfs pages count as file but stay).""" + vals = dict(l.split(" ", 1) for l in stat.splitlines() if " " in l) + def get(k): + return int(vals[k]) if vals.get(k, "").strip().isdigit() else 0 + return max(0, get("file") - get("shmem")) / GB + + +def gb_or_none(text: str) -> float | None: + return int(text) / GB if text.strip().isdigit() else None + + +# ---------------------------------------------------------------- formatting + +def fmt_gb(x: float) -> str: + return f"{x:.1f} GB" + + +def fmt_bytes(n: int) -> str: + return fmt_gb(n / GB) if n >= GB else f"{n / 1024 ** 2:.0f} MB" + + +def fmt_dur(minutes: int) -> str: + h, m = divmod(int(minutes), 60) + return (f"{h}h" if h else "") + (f"{m}m" if m or not h else "") + + +def hhmm(t: dt.datetime) -> str: + return t.strftime("%H:%M") + + +def rc_text(rc: int | None) -> str: + return "?" if rc is None else str(rc) + + +# ---------------------------------------------------------------- config + +@dataclass +class Config: + user_reserve_gb: float = 6.0 + user_reserve_cpus: int = 4 + game_reserve_gb: float = 12.0 + game_reserve_cpus: int = 8 + game_hours: float = 4.0 + small_headroom_gb: float = 2.0 + scratch_hours: float = 2.0 + session_mem_gb: float = 6.0 # MemoryHigh per Claude session scope (throttle, no kill); 0 = off + + +def load_config(text: str) -> Config: + try: + data = tomllib.loads(text) + except tomllib.TOMLDecodeError as e: + raise ResError(f"resources.toml: {e}") from None + cfg = Config() + unknown = sorted(set(data) - {f.name for f in fields(Config)}) + if unknown: + raise ResError("resources.toml: unknown key(s) " + ", ".join(unknown)) + for key, value in data.items(): + if isinstance(value, bool) or not isinstance(value, (int, float)) or value < 0: + raise ResError(f"resources.toml: {key} must be a number ≥ 0") + setattr(cfg, key, type(getattr(cfg, key))(value)) + return cfg + + +# ---------------------------------------------------------------- ledger + +TOP_KEYS = ("next", "game_until", "last_clean", "entries") +TIME_FIELDS = ("queued", "started", "expires", "ended") + + +@dataclass +class Entry: + id: str + project: str + owner: int + title: str + mem_gb: float + cpus: int + est_min: int + state: str # queued | running | note | done + cwd: str = "" + cmd: list[str] = field(default_factory=list) + unit: str = "" + log: str = "" + queued: dt.datetime | None = None + started: dt.datetime | None = None + expires: dt.datetime | None = None + ended: dt.datetime | None = None + rc: int | None = None + peak_gb: float | None = None + why: str = "" + env: dict = field(default_factory=dict) # caller variables the user manager lacks (env_diff) + by: dict = field(default_factory=dict) # who started it (owner_by): name, task, batch, address + lock: str = "" # '<key>@<main tree>': one queued/running job per lock (lock_for) + extra: dict = field(default_factory=dict, repr=False, compare=False) # unknown keys of a newer version: kept on rewrite + + def eta(self) -> dt.datetime | None: + return self.started + dt.timedelta(minutes=self.est_min) if self.started else None + + +@dataclass +class Ledger: + next: int = 1 + game_until: dt.datetime | None = None + last_clean: dt.datetime | None = None + entries: list[Entry] = field(default_factory=list) + extra: dict = field(default_factory=dict, repr=False, compare=False) # unknown top-level keys: kept on rewrite + + def get(self, id: str) -> Entry: + for e in self.entries: + if e.id == id: + return e + raise ResError(f"no entry '{id}'") + + def new_id(self) -> str: + id = f"r-{self.next}" + self.next += 1 + return id + + +def _time(v): + return dt.datetime.fromisoformat(v) if v else None + + +def _iso(t): + return t.isoformat(timespec="seconds") if t else None + + +def entry_dict(e: Entry) -> dict: + d = {f.name: getattr(e, f.name) for f in fields(Entry) if f.name != "extra"} + d.update({k: v for k, v in e.extra.items() if k not in d}) + for k in TIME_FIELDS: + d[k] = _iso(d[k]) + return d + + +def loads(text: str) -> Ledger: + if not text.strip(): + return Ledger() + try: + d = json.loads(text) + entries, known = [], {f.name for f in fields(Entry)} - {"extra"} + for raw in d.get("entries", []): + extra = {k: v for k, v in raw.items() if k not in known} # newer version's keys: kept, not fatal + raw = {k: v for k, v in raw.items() if k in known} + raw["extra"] = extra + for k in TIME_FIELDS: + raw[k] = _time(raw.get(k)) + entries.append(Entry(**raw)) + return Ledger(next=int(d.get("next", 1)), game_until=_time(d.get("game_until")), + last_clean=_time(d.get("last_clean")), entries=entries, + extra={k: v for k, v in d.items() if k not in TOP_KEYS}) + except (ValueError, TypeError, AttributeError) as e: + raise ResError(f"corrupt ledger: {e}") from None + + +def salvage_next(text: str) -> int: + """Id counter of an unreadable ledger, so a fresh one never reuses ids (logs/<id>.rc, wf-<id>.service).""" + nums = [int(n) for n in re.findall(r'"next":\s*(\d+)', text)] + \ + [int(n) + 1 for n in re.findall(r'"r-(\d+)"', text)] + return max(nums, default=1) + + +def dumps(led: Ledger) -> str: + return json.dumps({**{k: v for k, v in led.extra.items() if k not in TOP_KEYS}, "next": led.next, "game_until": _iso(led.game_until), "last_clean": _iso(led.last_clean), + "entries": [entry_dict(e) for e in led.entries]}, indent=1) + "\n" + + +# ---------------------------------------------------------------- facts, liveness + +@dataclass +class Unit: + active: bool + current_gb: float | None = None + stall: float | None = None # memory.pressure `some avg60`: % of the last minute stalled on memory + + +@dataclass +class Facts: + """Snapshot of the machine, gathered by wf_res under the lock.""" + now: dt.datetime + total_gb: float + available_gb: float + nproc: int + units: dict[str, Unit] = field(default_factory=dict) # unit name → state, for running entries + live: set[int] = field(default_factory=set) # note owner pids still alive + results: dict[str, tuple] = field(default_factory=dict) # id → (rc, peak_gb, why) from logs/, journal + slice_gb: float = 0.0 # agents.slice MemoryCurrent + slice_cache_gb: float = 0.0 # of it reclaimable file cache (cache_gb) + + +def _finish(e: Entry, now: dt.datetime, why: str) -> None: + e.state, e.ended, e.why = "done", now, why + + +def prune(led: Ledger, facts: Facts) -> list[str]: + """Liveness and expiry. Running jobs past their ETA are kept (never killed).""" + out = [] + now = facts.now + for e in led.entries: + if e.state == "running": + unit = facts.units.get(e.unit) + if unit is None or not unit.active: + e.rc, e.peak_gb, why = facts.results.get(e.id, (None, None, "")) + _finish(e, now, why or "exited") + out.append(f"{e.id} {why}" if why else f"{e.id} exited rc={rc_text(e.rc)}") + elif e.state == "note": + if e.owner not in facts.live: + _finish(e, now, "owner gone") + out.append(f"{e.id} freed (owner gone)") + elif e.expires and now >= e.expires: + _finish(e, now, "expired") + out.append(f"{e.id} freed (expired)") + led.entries = [e for e in led.entries if not (e.state == "done" and e.ended and now - e.ended > DONE_KEEP)] + if led.game_until and now >= led.game_until: + led.game_until = None + out.append("game off (expired)") + return out + + +# ---------------------------------------------------------------- capacity + +def gaming(led: Ledger, now: dt.datetime) -> bool: + return bool(led.game_until and now < led.game_until) + + +def reserve(cfg: Config, led: Ledger, now: dt.datetime) -> tuple[float, int]: + if gaming(led, now): + return cfg.game_reserve_gb, cfg.game_reserve_cpus + return cfg.user_reserve_gb, cfg.user_reserve_cpus + + +def _used(e: Entry, facts: Facts) -> float: + unit = facts.units.get(e.unit) + return unit.current_gb if unit and unit.current_gb is not None else 0.0 + + +def held_gb(e: Entry, facts: Facts) -> float: + """What an entry still claims beyond what MemAvailable already shows as used.""" + if e.state == "note": + return e.mem_gb + if e.state == "running": + return max(0.0, e.mem_gb - _used(e, facts)) + return 0.0 + + +def frees_gb(e: Entry, facts: Facts) -> float: + """Budget gained when the entry ends.""" + return max(e.mem_gb, _used(e, facts)) if e.state == "running" else e.mem_gb + + +def _claims(led: Ledger) -> list[Entry]: + return [e for e in led.entries if e.state in ("running", "note")] + + +def budget(cfg: Config, led: Ledger, facts: Facts) -> tuple[float, int]: + res_gb, res_cpus = reserve(cfg, led, facts.now) + claims = _claims(led) + mem = facts.available_gb - res_gb - cfg.small_headroom_gb - sum(held_gb(e, facts) for e in claims) + return round(mem, 6), facts.nproc - res_cpus - sum(e.cpus for e in claims) + + +def force_room(cfg: Config, led: Ledger, facts: Facts) -> tuple[float, int]: + """Budget for --force: what the machine really has free beyond the user reserve and headroom; ledger + claims (unused reservations, notes) and the queue are ignored.""" + res_gb, res_cpus = reserve(cfg, led, facts.now) + return round(facts.available_gb - res_gb - cfg.small_headroom_gb, 6), facts.nproc - res_cpus + + +def fits(mem_gb: float, cpus: int, b: tuple[float, int]) -> bool: + return mem_gb <= b[0] + EPS and cpus <= b[1] + + +def queued_total(led: Ledger) -> tuple[float, int]: + q = [e for e in led.entries if e.state == "queued"] + return sum(e.mem_gb for e in q), sum(e.cpus for e in q) + + +def never_fits(cfg: Config, facts: Facts, mem_gb: float, cpus: int) -> str | None: + """Larger than the whole agent budget on an empty machine (normal reserve).""" + max_gb = facts.total_gb - cfg.user_reserve_gb - cfg.small_headroom_gb + max_cpus = facts.nproc - cfg.user_reserve_cpus + if mem_gb > max_gb + EPS: + return f"{fmt_gb(mem_gb)} can never fit (max {fmt_gb(max_gb)} for agents)" + if cpus > max_cpus: + return f"{cpus} cpus can never fit (max {max_cpus} for agents)" + return None + + +def end_time(e: Entry) -> dt.datetime: + return (e.expires if e.state == "note" else e.eta()) or e.started + + +def until(e: Entry, now: dt.datetime) -> str: + t = end_time(e) + return f"~{hhmm(t)}" if t > now else "overdue" + + +def needed(led: Ledger, facts: Facts, need_gb: float, need_cpus: int) -> tuple[list[Entry], bool]: + """Claims that must end (earliest first) to free need_gb and need_cpus; and whether that is enough.""" + got_gb, got_cpus, out = 0.0, 0, [] + for e in sorted(_claims(led), key=end_time): + if got_gb >= need_gb - EPS and got_cpus >= need_cpus: + break + out.append(e) + got_gb += frees_gb(e, facts) + got_cpus += e.cpus + return out, got_gb >= need_gb - EPS and got_cpus >= need_cpus + + +def force_hint(cfg: Config, led: Ledger, facts: Facts, mem_gb: float, cpus: int) -> str: + """Busy-line suffix naming --force when the request fits the memory that is really free.""" + room = force_room(cfg, led, facts) + if not fits(mem_gb, cpus, room): + return "" + return (f"; {fmt_gb(room[0])} really free beyond the reserve: --force starts it past the ledger " + "(only if the holders will not use what they reserved)") + + +def busy_line(cfg: Config, led: Ledger, facts: Facts, mem_gb: float, cpus: int, hint: str = "") -> str: + b_gb, b_cpus = budget(cfg, led, facts) + need_gb, need_cpus = mem_gb - b_gb, cpus - b_cpus + holders, _ = needed(led, facts, need_gb, need_cpus) + if not holders: + return (f"busy: only {fmt_gb(max(0.0, b_gb))} free for agents and no agent job holds any; " + "other programs use the rest; retry later or work on something else" + hint) + by_mem = need_gb > EPS + + def held(e): + what = fmt_gb(frees_gb(e, facts)) if by_mem else f"{e.cpus} cpus" + return f'{what} held by {e.project} "{e.title}" ({e.id}) until {until(e, facts.now)}' + + free = fmt_gb(max(0.0, b_gb)) if by_mem else f"{max(0, b_cpus)} cpus" + last = max(end_time(e) for e in holders) + retry = f"after ~{hhmm(last)}" if last > facts.now else "later" + return (f"busy: {', '.join(held(e) for e in holders)}; {free} free for agents; " + f"retry {retry} or work on something else" + hint) + + +# ---------------------------------------------------------------- locks + +LOCK_TITLE_RE = re.compile(r"gate\b", re.I) # 'gate …' titles lock 'gate' unasked: one shared gate checkout + + +def lock_for(title: str, key: str, main: str) -> str: + """Ledger lock of a run: --lock KEY, else 'gate' for a title starting with the word gate; '' = none. + Scoped to the main tree (lane worktrees share their project's gate dir).""" + key = key or ("gate" if LOCK_TITLE_RE.match(title.strip()) else "") + return f"{key}@{main}" if key else "" + + +def lock_holder(led: Ledger, lock: str) -> Entry | None: + """The running (else first queued) entry holding this lock.""" + if not lock: + return None + mine = [e for e in led.entries if e.lock == lock and e.state in ("running", "queued")] + return next((e for e in mine if e.state == "running"), None) or next(iter(_queue_of(mine)), None) + + +def _queue_of(entries: list[Entry]) -> list[Entry]: + return sorted((e for e in entries if e.state == "queued"), key=lambda e: (e.queued, int(e.id[2:]))) + + +def lock_busy_line(h: Entry, facts: Facts) -> str: + key = h.lock.split("@", 1)[0] + when = f"ETA {until(h, facts.now)}" if h.state == "running" else "queued" + return (f"busy: lock '{key}' held by {h.id} \"{h.title}\" ({when}): one at a time (shared checkout); " + f"--queue waits for it (--force does not override a lock); never a hand-written copy without the lock") + + +# ---------------------------------------------------------------- run, status lines + +RES_ID_VAR = "WF_RES_ID" # set in every job's unit: wf res calls from inside it name it as their batch +ENV_SKIP = {"PWD", "OLDPWD", "SHLVL", "_", RES_ID_VAR} +SECRET_RE = re.compile(r"TOKEN|SECRET|PASSWORD|PASSWD|CREDENTIAL", re.I) # never written to the ledger + + +def env_diff(caller: dict[str, str], manager: dict[str, str]) -> dict[str, str]: + """Caller variables a systemd --user unit would not get as is (it starts from the manager's environment).""" + return {k: v for k, v in caller.items() + if manager.get(k) != v and k not in ENV_SKIP and not SECRET_RE.search(k)} + + +def parse_show_environment(text: str) -> dict[str, str]: + return dict(l.split("=", 1) for l in text.splitlines() if "=" in l) + + +def run_argv(e: Entry, logdir: str) -> list[str]: + """systemd-run argv; the sh wrapper records exit code and peak bytes (unit is collected on exit).""" + rc, peak = shlex.quote(f"{logdir}/{e.id}.rc"), shlex.quote(f"{logdir}/{e.id}.peak") + script = (f'"$@"; rc=$?; cat /sys/fs/cgroup$(cut -d: -f3 /proc/self/cgroup)/memory.peak > {peak} ' + f'2>/dev/null; echo $rc > {rc}') + return ["systemd-run", "--user", "--quiet", "--collect", f"--slice={JOBS_SLICE}", f"--unit={e.unit}", + f"--working-directory={e.cwd}", *(f"--setenv={k}={v}" for k, v in sorted(e.env.items())), + f"--setenv={RES_ID_VAR}={e.id}", + "-p", f"MemoryMax={int(e.mem_gb * GB)}", "-p", f"MemoryHigh={int(e.mem_gb * 0.9 * GB)}", + "-p", "MemorySwapMax=0", "-p", "Nice=10", + "-p", f"StandardOutput=append:{e.log}", "-p", f"StandardError=append:{e.log}", + "/bin/sh", "-c", script, "sh", *e.cmd] + + +TASK_BRANCH_RE = re.compile(r"^[\w.-]+/([\w.-]+)$") # <lane>/<task id> (wf worktree branches) + + +def owner_by(environ: dict[str, str], branch: str, records: list[dict], name: str = "") -> dict[str, str]: + """Who starts an entry, empty keys dropped. name: --by, else WF_SESSION_NAME, else '<lane> session' from the + wf lane session record (.wf/sessions) with this CLAUDE_PID; task: WF_TASK, else the <lane>/<id> branch; + batch: WF_RES_ID (the wf res job, e.g. wf batch, this runs in); address: uds:<CLAUDE_CODE_MESSAGING_SOCKET> + (a batch worker's = its orchestrator's: SendMessage there reaches the batch).""" + pid = environ.get("CLAUDE_PID", "") + name = name or environ.get("WF_SESSION_NAME", "") + if not name and pid: + name = next((f"{r['lane']} session" for r in records + if isinstance(r, dict) and r.get("lane") and str(r.get("pid")) == pid), "") + m = TASK_BRANCH_RE.match(branch) + sock = environ.get("CLAUDE_CODE_MESSAGING_SOCKET", "") + d = {"name": name, "task": environ.get("WF_TASK") or (m.group(1) if m else ""), + "batch": environ.get(RES_ID_VAR, ""), "address": f"uds:{sock}" if sock else ""} + return {k: v for k, v in d.items() if v} + + +def by_text(e: Entry) -> str: + """'slow session t-x batch r-3, message uds:/s' — empty when nothing is known.""" + b = e.by or {} + who = " ".join(x for x in (b.get("name"), b.get("task"), f"batch {b['batch']}" if b.get("batch") else "") if x) + msg = f"message {b['address']}" if b.get("address") else "" + return ", ".join(x for x in (who, msg) if x) + + +def position(led: Ledger, e: Entry) -> int: + return _queue(led).index(e) + 1 + + +def entry_line(e: Entry, led: Ledger, facts: Facts) -> str: + by = by_text(e) if e.state != "done" else "" + return _entry_line(e, led, facts) + (f" [by {by}]" if by else "") + + +def _entry_line(e: Entry, led: Ledger, facts: Facts) -> str: + head = f'{e.id} {e.project} "{e.title}"' + if e.state == "running": + unit = facts.units.get(e.unit) + used = fmt_gb(unit.current_gb) if unit and unit.current_gb is not None else "?" + return (f"{head} running {fmt_gb(e.mem_gb)} used {used} {e.cpus} cpu since {hhmm(e.started)} " + f"ETA {until(e, facts.now)}" + (" (still running, not killed)" if end_time(e) <= facts.now else "")) + if e.state == "queued": + return f"{head} queued #{position(led, e)} {fmt_gb(e.mem_gb)} {e.cpus} cpu" + if e.state == "note": + return f"{head} note {fmt_gb(e.mem_gb)} {e.cpus} cpu until {until(e, facts.now)}" + peak = fmt_gb(e.peak_gb) if e.peak_gb is not None else "?" + return f"{head} done ({e.why}) rc={rc_text(e.rc)} peak {peak} at {hhmm(e.ended)}" + + +def status_lines(cfg: Config, led: Ledger, facts: Facts) -> list[str]: + active = [e for e in led.entries if e.state != "done"] + done = [e for e in led.entries if e.state == "done"] + out = [entry_line(e, led, facts) for e in active + done] or ["no reservations"] + b_gb, b_cpus = budget(cfg, led, facts) + r_gb, r_cpus = reserve(cfg, led, facts.now) + out.append(f"agents may use {fmt_gb(max(0.0, b_gb))}, {max(0, b_cpus)} cpus now; " + f"reserve {fmt_gb(r_gb)}/{r_cpus} cpus") + if facts.slice_gb > 0: + jobs = sum(_used(e, facts) for e in led.entries if e.state == "running") + free = max(0.0, facts.slice_gb - jobs) + cache = min(free, facts.slice_cache_gb) + out.append(f"unreserved agent memory {fmt_gb(free)}" + + (f" ({fmt_gb(cache)} of it file cache, reclaimable)" if cache >= 0.05 else "")) + if gaming(led, facts.now): + left = int((led.game_until - facts.now).total_seconds() // 60) + out.append(f"game on until {hhmm(led.game_until)} ({fmt_dur(left)} left)") + out.extend(shortfall_lines(cfg, led, facts)) + return out + + +def status_json(cfg: Config, led: Ledger, facts: Facts) -> str: + b_gb, b_cpus = budget(cfg, led, facts) + return json.dumps({"budget_gb": b_gb, "budget_cpus": b_cpus, "gaming": gaming(led, facts.now), + "game_until": _iso(led.game_until), "entries": [entry_dict(e) for e in led.entries]}, + indent=1) + + +def detail_lines(e: Entry) -> list[str]: + d = entry_dict(e) + d["cmd"] = shlex.join(e.cmd) + return [f"{k}: {v}" for k, v in d.items() if v not in (None, "", [])] + + +def slice_unit_text() -> str: + return ("[Unit]\nDescription=wf agents: Claude sessions and wf res jobs\n\n" + "[Slice]\nCPUWeight=20\nIOWeight=20\n") + + +def done_line(e: Entry) -> str: + minutes = round((e.ended - e.started).total_seconds() / 60) if e.ended and e.started else 0 + peak = fmt_gb(e.peak_gb) if e.peak_gb is not None else "?" + why = f" ({e.why})" if e.why and e.why != "exited" else "" + return f"{e.id} done rc={rc_text(e.rc)} peak {peak} in {minutes} min{why}" + + +JOURNAL_RESULT = re.compile(r"Failed with result '([^']+)'") +JOURNAL_MAIN = re.compile(r"Main process exited, code=\w+, status=(\S+)") +JOURNAL_PEAK = re.compile(r"(\d+(?:\.\d+)?)([BKMGT]) memory peak") +UNIT_SCALE = {"B": 1 / GB, "K": 1 / 1024 ** 2, "M": 1 / 1024, "G": 1.0, "T": 1024.0} + + +def journal_reason(text: str, mem_gb: float) -> tuple[str, float | None]: + """Why a unit died, from `journalctl -u UNIT -o cat` (used when the job wrote no .rc): (why, peak_gb).""" + m = JOURNAL_PEAK.search(text) + peak = round(float(m.group(1)) * UNIT_SCALE[m.group(2)], 3) if m else None + m = JOURNAL_RESULT.search(text) + result = m.group(1) if m else "" + m = JOURNAL_MAIN.search(text) + status = m.group(1) if m else "?" + if result == "oom-kill": + by = "by systemd-oomd" if "systemd-oomd killed" in text else "at MemoryMax" + return f"killed: oom-kill {by}, limit {fmt_gb(mem_gb)}; raise --mem", peak + if result in ("signal", "core-dump"): + return f"killed: signal {status}", peak + if result == "exit-code": + return f"failed: exit {status}", peak + return (f"killed: {result}" if result else ""), peak + + +def _queue(led: Ledger) -> list[Entry]: + return _queue_of(led.entries) + + +def to_start(cfg: Config, led: Ledger, facts: Facts) -> list[Entry]: + """Strict FIFO: queued entries to start now, stopping at the first that does not fit; an entry whose lock a + running job (or one started in this pass) holds is skipped, not blocking the rest.""" + b_gb, b_cpus = budget(cfg, led, facts) + out, held = [], {e.lock for e in led.entries if e.state == "running" and e.lock} + for e in _queue(led): + if e.lock and e.lock in held: + continue + if not fits(e.mem_gb, e.cpus, (b_gb, b_cpus)): + break + out.append(e) + held.add(e.lock) + b_gb, b_cpus = b_gb - e.mem_gb, b_cpus - e.cpus + return out + + +def queue_estimate(cfg: Config, led: Ledger, facts: Facts, e: Entry) -> dt.datetime | None: + """When enough claims end for e and everything ahead of it; None if they never free enough.""" + q = _queue(led) + ahead = q[:q.index(e) + 1] + b_gb, b_cpus = budget(cfg, led, facts) + holders, enough = needed(led, facts, sum(x.mem_gb for x in ahead) - b_gb, sum(x.cpus for x in ahead) - b_cpus) + if not enough: + return None + lock = [end_time(x) for x in led.entries if x is not e and x.lock and x.lock == e.lock and x.state == "running"] + return max([*(end_time(h) for h in holders), *lock], default=facts.now) + + +def slice_props(cfg: Config, led: Ledger, facts: Facts) -> dict[str, str]: + """agents.slice values: normal, or gaming (MemoryHigh never below current use: slow, never squeeze).""" + if gaming(led, facts.now): + high, weight = max(facts.total_gb - cfg.game_reserve_gb, facts.slice_gb), "5" + else: + high, weight = facts.total_gb - cfg.user_reserve_gb, "20" + return {"CPUWeight": weight, "IOWeight": weight, "MemoryHigh": str(int(round(high * GB)))} + + +def shortfall_lines(cfg: Config, led: Ledger, facts: Facts) -> list[str]: + """Empty when the game reserve is free now; else who holds it and when it frees.""" + free = facts.available_gb - sum(held_gb(e, facts) for e in _claims(led)) + short = round(cfg.game_reserve_gb - free, 6) + if short <= EPS: + return [] + holders, enough = needed(led, facts, short, 0) + head = f"short {fmt_gb(short)} of {fmt_gb(cfg.game_reserve_gb)}: " + if not holders: + return [head + "other programs, not agent jobs", "agent jobs alone cannot free it; close other programs"] + parts = ", ".join(f'{e.id} {e.project} "{e.title}" {fmt_gb(frees_gb(e, facts))} {until(e, facts.now)}' + for e in holders) + if not enough: + return [head + parts, "agent jobs alone cannot free it; close other programs"] + last = max(end_time(e) for e in holders) + big = max(holders, key=lambda e: frees_gb(e, facts)) + when = f"~{hhmm(last)}" if last > facts.now else "soon" + return [head + parts, f"full reserve free {when} (est.); free now: wf res release {big.id} --stop"] + + +def game_on_lines(cfg: Config, led: Ledger, facts: Facts, minutes: int) -> list[str]: + return [f"game on until {hhmm(led.game_until)} ({fmt_dur(minutes)}); CPU/IO now yours", + *shortfall_lines(cfg, led, facts)] + + +def scratch_victims(tree: dict, now_ts: float, hours: float, keep: set[str] = frozenset()) -> list[str]: + """tree = {top path: (newest mtime, {child path: newest mtime})}. A stale top goes whole; under a fresh + top, its stale children go. Children named in keep (live Claude session ids) never go; a stale top + holding one is pruned per child. Sorted.""" + limit = now_ts - hours * 3600 + out = [] + for top, (newest, children) in tree.items(): + held = any(c.rsplit("/", 1)[-1] in keep for c in children) + if newest <= limit and not held: + out.append(top) + else: + out.extend(c for c, m in children.items() if m <= limit and c.rsplit("/", 1)[-1] not in keep) + return sorted(out) + + +DOTNET_PIPE_RE = re.compile(r"^(?:clr-debug-pipe|dotnet-diagnostic)-(\d+)-") + + +def litter_victims(entries: list[tuple[str, str, float]], now_ts: float, hours: float, + alive: Callable[[int], bool]) -> list[str]: + """Top-level temp-dir litter (own entries only): (path, kind emptydir|dir|file|other, mtime) → sorted paths. + Empty dirs idle for `hours`; .NET debug pipes/sockets of dead pids. Content is never swept.""" + limit = now_ts - hours * 3600 + out = [] + for path, kind, mtime in entries: + m = DOTNET_PIPE_RE.match(path.rsplit("/", 1)[-1]) + if (kind == "emptydir" and mtime <= limit) or (m and kind != "dir" and not alive(int(m.group(1)))): + out.append(path) + return sorted(out) + + +def live_session_id(text: str, proc_start: str) -> str | None: + """~/.claude/sessions/<pid>.json of a live pid → its sessionId (= its scratch dir name), unless the + recorded procStart shows the pid now belongs to another process.""" + try: + d = json.loads(text) + except ValueError: + return None + if not isinstance(d, dict) or not isinstance(d.get("sessionId"), str): + return None + if d.get("procStart") is not None and str(d["procStart"]) != proc_start: + return None + return d["sessionId"] + + +AGE_RE = re.compile(r"^(.*):(\d+)d$") + + +def cleanup_rule(text: str) -> tuple[str, int | None]: + """'out/logs/*.log:30d' → ('out/logs/*.log', 30); relative, no '..'.""" + m = AGE_RE.match(text) + pattern, days = (m.group(1), int(m.group(2))) if m else (text, None) + if not pattern or pattern.startswith("/") or ".." in pattern.split("/"): + raise ResError(f"cleanup pattern '{text}' (want a path relative to the project, no '..')") + return pattern, days + + +CLAUDE_BIN_RE = re.compile(r"/claude/versions/[^/]+$") + + +def proc_name(comm: str, argv0: str) -> str: + """`claude` for a Claude Code process, else comm. Background sessions run the versioned binary, so comm + is the version ('2.1.283'); argv0 is …/claude/versions/X, or a rewritten title 'claude bg-pty-host'. Tools it + re-execs (ugrep) keep their own argv0.""" + if argv0.split(" ", 1)[0].rsplit("/", 1)[-1] == "claude" or CLAUDE_BIN_RE.search(argv0): + return "claude" + return comm + + +def adopt_groups(procs: list[tuple[int, int, str, str]]) -> dict[int, list[int]]: + """procs = (pid, ppid, comm, cgroup). Claude session roots outside agents.slice (named `claude`, parent not + `claude`) → sorted pids of the root and its descendants that are outside the slice.""" + inside = f"/{SLICE}/" + by_pid = {p[0]: p for p in procs} + kids: dict[int, list[int]] = {} + for pid, ppid, _, _ in procs: + kids.setdefault(ppid, []).append(pid) + out = {} + for pid, ppid, comm, cgroup in procs: + parent = by_pid.get(ppid) + if comm != "claude" or inside in cgroup or (parent and parent[2] == "claude"): + continue + tree, stack = [], [pid] + while stack: + p = stack.pop() + if inside not in by_pid[p][3]: + tree.append(p) + stack.extend(kids.get(p, [])) + out[pid] = sorted(tree) + return out + + +THROTTLE_FILL = 0.85 # used ≥ this × --mem: at MemoryHigh (0.9 × --mem), where the kernel throttles +THROTTLE_STALL = 20.0 # % of the last minute stalled on memory: page cache alone at the limit stays near 0 + + +def throttle_lines(led: Ledger, facts: Facts) -> list[str]: + """Running jobs stuck at their own memory limit (they crawl, then systemd-oomd kills them).""" + out = [] + for e in led.entries: + u = facts.units.get(e.unit) + if (e.state == "running" and u and u.active and u.current_gb is not None and u.stall is not None + and u.current_gb >= THROTTLE_FILL * e.mem_gb and u.stall >= THROTTLE_STALL): + out.append(f"{e.id} throttled at its memory limit ({u.current_gb:.1f} of {fmt_gb(e.mem_gb)}, " + f"stalled {u.stall:.0f}% of the last minute): likely too small; " + f"`wf res release {e.id} --stop` and re-run with a bigger --mem" + + (f"; started by {by_text(e)}" if by_text(e) else "")) + return out + + +DESKTOP_UNIT = ("app-", "dbus-") # transient units the desktop starts (XDG app launch, dbus activation) + + +def bare_units(blocks: dict[str, dict[str, str]]) -> list[tuple[str, float | None]]: + """Running transient services outside agents.slice not from the desktop = jobs from a bare `systemd-run`: + sorted (unit, used GB).""" + return sorted((n, gb_or_none(b.get("MemoryCurrent", ""))) for n, b in blocks.items() + if b.get("Transient") == "yes" and not n.startswith(DESKTOP_UNIT) + and not b.get("Slice", "").startswith("agents")) + + +def session_caps(blocks: dict[str, dict[str, str]], cap_gb: float) -> list[tuple[str, str]]: + """Session scopes in agents.slice whose MemoryHigh differs from the cap → sorted (unit, value) to set. + MemoryHigh only: over it the session is throttled; a MemoryMax kill could pick `claude` itself.""" + want = str(int(cap_gb * GB)) if cap_gb > 0 else "infinity" + return sorted((n, want) for n, b in blocks.items() + if n.endswith(".scope") and b.get("Slice") == SLICE and b.get("MemoryHigh") != want) + + +def session_cap_lines(blocks: dict[str, dict[str, str]], stall: dict[str, float], cap_gb: float) -> list[str]: + """Sessions stuck at their cap: used ≥ 0.9 × cap and stalled ≥ THROTTLE_STALL %.""" + if cap_gb <= 0: + return [] + out = [] + for n, b in sorted(blocks.items()): + cur = gb_or_none(b.get("MemoryCurrent", "")) + if cur is not None and cur >= 0.9 * cap_gb and stall.get(n, 0.0) >= THROTTLE_STALL: + who = n.removeprefix("wf-claude-").removesuffix(".scope") + out.append(f"warning: claude session {who} at its memory cap ({cur:.1f} of {fmt_gb(cap_gb)}, " + f"stalled {stall[n]:.0f}% of the last minute): run big work with wf res run") + return out + + +def warning_lines(total_gb: float, tmp_gb: float, outside: int, + bare: list[tuple[str, float | None]] = ()) -> list[str]: + out = [] + if tmp_gb > 0.25 * total_gb: + out.append(f"warning: /tmp (RAM) holds {fmt_gb(tmp_gb)}; wf res clean") + if outside: + out.append(f"warning: {outside} claude sessions outside {SLICE} (wf res adopt)") + if bare: + units = ", ".join(f"{n} {'?' if gb is None else f'{gb:.1f}'} GB" for n, gb in bare) + out.append(f"warning: {len(bare)} jobs outside wf res (bare systemd-run): {units}; " + "start jobs with wf res run, also from project scripts") + return out + + +def adopt_argv(root: int, pids: list[int]) -> list[str]: + """Move live pids into a new scope in agents.slice (systemd StartTransientUnit with PIDs).""" + return ["busctl", "--user", "call", "org.freedesktop.systemd1", "/org/freedesktop/systemd1", + "org.freedesktop.systemd1.Manager", "StartTransientUnit", "ssa(sv)a(sa(sv))", + f"wf-claude-{root}.scope", "fail", "2", "PIDs", "au", str(len(pids)), *map(str, pids), + "Slice", "s", SLICE, "0"] + + +def service_text(python: str, wf: str) -> str: + return ("[Unit]\nDescription=wf res tick\n\n[Service]\nType=oneshot\n" + f"ExecStart={python} {wf} res tick\n") + + +def timer_text() -> str: + return ("[Unit]\nDescription=wf res tick every minute\n\n[Timer]\nOnBootSec=1min\nOnUnitActiveSec=1min\n\n" + "[Install]\nWantedBy=timers.target\n") + + +def hook_output(payload: dict, game_until: dt.datetime | None, now: dt.datetime) -> dict | None: + """Claude Code PreToolUse hook: while game mode is on, Bash commands run without a display (no windows pop + up over the game). None = no output, the command runs unchanged.""" + cmd = (payload.get("tool_input") or {}).get("command") + if payload.get("tool_name") != "Bash" or not cmd or not (game_until and now < game_until): + return None + return {"hookSpecificOutput": { + "hookEventName": "PreToolUse", + "updatedInput": {**payload["tool_input"], "command": f"unset DISPLAY WAYLAND_DISPLAY; {cmd}"}, + "additionalContext": f"wf res game mode until {hhmm(game_until)}: this command has no display, so GUI " + "windows fail. Run headless/offscreen, or do other work until game off."}} + + +def shell_init_line() -> str: + return f"alias claude='systemd-run --user --scope --quiet --slice={SLICE} claude'" + + +# ---------------------------------------------------------------- history (reservation sizes) + +HIST_MIN_RUNS = 3 +HIST_MEM_PCT, HIST_MEM_X, HIST_MEM_FLOOR = 0.95, 1.15, 0.2 # suggest mem = p95 peak ×1.15, ≥ 0.2 GB +HIST_DUR_PCT, HIST_DUR_X, HIST_DUR_FLOOR = 0.90, 1.5, 5 # suggest for = p90 duration ×1.5, ≥ 5 min +HIST_OVER = 2.0 # request > 2× suggestion → hint +HIST_KEEP = 2000 # history file: newest runs kept +_PAREN_TAIL = re.compile(r"\s*\([^()]*\)\s*$") +_HEX_TAIL = re.compile(r"\s+(?=[0-9a-f]*\d)[0-9a-f]{7,40}$") + + +def title_kind(title: str) -> str: + """A title without its trailing (…) and sha/hex words: 'gate 2c91cac7 (fix x)' → 'gate'.""" + t = title.strip() + while True: + s = _HEX_TAIL.sub("", _PAREN_TAIL.sub("", t)).strip() + if s == t or not s: + return t + t = s + + +def pct(xs: list[float], p: float) -> float: + """Nearest-rank percentile.""" + s = sorted(xs) + return s[max(0, math.ceil(round(p * len(s), 9)) - 1)] + + +def history_record(e: Entry) -> dict | None: + """A finished, successful run with a measured peak → history record; else None.""" + if e.state != "done" or e.rc != 0 or e.peak_gb is None or not e.started or not e.ended: + return None + return {"id": e.id, "project": e.project, "title": e.title, "mem_gb": e.mem_gb, "est_min": e.est_min, + "rc": e.rc, "peak_gb": e.peak_gb, "min": round((e.ended - e.started).total_seconds() / 60, 2)} + + +def all_runs(records: list[dict], led: Ledger) -> list[dict]: + """History file records + the ledger's finished runs not yet in it (ids never repeat).""" + seen = {r.get("id") for r in records} + return records + [r for r in map(history_record, led.entries) if r and r["id"] not in seen] + + +def _kind_runs(runs: list[dict], project: str, kind: str) -> list[dict]: + return [r for r in runs if r.get("project") == project and title_kind(r.get("title", "")) == kind + and r.get("rc") == 0 and r.get("peak_gb") is not None and r.get("min") is not None] + + +def _ceil1(x: float) -> float: + return math.ceil(round(x * 10, 6)) / 10 + + +def suggest(runs: list[dict], project: str, kind: str) -> tuple[float, int, int] | None: + """(mem GB, minutes, n) from ≥ 3 runs of this kind in this project, else None.""" + rs = _kind_runs(runs, project, kind) + if len(rs) < HIST_MIN_RUNS: + return None + mem = max(HIST_MEM_FLOOR, _ceil1(pct([r["peak_gb"] for r in rs], HIST_MEM_PCT) * HIST_MEM_X)) + minutes = max(HIST_DUR_FLOOR, math.ceil(round(pct([r["min"] for r in rs], HIST_DUR_PCT) * HIST_DUR_X, 6))) + return mem, minutes, len(rs) + + +def hint_line(mem_gb: float, est_min: int, sug: tuple[float, int, int] | None) -> str | None: + """Request far over the history (no auto-resize): one line, else None.""" + if not sug or (mem_gb <= HIST_OVER * sug[0] + EPS and est_min <= HIST_OVER * sug[1]): + return None + return f"hint: history says ~{sug[0]:g} GB / {sug[1]} min ({sug[2]} runs)" + + +def hist_lines(runs: list[dict], project: str | None) -> list[str]: + """wf res hist: per project and title kind: n, request, peak p50/p95, estimate, duration p50/p90, suggestion.""" + keys = sorted({(r.get("project", ""), title_kind(r.get("title", ""))) for r in runs + if project is None or r.get("project") == project}) + out = [] + for p, k in keys: + rs = _kind_runs(runs, p, k) + if not rs: + continue + peaks, durs = [r["peak_gb"] for r in rs], [r["min"] for r in rs] + sug = suggest(runs, p, k) + out.append(f"{p} · {k} · n {len(rs)} · req {fmt_gb(pct([r['mem_gb'] for r in rs], 0.5))} · " + f"peak {pct(peaks, 0.5):.1f}/{pct(peaks, HIST_MEM_PCT):.1f} GB · " + f"est {fmt_dur(pct([r['est_min'] for r in rs], 0.5))} · " + f"dur {fmt_dur(round(pct(durs, 0.5)))}/{fmt_dur(round(pct(durs, HIST_DUR_PCT)))} · " + + (f"suggest {sug[0]:g} GB {fmt_dur(sug[1])}" if sug else f"suggest - (< {HIST_MIN_RUNS} runs)")) + return out or ["no finished runs with a peak yet"] + + +# ---------------------------------------------------------------- batch fit (wf batch --left / --time-left) + +TASK_P90_DEFAULT, TASK_P90_MIN_RUNS = 30, 3 # minutes per task when out/wf-orch.log has < 3 timed rows +_ORCH_ROW = re.compile(r"^\d{4}-\d\d-\d\dT\S+ (\S+) \S+ \S+ (\S+) \S+ (?:(\d+)h(\d+)m|(\d+)m(\d+)s)(?:\s|$)") + + +def orch_durations(text: str) -> list[float]: + """Minutes of done*/handback rows of local lanes in out/wf-orch.log with a real (> 0) duration.""" + out = [] + for line in text.splitlines(): + m = _ORCH_ROW.match(line) + if not m or m.group(1) == "cloud" or not (m.group(2).startswith("done") or m.group(2) == "handback"): + continue + h, hm, mm, s = (int(x or 0) for x in m.groups()[2:]) + minutes = h * 60 + hm + mm + s / 60 + if minutes > 0: + out.append(round(minutes, 4)) + return out + + +def task_p90(text: str) -> tuple[int, int]: + """(p90 task minutes rounded up, runs) from out/wf-orch.log text; < 3 runs → (30, runs).""" + durs = orch_durations(text) + if len(durs) < TASK_P90_MIN_RUNS: + return TASK_P90_DEFAULT, len(durs) + return math.ceil(round(pct(durs, 0.90), 6)), len(durs) + + +def batch_fit(k: int, left_min: int, p90_min: int) -> int: + """Tasks that fit: min(K, floor(left / p90)).""" + return max(0, min(k, left_min // max(1, p90_min))) diff --git a/wflib/search.py b/wflib/search.py new file mode 100644 index 0000000..600902d --- /dev/null +++ b/wflib/search.py @@ -0,0 +1,82 @@ +"""Ranked word search over open tasks, doc sections and the archive. + +A word matches at the start of a word, any case ("guard" finds "Guards"). +Rank: more of the words found first, then the weight of where they were +found, then open tasks before docs before archive. No index: read per call. +""" +from __future__ import annotations + +import re +from dataclasses import dataclass + +from . import refs, tasks + +W_TITLE, W_HEADING, W_GOAL, W_ARCHIVE, W_TEXT = 5, 4, 3, 2, 1 +KINDS = ("task", "doc", "archive") + + +@dataclass +class Hit: + found: int # how many of the words + score: int + kind: str + where: str # id, or path:line + label: str + line: str # the best matching line + + +def _patterns(words: list[str]) -> list[re.Pattern]: + return [re.compile(r"(?<![A-Za-z0-9])" + re.escape(w), re.I) for w in words if w.strip()] + + +def _rate(patterns: list[re.Pattern], parts: list[tuple[int, list[str]]]) -> tuple[int, int, str]: + """parts = (weight, lines). Returns (words found, score, best line).""" + found = score = 0 + best, best_weight = "", 0 + for p in patterns: + hit = max(((w, l) for w, lines in parts for l in lines if p.search(l)), key=lambda x: x[0], default=None) + if hit: + found += 1 + score += hit[0] + if hit[0] > best_weight: + best_weight, best = hit + return found, score, " ".join(best.split())[:120] + + +def search(words: list[str], doc: tasks.Doc, archive_text: str, docs: dict[str, str], + kinds: set[str] | None = None, limit: int = 15, archive_name: str = "archive") -> list[Hit]: + patterns = _patterns(words) + if not patterns: + return [] + kinds = kinds or set(KINDS) + hits: list[Hit] = [] + + def add(kind, where, label, parts, header=None): + found, score, line = _rate(patterns, parts) + if found: + in_header = header and any(line == " ".join(l.split())[:120] for w, ls in parts[:2] for l in ls) + hits.append(Hit(found, score, kind, where, label, header if in_header else line)) + + if "task" in kinds: + for item in doc.all_items(): + if item.error: + add("task", item.id, "", [(W_TEXT, [item.raw, *item.body])]) + else: + add("task", item.id, item.title, + [(W_TITLE, [item.id, item.title]), (W_GOAL, [item.goal]), (W_TEXT, item.body)], + header=item.text) + if "doc" in kinds: + for path, text in docs.items(): + lines = text.replace("\r\n", "\n").split("\n") + starts = [s for s in refs.sections(text) if s.level] + for n, s in enumerate(starts): + end = starts[n + 1].start if n + 1 < len(starts) else len(lines) + add("doc", f"{path}:{s.start + 1}", s.heading, + [(W_HEADING, [s.heading]), (W_TEXT, lines[s.start + 1:end])]) + if "archive" in kinds: + for n, line in enumerate(archive_text.replace("\r\n", "\n").split("\n"), 1): + if line.startswith("- "): + add("archive", f"{archive_name}:{n}", "", [(W_ARCHIVE, [line[2:]])]) + hits.sort(key=lambda h: (-h.found, -h.score, KINDS.index(h.kind))) + return hits[:limit] + diff --git a/wflib/tasks.py b/wflib/tasks.py new file mode 100644 index 0000000..80264c7 --- /dev/null +++ b/wflib/tasks.py @@ -0,0 +1,728 @@ +"""TASKS.md as data: parse into sections and items, change, render back. + +Pure text in, text out; no file access. Format: docs/design.md (section 3). +""" +from __future__ import annotations + +import difflib +import re +from dataclasses import dataclass, field + +SECTIONS = { + "awaiting": "Awaiting your decision", + "pending": "Pending", + "human": "Needs human", + "deferred": "Deferred", +} +TASK_KEYS = ("pending", "human", "deferred") +EFFORTS = ("<1h", "1h", "5h", "10h", "100h") +INDENT = " " + +ID_RE = re.compile(r"^[ta]-[a-z0-9]+(?:-[a-z0-9]+)*$") +LINK_RE = re.compile(r"\[\[([^\]|#]+)\]\]") +HEADER_RE = re.compile( + r"^- \*\*(?P<id>[^*\s]+)\*\*" + r"(?: \[P(?P<prio>\d)\])?" + r"(?: \((?!in progress:|blocked:)(?P<effort>[^)]*)\))?" + r"(?: \((?P<status>(?:in progress|blocked):[^)]*)\))?" + r": (?P<text>.*)$" +) +ITEM_START = "- **" +AFTER_RE = re.compile(r"^\s*(?:-\s+)?After:\s*(.*)$") +REF_RE = re.compile(r"^\s*(?:-\s+)?Ref:\s*(.*)$") +SLICES_RE = re.compile(r"^\s*-\s+Slices:\s*(.*)$") +MODEL_RE = re.compile(r"^\s*(?:-\s+)?Model:\s*(.*)$") +MODELS = ("haiku", "sonnet", "opus") # low → high; no Model line = opus +DONE_RE = re.compile(r"^\s*(?:-\s+)?Done:\s*(.*)$") +HUMAN_DONE_RE = re.compile( + r"\bowner(?:'s)?\s+(?:approves?|approval|tests?|checks?|runs?|says?|confirms?|verif(?:y|ies)|decides?|reviews?|signs?|sees?|ok)\b" + r"|\breport to\b|\bconfirm with\b|\bplease confirm\b|^\s*confirm\b", re.I) +SESSIONS_RE = re.compile(r"^\s*(?:-\s+)?Sessions:\s*(.*)$") +SESSIONS = ("parallel", "solo", "owner") # no Sessions line = parallel +CLOUD_RE = re.compile(r"^\s*(?:-\s+)?Cloud:\s*(.*)$") +CLOUDS = ("yes", "no") # no Cloud line = decided by the fit rules (wflib/cloud.py) +LEADING_ID_RE = re.compile(r"^([ta]-[a-z0-9]+(?:-[a-z0-9]+)*):\s+(.*)$") + + +class TaskError(ValueError): + pass + + +@dataclass +class Item: + id: str + prio: int | None = None + effort: str | None = None + interactive: bool = False + status: str | None = None + text: str = "" + body: list[str] = field(default_factory=list) + error: str | None = None + raw: str | None = None + line: int | None = None # 1-based line of the header in the parsed text + + def header(self) -> str: + if self.error: + return self.raw or "" + out = f"- **{self.id}**" + if self.prio is not None: + out += f" [P{self.prio}]" + if self.effort is not None: + out += f" ({self.effort}{', interactive' if self.interactive else ''})" + if self.status: + out += f" ({self.status})" + return f"{out}: {self.text}" + + def lines(self) -> list[str]: + return [self.header(), *self.body] + + @property + def title(self) -> str: + return self.text.split(". ", 1)[0].rstrip(".").strip() + + @property + def goal(self) -> str: + return self.text.split(". ", 1)[1].strip() if ". " in self.text else "" + + @property + def blocked_on(self) -> str | None: + if self.status and self.status.startswith("blocked:"): + found = LINK_RE.findall(self.status) + return found[0] if found else None + return None + + def _line(self, pattern: re.Pattern) -> int | None: + return next((i for i, l in enumerate(self.body) if pattern.match(l)), None) + + @property + def after(self) -> list[str]: + i = self._line(AFTER_RE) + return LINK_RE.findall(self.body[i]) if i is not None else [] + + @property + def slices(self) -> list[str]: + i = self._line(SLICES_RE) + return LINK_RE.findall(self.body[i]) if i is not None else [] + + @property + def model(self) -> str: + i = self._line(MODEL_RE) + words = MODEL_RE.match(self.body[i]).group(1).split() if i is not None else [] + return words[0] if words else "opus" + + @property + def done_text(self) -> str | None: + """Text of the Done line, None without one.""" + i = self._line(DONE_RE) + return DONE_RE.match(self.body[i]).group(1) if i is not None else None + + @property + def runner_ready(self) -> bool: + """Headless worker can finish it: has a Done line, not owner-bound, Done not about the owner/confirming.""" + done = self.done_text + return bool(done and done.strip()) and self.sessions != "owner" and not HUMAN_DONE_RE.search(done) + + @property + def human_done_match(self) -> str | None: + """The Done phrase that makes it not runner-ready (human action), None if none.""" + m = HUMAN_DONE_RE.search(self.done_text or "") + return m.group(0) if m else None + + @property + def sessions(self) -> str: + """parallel | solo (no other live session) | owner (owner present); header ', interactive' = owner.""" + i = self._line(SESSIONS_RE) + words = SESSIONS_RE.match(self.body[i]).group(1).split() if i is not None else [] + if words and words[0] in SESSIONS: + return words[0] + return "owner" if self.interactive else "parallel" + + @property + def cloud(self) -> str | None: + """yes | no | None (no Cloud line, or an unknown word).""" + i = self._line(CLOUD_RE) + words = CLOUD_RE.match(self.body[i]).group(1).split() if i is not None else [] + return words[0] if words and words[0] in CLOUDS else None + + @property + def refs(self) -> list[tuple[str, str | None]]: + i = self._line(REF_RE) + if i is None: + return [] + out = [] + for part in split_refs(REF_RE.match(self.body[i]).group(1)): + target = re.sub(r"\s*\(.*$", "", part).strip() + if target: + path, _, anchor = target.partition("#") + out.append((path, anchor or None)) + return out + + +def split_refs(text: str) -> list[str]: + """Split on commas that are not inside parentheses.""" + parts, depth, cur = [], 0, "" + for ch in text: + if ch == "(": + depth += 1 + elif ch == ")": + depth = max(0, depth - 1) + if ch == "," and depth == 0: + parts.append(cur) + cur = "" + else: + cur += ch + return [p.strip() for p in parts + [cur] if p.strip()] + + +@dataclass +class Section: + heading: str + key: str | None + prefix: list[str] = field(default_factory=list) + items: list[Item] = field(default_factory=list) + suffix: list[str] = field(default_factory=list) + line: int | None = None # 1-based line of the heading + suffix_line: int | None = None + + def lines(self) -> list[str]: + out = [f"## {self.heading}", *self.prefix] + for item in self.items: + out += [*item.lines(), ""] + return out + self.suffix + + +@dataclass +class Doc: + preamble: list[str] + sections: list[Section] + newline: str = "\n" + + def section(self, key: str) -> Section: + for s in self.sections: + if s.key == key: + return s + raise TaskError(f"no '## {SECTIONS[key]}' section") + + def all_items(self) -> list[Item]: + return [i for s in self.sections for i in s.items] + + def ids(self) -> set[str]: + return {i.id for i in self.all_items() if i.id} + + def find(self, id: str) -> tuple[Section, int]: + for s in self.sections: + for n, item in enumerate(s.items): + if item.id == id: + return s, n + raise TaskError(unknown_id(id, self.ids())) + + def item(self, id: str) -> Item: + s, n = self.find(id) + return s.items[n] + + +def unknown_id(id: str, known: set[str]) -> str: + near = difflib.get_close_matches(id, sorted(known), n=3, cutoff=0.0) + return f"unknown id '{id}'" + (f" (nearest: {', '.join(near)})" if near else "") + + +def parse_header(line: str) -> Item: + m = HEADER_RE.match(line) + if not m: + found = re.match(r"^- \*\*([^*]+)\*\*", line) + return Item(id=found.group(1).strip() if found else "", raw=line, + error="bad header: want '- **id** [Pn] (effort) [(status)]: Title. Goal.'") + effort, interactive = m.group("effort"), False + if effort is not None: + parts = [p.strip() for p in effort.split(",")] + interactive = "interactive" in parts[1:] + effort = parts[0] + status = m.group("status") + return Item(id=m.group("id"), prio=int(m.group("prio")) if m.group("prio") else None, + effort=effort, interactive=interactive, + status=status.strip() if status else None, text=m.group("text").strip()) + + +def _strip_blank_tail(lines: list[str]) -> list[str]: + while lines and not lines[-1].strip(): + lines.pop() + return lines + + +def _parse_section(heading: str, lines: list[str], at: int = 0) -> Section: + """`at` = 1-based line number of the heading.""" + key = next((k for k, h in SECTIONS.items() if h == heading), None) + section = Section(heading, key, line=at) + if key is None: + section.prefix = lines + return section + for n, line in enumerate(lines, at + 1): + if section.suffix: + section.suffix.append(line) + elif line.startswith(ITEM_START): + section.items.append(parse_header(line)) + section.items[-1].line = n + elif not section.items: + section.prefix.append(line) + elif line.strip() and not line[0].isspace(): + section.suffix.append(line) + section.suffix_line = n + else: + section.items[-1].body.append(line) + for item in section.items: + _strip_blank_tail(item.body) + return section + + +def parse(text: str) -> Doc: + newline = "\r\n" if "\r\n" in text else "\n" + lines = _strip_blank_tail(text.replace("\r\n", "\n").split("\n")) + starts = [i for i, l in enumerate(lines) if l.startswith("## ")] + doc = Doc(lines[: starts[0]] if starts else lines, [], newline) + for n, start in enumerate(starts): + end = starts[n + 1] if n + 1 < len(starts) else len(lines) + doc.sections.append(_parse_section(lines[start][3:].strip(), lines[start + 1:end], start + 1)) + return doc + + +def render(doc: Doc) -> str: + lines = list(doc.preamble) + for s in doc.sections: + lines += s.lines() + _strip_blank_tail(lines) + return doc.newline.join(lines) + doc.newline + + +# ---------------------------------------------------------------- ids + +def make_id(title: str, taken: set[str], prefix: str = "t-") -> str: + words = [w for w in re.sub(r"[^a-z0-9]+", " ", title.lower().encode("ascii", "ignore").decode()).split() if w] + slug = "" + for w in words: + longer = f"{slug}-{w}" if slug else w + if len(longer) > 40: + break + slug = longer + if not slug and words: + slug = words[0][:40] + if not slug: + raise TaskError(f"no id can be made from title '{title}': give one with --id") + base, n, id = prefix + slug, 2, prefix + slug + while id in taken: + id = f"{base}-{n}" + n += 1 + return id + + +def slice_id(parent: str, taken: set[str]) -> str: + n = 1 + while f"{parent}-{n}" in taken: + n += 1 + return f"{parent}-{n}" + + +def open_slices(doc: Doc, id: str) -> list[str]: + """Open items listed in the Slices: line of `id`, then open `<id>-N` items.""" + pattern = re.compile(re.escape(id) + r"-\d+$") + open_ = doc.ids() + named = [s for s in doc.item(id).slices if s in open_] + return named + [i.id for i in doc.all_items() if pattern.match(i.id) and i.id not in named] + + +def parent_of(doc: Doc, id: str) -> str | None: + for item in doc.all_items(): + if id in item.slices: + return item.id + m = re.match(r"(.+)-\d+$", id) + return m.group(1) if m and m.group(1) in doc.ids() else None + + +def rename(doc: Doc, old: str, new: str, taken: set[str]) -> int: + """Change the id of an open item and every [[old]] link; returns the link count. + `taken` = ids used in the archive.""" + item = doc.item(old) + if not ID_RE.match(new): + raise TaskError(f"bad id '{new}' (want t-… or a-…, lowercase a-z 0-9 -)") + if new[:2] != old[:2]: + raise TaskError(f"a rename keeps the kind ({old[:2]}…)") + if new in doc.ids(): + raise TaskError(f"id '{new}' already exists") + if new in taken: + raise TaskError(f"id '{new}' already used in the archive (ids are never reused)") + link, n = re.compile(r"\[\[" + re.escape(old) + r"\]\]"), 0 + for other in doc.all_items(): + other.text, k = link.subn(f"[[{new}]]", other.text) + n += k + if other.status: + other.status, k = link.subn(f"[[{new}]]", other.status) + n += k + for i, line in enumerate(other.body): + other.body[i], k = link.subn(f"[[{new}]]", line) + n += k + item.id = new + return n + + +# ---------------------------------------------------------------- placing + +def _prio(item: Item) -> int: + return 9 if item.prio is None else item.prio + + +def place(section: Section, item: Item) -> int: + """Insert rule: after the last item of the same or higher priority, and + after every item named in the new item's After:.""" + at = max((n + 1 for n, other in enumerate(section.items) if _prio(other) <= _prio(item)), default=0) + deps = set(item.after) + return max([at, *(n + 1 for n, other in enumerate(section.items) if other.id in deps)]) + + +def _check_kind(item: Item, key: str) -> None: + if key not in SECTIONS: + raise TaskError(f"unknown section '{key}' (want {', '.join(SECTIONS)})") + if item.id.startswith("a-") != (key == "awaiting"): + raise TaskError("a- items belong in Awaiting, t- items in Pending / Needs human / Deferred") + if key != "awaiting" and item.prio is None: + raise TaskError(f"task '{item.id}' needs a priority [P0]-[P3]") + + +def insert(doc: Doc, item: Item, key: str) -> None: + if item.error: + raise TaskError(item.error) + if item.id in doc.ids(): + raise TaskError(f"id '{item.id}' already exists") + _check_kind(item, key) + section = doc.section(key) + section.items.insert(place(section, item), item) + + +def remove(doc: Doc, id: str) -> Item: + section, n = doc.find(id) + return section.items.pop(n) + + +def _section_key(doc: Doc, id: str) -> str: + return doc.find(id)[0].key + + +def set_prio(doc: Doc, id: str, prio: int) -> None: + if prio not in (0, 1, 2, 3): + raise TaskError("priority 0-3") + key = _section_key(doc, id) + item = doc.item(id) + if item.id.startswith("a-"): + raise TaskError("Awaiting items have no priority") + remove(doc, id) + item.prio = prio + insert(doc, item, key) + + +def move_to(doc: Doc, id: str, key: str) -> None: + item = doc.item(id) + _check_kind(item, key) + doc.section(key) + remove(doc, id) + insert(doc, item, key) + + +def move_rel(doc: Doc, id: str, other: str, before: bool, force: bool = False) -> None: + if id == other: + raise TaskError("cannot move an item relative to itself") + item = doc.item(id) + target, _ = doc.find(other) + _check_kind(item, target.key) + source, n = doc.find(id) + source.items.pop(n) + at = next(i for i, o in enumerate(target.items) if o.id == other) + (0 if before else 1) + target.items.insert(at, item) + + def undo(): + target.items.pop(at) + source.items.insert(n, item) + + later = {o.id for o in target.items[at + 1:]} + for dep in item.after: + if dep in later: + undo() + raise TaskError(f"'{id}' is After: [[{dep}]], it cannot go before it") + for o in target.items[:at]: + if id in o.after: + undo() + raise TaskError(f"'{o.id}' is After: [[{id}]], '{id}' cannot go after it") + prev = target.items[at - 1] if at else None + nxt = target.items[at + 1] if at + 1 < len(target.items) else None + if not force and ((prev and _prio(prev) > _prio(item)) or (nxt and _prio(nxt) < _prio(item))): + undo() + raise TaskError(f"moving '{id}' there breaks priority order (wf prio, or --force)") + + +# ---------------------------------------------------------------- fields + +def set_status(doc: Doc, id: str, status: str | None) -> None: + item = doc.item(id) + if item.id.startswith("a-"): + raise TaskError("Awaiting items have no status") + if status: + status = status.strip() + if status.startswith("blocked:"): + on = LINK_RE.findall(status) + open_awaiting = {i.id for s in doc.sections if s.key == "awaiting" for i in s.items} + if len(on) != 1 or on[0] not in open_awaiting: + raise TaskError(f"'{on[0] if on else status}' is not an open Awaiting item") + elif not status.startswith("in progress:"): + raise TaskError("status is 'in progress: <note>' or 'blocked: [[a-id]]'") + if ")" in status or "(" in status: + raise TaskError("status note cannot contain parentheses (the status line ends in one); use - or :") + item.status = status or None + + +def _tail_start(item: Item) -> int: + """Index where the trailing Sessions:/Model:/After:/Ref: lines of the body begin.""" + n = len(item.body) + while n and any(p.match(item.body[n - 1]) for p in (AFTER_RE, REF_RE, MODEL_RE, SESSIONS_RE, CLOUD_RE)): + n -= 1 + return n + + +def _set_line(item: Item, pattern: re.Pattern, line: str | None, at_end: bool) -> None: + i = item._line(pattern) + if i is not None: + if line is None: + item.body.pop(i) + else: + item.body[i] = line + elif line is not None: + if at_end: + item.body.append(line) + else: + ref = item._line(REF_RE) + item.body.insert(len(item.body) if ref is None else ref, line) + + +def set_fields(doc: Doc, id: str, *, title: str | None = None, effort: str | None = None, + interactive: bool | None = None, after: list[str] | None = None, + refs: list[str] | None = None, model: str | None = None, + sessions: str | None = None, done: str | None = None, + cloud: str | None = None) -> None: + item = doc.item(id) + if done is not None: + done = done.strip() + if (i := item._line(DONE_RE)) is not None: + if done: + item.body[i] = INDENT + "Done: " + done + else: + item.body.pop(i) + elif done: + item.body.insert(_tail_start(item), INDENT + "Done: " + done) + if sessions is not None: + if sessions not in ("", *SESSIONS): + raise TaskError(f"sessions '{sessions}' (want {', '.join(SESSIONS)})") + item.interactive = False + line = INDENT + "Sessions: " + sessions if sessions else None + if (i := item._line(SESSIONS_RE)) is not None: + if line: + item.body[i] = line + else: + item.body.pop(i) + elif line: + m = item._line(MODEL_RE) + item.body.insert(m if m is not None and m >= _tail_start(item) else _tail_start(item), line) + if cloud is not None: + if cloud not in ("", *CLOUDS): + raise TaskError(f"cloud '{cloud}' (want {', '.join(CLOUDS)})") + line = INDENT + "Cloud: " + cloud if cloud else None + if (i := item._line(CLOUD_RE)) is not None: + if line: + item.body[i] = line + else: + item.body.pop(i) + elif line: + item.body.insert(_tail_start(item), line) + if model: + if model not in MODELS: + raise TaskError(f"model '{model}' (want {', '.join(MODELS)})") + line = INDENT + "Model: " + model + if (i := item._line(MODEL_RE)) is not None: + item.body[i] = line + else: + item.body.insert(_tail_start(item), line) + elif model is not None and (i := item._line(MODEL_RE)) is not None: + item.body.pop(i) + if title is not None: + title = title.strip().rstrip(".") + if not title: + raise TaskError("empty title") + if ". " in title or title.endswith(("?", "!")): + item.text = title if title.endswith(("?", "!")) else f"{title}." + else: + item.text = f"{title}. {item.goal}" if item.goal else f"{title}." + if effort is not None: + if effort not in EFFORTS: + raise TaskError(f"effort '{effort}' (want {', '.join(EFFORTS)})") + item.effort = effort + if interactive is not None: + item.interactive = interactive + if after is not None: + line = INDENT + "- After: " + ", ".join(f"[[{a}]]" for a in after) if after else None + _set_line(item, AFTER_RE, line, at_end=False) + if refs is not None: + _set_line(item, REF_RE, INDENT + "Ref: " + ", ".join(refs) if refs else None, at_end=True) + + +def add_note(doc: Doc, id: str, line: str) -> None: + item = doc.item(id) + line = line.strip() + if not line: + raise TaskError("empty note") + item.body.insert(_tail_start(item), f"{INDENT}- {line}") + + +def _link_after(line: str) -> str: + """'After: t-a, t-b' -> 'After: [[t-a]], [[t-b]]' (bare ids accepted, written as links).""" + m = re.match(r"^(\s*(?:-\s+)?After:\s*)(.*)$", line) + if not m: + return line + parts = [w for w in re.split(r"[\s,;]+", m.group(2)) if w] + if not parts or not all(LINK_RE.fullmatch(w) or ID_RE.match(w) for w in parts): + return line + return m.group(1) + ", ".join(w if w.startswith("[[") else f"[[{w}]]" for w in parts) + + +def indent_body(lines: list[str]) -> list[str]: + import textwrap + text = textwrap.dedent("\n".join(_link_after(l.rstrip()) for l in lines)) + return _strip_blank_tail([INDENT + l if l.strip() else "" for l in text.split("\n")]) if text.strip() else [] + + +def set_body(doc: Doc, id: str, lines: list[str]) -> None: + item = doc.item(id) + new = indent_body(lines) + kept = [l for l in item.body[_tail_start(item):] + if not any(p.match(l) and any(p.match(n) for n in new) for p in (AFTER_RE, REF_RE))] + item.body = new + kept + + +BOX_RE = re.compile(r"^(\s*- \[)([ xX])(\] )(.*)$") + + +def tick(doc: Doc, id: str, which: str) -> str: + item = doc.item(id) + boxes = [(i, BOX_RE.match(l)) for i, l in enumerate(item.body) if BOX_RE.match(l)] + if which.isdigit(): + if not 1 <= int(which) <= len(boxes): + raise TaskError(f"no box {which} in '{id}' ({len(boxes)} boxes)") + hits = [boxes[int(which) - 1]] + else: + hits = [(i, m) for i, m in boxes if which.lower() in m.group(4).lower()] + if not hits: + raise TaskError(f"no box matches '{which}' in '{id}'") + if len(hits) > 1: + raise TaskError(f"'{which}' matches {len(hits)} boxes in '{id}'") + i, m = hits[0] + if m.group(2) != " ": + raise TaskError(f"box already ticked: {m.group(4)}") + item.body[i] = f"{m.group(1)}x{m.group(3)}{m.group(4)}" + return m.group(4) + + +def unblock(doc: Doc, a_id: str) -> list[str]: + out = [] + for item in doc.all_items(): + if item.blocked_on == a_id: + item.status = None + out.append(item.id) + return out + + +# ---------------------------------------------------------------- pick + +def why_not(doc: Doc, item: Item, archived: set[str], held: dict[str, str] | None = None, + others: int = 0, owner: bool = True) -> str | None: + """Why a Pending item can't be picked now; None = pickable. held: id → holder of a live + claim by another session; others: other live sessions in the project; owner: owner present.""" + if item.error: + return "bad header" + if held and item.id in held: + return f"in progress by {held[item.id]}" + if item.sessions == "owner" and not owner: + return "owner: needs the owner (wf next --owner)" + if item.sessions == "solo" and others: + return f"solo: {others} other live session{'s' if others > 1 else ''}" + if item.status and item.status.startswith("blocked:"): + return f"blocked: {item.blocked_on}" + if open_slices(doc, item.id): + return "open slices: " + ", ".join(open_slices(doc, item.id)) + if [a for a in item.after if a not in archived]: + return "after: " + ", ".join(a for a in item.after if a not in archived) + return None + + +def waiters(doc: Doc) -> dict[str, set[str]]: + """Open task id → every open task (Pending, Needs human) that waits on it, transitively: + via After:, and a parent waits on its open slices.""" + items = {i.id: i for s in doc.sections if s.key in ("pending", "human") for i in s.items if not i.error} + direct: dict[str, set[str]] = {id: set() for id in items} + for item in items.values(): + for dep in item.after: + if dep in direct: + direct[dep].add(item.id) + for s in item.slices: + if s in direct: + direct[s].add(item.id) + out: dict[str, set[str]] = {} + for id in direct: + seen, todo = set(), list(direct[id]) + while todo: + w = todo.pop() + if w not in seen and w != id: + seen.add(w) + todo += direct[w] + out[id] = seen + return out + + +def solo_running(doc: Doc, held: dict[str, str]) -> tuple[str, str] | None: + """(id, holder) of a solo task in progress by another live session; held: see why_not.""" + return next(((i.id, held[i.id]) for i in doc.section("pending").items + if i.id in held and not i.error and i.sessions == "solo"), None) + + +# ---------------------------------------------------------------- archive + +ARCHIVE_ID_RE = re.compile(r"^- .*?\*\*([ta]-[a-z0-9-]+)\*\*") + + +def archive_ids(text: str) -> set[str]: + return {m.group(1) for l in text.replace("\r\n", "\n").split("\n") if (m := ARCHIVE_ID_RE.match(l))} + + +def archive_line(date: str, item: Item, entry: str) -> str: + entry = " ".join(entry.split()) + return f"- {date} **{item.id}** {item.title}" + (f" — {entry}" if entry else "") + + +def archive_prepend(text: str, line: str) -> str: + newline = "\r\n" if "\r\n" in text else "\n" + lines = _strip_blank_tail(text.replace("\r\n", "\n").split("\n")) + at = next((i for i, l in enumerate(lines) if l.startswith("- ")), None) + if at is None: + lines += ["", line] + else: + lines.insert(at, line) + return newline.join(lines) + newline + + +# ---------------------------------------------------------------- blocks + +def parse_block(text: str) -> Item: + lines = _strip_blank_tail(text.replace("\r\n", "\n").strip("\n").split("\n")) + if not lines or not lines[0].strip(): + raise TaskError("empty item") + first = lines[0].strip() + if re.match(r"^\d+\. ", first): + raise TaskError("old numbered format: write '- **id** [Pn] (effort): Title. Goal.'") + item = parse_header(first) + if item.error: + raise TaskError(item.error) + item.body = indent_body(lines[1:]) + return item diff --git a/wflib/usage.py b/wflib/usage.py new file mode 100644 index 0000000..5b7c1ee --- /dev/null +++ b/wflib/usage.py @@ -0,0 +1,375 @@ +"""Token usage and API-price cost from Claude Code transcripts (jsonl), per model. + +One API request is streamed as several transcript entries (one per content block) that repeat its +usage: count each (message.id, requestId) once, taking its last entry (final output_tokens). +Subagent transcripts never get a final entry (stop_reason null, output_tokens = the message_start +placeholder): there output is estimated from the content (out_estimate), counted in Usage.est. +""" +from __future__ import annotations + +import json +import re +from dataclasses import dataclass, fields +from statistics import median + +# $ per million tokens: input, output, cache read. Cache writes: 1.25x input (5 min), 2x input (1 h). +# Source: Anthropic pricing as of 2026-09-25. Longest matching prefix of the model id wins. +PRICES = { + "claude-fable-5-1": (10.0, 50.0, 0.25), + "claude-fable-5": (10.0, 50.0, 1.0), + "claude-opus-5-5": (4.0, 20.0, 0.20), + "claude-opus-5": (5.0, 25.0, 0.50), + "claude-opus-4-8": (5.0, 25.0, 0.50), + "claude-sonnet-5-5": (2.0, 10.0, 0.20), + "claude-sonnet-5": (2.0, 10.0, 0.20), + "claude-sonnet-4-6": (3.0, 15.0, 0.30), + "claude-haiku-4-5": (1.0, 5.0, 0.10), +} + + +@dataclass +class Usage: + turns: int = 0 + inp: int = 0 + cw5: int = 0 + cw1h: int = 0 + cr: int = 0 + out: int = 0 + est: int = 0 # requests whose out is estimated from content + + @property + def cw(self) -> int: + return self.cw5 + self.cw1h + + def add(self, other: "Usage") -> None: + for f in fields(self): + setattr(self, f.name, getattr(self, f.name) + getattr(other, f.name)) + + +def price(model: str) -> tuple[float, float, float] | None: + keys = [k for k in PRICES if model == k or model.startswith(k + "-")] + return PRICES[max(keys, key=len)] if keys else None + + +def cost(model: str, u: Usage) -> float | None: + """API-price $ (subscription sessions: what it would cost on the API); None = unknown model.""" + p = price(model) + if p is None: + return None + inp, out, read = p + return (u.inp * inp + u.cw5 * inp * 1.25 + u.cw1h * inp * 2 + u.cr * read + u.out * out) / 1e6 + + +def _usage(raw: dict) -> Usage: + cw = raw.get("cache_creation_input_tokens") or 0 + cw1h = (raw.get("cache_creation") or {}).get("ephemeral_1h_input_tokens") or 0 + return Usage(1, raw.get("input_tokens") or 0, cw - cw1h, cw1h, + raw.get("cache_read_input_tokens") or 0, raw.get("output_tokens") or 0) + + +# Output-token estimate per content block, fitted on ~100k main-session requests with final usage +# (2026-10): per-session totals within ±5%. Thinking text is usually empty; its signature grows ~3.3 +# chars per thinking token over a ~800-char base. +EST_TEXT, EST_TOOL_CHAR, EST_TOOL, EST_SIG, EST_SIG_BASE = 0.3, 0.44, 30, 0.3, 800 + + +def out_estimate(blocks: list) -> int: + t = 0.0 + for b in blocks: + if not isinstance(b, dict): + continue + kind = b.get("type") + if kind == "text": + t += len(b.get("text") or "") * EST_TEXT + elif kind == "tool_use": + t += EST_TOOL + len(json.dumps(b.get("input") or {}, ensure_ascii=False)) * EST_TOOL_CHAR + elif kind == "thinking": + t += (len(b.get("thinking") or "") * EST_TEXT + + max(0, len(b.get("signature") or "") - EST_SIG_BASE) * EST_SIG) + return round(t) + + +def parse(text: str, since: str | None = None, until: str | None = None) -> dict[str, Usage]: + """Model id → summed Usage. since/until: ISO prefixes compared with the UTC timestamps.""" + last: dict[tuple, tuple[str, dict]] = {} + first_ts: dict[tuple, str] = {} + final: set[tuple] = set() + blocks: dict[tuple, list] = {} + seen: set[str] = set() + for line in text.split("\n"): + try: + d = json.loads(line) + except ValueError: + continue + if not isinstance(d, dict) or d.get("type") != "assistant": + continue + m = d.get("message") or {} + model, raw = m.get("model") or "?", m.get("usage") + if not raw or model == "<synthetic>": + continue + key = (m.get("id"), d.get("requestId")) + first_ts.setdefault(key, d.get("timestamp") or "") + last[key] = (model, raw) + if m.get("stop_reason"): + final.add(key) + uid = d.get("uuid") + if uid not in seen and isinstance(m.get("content"), list): + blocks.setdefault(key, []).extend(m["content"]) + if uid: + seen.add(uid) + out: dict[str, Usage] = {} + for key, (model, raw) in last.items(): + ts = first_ts[key] + if (since and ts < since) or (until and ts >= until): + continue + u = _usage(raw) + if key not in final: + guess = out_estimate(blocks.get(key, [])) + if guess > u.out: + u.out, u.est = guess, 1 + out.setdefault(model, Usage()).add(u) + return out + + +def context_tokens(text: str) -> int | None: + """Prompt size of the last main-thread request (input + cache write + cache read); None = no request.""" + size = None + for line in text.split("\n"): + try: + d = json.loads(line) + except ValueError: + continue + if not isinstance(d, dict) or d.get("type") != "assistant" or d.get("isSidechain"): + continue + m = d.get("message") or {} + raw = m.get("usage") + if raw and m.get("model") != "<synthetic>": + u = _usage(raw) + size = u.inp + u.cw + u.cr + return size + + +def ctx_hint(tokens: int | None, limit: int) -> str | None: + if not limit or tokens is None or tokens <= limit: + return None + return (f"context ~{tokens // 1000}k tokens (> {limit // 1000}k): ask the owner to /clear, then continue " + "(subagent: ignore, this is the main session)") + + +def short(model: str) -> str: + return re.sub(r"-\d{8}$", "", model.removeprefix("claude-")) + + +def fmt(n: int) -> str: + if n < 1000: + return str(n) + return f"{n / 1e3:.1f}k" if n < 1e6 else f"{n / 1e6:.2f}M" + + +# ------------------------------------------------------------------ cost log (out/wf-cost.log) + +def log_line(time: str, project: str, task: str, effort: str, outcome: str, agent: str, + by_model: dict[str, Usage], lane: str | None = None, dur: int | None = None) -> str: + """One key=value line per worker; model = the model with most turns, lane = given or model; usd = priced models only.""" + total = Usage() + for u in by_model.values(): + total.add(u) + model = short(max(by_model, key=lambda m: by_model[m].turns)).split("-")[0] if by_model else "?" + usd = sum(cost(m, u) or 0 for m, u in by_model.items()) + return (f"{time} project={project} task={task} lane={lane or model} model={model} effort={effort} outcome={outcome} " + f"turns={total.turns} in={total.inp} cw={total.cw} cr={total.cr} out={total.out} " + + (f"est={total.est} " if total.est else "") + f"usd={usd:.4f} " + (f"dur={dur} " if dur is not None else "") + f"agent={agent}") + + +def parse_log(text: str, since: str | None = None) -> list[dict]: + out = [] + for line in text.split("\n"): + time, _, rest = line.partition(" ") + fields_ = dict(f.split("=", 1) for f in rest.split() if "=" in f) + if not re.match(r"\d{4}-\d\d-\d\dT", time) or "lane" not in fields_ or (since and time < since): + continue + try: + out.append({**fields_, "time": time, "usd": float(fields_.get("usd", 0)), + "turns": int(fields_.get("turns", 0)), + **({"dur": int(fields_["dur"])} if "dur" in fields_ else {})}) + except ValueError: + continue + return out + + +def _num(x: float) -> float | int: + x = round(x, 4) + return int(x) if x == int(x) and not isinstance(x, bool) else x + + +def report(entries: list[dict], efforts: tuple[str, ...]) -> list[tuple]: + """Per lane, then per lane+effort: (lane, effort, n, done, other, median $, total $, $/done, median turns, median dur s or None).""" + def row(lane, effort, es): + done = sum(e.get("outcome") in ("done", "done+gate-red") for e in es) + total = sum(e["usd"] for e in es) + return (lane, effort, len(es), done, len(es) - done, round(median(e["usd"] for e in es), 4), + round(total, 4), round(total / done, 4) if done else None, _num(median(e["turns"] for e in es)), + _num(median(durs)) if (durs := [e["dur"] for e in es if "dur" in e]) else None) + + order = {e: i for i, e in enumerate(efforts)} + rows = [] + key = lambda e: f"{e['lane']}/{e['model']}" if "model" in e and e["model"] != e["lane"] else e["lane"] + for lane in sorted({key(e) for e in entries}): + es = [e for e in entries if key(e) == lane] + rows.append(row(lane, "all", es)) + for effort in sorted({e.get("effort", "-") for e in es}, key=lambda x: (order.get(x, len(order)), x)): + rows.append(row(lane, effort, [e for e in es if e.get("effort", "-") == effort])) + return rows + + +# --- explore: where an agent's billed input goes (exploring vs editing/running/bookkeeping) --------------------- +_MUT = re.compile(r"sed -i|python3 - <<|cat > |tee |expect-init|>> [a-zA-Z]|apply_patch|patch ") +_RUN = re.compile(r"dotnet (test|build|run)|own-fx\.py (check|measure|dump)|npx |npm |gate-bg|check\.sh|wf res run|pytest|unittest") +_BOOK = re.compile(r"wf\.py (done|finish|add|status|note|set|merge|check|tick|body)" + r"|git (commit|add|switch|rebase|merge|push)|git worktree (add|remove)|git branch -[dDmM]") +_FILE = re.compile(r"(?<![\w.-])((?:\./|/)?[\w.-]+(?:/[\w.-]+)+\.\w+)(?![\w/-])") +_RESULT_KINDS = ( # first match wins; Bash result kinds + ("git-hist", re.compile(r"git (show|log|diff)")), + ("wf", re.compile(r"wf\.py")), + ("build/test", re.compile(r"dotnet|npx|npm")), + ("grep", re.compile(r"\bgrep|\brg\b|find ")), + ("sed-cat", re.compile(r"sed -n|cat |head|tail")), +) + + +def call_kind(tool: dict) -> str: + """mut (edit/write) | book (wf, git writes) | run (build/test) | exp (everything else: exploring).""" + if tool.get("name") in ("Edit", "Write", "NotebookEdit"): + return "mut" + if tool.get("name") != "Bash": + return "exp" + c = (tool.get("input") or {}).get("command") or "" + return "book" if _BOOK.search(c) else "run" if _RUN.search(c) else "mut" if _MUT.search(c) else "exp" + + +def result_kind(tool: dict) -> str: + if tool.get("name") == "Bash": + c = (tool.get("input") or {}).get("command") or "" + return next((k for k, rx in _RESULT_KINDS if rx.search(c)), "other") + return tool.get("name") or "other" + + +def tool_files(tool: dict, root: str = "") -> list[str]: + """Files a Read / Bash call touches, relative to the project (.worktrees/<x>/ and root stripped).""" + inp = tool.get("input") or {} + if tool.get("name") == "Read": + paths = [inp.get("file_path") or ""] + elif tool.get("name") == "Bash": + paths = [p for p in _FILE.findall(inp.get("command") or "") if not p.startswith("/") or p.startswith(root + "/")] + else: + return [] + out = [] + for p in paths: + p = re.sub(r"^.*?/\.worktrees/[^/]+/", "", p) + if root and p.startswith(root.rstrip("/") + "/"): + p = p[len(root.rstrip("/")) + 1:] + if p and p not in out: + out.append(p) + return out + + +def _ctx(u: dict) -> int: + return (u.get("input_tokens") or 0) + (u.get("cache_read_input_tokens") or 0) + (u.get("cache_creation_input_tokens") or 0) + + +@dataclass +class Explore: + calls: int = 0 + first_edit: int | None = None # calls before the first edit; None = never edited + cost: dict = None # kind → billed input tokens + pre: int = 0 # billed input before the first edit + ctx_growth: int = 0 # context growth up to the first edit + results: dict = None # result kind → tokens + files: dict = None # file → result tokens (read by this agent) + + @property + def total(self) -> int: + return sum(self.cost.values()) + + @property + def edited(self) -> bool: + return self.first_edit is not None + + +def explore(text: str, since: str | None = None, root: str = "") -> Explore: + """One transcript → billed-input split by call kind, result tokens per tool kind, files read.""" + calls: dict[str, dict] = {} + for line in text.split("\n"): + try: + d = json.loads(line) + except ValueError: + continue + if not isinstance(d, dict) or d.get("type") != "assistant": + continue + m = d.get("message") or {} + key = m.get("id") or d.get("uuid") + c = calls.setdefault(key, {"u": {}, "t": [], "ts": d.get("timestamp") or ""}) + c["u"] = m.get("usage") or c["u"] + c["t"] += [t for t in m.get("content") or [] if isinstance(t, dict) and t.get("type") == "tool_use"] + kept = [c for c in calls.values() if not since or c["ts"] >= since] + tools = {t["id"]: t for c in kept for t in c["t"] if "id" in t} + ex = Explore(calls=len(kept), cost={}, results={}, files={}) + for i, c in enumerate(kept): + ks = [call_kind(t) for t in c["t"]] or ["exp"] + k = "mut" if "mut" in ks else ks[0] + if k == "mut" and ex.first_edit is None: + ex.first_edit = i + ex.pre = sum(_ctx(x["u"]) for x in kept[:i]) + ex.ctx_growth = _ctx(c["u"]) - _ctx(kept[0]["u"]) + ex.cost[k] = ex.cost.get(k, 0) + _ctx(c["u"]) + for line in text.split("\n"): + try: + d = json.loads(line) + except ValueError: + continue + content = (d.get("message") or {}).get("content") if isinstance(d, dict) and d.get("type") == "user" else None + for c in content if isinstance(content, list) else []: + if not isinstance(c, dict) or c.get("type") != "tool_result" or c.get("tool_use_id") not in tools: + continue + body = c.get("content") + size = len(body if isinstance(body, str) else json.dumps(body)) // 4 + t = tools[c["tool_use_id"]] + k = result_kind(t) + ex.results[k] = ex.results.get(k, 0) + size + for f in tool_files(t, root): + ex.files[f] = ex.files.get(f, 0) + size + return ex + + +def explore_report(agents: list[tuple[str, Explore]], min_calls: int = 4, min_pre: int = 10, top: int = 20) -> dict: + """agents: (label, Explore). Agents under min_calls are skipped; the pre-edit median uses agents with an + edit and >= min_pre calls. Returns rows, totals, results, files (≥2 agents: (file, agents, tokens)).""" + use = [(l, e) for l, e in agents if e.calls >= min_calls] + rows = [] + for label, e in use: + t = e.total or 1 + rows.append((label, e.calls, e.first_edit if e.edited else e.calls, e.cost.get("exp", 0) / t, e.pre / t)) + cost: dict[str, int] = {} + results: dict[str, int] = {} + for _, e in use: + for k, v in e.cost.items(): + cost[k] = cost.get(k, 0) + v + for k, v in e.results.items(): + results[k] = results.get(k, 0) + v + pre = [e for _, e in use if e.edited and e.calls >= min_pre] + seen: dict[str, list[int]] = {} + for _, e in use: + for f, n in e.files.items(): + s = seen.setdefault(f, [0, 0]) + s[0] += 1 + s[1] += n + files = sorted(((f, a, n) for f, (a, n) in seen.items() if a >= 2), key=lambda x: (-x[1], -x[2], x[0]))[:top] + total = sum(cost.values()) + return { + "rows": rows, "agents": len(use), "total": total, + "cost": cost, "results": results, "files": files, + "pre_n": len(pre), + "pre_share": median(e.pre / (e.total or 1) for e in pre) if pre else None, + "pre_calls": median(e.first_edit for e in pre) if pre else None, + "pre_ctx": median(e.ctx_growth for e in pre) if pre else None, + } |
