diff options
Diffstat (limited to 'docs/resource-ledger.md')
| -rw-r--r-- | docs/resource-ledger.md | 304 |
1 files changed, 304 insertions, 0 deletions
diff --git a/docs/resource-ledger.md b/docs/resource-ledger.md new file mode 100644 index 0000000..17a0443 --- /dev/null +++ b/docs/resource-ledger.md @@ -0,0 +1,304 @@ +# wf res — shared resource ledger for agents — design + +Status: approved and implemented 2026-10-01. + +## 1. Goal + +1. No agent job is OOM-killed because another agent started a big job at the same time. +2. The user keeps a guaranteed share of RAM and CPU (gaming), switchable on demand. +3. An agent that cannot get resources now gets one clear line (who holds them, until when) and works on + something else instead of waiting or retrying blindly. +4. Agent-made waste (tmpfs scratch, failed units) is cleaned routinely. +5. Python ≥ 3.11 stdlib, Linux systemd user manager, no daemon. Works in any directory, wf project or not. + +Machine facts (2026-10-01): 16 cores, 30 GB RAM, 7 GB swap, `/tmp` is tmpfs (16 GB), user manager +delegates `cpu io memory pids`; `systemd-run --user --scope --slice=agents.slice` and transient units with +`MemoryMax`/`CPUWeight` verified working. + +## 2. User decisions (2026-10-01) + +| # | Decision | +|---|----------| +| D1 | Default user reserve 6 GB RAM, 4 CPUs. | +| D2 | Gaming reserve 12 GB RAM, 8 CPUs. | +| D3 | Gaming mode slows agents; never freezes or kills them. | +| D4 | Scratch under `/tmp/claude-<uid>/` untouched for 2 h (newest mtime in the entry) may be deleted automatically by any agent, whoever made it. Except a live session's own dir (`<project>/<session-id>`, live = `~/.claude/sessions/<pid>.json` with that `sessionId`, pid alive, `procStart` matching; added 2026-10-02 after an idle session lost its scratchpad). Same automatic clean sweeps top-level litter in `/tmp` and `/var/tmp` (own entries only): empty dirs idle 2 h (`scratch_hours`) and `clr-debug-pipe-<pid>-*` / `dotnet-diagnostic-<pid>-*` of dead pids; entries with content are never swept. Per-project `out/` patterns only list unless `--yes`. | +| D5 | v1 = everything in the proposal: run, status, wait, release, note, queue, game, clean, timer, shared rule. | +| D6 | `game on [--for D]`, default 4 h, turns itself off. | +| D7 | Install a 1-minute user timer via an explicit `wf res timer on`. | +| D8 | `game on` never fails and never squeezes running jobs; it switches budget and weights at once and prints the shortfall and the estimated time until the full reserve is free. | +| D9 | Small unreserved work: headroom (2 GB) in the budget **and** all agent processes inside `agents.slice`: running Claude sessions are moved there live by `wf res adopt` (run by every timer tick; no restart), new ones optionally start there via the `shell-init` alias. | + +## 3. Architecture + +### 3.1 Cgroup layout + +``` +user@<uid>.service +└── agents.slice ← MemoryHigh, CPUWeight, IOWeight (normal / gaming) + ├── run-…scope ← each Claude session (started via the alias), its shells, small jobs + └── agents-jobs.slice + └── wf-r-N.service ← each `wf res run` job: MemoryMax, MemoryHigh, MemorySwapMax=0, Nice=10 +``` + +- `agents.slice` caps the sum of everything agents do. The ledger decides who may start big work; the slice + is the kernel backstop for wrong estimates and unreserved small work. +- Slice properties are set with `systemctl --user set-property --runtime agents.slice …`. The slice unit file + `~/.config/systemd/user/agents.slice` is written on first use (then `daemon-reload`), so the slice exists + before any set-property. +- Normal: `CPUWeight=20`, `IOWeight=20`, `MemoryHigh = RAM − user_reserve_gb`. +- Gaming: `CPUWeight=5`, `IOWeight=5`, `MemoryHigh = max(RAM − game_reserve_gb, agents.slice MemoryCurrent)`; + re-lowered towards the target on every `wf res` call / tick (D8: no squeezing below current use). +- Weights only matter under contention; on an idle machine agents get full speed. + +### 3.2 Code layout + +- `wflib/res.py` — pure, text in / data out: parse `/proc/meminfo`, parse `systemctl show` output, durations + and sizes (`10G`, `40m`, `4h`), prune, capacity rule, queue start order, game slice values and shortfall, + stale-scratch selection (from a given list of (path, newest mtime)), all output lines. +- `wf_res.py` — I/O: lock, atomic ledger write, `/proc` and `/sys/fs/cgroup` reads, filesystem walk, the + runner (`subprocess.run` wrapper, injectable for tests), unit/timer file writing, printing. +- `wf.py` — only registers `wf res …` and dispatches to `wf_res.main(argv)`. `wf res` needs no + `workflow.toml`. + +### 3.3 Files + +- State dir `~/.local/state/wf/` (`$XDG_STATE_HOME/wf` if set): `resources.json`, `resources.lock`, + `logs/r-N.log`, `logs/r-N.rc`, `logs/r-N.peak`, `resources-history.jsonl` (finished rc=0 runs with a peak, + appended when pruned from the ledger after 24 h; newest 2000 kept; read by `wf res hist`). +- Config `~/.config/wf/resources.toml` (`$XDG_CONFIG_HOME`), all keys optional: + ```toml + user_reserve_gb = 6 + user_reserve_cpus = 4 + game_reserve_gb = 12 + game_reserve_cpus = 8 + game_hours = 4 + small_headroom_gb = 2 + scratch_hours = 2 + ``` +- Unit files under `~/.config/systemd/user/`: `agents.slice`, and with `timer on`: `wf-res.service`, + `wf-res.timer`. + +### 3.4 Ledger + +`resources.json`: `{"next": N, "game_until": iso|null, "last_clean": iso|null, "entries": [...]}`. + +Entry fields: `id` (`r-N`, never reused), `project` (git toplevel basename, else cwd basename), `owner` (pid of +the nearest ancestor process named `claude`, else the parent pid of `wf`), `title`, `mem_gb`, `cpus` +(default 1), `est_min`, `state` (`queued` | `running` | `note` | `done`), `unit` (`wf-r-N.service`, run only), +`cmd` (argv list), `cwd`, `log`, `queued` (iso), `started` (iso), `expires` (iso = started + 2 × est_min), +and for `done`: `ended`, `rc` (int or null), `peak_gb`, `why` (`exited` | `expired` | `owner gone` | `released` | a kill reason, e.g. +`killed: oom-kill by systemd-oomd, limit 4.0 GB; raise --mem`). +`env` (caller variables the user manager lacks; `WF_RES_ID` and secrets never), `by` (who to tell, run/note; +keys dropped when unknown): `name` (`--by`, else `WF_SESSION_NAME`, else `<lane> session` from the wf lane +record `.wf/sessions/*.json` with this `CLAUDE_PID`, main tree or toplevel), `task` (`WF_TASK`, else the +`<lane>/<id>` branch of cwd), `batch` (`WF_RES_ID`: every job unit gets `WF_RES_ID=r-N`, so a `wf batch` +worker's entries name the batch), `address` (`uds:$CLAUDE_CODE_MESSAGING_SOCKET`; a batch worker's = its +orchestrator's, reachable via SendMessage). Status lines of live entries end `[by <name> <task> batch r-N, +message uds:…]`; the throttle warning ends `; started by …`. + +Every command takes the lock (`fcntl.flock` exclusive on `resources.lock`), reads, then **prunes**, acts, +writes atomically (temp file in the same dir + `os.replace`), releases. + +Prune (pure, given unit states, live pids, now): +- `running` whose unit is inactive/failed/unknown → `done` (`rc` from `logs/r-N.rc`, `peak_gb` from + `logs/r-N.peak`). No `.rc` (the sh wrapper died with the job) → `rc` null and `why`/`peak_gb` from + `journalctl --user -u wf-r-N.service --since @started` (`Failed with result '…'`, `systemd-oomd killed`, + `… memory peak`), shown in `status` and `wait`; journal silent → `why` points at that journalctl command. +- `running` past `expires` → stays running (job not killed; `run` says so on start, status shows + `ETA overdue (still running, not killed)`), flagged `overdue` in status, and counted at + `max(reserved, used)`. +- `note` whose owner pid is gone or past `expires` → `done`. +- `done` older than 24 h → removed. +- `game_until` in the past → game off (slice back to normal values). + +### 3.5 Capacity rule + +``` +budget_mem = MemAvailable − reserve_gb − small_headroom_gb + − Σ over running: max(0, mem_gb − MemoryCurrent(unit)) + − Σ over notes: mem_gb +budget_cpus = nproc − reserve_cpus − Σ over running and notes: cpus +fits(req) = req.mem_gb ≤ budget_mem and req.cpus ≤ budget_cpus +``` + +- `reserve_*` is the user or the gaming value depending on game mode. +- A note counts its full reservation (its memory cannot be measured; overcounting briefly is safe). +- A running job's memory already in use is inside `MemAvailable`, so only its unused part is subtracted. + +## 4. Commands + +All print short lines; `--json` where stated. Exit: 0 ok · 3 busy (refused) · 1 `wf: …` one line (unknown +id, systemd failure) · 2 usage. + +### 4.1 `wf res run --mem 10G [--cpus N] --for 40m --title "…" [--queue] -- <cmd …>` + +- Fits → `systemd-run --user --slice=agents-jobs.slice --unit=wf-r-N --collect + -p MemoryMax=<mem> -p MemoryHigh=<0.9·mem> -p MemorySwapMax=0 -p Nice=10 + -p WorkingDirectory=<cwd> -p StandardOutput=append:<log> -p StandardError=append:<log> + /bin/sh -c '"$@"; rc=$?; cat /sys/fs/cgroup$(cut -d: -f3 /proc/self/cgroup)/memory.peak > <peak>; echo $rc > <rc>' sh <cmd …>` + (the unit is collected on exit, so the wrapper records exit code and peak bytes itself). Prints `r-N started; log <path>; ETA ~HH:MM`. Exit 0. +- Environment: a user unit starts from the user manager's environment, so the caller's variables that differ + from `systemctl --user show-environment` go along as `--setenv=NAME=VALUE` (stored in the entry, so a queued + job gets them too); skipped: `PWD OLDPWD SHLVL _` and names with TOKEN/SECRET/PASSWORD/PASSWD/CREDENTIAL. +- Does not fit, no `--queue` → exit 3, one line: + `busy: 7.5 GB held by proj-a "dotnet e2e" (r-4) until ~14:40; 2.1 GB free for agents; retry after ~14:40 or work on something else` + (names the holders whose release would make it fit, earliest ETA first; CPU shortage named the same way). +- Does not fit, `--queue` → `queued`, prints `r-N queued, position P; est. start ~HH:MM; cancel: wf res release r-N`. Exit 0. +- A request larger than the whole agent budget on an empty machine → exit 1 `wf: 20 GB can never fit (max …)`. +- `--force` (memory really free though the ledger says busy: unused reservations, stale notes): fits against + MemAvailable − user reserve (game reserve in game mode) − headroom only; ledger claims and the queue are ignored; + the entry is still recorded. Does not fit even so → exit 3 `busy even with --force: only X GB really free beyond + the reserve; …`. The normal busy line ends with `; X GB really free beyond the reserve: --force starts it past + the ledger (only if the holders will not use what they reserved)` when `--force` would fit. +- `systemd-run` failure → entry removed, exit 1 with its first stderr line. +- `--lock KEY`: one queued/running job per KEY and main tree (git common dir's parent: lane worktrees share it) + at a time, e.g. jobs sharing one checkout (a-game-project-builder `../.<project>-gate`). A title starting with the word + `gate` locks `gate` unasked (2026-10-06: two hand-started gates overlapped in one gate checkout). Held → exit 3 + `busy: lock 'KEY' held by r-N "title" (ETA ~HH:MM|queued): …; --queue waits for it (--force does not override + a lock)`; `--queue` → queued `(lock held by r-N)`. Ledger field `lock` = `KEY@<main tree>`. +- The command reaches the job verbatim (no systemd `%` specifier expansion: `--format='%h %s'` is safe). + +### 4.2 Queue + +Strict FIFO by `queued`: the head starts when it fits; entries behind a waiting head wait (a big job is never +starved); an entry whose lock a running job holds is skipped (it neither starts nor blocks the rest). Starting happens inside every `wf res` call (after prune) and every timer tick. A started queued job +runs with the cwd and argv it was queued with; its ETA counts from the actual start. + +### 4.3 `wf res status [r-N] [--json]` + +Without id: one line per entry (id, project, title, state, reserved / used / peak GB, cpus, started, ETA or +`overdue`), then the budget line (`agents may use X GB, Y cpus now; reserve 6 GB/4 cpus`), game mode with +time left and shortfall line (§4.6), unreserved agent memory (`agents.slice` MemoryCurrent − jobs; `(X GB of it file cache, reclaimable)` = +slice `memory.stat` `file` − `shmem`, capped at that figure), warnings: +- running job throttled at its own limit: used ≥ 0.85 × `--mem` **and** its cgroup `memory.pressure` + `some avg60` ≥ 20 % → `r-N throttled at its memory limit (…): likely too small; release --stop, re-run + with a bigger --mem` (page cache alone at the limit does not stall, so no false alarm; the timer tick + has no reader, so it does not warn); +- tmpfs `/tmp` above 25 % of RAM; +- `N claude sessions outside agents.slice (wf res adopt)`. +- `N jobs outside wf res (bare systemd-run): <unit> <used> GB, …`: running transient user services outside + `agents*` slices, except desktop ones (`app-*`, `dbus-*`). Warning only: a service cannot change slice live. + +With id: that entry in full, including `rc`, `peak_gb`, `why` for done entries. + +### 4.4 `wf res wait r-N [--timeout 2h]` + +Polls every 15 s (each poll is a normal locked call, so it also prunes and starts queued work) until the entry +is `done`; prints `r-N done rc=0 peak 9.4 GB in 37 min`. The throttle warning (§4.3) goes to stderr once. Exit 0 when done (whatever `rc`), 1 on timeout. If +reaped, `status r-N` answers from the ledger. + +### 4.5 `wf res release r-N [--stop]` + +- `queued` / `note` → removed. Running unit → refused with `wf: r-N is running; --stop to kill it` unless + `--stop`, then `systemctl --user stop`, entry → done (`why=released`). + +### 4.6 `wf res game on [--for 4h] | off` + +- `on`: `game_until = now + for`; set gaming slice values (§3.1); prints + ``` + game on until 22:40 (4h); CPU/IO now yours + short 3.1 GB of 12 GB: r-4 proj-a "dotnet e2e" 2.5 GB ~21:05, r-7 proj-b "r5 rebuild" 9 GB ~21:30 + full reserve free ~21:30 (est.); free now: wf res release r-7 --stop + ``` + The shortfall/estimate lines appear only when the reserve is not free now. Estimate = ETAs of the running + jobs that must end, in order (jobs are never stopped by game mode). New requests that do not fit are refused + or queued as usual. `on` while on → extends `game_until`. +- `off` or expiry: normal slice values, `game_until = null`. +- No windows over the game: `wf res hook` (Claude Code `PreToolUse` hook, matcher `Bash`, in + `~/.claude/settings.json`; the user adds it, wf never edits dotfiles) prefixes every agent Bash command with + `unset DISPLAY WAYLAND_DISPLAY;` while `game_until > now`, and adds context telling the agent why GUI apps + fail (run headless/offscreen or later). Reads the ledger only (no lock, no systemctl, ~60 ms); off, expired, + other tools or any error → no output, command unchanged. Live for running sessions: state is read per call. + ```json + "hooks": {"PreToolUse": [{"matcher": "Bash", "hooks": [{"type": "command", + "command": "python3 /projects/public/workflow/wf_res.py hook"}]}]} + ``` + +### 4.7 `wf res note --mem 3G [--cpus N] --for 20m [--force] "title"` + +Reserves for foreground work (no unit). Freed when `owner` exits, at `expires`, or by `release`. Refused like +`run` (exit 3) when it does not fit; `--force` as in `run`. + +### 4.8 `wf res clean [--yes]` + +- Automatic part (no `--yes` needed; also run by any `wf res` call when `last_clean` > 10 min ago, and by every + tick): + - every direct and second-level entry under `/tmp/claude-<uid>/` whose newest mtime (recursive) is older than + `scratch_hours` → deleted; lines `freed 1.2 GB: /tmp/claude-1000/-projects-x/<session>` (only printed by + `clean` itself; silent from other commands); + - `systemctl --user reset-failed 'wf-r-*'`. +- Listed only, deleted with `--yes`: per-project patterns from the current project's `workflow.toml` + `cleanup = ["out/prof", "out/history-logs/*.log:30d"]` (glob relative to the project root, optional `:Nd` + minimum age by mtime). Never follows symlinks; never leaves the project root. + +### 4.9 `wf res timer on | off`, `wf res tick` + +- `timer on` writes `wf-res.service` (`ExecStart=python3 /projects/public/workflow/wf.py res tick`) and + `wf-res.timer` (`OnBootSec=1min`, `OnUnitActiveSec=1min`), `daemon-reload`, `enable --now wf-res.timer`. + `off` disables and removes them. +- `tick` = lock, prune, start queued, game expiry and slice re-lowering, automatic clean; then adopt., then + session caps: every running scope in `agents.slice` (a Claude session) gets `MemoryHigh = session_mem_gb` + (config, default 6, 0 = off → `infinity`) via `systemctl --user set-property --runtime`, only where it differs. + MemoryHigh only: a session over it is throttled, never killed (a MemoryMax OOM kill could pick `claude`, whose + oom_score_adj is 200). `status` warns `claude session N at its memory cap (…)` at ≥ 0.9 × cap and ≥ 20 % stall. + `wf res run` jobs live in `agents-jobs.slice`, not in a session, so the cap never limits them. Silent. + +### 4.10 `wf res adopt` + +Moves running Claude sessions into `agents.slice` without a restart (verified 2026-10-01: a live process moved +by `StartTransientUnit` with `PIDs`). Session roots = processes named `claude` whose parent is not `claude`, +outside the slice; each root and its descendants outside the slice go into a new scope +`wf-claude-<root pid>.scope` via +`busctl --user call org.freedesktop.systemd1 /org/freedesktop/systemd1 org.freedesktop.systemd1.Manager +StartTransientUnit 'ssa(sv)a(sa(sv))' wf-claude-<pid>.scope fail 2 PIDs au <n> <pids…> Slice s agents.slice 0`. +Children forked later inherit the scope. Prints one line per session; errors (process gone) are reported, and +the timer retries next minute. + +### 4.11 `wf res shell-init` + +Prints the alias line for `~/.bashrc` (the user adds it; wf never edits dotfiles): +`alias claude='systemd-run --user --scope --quiet --slice=agents.slice claude'`. + +### 4.12 `wf res hist [--project P]` + +Reservation sizes from history: history file + the ledger's finished runs (rc=0, peak known). Grouped by +project and title kind = title minus trailing `(…)` and trailing sha/hex words (7–40 hex chars, ≥ 1 digit): +`gate 2c91cac7 (fix x)` → `gate`. One line per kind: `n`, median request, peak p50/p95, median estimate, +duration p50/p90, suggestion. Suggestion (≥ 3 runs): mem = p95 peak ×1.15 (up to 0.1 GB, ≥ 0.2 GB), for = p90 +duration ×1.5 (≥ 5 min); percentiles nearest-rank. `run`/`note` asking > 2× the suggestion (mem or for) print +`hint: history says ~X GB / Y min (n runs)` on stderr; nothing is resized. `wf batch` without `--mem`/`--for` +uses the `wf-batch` suggestion of the project (not `--prep`; the bg-wait ceiling stays 6h unless `--for`). + +## 5. Shared rule (`shared/CLAUDE.md`, ~4 lines) + +- Job expected > 2 GB RAM, > 4 cores or > 10 min → `wf res run` (never bare `systemd-run`, never a long + foreground command); big foreground step → `wf res note`. +- Exit 3 = busy: note it on the task, do other work, retry after the printed time. No tight polling. + The line names `--force` when the memory is really free (claims unused): rerun with it if the holders will not grow. +- No big scratch in `/tmp` (RAM); use the project's git-ignored `out/`. +- Session end: release your own entries (`wf res status`). + +## 6. Errors + +- Not Linux / no systemd user manager / no cgroup v2 → `wf: wf res needs a systemd user session` exit 1. +- Corrupt ledger → renamed to `resources.json.bad-<ts>`, start empty (id counter salvaged from its text: ids never reused), one warning line. Unknown entry keys (written by a newer version) are ignored, never corrupt: a `wf res wait` started before a release must not wipe the ledger. Starting a job deletes stale `logs/<id>.rc/.peak`. +- Lock is held at most for one command's critical section (no subprocess wait longer than a `systemctl` + call inside the lock; `wait` polls with the lock released between polls). + +## 7. Tests + +- Pure (`tests/test_res.py`, hand-written literals): meminfo parse; size/duration parse; capacity with mixed + running/note entries, used figures, headroom and game reserve; prune (unit gone, pid gone, expired note, + overdue run, 24 h done removal, game expiry); FIFO start order incl. blocked head; busy line holder choice and + wording; game slice values and shortfall/estimate lines; stale scratch selection; cleanup pattern + age. +- I/O (`tests/test_res_io.py`): fake runner records argv for run/stop/set-property/timer; atomic write; corrupt + ledger recovery; two real processes reserving at once under the real lock never overbook (budget faked). +- Smoke (manual, release step): real `wf res run --mem 100M --for 1m -- true`, `wait`, `status`. + +## 8. Release + +Slices (≈1 h each): pure core · ledger+lock+run/status/release · wait/rc/peak · queue · note · game+slice · +clean · timer+tick+shell-init · shared rule + CHANGES + release. Release rules of the workflow repo apply; +after release the workflow session runs `wf res timer on` once and tells the user to add the +`shell-init` alias. |
