workflow

git clone https://git.godosa.eu/workflow

master

raw · 22694 bytes

wf res — shared resource ledger for agents — design

Status: approved and implemented 2026-10-01.

1. Goal

  1. No agent job is OOM-killed because another agent started a big job at the same time.
  2. The user keeps a guaranteed share of RAM and CPU (gaming), switchable on demand.
  3. An agent that cannot get resources now gets one clear line (who holds them, until when) and works on something else instead of waiting or retrying blindly.
  4. Agent-made waste (tmpfs scratch, failed units) is cleaned routinely.
  5. Python ≥ 3.11 stdlib, Linux systemd user manager, no daemon. Works in any directory, wf project or not.

Machine facts (2026-10-01): 16 cores, 30 GB RAM, 7 GB swap, /tmp is tmpfs (16 GB), user manager delegates cpu io memory pids; systemd-run --user --scope --slice=agents.slice and transient units with MemoryMax/CPUWeight verified working.

2. User decisions (2026-10-01)

# Decision
D1 Default user reserve 6 GB RAM, 4 CPUs.
D2 Gaming reserve 12 GB RAM, 8 CPUs.
D3 Gaming mode slows agents; never freezes or kills them.
D4 Scratch under /tmp/claude-<uid>/ untouched for 2 h (newest mtime in the entry) may be deleted automatically by any agent, whoever made it. Except a live session's own dir (<project>/<session-id>, live = ~/.claude/sessions/<pid>.json with that sessionId, pid alive, procStart matching; added 2026-10-02 after an idle session lost its scratchpad). Same automatic clean sweeps top-level litter in /tmp and /var/tmp (own entries only): empty dirs idle 2 h (scratch_hours) and clr-debug-pipe-<pid>-* / dotnet-diagnostic-<pid>-* of dead pids; entries with content are never swept. Per-project out/ patterns only list unless --yes.
D5 v1 = everything in the proposal: run, status, wait, release, note, queue, game, clean, timer, shared rule.
D6 game on [--for D], default 4 h, turns itself off.
D7 Install a 1-minute user timer via an explicit wf res timer on.
D8 game on never fails and never squeezes running jobs; it switches budget and weights at once and prints the shortfall and the estimated time until the full reserve is free.
D9 Small unreserved work: headroom (2 GB) in the budget and all agent processes inside agents.slice: running Claude sessions are moved there live by wf res adopt (run by every timer tick; no restart), new ones optionally start there via the shell-init alias.

3. Architecture

3.1 Cgroup layout

user@<uid>.service
└── agents.slice              ← MemoryHigh, CPUWeight, IOWeight (normal / gaming)
    ├── run-…scope            ← each Claude session (started via the alias), its shells, small jobs
    └── agents-jobs.slice
        └── wf-r-N.service    ← each `wf res run` job: MemoryMax, MemoryHigh, MemorySwapMax=0, Nice=10
  • agents.slice caps the sum of everything agents do. The ledger decides who may start big work; the slice is the kernel backstop for wrong estimates and unreserved small work.
  • Slice properties are set with systemctl --user set-property --runtime agents.slice …. The slice unit file ~/.config/systemd/user/agents.slice is written on first use (then daemon-reload), so the slice exists before any set-property.
  • Normal: CPUWeight=20, IOWeight=20, MemoryHigh = RAM − user_reserve_gb.
  • Gaming: CPUWeight=5, IOWeight=5, MemoryHigh = max(RAM − game_reserve_gb, agents.slice MemoryCurrent); re-lowered towards the target on every wf res call / tick (D8: no squeezing below current use).
  • Weights only matter under contention; on an idle machine agents get full speed.

3.2 Code layout

  • wflib/res.py — pure, text in / data out: parse /proc/meminfo, parse systemctl show output, durations and sizes (10G, 40m, 4h), prune, capacity rule, queue start order, game slice values and shortfall, stale-scratch selection (from a given list of (path, newest mtime)), all output lines.
  • wf_res.py — I/O: lock, atomic ledger write, /proc and /sys/fs/cgroup reads, filesystem walk, the runner (subprocess.run wrapper, injectable for tests), unit/timer file writing, printing.
  • wf.py — only registers wf res … and dispatches to wf_res.main(argv). wf res needs no workflow.toml.

3.3 Files

  • State dir ~/.local/state/wf/ ($XDG_STATE_HOME/wf if set): resources.json, resources.lock, logs/r-N.log, logs/r-N.rc, logs/r-N.peak, resources-history.jsonl (finished rc=0 runs with a peak, appended when pruned from the ledger after 24 h; newest 2000 kept; read by wf res hist).
  • Config ~/.config/wf/resources.toml ($XDG_CONFIG_HOME), all keys optional: toml user_reserve_gb = 6 user_reserve_cpus = 4 game_reserve_gb = 12 game_reserve_cpus = 8 game_hours = 4 small_headroom_gb = 2 scratch_hours = 2
  • Unit files under ~/.config/systemd/user/: agents.slice, and with timer on: wf-res.service, wf-res.timer.

3.4 Ledger

resources.json: {"next": N, "game_until": iso|null, "last_clean": iso|null, "entries": [...]}.

Entry fields: id (r-N, never reused), project (git toplevel basename, else cwd basename), owner (pid of the nearest ancestor process named claude, else the parent pid of wf), title, mem_gb, cpus (default 1), est_min, state (queued | running | note | done), unit (wf-r-N.service, run only), cmd (argv list), cwd, log, queued (iso), started (iso), expires (iso = started + 2 × est_min), and for done: ended, rc (int or null), peak_gb, why (exited | expired | owner gone | released | a kill reason, e.g. killed: oom-kill by systemd-oomd, limit 4.0 GB; raise --mem). env (caller variables the user manager lacks; WF_RES_ID and secrets never), by (who to tell, run/note; keys dropped when unknown): name (--by, else WF_SESSION_NAME, else <lane> session from the wf lane record .wf/sessions/*.json with this CLAUDE_PID, main tree or toplevel), task (WF_TASK, else the <lane>/<id> branch of cwd), batch (WF_RES_ID: every job unit gets WF_RES_ID=r-N, so a wf batch worker's entries name the batch), address (uds:$CLAUDE_CODE_MESSAGING_SOCKET; a batch worker's = its orchestrator's, reachable via SendMessage). Status lines of live entries end [by <name> <task> batch r-N, message uds:…]; the throttle warning ends ; started by ….

Every command takes the lock (fcntl.flock exclusive on resources.lock), reads, then prunes, acts, writes atomically (temp file in the same dir + os.replace), releases.

Prune (pure, given unit states, live pids, now): - running whose unit is inactive/failed/unknown → done (rc from logs/r-N.rc, peak_gb from logs/r-N.peak). No .rc (the sh wrapper died with the job) → rc null and why/peak_gb from journalctl --user -u wf-r-N.service --since @started (Failed with result '…', systemd-oomd killed, … memory peak), shown in status and wait; journal silent → why points at that journalctl command. - running past expires → stays running (job not killed; run says so on start, status shows ETA overdue (still running, not killed)), flagged overdue in status, and counted at max(reserved, used). - note whose owner pid is gone or past expires → done. - done older than 24 h → removed. - game_until in the past → game off (slice back to normal values).

3.5 Capacity rule

budget_mem  = MemAvailable − reserve_gb − small_headroom_gb
              − Σ over running: max(0, mem_gb − MemoryCurrent(unit))
              − Σ over notes:   mem_gb
budget_cpus = nproc − reserve_cpus − Σ over running and notes: cpus
fits(req)   = req.mem_gb ≤ budget_mem and req.cpus ≤ budget_cpus
  • reserve_* is the user or the gaming value depending on game mode.
  • A note counts its full reservation (its memory cannot be measured; overcounting briefly is safe).
  • A running job's memory already in use is inside MemAvailable, so only its unused part is subtracted.

4. Commands

All print short lines; --json where stated. Exit: 0 ok · 3 busy (refused) · 1 wf: … one line (unknown id, systemd failure) · 2 usage.

4.1 wf res run --mem 10G [--cpus N] --for 40m --title "…" [--queue] -- <cmd …>

  • Fits → systemd-run --user --slice=agents-jobs.slice --unit=wf-r-N --collect -p MemoryMax=<mem> -p MemoryHigh=<0.9·mem> -p MemorySwapMax=0 -p Nice=10 -p WorkingDirectory=<cwd> -p StandardOutput=append:<log> -p StandardError=append:<log> /bin/sh -c '"$@"; rc=$?; cat /sys/fs/cgroup$(cut -d: -f3 /proc/self/cgroup)/memory.peak > <peak>; echo $rc > <rc>' sh <cmd …> (the unit is collected on exit, so the wrapper records exit code and peak bytes itself). Prints r-N started; log <path>; ETA ~HH:MM. Exit 0.
  • Environment: a user unit starts from the user manager's environment, so the caller's variables that differ from systemctl --user show-environment go along as --setenv=NAME=VALUE (stored in the entry, so a queued job gets them too); skipped: PWD OLDPWD SHLVL _ and names with TOKEN/SECRET/PASSWORD/PASSWD/CREDENTIAL.
  • Does not fit, no --queue → exit 3, one line: busy: 7.5 GB held by proj-a "dotnet e2e" (r-4) until ~14:40; 2.1 GB free for agents; retry after ~14:40 or work on something else (names the holders whose release would make it fit, earliest ETA first; CPU shortage named the same way).
  • Does not fit, --queue → queued, prints r-N queued, position P; est. start ~HH:MM; cancel: wf res release r-N. Exit 0.
  • A request larger than the whole agent budget on an empty machine → exit 1 wf: 20 GB can never fit (max …).
  • --force (memory really free though the ledger says busy: unused reservations, stale notes): fits against MemAvailable − user reserve (game reserve in game mode) − headroom only; ledger claims and the queue are ignored; the entry is still recorded. Does not fit even so → exit 3 busy even with --force: only X GB really free beyond the reserve; …. The normal busy line ends with ; X GB really free beyond the reserve: --force starts it past the ledger (only if the holders will not use what they reserved) when --force would fit.
  • systemd-run failure → entry removed, exit 1 with its first stderr line.
  • --lock KEY: one queued/running job per KEY and main tree (git common dir's parent: lane worktrees share it) at a time, e.g. jobs sharing one checkout (a-game-project-builder ../.<project>-gate). A title starting with the word gate locks gate unasked (2026-10-06: two hand-started gates overlapped in one gate checkout). Held → exit 3 busy: lock 'KEY' held by r-N "title" (ETA ~HH:MM|queued): …; --queue waits for it (--force does not override a lock); --queue → queued (lock held by r-N). Ledger field lock = KEY@<main tree>.
  • --commit REV: the commit the job tests (ledger commit, full sha); default for a locked job: a 7–40 hex word of the title that resolves to a commit in the cwd (gate 1a2b3c4). Coalescing, §4.2.
  • The command reaches the job verbatim (no systemd % specifier expansion: --format='%h %s' is safe).

4.2 Queue

Strict FIFO by queued: the head starts when it fits; entries behind a waiting head wait (a big job is never starved); an entry whose lock a running job holds is skipped (it neither starts nor blocks the rest). Starting happens inside every wf res call (after prune) and every timer tick. A started queued job runs with the cwd and argv it was queued with; its ETA counts from the actual start.

Batch loan: a job started from inside a live job (by.batch = its WF_RES_ID, e.g. a wf batch worker's gate) fits against budget + the parent's claim (mem and cpus) minus what the parent's other running children hold, so a batch waiting for its own gate never deadlocks (direct run and queue alike; not with --force). A queued entry whose parent job ended (done or gone) is dropped at prune (r-N dropped (batch r-M gone), why owner gone): nobody waits for it, and it must not head the queue.

Coalescing (ruling 2026-10-07, t-coalesce-batch-gates; chosen over "batch queues one gate at its end": works for every caller, batch or not, needs no batch-end hook): when a locked job with a commit is queued, same-lock queued entries are compared by ancestry (git merge-base --is-ancestor in the cwd). Queued entries on ancestor commits end superseded by r-N (ledger superseded_by); the new entry takes the earliest queued time of those (keeps its turn) and lists their commits in covers (oldest queued first). A queued entry on the same or a descendant commit already exists → nothing is queued: r-M queued, position … ; covers <sha7> (queued on descendant <sha7>, nothing new queued): wf res wait r-M (r-M gains the commit in covers). Running entries and unrelated commits (side branches) are never touched. So a batch of N commits yields ≤ 1 queued gate per lock. wf res wait on a superseded id prints r-N superseded by r-M …; waiting for it and follows the chain. Done line of a coalesced entry: …; covers a,b; red (rc ≠ 0): ; red: culprit is any commit in <oldest>^..<sha> -> bisect (git bisect start <sha> <oldest>^, rerun the job per step) or name that range in the P0 fix task.

4.3 wf res status [r-N] [--json]

Without id: one line per entry (id, project, title, state, reserved / used / peak GB, cpus, started, ETA or overdue), then the budget line (agents may use X GB, Y cpus now; reserve 6 GB/4 cpus), game mode with time left and shortfall line (§4.6), unreserved agent memory (agents.slice MemoryCurrent − jobs; (X GB of it file cache, reclaimable) = slice memory.stat file − shmem, capped at that figure), warnings: - running job throttled at its own limit: used ≥ 0.85 × --mem and its cgroup memory.pressure some avg60 ≥ 20 % → r-N throttled at its memory limit (…): likely too small; release --stop, re-run with a bigger --mem (page cache alone at the limit does not stall, so no false alarm; the timer tick has no reader, so it does not warn); - tmpfs /tmp above 25 % of RAM; - N claude sessions outside agents.slice (wf res adopt). - N jobs outside wf res (bare systemd-run): <unit> <used> GB, …: running transient user services outside agents* slices, except desktop ones (app-*, dbus-*). Warning only: a service cannot change slice live.

With id: that entry in full, including rc, peak_gb, why for done entries.

4.4 wf res wait r-N [--timeout 2h]

Polls every 15 s (each poll is a normal locked call, so it also prunes and starts queued work) until the entry is done; prints r-N done rc=0 peak 9.4 GB in 37 min (coalesced: superseded ids followed, §4.2). The throttle warning (§4.3) goes to stderr once. Exit 0 when done (whatever rc), 1 on timeout. If reaped, status r-N answers from the ledger.

4.5 wf res release r-N [--stop]

  • queued / note → removed. Running unit → refused with wf: r-N is running; --stop to kill it unless --stop, then systemctl --user stop, entry → done (why=released).

4.6 wf res game on [--for 4h] | off

  • on: game_until = now + for; set gaming slice values (§3.1); prints game on until 22:40 (4h); CPU/IO now yours short 3.1 GB of 12 GB: r-4 proj-a "dotnet e2e" 2.5 GB ~21:05, r-7 proj-b "r5 rebuild" 9 GB ~21:30 full reserve free ~21:30 (est.); free now: wf res release r-7 --stop The shortfall/estimate lines appear only when the reserve is not free now. Estimate = ETAs of the running jobs that must end, in order (jobs are never stopped by game mode). New requests that do not fit are refused or queued as usual. on while on → extends game_until.
  • off or expiry: normal slice values, game_until = null.
  • No windows over the game: wf res hook (Claude Code PreToolUse hook, matcher Bash, in ~/.claude/settings.json; the user adds it, wf never edits dotfiles) prefixes every agent Bash command with unset DISPLAY WAYLAND_DISPLAY; while game_until > now, and adds context telling the agent why GUI apps fail (run headless/offscreen or later). Reads the ledger only (no lock, no systemctl, ~60 ms); off, expired, other tools or any error → no output, command unchanged. Live for running sessions: state is read per call. json "hooks": {"PreToolUse": [{"matcher": "Bash", "hooks": [{"type": "command", "command": "python3 /projects/public/workflow/wf_res.py hook"}]}]}

4.7 wf res note --mem 3G [--cpus N] --for 20m [--force] "title"

Reserves for foreground work (no unit). Freed when owner exits, at expires, or by release. Refused like run (exit 3) when it does not fit; --force as in run.

4.8 wf res clean [--yes]

  • Automatic part (no --yes needed; also run by any wf res call when last_clean > 10 min ago, and by every tick):
  • every direct and second-level entry under /tmp/claude-<uid>/ whose newest mtime (recursive) is older than scratch_hours → deleted; lines freed 1.2 GB: /tmp/claude-1000/-projects-x/<session> (only printed by clean itself; silent from other commands);
  • systemctl --user reset-failed 'wf-r-*'.
  • Listed only, deleted with --yes: per-project patterns from the current project's workflow.toml cleanup = ["out/prof", "out/history-logs/*.log:30d"] (glob relative to the project root, optional :Nd minimum age by mtime). Never follows symlinks; never leaves the project root.

4.9 wf res timer on | off, wf res tick

  • timer on writes wf-res.service (ExecStart=python3 /projects/public/workflow/wf.py res tick) and wf-res.timer (OnBootSec=1min, OnUnitActiveSec=1min), daemon-reload, enable --now wf-res.timer. off disables and removes them.
  • tick = lock, prune, start queued, game expiry and slice re-lowering, automatic clean; then adopt., then session caps: every running scope in agents.slice (a Claude session) gets MemoryHigh = session_mem_gb (config, default 6, 0 = off → infinity) via systemctl --user set-property --runtime, only where it differs. MemoryHigh only: a session over it is throttled, never killed (a MemoryMax OOM kill could pick claude, whose oom_score_adj is 200). status warns claude session N at its memory cap (…) at ≥ 0.9 × cap and ≥ 20 % stall. wf res run jobs live in agents-jobs.slice, not in a session, so the cap never limits them. Silent.

4.10 wf res adopt

Moves running Claude sessions into agents.slice without a restart (verified 2026-10-01: a live process moved by StartTransientUnit with PIDs). Session roots = processes named claude whose parent is not claude, outside the slice; each root and its descendants outside the slice go into a new scope wf-claude-<root pid>.scope via busctl --user call org.freedesktop.systemd1 /org/freedesktop/systemd1 org.freedesktop.systemd1.Manager StartTransientUnit 'ssa(sv)a(sa(sv))' wf-claude-<pid>.scope fail 2 PIDs au <n> <pids…> Slice s agents.slice 0. Children forked later inherit the scope. Prints one line per session; errors (process gone) are reported, and the timer retries next minute.

4.11 wf res shell-init

Prints the alias line for ~/.bashrc (the user adds it; wf never edits dotfiles): alias claude='systemd-run --user --scope --quiet --slice=agents.slice claude'.

4.12 wf res hist [--project P]

Reservation sizes from history: history file + the ledger's finished runs (rc=0, peak known). Grouped by project and title kind = title minus trailing (…) and trailing sha/hex words (7–40 hex chars, ≥ 1 digit): gate 2c91cac7 (fix x) → gate. One line per kind: n, median request, peak p50/p95, median estimate, duration p50/p90, suggestion. Suggestion (≥ 3 runs): mem = p95 peak ×1.15 (up to 0.1 GB, ≥ 0.2 GB), for = p90 duration ×1.5 (≥ 5 min); percentiles nearest-rank. run/note asking > 2× the suggestion (mem or for) print hint: history says ~X GB / Y min (n runs) on stderr; nothing is resized. wf batch without --mem/--for uses the wf-batch suggestion of the project (not --prep; the bg-wait ceiling stays 6h unless --for).

5. Shared rule (shared/CLAUDE.md, ~4 lines)

  • Job expected > 2 GB RAM, > 4 cores or > 10 min → wf res run (never bare systemd-run, never a long foreground command); big foreground step → wf res note.
  • Exit 3 = busy: note it on the task, do other work, retry after the printed time. No tight polling. The line names --force when the memory is really free (claims unused): rerun with it if the holders will not grow.
  • No big scratch in /tmp (RAM); use the project's git-ignored out/.
  • Session end: release your own entries (wf res status).

6. Errors

  • Not Linux / no systemd user manager / no cgroup v2 → wf: wf res needs a systemd user session exit 1.
  • Corrupt ledger → renamed to resources.json.bad-<ts>, start empty (id counter salvaged from its text: ids never reused), one warning line. Unknown entry keys (written by a newer version) are ignored, never corrupt: a wf res wait started before a release must not wipe the ledger. Starting a job deletes stale logs/<id>.rc/.peak.
  • Lock is held at most for one command's critical section (no subprocess wait longer than a systemctl call inside the lock; wait polls with the lock released between polls).

7. Tests

  • Pure (tests/test_res.py, hand-written literals): meminfo parse; size/duration parse; capacity with mixed running/note entries, used figures, headroom and game reserve; prune (unit gone, pid gone, expired note, overdue run, 24 h done removal, game expiry); FIFO start order incl. blocked head; busy line holder choice and wording; game slice values and shortfall/estimate lines; stale scratch selection; cleanup pattern + age.
  • I/O (tests/test_res_io.py): fake runner records argv for run/stop/set-property/timer; atomic write; corrupt ledger recovery; two real processes reserving at once under the real lock never overbook (budget faked).
  • Smoke (manual, release step): real wf res run --mem 100M --for 1m -- true, wait, status.

8. Release

Slices (≈1 h each): pure core · ledger+lock+run/status/release · wait/rc/peak · queue · note · game+slice · clean · timer+tick+shell-init · shared rule + CHANGES + release. Release rules of the workflow repo apply; after release the workflow session runs wf res timer on once and tells the user to add the shell-init alias.