1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
|
# wf res — shared resource ledger for agents — design
Status: approved and implemented 2026-10-01.
## 1. Goal
1. No agent job is OOM-killed because another agent started a big job at the same time.
2. The user keeps a guaranteed share of RAM and CPU (gaming), switchable on demand.
3. An agent that cannot get resources now gets one clear line (who holds them, until when) and works on
something else instead of waiting or retrying blindly.
4. Agent-made waste (tmpfs scratch, failed units) is cleaned routinely.
5. Python ≥ 3.11 stdlib, Linux systemd user manager, no daemon. Works in any directory, wf project or not.
Machine facts (2026-10-01): 16 cores, 30 GB RAM, 7 GB swap, `/tmp` is tmpfs (16 GB), user manager
delegates `cpu io memory pids`; `systemd-run --user --scope --slice=agents.slice` and transient units with
`MemoryMax`/`CPUWeight` verified working.
## 2. User decisions (2026-10-01)
| # | Decision |
|---|----------|
| D1 | Default user reserve 6 GB RAM, 4 CPUs. |
| D2 | Gaming reserve 12 GB RAM, 8 CPUs. |
| D3 | Gaming mode slows agents; never freezes or kills them. |
| D4 | Scratch under `/tmp/claude-<uid>/` untouched for 2 h (newest mtime in the entry) may be deleted automatically by any agent, whoever made it. Except a live session's own dir (`<project>/<session-id>`, live = `~/.claude/sessions/<pid>.json` with that `sessionId`, pid alive, `procStart` matching; added 2026-10-02 after an idle session lost its scratchpad). Same automatic clean sweeps top-level litter in `/tmp` and `/var/tmp` (own entries only): empty dirs idle 2 h (`scratch_hours`) and `clr-debug-pipe-<pid>-*` / `dotnet-diagnostic-<pid>-*` of dead pids; entries with content are never swept. Per-project `out/` patterns only list unless `--yes`. |
| D5 | v1 = everything in the proposal: run, status, wait, release, note, queue, game, clean, timer, shared rule. |
| D6 | `game on [--for D]`, default 4 h, turns itself off. |
| D7 | Install a 1-minute user timer via an explicit `wf res timer on`. |
| D8 | `game on` never fails and never squeezes running jobs; it switches budget and weights at once and prints the shortfall and the estimated time until the full reserve is free. |
| D9 | Small unreserved work: headroom (2 GB) in the budget **and** all agent processes inside `agents.slice`: running Claude sessions are moved there live by `wf res adopt` (run by every timer tick; no restart), new ones optionally start there via the `shell-init` alias. |
## 3. Architecture
### 3.1 Cgroup layout
```
user@<uid>.service
└── agents.slice ← MemoryHigh, CPUWeight, IOWeight (normal / gaming)
├── run-…scope ← each Claude session (started via the alias), its shells, small jobs
└── agents-jobs.slice
└── wf-r-N.service ← each `wf res run` job: MemoryMax, MemoryHigh, MemorySwapMax=0, Nice=10
```
- `agents.slice` caps the sum of everything agents do. The ledger decides who may start big work; the slice
is the kernel backstop for wrong estimates and unreserved small work.
- Slice properties are set with `systemctl --user set-property --runtime agents.slice …`. The slice unit file
`~/.config/systemd/user/agents.slice` is written on first use (then `daemon-reload`), so the slice exists
before any set-property.
- Normal: `CPUWeight=20`, `IOWeight=20`, `MemoryHigh = RAM − user_reserve_gb`.
- Gaming: `CPUWeight=5`, `IOWeight=5`, `MemoryHigh = max(RAM − game_reserve_gb, agents.slice MemoryCurrent)`;
re-lowered towards the target on every `wf res` call / tick (D8: no squeezing below current use).
- Weights only matter under contention; on an idle machine agents get full speed.
### 3.2 Code layout
- `wflib/res.py` — pure, text in / data out: parse `/proc/meminfo`, parse `systemctl show` output, durations
and sizes (`10G`, `40m`, `4h`), prune, capacity rule, queue start order, game slice values and shortfall,
stale-scratch selection (from a given list of (path, newest mtime)), all output lines.
- `wf_res.py` — I/O: lock, atomic ledger write, `/proc` and `/sys/fs/cgroup` reads, filesystem walk, the
runner (`subprocess.run` wrapper, injectable for tests), unit/timer file writing, printing.
- `wf.py` — only registers `wf res …` and dispatches to `wf_res.main(argv)`. `wf res` needs no
`workflow.toml`.
### 3.3 Files
- State dir `~/.local/state/wf/` (`$XDG_STATE_HOME/wf` if set): `resources.json`, `resources.lock`,
`logs/r-N.log`, `logs/r-N.rc`, `logs/r-N.peak`, `resources-history.jsonl` (finished rc=0 runs with a peak,
appended when pruned from the ledger after 24 h; newest 2000 kept; read by `wf res hist`).
- Config `~/.config/wf/resources.toml` (`$XDG_CONFIG_HOME`), all keys optional:
```toml
user_reserve_gb = 6
user_reserve_cpus = 4
game_reserve_gb = 12
game_reserve_cpus = 8
game_hours = 4
small_headroom_gb = 2
scratch_hours = 2
```
- Unit files under `~/.config/systemd/user/`: `agents.slice`, and with `timer on`: `wf-res.service`,
`wf-res.timer`.
### 3.4 Ledger
`resources.json`: `{"next": N, "game_until": iso|null, "last_clean": iso|null, "entries": [...]}`.
Entry fields: `id` (`r-N`, never reused), `project` (git toplevel basename, else cwd basename), `owner` (pid of
the nearest ancestor process named `claude`, else the parent pid of `wf`), `title`, `mem_gb`, `cpus`
(default 1), `est_min`, `state` (`queued` | `running` | `note` | `done`), `unit` (`wf-r-N.service`, run only),
`cmd` (argv list), `cwd`, `log`, `queued` (iso), `started` (iso), `expires` (iso = started + 2 × est_min),
and for `done`: `ended`, `rc` (int or null), `peak_gb`, `why` (`exited` | `expired` | `owner gone` | `released` | a kill reason, e.g.
`killed: oom-kill by systemd-oomd, limit 4.0 GB; raise --mem`).
`env` (caller variables the user manager lacks; `WF_RES_ID` and secrets never), `by` (who to tell, run/note;
keys dropped when unknown): `name` (`--by`, else `WF_SESSION_NAME`, else `<lane> session` from the wf lane
record `.wf/sessions/*.json` with this `CLAUDE_PID`, main tree or toplevel), `task` (`WF_TASK`, else the
`<lane>/<id>` branch of cwd), `batch` (`WF_RES_ID`: every job unit gets `WF_RES_ID=r-N`, so a `wf batch`
worker's entries name the batch), `address` (`uds:$CLAUDE_CODE_MESSAGING_SOCKET`; a batch worker's = its
orchestrator's, reachable via SendMessage). Status lines of live entries end `[by <name> <task> batch r-N,
message uds:…]`; the throttle warning ends `; started by …`.
Every command takes the lock (`fcntl.flock` exclusive on `resources.lock`), reads, then **prunes**, acts,
writes atomically (temp file in the same dir + `os.replace`), releases.
Prune (pure, given unit states, live pids, now):
- `running` whose unit is inactive/failed/unknown → `done` (`rc` from `logs/r-N.rc`, `peak_gb` from
`logs/r-N.peak`). No `.rc` (the sh wrapper died with the job) → `rc` null and `why`/`peak_gb` from
`journalctl --user -u wf-r-N.service --since @started` (`Failed with result '…'`, `systemd-oomd killed`,
`… memory peak`), shown in `status` and `wait`; journal silent → `why` points at that journalctl command.
- `running` past `expires` → stays running (job not killed; `run` says so on start, status shows
`ETA overdue (still running, not killed)`), flagged `overdue` in status, and counted at
`max(reserved, used)`.
- `note` whose owner pid is gone or past `expires` → `done`.
- `done` older than 24 h → removed.
- `game_until` in the past → game off (slice back to normal values).
### 3.5 Capacity rule
```
budget_mem = MemAvailable − reserve_gb − small_headroom_gb
− Σ over running: max(0, mem_gb − MemoryCurrent(unit))
− Σ over notes: mem_gb
budget_cpus = nproc − reserve_cpus − Σ over running and notes: cpus
fits(req) = req.mem_gb ≤ budget_mem and req.cpus ≤ budget_cpus
```
- `reserve_*` is the user or the gaming value depending on game mode.
- A note counts its full reservation (its memory cannot be measured; overcounting briefly is safe).
- A running job's memory already in use is inside `MemAvailable`, so only its unused part is subtracted.
## 4. Commands
All print short lines; `--json` where stated. Exit: 0 ok · 3 busy (refused) · 1 `wf: …` one line (unknown
id, systemd failure) · 2 usage.
### 4.1 `wf res run --mem 10G [--cpus N] --for 40m --title "…" [--queue] -- <cmd …>`
- Fits → `systemd-run --user --slice=agents-jobs.slice --unit=wf-r-N --collect
-p MemoryMax=<mem> -p MemoryHigh=<0.9·mem> -p MemorySwapMax=0 -p Nice=10
-p WorkingDirectory=<cwd> -p StandardOutput=append:<log> -p StandardError=append:<log>
/bin/sh -c '"$@"; rc=$?; cat /sys/fs/cgroup$(cut -d: -f3 /proc/self/cgroup)/memory.peak > <peak>; echo $rc > <rc>' sh <cmd …>`
(the unit is collected on exit, so the wrapper records exit code and peak bytes itself). Prints `r-N started; log <path>; ETA ~HH:MM`. Exit 0.
- Environment: a user unit starts from the user manager's environment, so the caller's variables that differ
from `systemctl --user show-environment` go along as `--setenv=NAME=VALUE` (stored in the entry, so a queued
job gets them too); skipped: `PWD OLDPWD SHLVL _` and names with TOKEN/SECRET/PASSWORD/PASSWD/CREDENTIAL.
- Does not fit, no `--queue` → exit 3, one line:
`busy: 7.5 GB held by proj-a "dotnet e2e" (r-4) until ~14:40; 2.1 GB free for agents; retry after ~14:40 or work on something else`
(names the holders whose release would make it fit, earliest ETA first; CPU shortage named the same way).
- Does not fit, `--queue` → `queued`, prints `r-N queued, position P; est. start ~HH:MM; cancel: wf res release r-N`. Exit 0.
- A request larger than the whole agent budget on an empty machine → exit 1 `wf: 20 GB can never fit (max …)`.
- `--force` (memory really free though the ledger says busy: unused reservations, stale notes): fits against
MemAvailable − user reserve (game reserve in game mode) − headroom only; ledger claims and the queue are ignored;
the entry is still recorded. Does not fit even so → exit 3 `busy even with --force: only X GB really free beyond
the reserve; …`. The normal busy line ends with `; X GB really free beyond the reserve: --force starts it past
the ledger (only if the holders will not use what they reserved)` when `--force` would fit.
- `systemd-run` failure → entry removed, exit 1 with its first stderr line.
- `--lock KEY`: one queued/running job per KEY and main tree (git common dir's parent: lane worktrees share it)
at a time, e.g. jobs sharing one checkout (a-game-project-builder `../.<project>-gate`). A title starting with the word
`gate` locks `gate` unasked (2026-10-06: two hand-started gates overlapped in one gate checkout). Held → exit 3
`busy: lock 'KEY' held by r-N "title" (ETA ~HH:MM|queued): …; --queue waits for it (--force does not override
a lock)`; `--queue` → queued `(lock held by r-N)`. Ledger field `lock` = `KEY@<main tree>`.
- The command reaches the job verbatim (no systemd `%` specifier expansion: `--format='%h %s'` is safe).
### 4.2 Queue
Strict FIFO by `queued`: the head starts when it fits; entries behind a waiting head wait (a big job is never
starved); an entry whose lock a running job holds is skipped (it neither starts nor blocks the rest). Starting happens inside every `wf res` call (after prune) and every timer tick. A started queued job
runs with the cwd and argv it was queued with; its ETA counts from the actual start.
Batch loan: a job started from inside a live job (`by.batch` = its `WF_RES_ID`, e.g. a `wf batch` worker's gate)
fits against budget + the parent's claim (mem and cpus) minus what the parent's other running children hold, so a
batch waiting for its own gate never deadlocks (direct run and queue alike; not with `--force`). A queued entry
whose parent job ended (done or gone) is dropped at prune (`r-N dropped (batch r-M gone)`, why `owner gone`):
nobody waits for it, and it must not head the queue.
### 4.3 `wf res status [r-N] [--json]`
Without id: one line per entry (id, project, title, state, reserved / used / peak GB, cpus, started, ETA or
`overdue`), then the budget line (`agents may use X GB, Y cpus now; reserve 6 GB/4 cpus`), game mode with
time left and shortfall line (§4.6), unreserved agent memory (`agents.slice` MemoryCurrent − jobs; `(X GB of it file cache, reclaimable)` =
slice `memory.stat` `file` − `shmem`, capped at that figure), warnings:
- running job throttled at its own limit: used ≥ 0.85 × `--mem` **and** its cgroup `memory.pressure`
`some avg60` ≥ 20 % → `r-N throttled at its memory limit (…): likely too small; release --stop, re-run
with a bigger --mem` (page cache alone at the limit does not stall, so no false alarm; the timer tick
has no reader, so it does not warn);
- tmpfs `/tmp` above 25 % of RAM;
- `N claude sessions outside agents.slice (wf res adopt)`.
- `N jobs outside wf res (bare systemd-run): <unit> <used> GB, …`: running transient user services outside
`agents*` slices, except desktop ones (`app-*`, `dbus-*`). Warning only: a service cannot change slice live.
With id: that entry in full, including `rc`, `peak_gb`, `why` for done entries.
### 4.4 `wf res wait r-N [--timeout 2h]`
Polls every 15 s (each poll is a normal locked call, so it also prunes and starts queued work) until the entry
is `done`; prints `r-N done rc=0 peak 9.4 GB in 37 min`. The throttle warning (§4.3) goes to stderr once. Exit 0 when done (whatever `rc`), 1 on timeout. If
reaped, `status r-N` answers from the ledger.
### 4.5 `wf res release r-N [--stop]`
- `queued` / `note` → removed. Running unit → refused with `wf: r-N is running; --stop to kill it` unless
`--stop`, then `systemctl --user stop`, entry → done (`why=released`).
### 4.6 `wf res game on [--for 4h] | off`
- `on`: `game_until = now + for`; set gaming slice values (§3.1); prints
```
game on until 22:40 (4h); CPU/IO now yours
short 3.1 GB of 12 GB: r-4 proj-a "dotnet e2e" 2.5 GB ~21:05, r-7 proj-b "r5 rebuild" 9 GB ~21:30
full reserve free ~21:30 (est.); free now: wf res release r-7 --stop
```
The shortfall/estimate lines appear only when the reserve is not free now. Estimate = ETAs of the running
jobs that must end, in order (jobs are never stopped by game mode). New requests that do not fit are refused
or queued as usual. `on` while on → extends `game_until`.
- `off` or expiry: normal slice values, `game_until = null`.
- No windows over the game: `wf res hook` (Claude Code `PreToolUse` hook, matcher `Bash`, in
`~/.claude/settings.json`; the user adds it, wf never edits dotfiles) prefixes every agent Bash command with
`unset DISPLAY WAYLAND_DISPLAY;` while `game_until > now`, and adds context telling the agent why GUI apps
fail (run headless/offscreen or later). Reads the ledger only (no lock, no systemctl, ~60 ms); off, expired,
other tools or any error → no output, command unchanged. Live for running sessions: state is read per call.
```json
"hooks": {"PreToolUse": [{"matcher": "Bash", "hooks": [{"type": "command",
"command": "python3 /projects/public/workflow/wf_res.py hook"}]}]}
```
### 4.7 `wf res note --mem 3G [--cpus N] --for 20m [--force] "title"`
Reserves for foreground work (no unit). Freed when `owner` exits, at `expires`, or by `release`. Refused like
`run` (exit 3) when it does not fit; `--force` as in `run`.
### 4.8 `wf res clean [--yes]`
- Automatic part (no `--yes` needed; also run by any `wf res` call when `last_clean` > 10 min ago, and by every
tick):
- every direct and second-level entry under `/tmp/claude-<uid>/` whose newest mtime (recursive) is older than
`scratch_hours` → deleted; lines `freed 1.2 GB: /tmp/claude-1000/-projects-x/<session>` (only printed by
`clean` itself; silent from other commands);
- `systemctl --user reset-failed 'wf-r-*'`.
- Listed only, deleted with `--yes`: per-project patterns from the current project's `workflow.toml`
`cleanup = ["out/prof", "out/history-logs/*.log:30d"]` (glob relative to the project root, optional `:Nd`
minimum age by mtime). Never follows symlinks; never leaves the project root.
### 4.9 `wf res timer on | off`, `wf res tick`
- `timer on` writes `wf-res.service` (`ExecStart=python3 /projects/public/workflow/wf.py res tick`) and
`wf-res.timer` (`OnBootSec=1min`, `OnUnitActiveSec=1min`), `daemon-reload`, `enable --now wf-res.timer`.
`off` disables and removes them.
- `tick` = lock, prune, start queued, game expiry and slice re-lowering, automatic clean; then adopt., then
session caps: every running scope in `agents.slice` (a Claude session) gets `MemoryHigh = session_mem_gb`
(config, default 6, 0 = off → `infinity`) via `systemctl --user set-property --runtime`, only where it differs.
MemoryHigh only: a session over it is throttled, never killed (a MemoryMax OOM kill could pick `claude`, whose
oom_score_adj is 200). `status` warns `claude session N at its memory cap (…)` at ≥ 0.9 × cap and ≥ 20 % stall.
`wf res run` jobs live in `agents-jobs.slice`, not in a session, so the cap never limits them. Silent.
### 4.10 `wf res adopt`
Moves running Claude sessions into `agents.slice` without a restart (verified 2026-10-01: a live process moved
by `StartTransientUnit` with `PIDs`). Session roots = processes named `claude` whose parent is not `claude`,
outside the slice; each root and its descendants outside the slice go into a new scope
`wf-claude-<root pid>.scope` via
`busctl --user call org.freedesktop.systemd1 /org/freedesktop/systemd1 org.freedesktop.systemd1.Manager
StartTransientUnit 'ssa(sv)a(sa(sv))' wf-claude-<pid>.scope fail 2 PIDs au <n> <pids…> Slice s agents.slice 0`.
Children forked later inherit the scope. Prints one line per session; errors (process gone) are reported, and
the timer retries next minute.
### 4.11 `wf res shell-init`
Prints the alias line for `~/.bashrc` (the user adds it; wf never edits dotfiles):
`alias claude='systemd-run --user --scope --quiet --slice=agents.slice claude'`.
### 4.12 `wf res hist [--project P]`
Reservation sizes from history: history file + the ledger's finished runs (rc=0, peak known). Grouped by
project and title kind = title minus trailing `(…)` and trailing sha/hex words (7–40 hex chars, ≥ 1 digit):
`gate 2c91cac7 (fix x)` → `gate`. One line per kind: `n`, median request, peak p50/p95, median estimate,
duration p50/p90, suggestion. Suggestion (≥ 3 runs): mem = p95 peak ×1.15 (up to 0.1 GB, ≥ 0.2 GB), for = p90
duration ×1.5 (≥ 5 min); percentiles nearest-rank. `run`/`note` asking > 2× the suggestion (mem or for) print
`hint: history says ~X GB / Y min (n runs)` on stderr; nothing is resized. `wf batch` without `--mem`/`--for`
uses the `wf-batch` suggestion of the project (not `--prep`; the bg-wait ceiling stays 6h unless `--for`).
## 5. Shared rule (`shared/CLAUDE.md`, ~4 lines)
- Job expected > 2 GB RAM, > 4 cores or > 10 min → `wf res run` (never bare `systemd-run`, never a long
foreground command); big foreground step → `wf res note`.
- Exit 3 = busy: note it on the task, do other work, retry after the printed time. No tight polling.
The line names `--force` when the memory is really free (claims unused): rerun with it if the holders will not grow.
- No big scratch in `/tmp` (RAM); use the project's git-ignored `out/`.
- Session end: release your own entries (`wf res status`).
## 6. Errors
- Not Linux / no systemd user manager / no cgroup v2 → `wf: wf res needs a systemd user session` exit 1.
- Corrupt ledger → renamed to `resources.json.bad-<ts>`, start empty (id counter salvaged from its text: ids never reused), one warning line. Unknown entry keys (written by a newer version) are ignored, never corrupt: a `wf res wait` started before a release must not wipe the ledger. Starting a job deletes stale `logs/<id>.rc/.peak`.
- Lock is held at most for one command's critical section (no subprocess wait longer than a `systemctl`
call inside the lock; `wait` polls with the lock released between polls).
## 7. Tests
- Pure (`tests/test_res.py`, hand-written literals): meminfo parse; size/duration parse; capacity with mixed
running/note entries, used figures, headroom and game reserve; prune (unit gone, pid gone, expired note,
overdue run, 24 h done removal, game expiry); FIFO start order incl. blocked head; busy line holder choice and
wording; game slice values and shortfall/estimate lines; stale scratch selection; cleanup pattern + age.
- I/O (`tests/test_res_io.py`): fake runner records argv for run/stop/set-property/timer; atomic write; corrupt
ledger recovery; two real processes reserving at once under the real lock never overbook (budget faked).
- Smoke (manual, release step): real `wf res run --mem 100M --for 1m -- true`, `wait`, `status`.
## 8. Release
Slices (≈1 h each): pure core · ledger+lock+run/status/release · wait/rc/peak · queue · note · game+slice ·
clean · timer+tick+shell-init · shared rule + CHANGES + release. Release rules of the workflow repo apply;
after release the workflow session runs `wf res timer on` once and tells the user to add the
`shell-init` alias.
|