workflow

git clone https://git.godosa.eu/workflow

master

raw · 17296 bytes

wf — owner's manual

How to set up and run wf day to day. The README tells you what it is and why it exists. The reference docs give the details: design.md (format, config, every command), orchestrator.md, lanes.md, multi-session.md, resource-ledger.md. Every command below also has wf <command> -h.

This is still a personal setup. The manual describes how it runs on the author's machine. It is not a supported product.

1. Setup

Requirements

  • Linux with systemd. wf res needs systemd user units and cgroup v2. Everything else only needs a shell.
  • Python ≥ 3.11 (standard library only, nothing to pip install).
  • git, and Claude Code (claude on PATH).
  • Optional: the superpowers Claude Code plugin. The shared rules refer to its brainstorming, debugging and plan-writing skills. Without it, agents just follow the plain rules.

One root folder

All projects live under one folder. The shipped rules, skills and agent file use the absolute path /projects (for example python3 /projects/public/workflow/wf.py). Keep that path: create it once and own it, or point it at a folder in your home directory.

sudo mkdir /projects && sudo chown "$USER": /projects     # or: sudo ln -s ~/projects /projects
git clone <this repo> /projects/public/workflow
ln -s workflow/shared/CLAUDE.md /projects/CLAUDE.md        # every agent under /projects loads it
echo "alias wf='python3 /projects/public/workflow/wf.py'" >> ~/.bashrc

Claude Code reads CLAUDE.md from the working directory and from every parent folder, so every session in /projects/<anything> gets the shared rules. If the symlink is not picked up, use a one-line /projects/CLAUDE.md with the content @workflow/shared/CLAUDE.md instead. wf projects scans the parent folder of the tool repo, or its grandparent when the parent holds no project (tool under e.g. public/; override: WF_ROOT). Workflow reports go to /projects/public/workflow/inbox.md (override: WF_INBOX).

Agent and skills

The orchestrator spawns the worker subagent. The skills are the session modes in section 2. Install them as symlinks, so that git pull updates them:

mkdir -p ~/.claude/agents ~/.claude/skills
ln -s /projects/public/workflow/shared/agents/wf-worker.md ~/.claude/agents/wf-worker.md
for s in wf-orchestrate wf-pilot wf-design; do ln -s /projects/public/workflow/shared/skills/$s ~/.claude/skills/; done

Agents and skills appear only in sessions started after the install, so restart any running session.

A project

cd /projects/games/mygame        # any depth ≤ 3 under /projects; a git repo is recommended
wf init                          # writes workflow.toml, TASKS.md, tasks/archive.md, CLAUDE.md (keeps existing files)
wf check                         # 0 errors

If the project already has a numbered task list, wf migrate shows the conversion (dry run) and wf migrate --write applies it.

Edit the generated CLAUDE.md: layout, verify command, gotchas and one ### <area> block per code area (code map with grep anchors, test recipe, paths). Agents read these area notes before they read code.

workflow.toml (the template lists every key with a comment; unknown keys are errors):

Key What it does
tasks, archive file names (defaults TASKS.md, tasks/archive.md)
docs files/folders whose [[id]] links wf check validates
verify commands wf done reminds of (it never runs them) and wf ctx prints as Verify
done extra checklist lines wf done prints
quick_gate fast regression check; wf gate and wf finish run it and refuse when it is red
worktree_setup idempotent commands wf setup runs in a fresh lane worktree (link git-ignored data, venv…; env WF_MAIN = main tree)
slice_above, [lanes.*] task sizes per lane and when a task must be sliced (default lanes: fast <1h, slow 1h); see lanes.md
areas, code_root, area_stale_commits, area_ignore where the area notes live and when they count as stale
ledgers, ctx_hint, [anchors] in-flight plan ledgers, /clear hint, required Ref anchors
cloud, cloud_include, cloud_note cloud lane opt-in (below)

Add out/ to the project's .gitignore: logs, batch summaries and scratch go there.

Resource ledger (wf res)

Optional, but needed once several agents build or test at the same time.

wf res shell-init                # prints an alias that starts claude inside agents.slice: add it to ~/.bashrc
wf res timer on                  # 1-minute timer: starts queued jobs, ends game mode, cleans scratch, adopts sessions
wf res status                    # budget, reservations, warnings

Reserves and limits go in ~/.config/wf/resources.toml (all keys optional: user_reserve_gb, user_reserve_cpus, game_reserve_gb, game_reserve_cpus, game_hours, small_headroom_gb, scratch_hours; see resource-ledger.md §3.3). Optional hook that hides the display from agent commands while game mode is on, in ~/.claude/settings.json:

"hooks": {"PreToolUse": [{"matcher": "Bash", "hooks": [{"type": "command",
  "command": "python3 /projects/public/workflow/wf_res.py hook"}]}]}

wf batch (and /wf-pilot) need one allow rule in ~/.claude/settings.json, because auto mode does not let an agent start another unattended agent: Bash(python3 /projects/public/workflow/wf.py batch:*).

Cloud lane (optional)

Tasks can also run as claude --cloud sessions, which are billed to your cloud budget, not your machine. You need a Claude Code version with --cloud / --teleport and a claude.ai login. Per project: cloud = true in workflow.toml, plus cloud_include for git-ignored inputs the task needs. Set the spending cap once with wf cloud ledger --budget 100 and check it with wf cloud ledger. Only opus tasks marked Cloud: yes (or passing the fit rules) are sent. See orchestrator.md "Cloud lane".

2. Daily use

There are four kinds of session. Start each one in the project folder with claude (the agents.slice alias if you set it up).

Worker session: "continue"

Type continue (or "next task"). The agent runs wf next --as <its model>. It shows open questions, items that need you, plans in flight and the next task, then works on that task: test red → implement → green → wf done → one commit. It does not ask before it starts. Several sessions in one project: give each a lane ("continue in the slow lane"). wf next then switches them to worktree mode (one branch per task, wf finish merges it). One session per lane.

Orchestrator: /wf-orchestrate

After /clear, type /wf-orchestrate (go = start the proposed run without asking, batch N = prepare an unattended batch). The orchestrator holds no lane and never reads code. It spawns one fresh wf-worker per task (wf orch pick / wf orch post), one per lane in parallel, on the task's model, and parks a task that ends awaiting, needs-owner (waits After: a Needs-human task), handback or red post-check (blocked / Sessions: owner, alert) while its lane keeps going. Use it when the backlog is runner-ready (section 3) and you want to steer.

Overnight: /wf-pilot

/wf-pilot [--batch 4] [--lanes a,b] [--for 8h] in an idle session, which can be one you opened from the phone. The pilot starts headless wf batch runs one after another until the deadline. The session stays small, and you get push notifications on new questions, on stops and at the end. Steering: wf batch --stop (graceful, running workers finish), wf move <id> deferred (skip a task), wf batch --status (newest batch and its summary). A one-off short run without a pilot is wf batch N [--lanes …]. wf batch N --prep makes tasks without a Done line runner-ready instead.

Design: /wf-design

/wf-design [topic | id] is for brainstorming, specs and rulings with you. Flow: brainstorm → spec (committed) → your approval → wf add slices ≤ 1h with Done, Model and Ref → the orchestrator picks them up. It answers open questions with you and records each ruling where it binds.

Your side

  • Awaiting (wf list -s awaiting): questions agents could not decide. Answer them in a design or orchestrator session. The session records the ruling and runs wf done <a-id>, which unblocks the waiting tasks.
  • Needs human (wf list -s human): things only you can do (play-test, hardware, accounts). Check boxes with wf tick <id> <n>. Done → tell any session ("t-x is done") or run wf done. Tasks marked Sessions: owner are picked only when you say you are present (wf next --owner).
  • Stop file: wf batch --stop writes out/wf-batch.stop. Batches, pilots and wf lanes --wait stop before the next spawn.
  • Game mode: wf res game on [--for 4h] before you play. Agents keep running but get less CPU and IO, and big jobs that do not fit wait. wf res game off when you are done.
  • Overview: wf projects (all projects), wf lanes, wf log -n 10, wf res status.
  • Workflow inbox: agents file problems with wf report. Triage /projects/public/workflow/inbox.md now and then: fix the tool, or turn entries into tasks.

3. Writing good tasks

wf add "Cave seams. Close the slit beside the lintel." -p 1 -e '<1h' --model sonnet \
  --done "tests/test_cave.py::test_seam passes; render of fixture shows no gap" --ref DESIGN.md#terrain
  • Title. Goal. One line: the title, then the outcome in one sentence.
  • Done line (--done): a check a worker can run without you, such as a file, a test result or an artifact. Never "owner confirms". Only tasks with a Done line are runner-ready: the orchestrator, batch and pilot skip the others. (wf batch --prep drafts them.)
  • Model line (--model): the cheapest model that fits. haiku = mechanical, exact steps. sonnet = exact Done, an existing pattern to copy, and a real-data pass/fail. opus (the default when the line is missing) = design, debugging, reverse engineering, anything unclear. A worker that finds the task too hard raises it itself and hands it back.
  • Effort (-e): <1h, 1h, 5h, 10h, 100h. Anything above slice_above (1h) becomes a slice job: an agent splits it into --parent slices, each with its own Done/Model, and never implements it whole. Slice big work yourself when you already know the steps.
  • Code anchors: when you know them, add body lines like Code: src/terrain.py build_cave; test to copy: tests/test_cave.py::test_floor. Write anchors (function or string names), not line numbers. They save the worker most of its exploring.
  • Refs and dependencies: --ref doc.md#heading makes wf ctx print that section; --after t-other keeps it unpicked until t-other is done.
  • Areas: keep each area's code map and test recipe in the project CLAUDE.md. wf ctx prints the areas a task names. When code drifts, wf done adds a t-map-… refresh task automatically.
  • Cloud (cloud projects): --cloud yes only when the Done line can be checked from the repo alone (no GUI, live server, LAN, local installs). When unsure, no.
  • Sessions: --sessions solo for repo-wide changes (nothing else runs meanwhile), --sessions owner when you must be there.

4. Best practices and pitfalls

  • One lane per session, at most one worker per lane worktree. Parallel lanes are safe (project lock, branch per task). Two sessions in one lane pick against each other.
  • Fresh context per task. Cost is mostly cache reads, which grow with context. Short-lived workers on the cheapest fitting model are cheaper than one long session. /clear the orchestrator after a big design talk, but only when no worker is running (a /clear kills its running command).
  • Big jobs through wf res: anything over 2 GB, 4 cores or 10 minutes goes through wf res run --mem … --for … --title … -- cmd. Exit 3 = busy: do other work and retry later, or use --queue. wf res hist suggests sizes from past runs.
  • Never hand-edit TASKS.md while agents run. Use wf add/set/note/move/done. If you must edit by hand, run wf check afterwards.
  • Commit ≠ merge ≠ push. Agents commit their own task. Workers ff-merge only their own task branch. Pushing goes only to a remote named home (git push home --all), and nothing else is pushed unless you ask.
  • Close the loop on problems. Agents work around a workflow problem and wf report it. Triaging the inbox is how the rules get better. Put project-only needs in the project's CLAUDE.md or workflow.toml, not in the shared rules.
  • Watch cost: wf usage (this session per agent), wf usage --report (per lane, model and effort across every project's out/wf-cost.log). Expensive lanes usually mean vague Done lines or a model that is too large.
  • Model choice: run the orchestrator and design sessions on opus. Tasks get the cheapest model that fits. A sonnet task without an exact Done line comes back as a handback.
  • Updating the tool: git -C /projects/public/workflow pull. Read the top of CHANGES.md: each line ends with "Projects: …", which says what a project must do (usually nothing).
  • Public by default: the tool repo's history is public, so never put private names (project names, home paths, hosts) in files or commit messages. A denylist guard checks each commit: put your private words, one per line (# comments, !token = exemption), in publish-denylist.local in the tool repo (git-ignored; extra lists via git config --add denylist.file <path>) and enable the hooks once per clone with git config core.hooksPath .githooks. The pre-commit hook scans staged lines and paths, the commit-msg hook the message, both case-insensitive; a hit names the place and word #<line> only and refuses the commit (git commit --no-verify for a false hit). No list → a warning, the commit goes through. Audit the whole tree with python3 scripts/denylist_check.py tree.

5. Troubleshooting

Symptom Cause / fix
busy: … and exit 3 from wf res run/note Not enough memory or CPUs beyond your reserve. Retry after the printed time, add --queue, or check wf res status for stale entries (wf res release r-N). If memory really is free (unused reservations), --force fits against real free memory.
wf orch post says post-check red The worker said done, but the archive line is missing, the branch is still there, the worktree is dirty, or wf check has errors. The message names the problem. Fix it in the worktree (wf merge, commit or clean), then post again.
Task stuck in progress though no one works on it A session or worker died. wf list --stale shows them, and wf status --clear-stale clears every claim without a live session (or wf status <id> clear). Re-run the worker with wf orch pick <lane> --id <id> --recovery "<why>" to keep its WIP.
wf start refuses: worktree has uncommitted changes A dead worker left WIP. --recovery keeps it and shows it. Otherwise commit or reset it in that worktree.
branch <lane>/<id> exists (local WIP?) (cloud pull) A local branch for that task already exists. Merge or delete it, then wf cloud pull <id> again. The cloud session is kept.
wf cloud pull exit 4 The session is still running. Pull again later; after 24 h without a result it counts as lost.
wf cloud pull handback / bad patch The task goes back to its local lane with a Recovery: cloud attempt … note. The patch is kept in out/cloud/<id>.patch.
wf merge: rebase conflict The branch is untouched. git rebase master, resolve, run verify, wf merge again.
wf finish refused The quick gate is red, files outside the given paths are uncommitted, or a path is outside the repo, missing or unchanged. Fix it and rerun. If it fails after done:, a rerun resumes.
A worker or batch does nothing wf list --runner is empty: tasks lack a Done line (wf batch --prep), are blocked on an awaiting item, or wait on After:. wf lanes shows counts per lane.
wf check errors after a hand edit Read the message (duplicate id, broken [[link]], bad section) and fix the item. Ids are never reused.
project format N is newer than this wf Update the tool: git -C /projects/public/workflow pull.

Split projects (private home repo + public code repo)

code_root in workflow.toml points at another git repo: TASKS/archive/specs stay in the private project, code goes to the code repo. wf start makes code worktrees (branch in the code repo, .wf-home file pointing back); wf finish/wf merge/ wf wip commit private paths in the private repo and code paths in the code repo. push = [cmd…] runs in the private root after a merge (env WF_MAIN); wf push runs it by hand; failure writes .wf/push-failed. Add .wf-home to the code repo's .gitignore. The cloud lane is refused. wf check warns: no push, .wf-home not ignored, merged leftover code worktrees.