aboutsummaryrefslogtreecommitdiffziptar.gz
path: root/docs/cloud-lane.md
blob: d48c1ed0a7a357be1fd589dfccf7e9723a8be39c (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
# wf cloud — cloud worker lane on promo credits — design

Status: owner decisions 2026-10-06 (design session).
Tool code goes to `/projects/public/workflow` (public tool repo); this spec was originally developed locally,
then a copy `docs/cloud-lane.md` in the tool repo becomes the live one (same as resource-ledger).

## 1. Goal

1. Use a promo cloud credit for worker tasks, so the weekly subscription limit goes to design/orchestration/local-only work.
2. No GitHub, no hosting: code goes up as a `claude --cloud` bundle upload from a local folder, results come
   back through the session transcript. Push stays local-only.
3. A local ledger keeps cloud spend under a cap; the orchestrator routes fitting work to the cloud only while
   the ledger is positive; otherwise everything runs locally as today.
4. Cloud work never touches TASKS.md / archive: bookkeeping stays local (`wf` is not in the VM).

## 2. Owner decisions (2026-10-06)

| # | Decision |
|---|----------|
| C1 | No GitHub (no mirror, no GitHub App). Bundle upload only. |
| C2 | Cloud ledger with budget set by owner (balance minus margin against overrun). |
| C3 | The orchestrator may put fitting work in the cloud on its own while the ledger balance is positive; at ≤ 0 it runs locally, no question asked. |
| C4 | Opus in the cloud is fine (most work is opus). |
| C5 | Purpose-built snapshot repos (code + the resource files a task needs) are the unit sent up; no hosted git. |

<a id="facts"></a>
## 3. Facts (measured 2026-10-06, CLI 2.1.290)

| # | Fact |
|---|------|
| F1 | `claude --cloud "<prompt>"` in a folder with no git remote uploads a bundle (history + tracked files; untracked and `.env`-like files left out; ≤ 100 MB, else branch-only, else squashed snapshot). It prints `Created cloud session: …`, `View: https://claude.ai/code/session_…`, `Resume with: claude --teleport session_…` and **exits** — scriptable. |
| F2 | `--cloud` needs a TTY (`-p` refused; piped stdout refused). A pty (tmux, python `pty`) works. |
| F3 | A never-seen folder shows the trust dialog first (blocks; "Yes" = Down, Enter). Pre-accepting it (`~/.claude.json` `projects[<path>].hasTrustDialogAccepted = true`, atomic rewrite) works — verified 2026-10-06, no dialog. |
| F4 | Cloud sessions run **Opus 5.5 at effort medium** by default (commit trailer, status line). |
| F5 | The VM can't reach local networks behind a firewall → it can't push to local git. Custom allowlist = claude.ai environment settings (owner action). |
| F6 | Large projects with native build tools may require environment setup; e.g., .NET SDK not in the base image → download from package repos (minutes, cached with environment setup script). |
| F7 | `claude --teleport <sid>` for a bundle session: validates, fetches logs, then **fails silently at "Checking out branch"** (branch only in the VM) — sometimes exits, sometimes still resumes the conversation. It never applies the branch. Teleporting a *running* session does not interrupt it (local copy shows "Interrupted"; cloud kept working). |
| F8 | After a teleport resumes, `/export <file>` writes the rendered conversation (80-col wrapped; tool output truncated `… +N lines`; **assistant text complete**) with **no model call**. The local `.jsonl` transcript only appears after a local message (= a full-context model call, avoid). |
| F9 | So the result channel is the session's **final assistant text**: a gzip+base64 patch between markers with a sha256, whitespace-stripped on read. A patch inside tool output is truncated and useless. |
| F10 | `claude -p "<msg>" --cloud <sid>` queues a follow-up into a running/idle session and exits (no TTY needed). |
| F12 | The export contains the prompt too: marker parsing must take only assistant lines (`● WF-RESULT …` = bullet prefix) after the prompt, never a bare substring match. |
| F13 | Environment configuration (owner, 2026-10-06): custom network allowlist for required services; environment variables for telemetry and timeouts; setup script for large dependencies. The exact setup depends on project needs and should be configured in claude.ai environment settings. |
| F11 | Cloud sessions share the subscription rate limits only after the promo credit is used up; until then they draw the credit (promo terms). No CLI shows the balance: owner reads it on claude.ai. |

Baseline local cost (`wf usage --report`, API-price $, 2026-10-06):
opus worker ≈ $1.10–1.32 per done task (median $0.58–0.96, n=64), sonnet ≈ $0.14–0.30, orchestrator
≈ $0.25 per dispatched task. Transcript output counts are estimates.

## 4. Design

<a id="snapshot"></a>
### 4.1 Unit of work: snapshot repo

`wf cloud send <id>` builds `<project>/out/cloud/<id>/` (git-ignored, deleted by `pull` after apply):
- `git archive <master HEAD>` of the project → plain files; plus every path in workflow.toml
  `cloud_include` (git-ignored inputs, e.g. large data files), force-added; one commit "base <sha>".
- Packed size > 90 MB → refuse (`wf: snapshot <n> MB > 90 MB; trim cloud_include`).
- Never in the snapshot: `TASKS.md`/archive/`.wf`/`.worktrees`/`out/` except `cloud_include`.
- Pre-accept the trust dialog for that path (F3), then run `claude --cloud "<prompt>"` under a python pty with
  a timeout (5 min); parse `session_…` from the output. No session id → exit 1 with the last 5 output lines.

<a id="prompt"></a>
### 4.2 Prompt (template `templates/cloud-prompt.md`, filled by `send`)

Fixed text (≈ 25 lines) + `wf show <id>` body + the project CLAUDE.md test recipe of the areas the task names
(`wf ctx` areas section). Rules in the template:
- `wf` is absent: ignore wf/TASKS/lane/gate rules in CLAUDE.md; never edit TASKS.md/archive.
- Never ask; blocked → final text `WF-AWAITING: <one-line question>` and stop.
- Test red → implement → green; commit on `main` (snapshot repo) with one-line messages.
- Network blocked for a needed source → say so in the report line, continue with what the repo has.
- **Final message text** (not a tool call), exactly:
  ```
  WF-RESULT done|awaiting|handback
  WF-REPORT <≤2-line archive entry>
  WF-USAGE in=<n> cw=<n> cr=<n> out=<n> model=<id>   (summed from its own transcript, script given in the template)
  WF-PATCH-BEGIN sha256=<hex of the gzip> bytes=<n>
  <base64 of `git format-patch --binary <base>..HEAD --stdout | gzip -9`>
  WF-PATCH-END
  ```
- Built (cloud send task): the template writes each marker line in brackets (`[WF-RESULT R]`) so no prompt line
  starts with a marker; `cloud.final_message(export)` = the F12 parser: lines of the last message starting
  `● WF-RESULT` (assistant bullet; the prompt renders as `❯ `/`  `, never `● `), bullet/indent stripped,
  ends at the next non-indented line. `pull` parses fields from that list only. Snapshot record
  `.wf/cloud/<id>.json`: id, sid, project, lane, master (sha for `git am`), base (snapshot commit = patch base),
  folder, sent, model, bytes. Ledger row reserved before `claude --cloud` (sid `pending:…`), removed on failure.

<a id="pull"></a>
### 4.3 Return: `wf cloud pull <id> | --all`

- Teleport under a pty in the snapshot folder, `/export <tmp>`, `/exit` (F7, F8); no model call.
- Parse the export: last line starting `● WF-RESULT` (F12); none → `running` (exit 4, prints age since send).
  None, but the last assistant message (no non-empty prompt after it) ends in a bare `WF-PATCH-END` line
  (key words dropped) → handback `bad result: no WF-RESULT key`, patch kept when a `sha256=… bytes=…` header decodes.
- `done`: join base64 between markers (strip whitespace and the 2-space indent), check sha256 and bytes,
  gunzip → patch. Apply in the task's lane worktree on branch `<lane>/<id>` from the recorded base:
  `git am --3way`; files outside the project's tracked tree or in `cloud_include` → refuse.
  Then `quick_gate`; green → `wf done <id> -m "<WF-REPORT>"` + commit + `wf merge` (the `wf finish` path, run by
  `pull` itself, so a pull costs the orchestrator one Bash call). Red / am conflict → handback (below).
- `awaiting` → `wf add -s awaiting "<q>"` + `wf status <id> blocked <a-id>`.
- `handback` / red gate / bad patch → `wf note <id> "cloud <sid>: <why>"`, `wf status <id> clear`, task goes back
  to the local queue with `Recovery: cloud attempt <sid> — <why>` in its note (a local worker may take the patch
  from `out/cloud/<id>.patch`, kept on handback).
- Every pull that ends the session charges the ledger (4.4) and deletes the snapshot folder.
- Archive (owner ruling 2026-10-06, built as part of cloud infrastructure): each pull that ends a session then archives it in the
  app (`POST https://api.anthropic.com/v1/code/sessions/<sid>/archive`, the CLI's internal archiveRemoteSession,
  CLI 2.1.291; 200/409 ok). Failure → one `archive <sid> failed: …` line, pull exit unchanged; ledger row
  `archived: true`; `wf cloud archive <sid>|--ended` retries.
- Built (cloud pull task): sections of the final message open at a key line, other lines continue it (export
  wraps); the patch header + base64 are joined with all whitespace removed (`sha256=HEXbytes=N` then base64:
  gzip base64 starts `H4sI`, never a digit). Refused paths: outside the tree / project folder, TASKS, archive,
  `.wf`, `.worktrees`, `out`, `cloud_include`. One handback note `Recovery: cloud attempt <sid> — <why>[; patch …]`.
  Lost = no result (or no export) and > 24 h since send. Branch `<lane>/<id>` already there / worktree_setup red →
  exit 1, session kept (pull again). `--export FILE` parses a given export (tests, manual). Exit 0 ended, 4 running.

<a id="ledger"></a>
### 4.4 Ledger: `~/.local/state/wf/cloud.json` (same state dir and lock style as `wf res`)

```
{"budget": 240.0, "spent": 0.0, "reserve_per_task": 4.0, "max_parallel": 3,
 "entries": [{"id", "project", "sid", "model", "sent", "state": "running|done|awaiting|handback|lost",
              "usd": null|float, "usd_source": "self|est|owner"}]}
```
- balance = budget − spent − reserve_per_task × running. `send` refuses (exit 3, one line) when
  balance < reserve_per_task or running ≥ max_parallel.
- Charge on end: `WF-USAGE` × `usage.PRICES` (reuse `wflib/usage.py`), plus a 15 % overhead factor for
  what the session can't see (title generation, setup); missing WF-USAGE → charge reserve_per_task,
  `usd_source=est`.
- `wf cloud ledger` prints budget/spent/balance/running; `--set-balance <usd from claude.ai>` sets
  spent = budget − balance (owner reconcile, `usd_source=owner` row); `--budget N` changes the cap.
- A running entry older than 24 h without `WF-RESULT` → `lost` (charged reserve), task back to local.

<a id="fit"></a>
### 4.5 What fits the cloud (routing)

Project opt-in: workflow.toml `cloud = true` (+ `cloud_include = [...]`, optional `cloud_note = "..."` appended
to the prompt) only enables the lane; routing is per task (owner ruling 2026-10-06, amended: opus only).
Whoever adds a task (design session, orchestrator, worker follow-ups) classifies it: `Cloud: yes|no`
(`wf add --cloud` / `wf set --cloud`); yes = Done checkable from the repo alone (code, synthetic data,
`cloud_include` data), no GUI/display, live game/server/capture, LAN, Wine or local installs, `wf`/`wf res` steps,
owner; opus only, prefer larger (1h); unsure → no. Task fits when all hold:
- project opted in; task runner-ready (Done + Model), not `Sessions: owner|solo`, no `Cloud: no` line;
- effort ≤ `slice_above` (no slice jobs); Model opus (sonnet/haiku never, also with `Cloud: yes`: overhead ≈ their cost, §5);
- `Cloud: yes` → fits (regex skipped); no Cloud line → body must not name `wf ` commands as steps
  (workflow-upgrade / bookkeeping tasks need `wf`), GUI, LAN hosts, `wf res`, or a live server — checked by a small
  regex list in `wflib/cloud.py`.
`wf orch pick cloud` order (`cloud.pick_key`): `Cloud: yes` first, then prio, then effort desc (1h before <1h).
A worker that finds a task unfit for the cloud adds `Cloud: no` (`wf set <id> --cloud no`).

<a id="orch"></a>
### 4.6 Orchestrator

`/wf-orchestrate` and `wf batch` gain a virtual lane `cloud`:
- `wf orch pick cloud` = first fitting task across the project's lanes, ledger allows → `wf cloud send`, claim
  `status progress cloud:<sid>`. Prints `stop lane cloud: <ledger|none fit|max parallel>` otherwise.
- The orchestrator calls `wf cloud pull --all` on each wake-up (no faster than every 10 min; a teleport
  takes ~30 s and no tokens). Results feed `wf orch post` like a worker report.
- Local lanes skip tasks claimed `cloud:`; the cloud lane takes opus tasks only (the expensive ones locally), `Cloud: yes` first (§4.5).
- Ledger ≤ reserve → the cloud lane stops; the local lanes run as today (C3).

## 5. Economics (measured, see §6)

Value of local work = its API-price $ (what the weekly limit is spent on). Per opus task:

| Item | $ | Source |
|---|---|---|
| Local opus worker avoided | 1.20 | Historical API-price per task (range 1.10–1.32, n=64) |
| Local overhead of a cloud task | 0.10 | send + pull = 2 orchestrator calls at ~150k cached ctx (~0.05 each); poll wakes amortised ~0.03; gate = CPU only |
| Failure (handback → local redo) | 0.24 | assumed p_fail 20 % × 1.20 |
| **Local $ saved** | **0.86** | |
| Cloud credit used | 1.40 | typical replay scenario with setup overhead |

Net per opus task = 0.86 − v × 1.40:

| Credit value v | Net per task | $240 credit → tasks / net |
|---|---|---|
| 0 % (use-or-lose, time-limited) | **+0.86** | ~150 tasks / **+$129 local** |
| 50 % | +0.16 | ~150 / +$24 |
| 100 % (as paid usage) | −0.54 | loss: only worth it when the local limit is hit |

The analysis shows that using cloud credits for opus-only tasks is economical only when the credit value is low (time-limited or not yet consumed).
Sonnet tasks: local $0.14–0.30 vs overhead + failure ≈ $0.15 → ≈ 0 net even at v = 0 → **cloud lane takes opus
only** (sonnet only via `Cloud: yes`). Build cost: 6 slices worth of engineering.

## 6. Validation

Cloud worker send/pull was validated with:
- Manual pty driver tests against the real CLI on throwaway repos (no project content).
- Round-trip end-to-end tests from send → export → pull with real snapshot repos.
- Synthetic test fixtures with the same message shape as real cloud sessions (for unit tests in the tool repo).

Key findings:
- Marker parsing must handle export word wrapping (spaces, indents).
- ANSI codes in the TUI output must be stripped before marker matching.
- Trust dialog pre-acceptance prevents interactive blocks.
- Export `/exit` may take multiple Enter presses to ensure the file is written.