返回 CodeWhale
runtime-contract-budget.json
根目录 / scripts / runtime-contract-budget.json
1 {
2 "_comment": "One-way numeric ceilings and exact structural identities for the provider-free runtime contract. Decreases pass; increases or identity changes fail. Lock in decreases with: python3 scripts/check-runtime-contract-budget.py --update The v0.9.8 child-receipt restore grew every production tool surface by 1496 schema bytes / 374 estimated tokens (agent tool). The v0.9.8 workshop read/tool-result byte fields then grew every production tool surface by 371 schema bytes / 93 estimated tokens. Both raises are explicit maintainer decisions; identities stay on the pre-raise digests only if the name set is unchanged \u2014 re-measure on Linux CI if Lint reports identity drift. The v0.9.8 pinned session prefix added the <context_update> sentence to the base prompt (5848 -> 6084 bytes, every representative stage re-hashed), and the host-side Workflow/Goal verbs plus honest child posture grew the tool catalog (active 16531 -> 16602 bytes, full 71473 -> 72371); both are explicit v0.9.8 maintainer decisions measured from the release train. The v0.9.9 configured-skills change hides only custom configured-root paths, preserves discoverable default-root paths, normalizes Windows prompt separators, and trims 50 redundant skills-prompt bytes. The skill/memory/goal/handoff identities were re-measured without raising any ceiling. Explicit maintainer decision for #5473/#5492. The v0.9.10 full surfaces intentionally add the safe read_media tool; their measured schemas remain below the prior byte/token ceilings. Representative prompt byte metrics now use the same host-independent normalized text as their identities; the normalized base is 6089 bytes. The v0.9.11 model-visible sub-agent surface intentionally retires six legacy agents/* tools in favor of the canonical agent tool; all affected schema and prompt metrics decrease. The v0.9.12 plugin prompt-match slice intentionally adds the request_plugin_install tool to the full tool surfaces (plan full: +518 schema bytes / +130 estimated tokens / 29 -> 30 tools) so a strong prompt match can surface the human review CTA; explicit maintainer decision for #5663/#5579. The v0.9.13 profile pins a non-executed bare bash shell so interpreter guidance is reproducible across hosts. The duplicate tts catalog entry is intentionally hidden; speech remains canonical and the alias remains available for saved-transcript dispatch. Explicit v0.9.13 maintainer decision (2026-09-08): after removing 1426 repeated guidance bytes and pinning the bash-v2 fixture, accept only the measured tool byte/token ceilings from all-features macOS source e27735bb63c897f88061c71567701fd971f5d396, verified libtest SHA-256 5e8cbe213f32c4ecdec63494c4de5e31857b4a40134edf7b21a55bca926b1b38: active 13274/3319 in every mode, Plan full 39885/9972, Act/Operate full 67603/16901, with no margin. Against the prior budget, active +390 bytes is agent -41 plus retained bash command syntax +431. Plan full also retains Git commit_plan +253, update_goal progress +583, github bounded local-report guidance +127, review complete-input refusal +35, and send_later dispatching status +14. Act/Operate full instead has github +2151 and additionally speech +230, hidden tts -2120, and tasks/automation exact model-route fields +274 each. The older budget predates v0.9.12: that tag had already removed 361 agent bytes and added the two 274-byte route fields; the retained initial increase versus the tag is 751 source-attributed bytes (agent +320, bash +431), not the +390 budget delta. Only the seven active definitions form the initial request; full catalogs include deferred tools. Estimated tokens use the existing bytes/4 heuristic, not provider usage or billing. Prompt, representative-context, skill-discovery and tool-name identities/ceilings are unchanged. Explicit v0.9.13 maintainer decision (2026-09-09): source ccc5dadfa2279545bf084d37cff3617e41ceaae2 intentionally exposes create_goal, get_goal and update_goal before continuation, so all three initial surfaces now contain ten tools. Measure exact source 4648d148eea64782be857eda6952af2c539cbfcc with the hosted macOS all-features libtest SHA-256 4485c88c7a8b66b8bb9a135807321e417a1e266457ecb74cffc7dfb92f850fc4: four exact provider-free metric tests pass. Active schemas are exactly 17847 bytes / 4462 estimated tokens (+4573 / +1143 for the three eager goal definitions); Plan full is 40597 / 10150 and Act/Operate full is 68315 / 17079, with no margin. The +712 full-catalog bytes are request_user_input guidance +358, update_goal state-change guidance +98, list_dir home-relative path guidance +56, explicit review max_passes schema +197, and three defer_loading true-to-false values +3. The three active name sets/digests, their counts, and measured active/full byte/token ceilings change; full name identities and all prompt, representative-context and skill-discovery measurements remain unchanged. This updates the earlier seven-tool initial-request receipt; bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.13 maintainer decision (2026-09-10): the +765 active/full tool-schema bytes since source 4648d148ee are exactly source-attributed \u2014 the agent tool's followup/parked-child continuation guidance (b6fad79373: +144 action description, +43 message parameter, +169 resume_from parameter) and the read tool's real output budget (e7f7c71e2c: +164 description, +245 for the new max_bytes parameter). No tool enters or leaves any surface: every name-set identity, count, and all prompt, representative-context and skill-discovery measurements are unchanged; only the measured byte/token ceilings move, to active 18612/4653 in every mode, Plan full 41362/10341, and Act/Operate full 69080/17270, with no margin. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.13 maintainer decision (2026-09-13): source 5bfe88c8e8f45266ea1f9abce87c76b2eed894af intentionally keeps the native workflow tool eager on Plan/Act/Operate first-turn surfaces (DEFAULT_ACTIVE_NATIVE_TOOLS; commit 9e49d0918). Active name sets gain `workflow` (10\u219211 tools); active schema ceilings move to 29402/7351 with margin pending exact Linux --update lock-in. Full catalogs already advertised workflow; only a small defer_loading true\u2192false spelling bump is reserved (+32 bytes / +8 tokens). Prompt, representative-context, and skill-discovery measurements are unchanged. Do not remove workflow from Plan. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.13 maintainer decision (2026-09-13, CI lock-in after b4d48e9a4): Linux Lint run 34766373771 measured the intentional eager `workflow` surface at active 31438/7860 in every mode, Plan full 50801/12701, and Act/Operate full 78513/19629. Prior ceilings (29402/7351 active, Plan full 41394/10349, Act/Operate full 69112/17278) under-counted the workflow schema body plus defer_loading true\u2192false on full catalogs; name-set identities are unchanged and `workflow` stays on Plan. Lock ceilings to those measured values with no margin. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.14 maintainer decision (2026-09-16): #5715 intentionally adds the two bounded read-only recall tools `session_get` and `session_search` to the Act/Operate full catalogs (50 -> 52 tools), so both full name-set identities and their digests move to 1203d192385fd2b02227ef9e4212e5379bc5cbb813a373f12539406b5958aef1. Plan full is unchanged: the session tools are not offered there. Measured on macOS aarch64 all-features from source 55a9e1b778fa; per the 2026-09-13 precedent the exact byte/token ceilings must be re-locked from a Linux Lint run if CI reports drift. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.14 maintainer decision (2026-09-17, #6319): re-measure from source b0866943b4 on macOS aarch64 all-features; per the 2026-09-13 precedent the exact byte/token ceilings must be re-locked from a Linux Lint run if CI reports drift. No tool enters or leaves any surface (52 tools; every name-set identity unchanged). Representative stages and system prompt grow exactly +742 bytes per stage/mode (base 6089 -> 6831, prompt 6084 -> 6826) for the deliberate dc32272f15 Bearing article plus mandate-first scope law; every stage re-hashed. Tool growth is exactly source-attributed per tool (measured per-tool at 55a9e1b778fa vs HEAD): agent +670 (2ca54ea8da #6282 output-token-cap parameter, 3c9c62571a spawn-requirement docs, 85d8dc7501 #6278 exact_files sentence, less 3dcb41f5d2 #6272 release-clause trim and the a7a8bdb338 token-allowance description removal), workflow +424 (a035914336 #6232 plan-child cwd property, serialized twice via phases/items and top-level children/items), Git +867 (b89349286f #6298 merge_tree verify surface), Run +158 (233da9fb76 #6296 bounded cwd), read +91 (f6fb5f42d1 #6283 size/truncated/line_count response fields), create_goal +76 (4de9e9e281 model-decides-goals description rewrite). Active 31453 -> 32714 (+1261 = agent +670, workflow +424, read +91, create_goal +76); Plan full 50816 -> 52944 (+2128 = active +1261 plus Git +867); Act/Operate full 79531 -> 81817 (+2286 = Plan full +2128 plus Run +158). bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.14 maintainer decision (2026-09-19 overnight): Code Mode Phase-1 advertises `execute_tools` on Act/Operate full catalogs (52 -> 53 tools; Plan full unchanged). Default Direct mode keeps it deferred (`defer_loading: true`), so active surfaces are unchanged. Name-set identity digests move with the sorted name list; measured tool JSON is 884 bytes (+885 including the catalog comma) so Act/Operate full ceilings lock to 82702/20676. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.10.0 source reconciliation (2026-09-19): retain the existing agent cwd parameter (+256 serialized bytes in every surface) and tasks create name parameter (+121 bytes in Act/Operate full only), both already present before the preceding execute_tools lock-in. Linux CI run 35490467762 measured active 32970/8243, Plan full 53200/13300, and Act/Operate full 83079/20770. Each increase is attributed to exactly one schema property; no margin, tool identity, prompt or representative-context change. These are schema byte/4 estimates, not billed tokens.",
3 "document_kind": "codewhale.runtime_contract_budget",
4 "representative_context": {
5 "fixture_id": "representative-v1",
6 "stages": {
7 "base": {
8 "bytes": 6831,
9 "identity_sha256": "47ec39be875d57bc86c0a490be4a7a31bee577011eb3707df51b706901d17b22"
10 },
11 "goal": {
12 "bytes": 8878,
13 "delta_bytes": 81,
14 "identity_sha256": "4e079faa30af095d903a630608174198ad72a00901900fd55a645e8d95b352b9"
15 },
16 "handoff": {
17 "bytes": 9266,
18 "delta_bytes": 388,
19 "identity_sha256": "05430757b2d2cd9627c6521c55706882261cb0d7bc33cbef1cca347ddd6a8276"
20 },
21 "instructions": {
22 "bytes": 7217,
23 "delta_bytes": 131,
24 "identity_sha256": "c6f4d19bb203f82f85c7852ee31310fe426441371850126682cd86d03596289c"
25 },
26 "memory": {
27 "bytes": 8797,
28 "delta_bytes": 963,
29 "identity_sha256": "45284c41712d97cca11e13f90c12c757bcd3111f7421c67661a88f04dbaa9b23"
30 },
31 "project": {
32 "bytes": 7086,
33 "delta_bytes": 255,
34 "identity_sha256": "7e2993b93e1dcadef95c70d3c502984423e9800ee037743ab6101c67831dd83e"
35 },
36 "skill": {
37 "bytes": 7834,
38 "delta_bytes": 617,
39 "identity_sha256": "28a1b91473b992440d0d2ffcd0db79cb9a7b5582c521e527bbb855b5fa506dbb"
40 }
41 },
42 "system_prompt_blocks": 6,
43 "total_bytes": 9266,
44 "total_tokens_est": 2317
45 },
46 "schema_version": 1,
47 "skill_discovery": {
48 "first_delta": {
49 "directories_visited": 1,
50 "root_discovery_calls": 1,
51 "skill_md_read_attempts": 1
52 },
53 "second_delta": {
54 "directories_visited": 0,
55 "root_discovery_calls": 0,
56 "skill_md_read_attempts": 0
57 }
58 },
59 "system_prompt": {
60 "modes": {
61 "act": {
62 "mode_instructions_bytes": 0,
63 "mode_instructions_tokens_est": 0,
64 "system_prompt_blocks": 4,
65 "system_prompt_bytes": 6826,
66 "system_prompt_tokens_est": 1707
67 },
68 "operate": {
69 "mode_instructions_bytes": 0,
70 "mode_instructions_tokens_est": 0,
71 "system_prompt_blocks": 4,
72 "system_prompt_bytes": 6826,
73 "system_prompt_tokens_est": 1707
74 },
75 "plan": {
76 "mode_instructions_bytes": 0,
77 "mode_instructions_tokens_est": 0,
78 "system_prompt_blocks": 4,
79 "system_prompt_bytes": 6826,
80 "system_prompt_tokens_est": 1707
81 }
82 }
83 },
84 "tool_catalog": {
85 "execution_shell": "bash",
86 "modes": {
87 "act": {
88 "active": {
89 "bytes": 32970,
90 "identity_sha256": "cf523fcd7528ab2e14efffd7fe6b0916a370d6a8b8426d15a159aa823a7b86ce",
91 "tokens_est": 8243,
92 "tool_names": [
93 "agent",
94 "bash",
95 "create_goal",
96 "edit",
97 "get_goal",
98 "read",
99 "todo_write",
100 "tool_search",
101 "update_goal",
102 "workflow",
103 "write"
104 ],
105 "tools": 11
106 },
107 "full": {
108 "bytes": 83079,
109 "identity_sha256": "45e989bbe5ac0bb1f2d9084361c021009f30a06539f132ebd4fd2331a1bb1954",
110 "tokens_est": 20770,
111 "tool_names": [
112 "Git",
113 "Run",
114 "Web",
115 "agent",
116 "apply_patch",
117 "automation",
118 "bash",
119 "create_goal",
120 "diagnostics",
121 "edit",
122 "execute_tools",
123 "file_search",
124 "fim_edit",
125 "finance",
126 "get_goal",
127 "github",
128 "grep_files",
129 "handle_read",
130 "harness",
131 "list_dir",
132 "load_skill",
133 "lsp",
134 "note",
135 "notify",
136 "project_map",
137 "read",
138 "read_media",
139 "request_plugin_install",
140 "request_user_input",
141 "retrieve_tool_result",
142 "revert_turn",
143 "review",
144 "send_later",
145 "session_get",
146 "session_search",
147 "speech",
148 "task_shell_start",
149 "task_shell_wait",
150 "tasks",
151 "terminal/cancel",
152 "terminal/reset",
153 "terminal/run",
154 "terminal/send",
155 "terminal/wait",
156 "todo_write",
157 "tool_search",
158 "tui_help",
159 "update_goal",
160 "validate_data",
161 "verify",
162 "web.run",
163 "workflow",
164 "write"
165 ],
166 "tools": 53
167 }
168 },
169 "operate": {
170 "active": {
171 "bytes": 32970,
172 "identity_sha256": "cf523fcd7528ab2e14efffd7fe6b0916a370d6a8b8426d15a159aa823a7b86ce",
173 "tokens_est": 8243,
174 "tool_names": [
175 "agent",
176 "bash",
177 "create_goal",
178 "edit",
179 "get_goal",
180 "read",
181 "todo_write",
182 "tool_search",
183 "update_goal",
184 "workflow",
185 "write"
186 ],
187 "tools": 11
188 },
189 "full": {
190 "bytes": 83079,
191 "identity_sha256": "45e989bbe5ac0bb1f2d9084361c021009f30a06539f132ebd4fd2331a1bb1954",
192 "tokens_est": 20770,
193 "tool_names": [
194 "Git",
195 "Run",
196 "Web",
197 "agent",
198 "apply_patch",
199 "automation",
200 "bash",
201 "create_goal",
202 "diagnostics",
203 "edit",
204 "execute_tools",
205 "file_search",
206 "fim_edit",
207 "finance",
208 "get_goal",
209 "github",
210 "grep_files",
211 "handle_read",
212 "harness",
213 "list_dir",
214 "load_skill",
215 "lsp",
216 "note",
217 "notify",
218 "project_map",
219 "read",
220 "read_media",
221 "request_plugin_install",
222 "request_user_input",
223 "retrieve_tool_result",
224 "revert_turn",
225 "review",
226 "send_later",
227 "session_get",
228 "session_search",
229 "speech",
230 "task_shell_start",
231 "task_shell_wait",
232 "tasks",
233 "terminal/cancel",
234 "terminal/reset",
235 "terminal/run",
236 "terminal/send",
237 "terminal/wait",
238 "todo_write",
239 "tool_search",
240 "tui_help",
241 "update_goal",
242 "validate_data",
243 "verify",
244 "web.run",
245 "workflow",
246 "write"
247 ],
248 "tools": 53
249 }
250 },
251 "plan": {
252 "active": {
253 "bytes": 32970,
254 "identity_sha256": "cf523fcd7528ab2e14efffd7fe6b0916a370d6a8b8426d15a159aa823a7b86ce",
255 "tokens_est": 8243,
256 "tool_names": [
257 "agent",
258 "bash",
259 "create_goal",
260 "edit",
261 "get_goal",
262 "read",
263 "todo_write",
264 "tool_search",
265 "update_goal",
266 "workflow",
267 "write"
268 ],
269 "tools": 11
270 },
271 "full": {
272 "bytes": 53200,
273 "identity_sha256": "ac8af1f4988199825be7b00b054c258724a44074b1d4e6de6c92ade7c1cffe63",
274 "tokens_est": 13300,
275 "tool_names": [
276 "Git",
277 "Web",
278 "agent",
279 "automation",
280 "bash",
281 "create_goal",
282 "diagnostics",
283 "edit",
284 "file_search",
285 "get_goal",
286 "github",
287 "grep_files",
288 "handle_read",
289 "list_dir",
290 "load_skill",
291 "notify",
292 "read",
293 "read_media",
294 "request_plugin_install",
295 "request_user_input",
296 "review",
297 "send_later",
298 "tasks",
299 "todo_write",
300 "tool_search",
301 "update_goal",
302 "validate_data",
303 "web.run",
304 "workflow",
305 "write"
306 ],
307 "tools": 30
308 }
309 }
310 },
311 "surface_profile": "production-default-builtins-no-mcp-no-host-interpreters-bash-v2"
312 }
313 }
314
314 lines JSON