The workbench
Search, read, edit, diff, shell, build, tags, task memory, and caller-bound verification surfaces.
Hands and controlsThe workbench gives the hands. Evidence gives the proof. Memory preserves context. Playbooks decide when to route, how to review, what to preserve, and where boundaries stay hard.
Playbooks are not a list of prompts. They are the operating layer that makes Codex and J.A.R.V.I.S behave consistently across reviews, bug hunts, wiki work, Source of Truth updates, contracts, and handoffs.
Search, read, edit, diff, shell, build, tags, task memory, and caller-bound verification surfaces.
Hands and controlsFetched corpus proof stays separate from durable wiki memory, with authority boundaries kept explicit.
Proof and recallProcedures, trigger routing, helper review, safety rules, and reusable behavioral patterns load at the moment of need.
Repeatable behaviorThe playbook catalog makes a broad agent feel like a team. Each request can land in a concrete workflow instead of relying on memory, vibes, or a one-off answer.
`review-helper` runs the local repo-grounded pass and routes bounded helper opinions through AGY, Codex, or ChatGPT handoff when useful. Findings stay hypotheses until locally verified.
The current custom catalog is focused on durable operator workflows, not one-off preferences. Each mature playbook gets wiki-backed recreation guidance so the behavior can survive beyond one machine.
A skill can change how every future agent acts. The skill-wiki-manager model keeps that power attractive without letting it become casual self-modification.
The active skill package owns behavior. The skills wiki owns durable explanation and recreation notes. The coding workbench may report readiness, but it must not mutate installed skills. Long-term memory stores context, not proof.
The accepted direction is not autonomous self-improvement. It is proposal-only distillation: useful session residue is classified, labeled, verified, and promoted only through explicit owner-approved work.
A session reveals repeated success, repeated failure, a durable decision, or a missing validation pattern.
The residue becomes a skill patch candidate, wiki memory candidate, SoT update candidate, follow-up task, or no action.
Evidence, uncertainty, authority surface, safety boundary, and rejected paths stay visible.
Repo evidence, local checks, helper review, and normal gates decide whether the candidate is real.
Only explicit operator-approved work mutates skills, wiki, SoT, task state, git state, or runtime behavior.
Playbooks complete the operating stack: the coding workbench provides controlled action, the evidence engine provides grounded knowledge, and durable skill guidance turns repeated expert behavior into portable operating doctrine.