Cross-Harness Model Routing
Route agent work by capability axis: use the model with the strongest taste and judgment to plan, judge, and orchestrate; use Codex for well-specified execution, computer use, and UI verification when it is the better substrate.
Trigger
Use this workflow when:
- A Claude Code workflow or subagent needs work that Codex handles better.
- A task can be split into taste/judgment and mechanical execution.
- A high-taste model should not spend tokens on bulk edits, migrations, or repeated verification.
- User-facing UI, copy, or API design needs a stronger taste pass than the executor can provide.
- End-to-end agent PRs are producing too many branches that get closed after review.
Core rule
Do not pick one model for the whole run. Split the run by role.
| Role | Route |
|---|---|
| Taste, judgment, task decomposition, plan review | Fable/Opus-class Claude model, or the best current equivalent |
| User-facing UI, copy, API design | Taste model or explicit human review before shipping |
| Bulk implementation, migrations, data transforms | Codex via a self-contained codex exec prompt when the spec is clear |
| Computer use and UI/UX verification | Codex/computer-use route with browser or desktop receipts |
| Independent review | A second model or harness, preferably a different family |
Cost is a tie-breaker, not the primary objective. For work that ships, correctness and taste outrank cheapness.
Cursor's SQLite swarm is a useful cost counterexample: planner choice and worker choice interacted, workers consumed most tokens, and a cheaper hybrid worker fleet beat a much more expensive uniform fleet on the reported run. Route and price the full planner + workers + reconciliation + verification tuple. Do not infer fleet economics from the planner model, a single token price, or agent count. Source: Cursor, “Agent swarm model economics,” 2026-08-07
Candidate admission gate
A launch post, leaderboard, context-window claim, open-weight release, or low token price may add a candidate; it cannot change the default route. Before a new model/provider/effort setting enters a workflow:
- Freeze the representative task set, repository revision, full harness, system/project instructions, skills, tools, context and compaction policy, sandbox, retry policy, and verifier.
- Pin the model or checkpoint, provider, API behavior, effort, temperature, maximum output, and input/output/cost/wall-time budgets.
- Run the incumbent and candidate under the same condition. Record per-task outcomes, failures, retries, tokens, latency, cost, variance, and artifacts.
- Change one declared axis for causal ablations. Treat model × harness or effort × context interactions as separate experiments.
- Promote the measured tuple—workflow + model + provider + effort + harness configuration—not a model name in isolation. Add rollback and revalidation dates because provider behavior, price, and availability drift.
GLM-5.2's complete launch chart is the concrete warning. Its score improves from non-thinking to high and then only modestly at max while average output roughly doubles toward 84K tokens per task; the chart also lacks variance, latency, and cost axes. That is enough to schedule a bounded pilot and define the required receipt, not enough to replace an incumbent. Source: X/@Zai_org, current replay 2026-08-11; GLM-5.2
Kimi K3, GLM-5.2, public design arenas, and planner–coder–judge anecdotes remain
candidate inputs under the same rule. Pin the released weights or exact hosted
endpoint, license, provider, arena submission condition, visual tool loop, and
complete cost; a text-only builder paired with a browser critic is one routed
tuple, not evidence that either model won alone. As of this review Kimi K3 has a
current public paper and source tree, but its social arena rank and individual
one-shots still do not establish Kevin's default without the frozen local suite.
Source: Kimi K3 technical report; MoonshotAI/Kimi-K3@3cb39df; saved GLM,
Kimi, Fable/GPT routing and visual-arena signals, reviewed 2026-08-12
Production router contract
A production router may optimize cost, time-to-first-token, throughput, deadline reliability, or task quality, but provider telemetry can only optimize the dimensions it measures. Cost/TTFT/TPS sorting is not semantic-quality routing. The route manifest therefore declares task class, allowed models/providers/service tiers, privacy/data-retention boundary, primary quality gate, latency deadline, price snapshot, retry/fallback policy, exploration budget, and rollback owner. Source: X/@vercel_dev AI Gateway sort release, 2026-05-18; X/@RampLabs Router sequence, 2026-07-20
Before live use, replay a frozen set and shadow the incumbent. Capture both the chosen and counterfactual route, prediction, quality result, tokens, latency, cost, retries, provider failure, and terminal artifact. Promote only task classes that meet the quality floor and SLO with uncertainty bounds; retain an incumbent fallback. Revalidate after model, provider, price, prompt, tool, grader, or harness changes. Short-lived free access, launch pricing, internal benchmarks, context-window size, and one-shot demos can admit candidates but never overwrite this contract.
Security profiles are orthogonal to intelligence profiles. A copied Codex
configuration that combines approval_policy = "never" with
danger-full-access is not a performance preset; it is an authority decision
that must be scoped to an intentionally isolated disposable environment. Model
and effort settings must be verified against the installed runtime's supported
config keys, while approval and sandbox policy follow the task's risk and
mutation boundary.
Guarded eval-authoring operator
OpenRouter's Ori Eval skill is retained as a procedure candidate, not an installed default. Its useful sequence is to locate model-call surfaces, ask a small set of sequential product questions, pin run and judge conditions, evaluate tool-called, tool-avoided, and answer/terminal quality, include the incumbent, preserve prompt state, rank candidates, regression-test the winner, report timing and cost, and permit “no change” as the result.
The fetched distribution does not meet the active-install bar. The skill is a
mutable remote URL; it expects an automatic closed-binary install; the install
script selects mutable release channels, sends installation telemetry unless
disabled, edits shell startup files, may create /usr/local/bin/ori through
sudo, and obtains checksums from the same release origin. Running an eval can
scan code, require OpenRouter authentication and spend, write scratch state,
and use a correlated run/judge model. Any future pilot must pin the skill hash,
installer and binary version/checksum; disable telemetry; prohibit shell-rc,
global-path, and sudo mutation; isolate code and data; obtain explicit budget
approval; use deterministic checks or an independent judge; include the
incumbent; and preserve all failures, retries, costs, and artifacts. No Ori
binary was installed or executed during this review. Source: OpenRouter Ori
skill SHA-256 24ee54ab75b1c62d0306d75d0fc9196d4095013fc33015495e85ba3d307d7a29
and installer SHA-256
9134c83cd1029c80bda688b8bfb873a7d446aa0826d3b9493da2b66ab34fcda7,
fetched and read 2026-08-12
Operating pattern
- The strong orchestrator reads the spec and writes a self-contained work packet.
- If the work is mechanical or verification-heavy, the orchestrator spawns a thin wrapper/subagent.
- The wrapper calls Codex CLI (
codex execorcodex review) with the packet. - Codex returns a patch, review, or verification receipt, not vague status.
- The orchestrator judges the output against acceptance criteria.
- If the output misses the bar, rerun or escalate instead of accepting mediocre work.
Thin wrapper contract
The wrapper is not the worker. It writes the prompt, calls Codex through the shell, and returns the result.
The prompt should include:
- repo path and branch
- files or subsystems in scope
- constraints and out-of-scope boundaries
- expected output format
- validation commands
- sandbox mode
- what evidence to return
Use read-only Codex runs for investigation and analysis. Use write-capable runs only when a patch or commit is intended. Always return receipts: commands run, files changed, screenshots or URLs for UI verification, and any failures.
Kevin stack defaults
Default to Codex for computer use, UI verification, and well-specified implementation when available. Default to Fable/Opus-class Claude models for plan review, implementation review, user-facing taste, API design, and copy. User-facing surfaces need a taste model or a human review gate.
Treat exact model names as dated defaults, not constants. Availability, pricing, and model quality change; the durable invariant is the split between judgment, execution, and verification.
Source snapshot
Theo's July 2026 CLAUDE.md example ranked models by cost, intelligence, and taste, then routed work accordingly. The table was an operator ranking, not a universal benchmark.
| Model | Cost | Intelligence | Taste |
|---|---|---|---|
| gpt-5.5 | 9 | 8 | 5 |
| sonnet-5 | 5 | 5 | 7 |
| opus-4.8 | 4 | 7 | 8 |
| fable-5 | 2 | 9 | 9 |
The useful procedure from the example:
- Bulk/mechanical work goes to GPT/Codex when it is effectively cheap.
- User-facing UI, copy, and API design need taste >= 7.
- Plans and implementations should be reviewed by a high-taste model; a second independent perspective is useful.
- When a workflow can only select Claude models directly, use a thin Claude wrapper to call
codex exec. - If the cheap run misses the bar, rerun or escalate.
Kevin's observation after adopting this workflow: Codex is still materially better for computer use, UI/UX verification, and efficient execution of well-specified work; routing that way reduced discarded end-to-end agent PRs from roughly half to zero in a day. Source: User, 2026-07-06
Failure modes
- Delegating vague work to a cheap executor produces bad PRs.
- Letting a wrapper summarize without receipts hides failures.
- Making cost primary ships mediocre work.
- Spending a taste model on mechanical chores burns scarce tokens.
- Using one model family to judge itself creates correlated blind spots.
Relationship to other pages
Borrowed Intelligence (Drain It into Plans) banks judgment into plans. This workflow drains judgment into live delegation packets. Claude Code Dynamic Workflows supplies the subagent/workflow dispatch layer. Cross-Agent Harness Portability keeps instructions, skills, and receipts portable so delegation does not drift per agent.
Timeline
-
2026-08-12 | Added the production-router manifest, replay/shadow/promotion sequence, counterfactual receipts, incumbent fallback, and risk-separated Codex configuration rule. Provider cost/TTFT/TPS sorting is operational telemetry, not semantic-quality routing. Source: Vercel AI Gateway sort, Ramp Router, free-access and Codex-config source cohort
-
2026-08-12 | Reconciled Kimi K3, GLM-5.2, design-arena, Browser Use, OrcaRouter, and Fable/GPT planner–coder–judge signals through the existing candidate gate. The admissible unit is the pinned model/provider/effort/tools/ verifier tuple; no social rank, one-shot, or anecdotal cost changed the default. Source: Kimi K3 report/current source; 14 saved routing and visual signals
-
2026-08-12 | Added planner/worker fleet economics: price and evaluate the complete routed tuple because worker tokens, context allocation, and reconciliation dominate some swarm runs. Source: Cursor agent-swarm model economics
-
2026-08-11 | Added the machine-checkable candidate-admission shape from the GLM-5.2 replay: freeze the complete condition, capture quality and run-economics receipts, ablate declared axes, and promote only the measured model/provider/effort/harness tuple. Source: X/@Zai_org; GLM-5.2
-
2026-07-06 | Created from Theo's CLAUDE.md routing section and Kevin's Codex/Fable operating note. Durable rule: high-taste model orchestrates and judges; Codex executes and verifies well-specified work where it is stronger. Source: User-provided screenshot and X/@theo, 2026-07-06