A local control plane that routes engineering tasks across Claude Code, OpenCode, and Codex — capability-tier routing, deterministic fallback, durable task state. A real tool used daily, not a demo built for this portfolio.
AI-assisted development means juggling several CLIs (Claude Code, OpenCode, Codex), each with its own agents, skills, and models. Left manual, every task means re-deciding which tool to use, which model fits, what happens if a provider fails mid-task, and how not to lose progress on interruption.
Without a routing policy, tasks get sent to the wrong model for their difficulty (wasting cost on easy tasks, or worse, under-powering hard ones), failures aren't retried consistently, and there's no shared record of what's actually installed across environments.
A Python control plane that classifies a task, selects skills/agent/model against real capability tiers, dispatches to whichever CLI backs that model, falls back deterministically on failure, and persists task state to SQLite so nothing is lost on interruption.
TASK -> CLARITY GATE (local, deterministic, skips triage for obvious tasks) -> TRIAGE (cheap-model classification for ambiguous tasks) -> SKILL SELECTOR (ranks 3-10 relevant skills from a 3,812-skill registry) -> AGENT SELECTOR (301 agents, 164 live + 137 offline) -> MODEL SELECTOR (enforces capability-tier floor, never silently downgrades) -> DISPATCH (real CLI call: claude -p / codex exec / opencode run) -> FALLBACK (classify error, retry next candidate, state preserved) -> TASK STATE (SQLite, durable across restarts) -> USAGE LOG (real latency, routing decisions, fallback events)
$ uv run python orchestrator/repl.py > fix the bug in src/auth.ts line 42 [clarity gate] clear=True — triage skipped [skill selector] -> debugging-strategies, typescript-expert, error-handling [model selector] -> tier B floor -> kimi-k2.7-code (opencode-go) [dispatch] opencode run -m kimi-k2.7-code ... [usage] logged: latency=4.8s, provider=opencode-go, fallback_used=false
No subagent spawning. A deliberate choice: all agent invocation is inline rather than spawned, to avoid the token overhead of cold-start subagents.
Two-stage triage. A free, local, deterministic "clarity gate" skips the cheap-model classification call for obviously unambiguous tasks, with an asymmetric-cost design — a false "clear" is dangerous, so every rule is deliberately narrow.
Real, benchmarked routing. Models aren't tier-rated by assumption — 9 of 26 available models were benchmarked against real tasks and judged from actual output; two real failures were found this way and correctly excluded from routing.
Untrusted-by-default trust boundary for web pages, downloaded documents, and skill files. Secret scanning reports only FOUND/MISSING/INVALID/UNVERIFIED, never values. Destructive operations (uninstalls, force push, mass deletion) are ASK/DENY by default, not ALLOW.
uv sync before this case study was written.