Contents

A model router for Claude Code, and why it should live on my GPU

For people running Claude Code who wonder whether a small classifier can send their prompts to cheaper models. Short answer: a little, with caveats. The code lives in my private infra repo; the numbers below are from its log.


My timeline has been full of “what Jev can do for you” posts. Jev sizes your request, Jev picks the right model, Jev saves you money. Fine. But every one of those setups has the same shape: before any work happens, your prompt goes to someone else’s model, just so that model can answer with one word.

That bothered me enough to go looking for alternatives. Open weights, hosted locally, on hardware I already own. This post is what came out of the first day: a model router for Claude Code, first built the way everybody builds it, with Jev in the cloud, and then moved onto a local classifier called Kev.

Before anything else, the honest part:

  • The savings are small. Claude Code can’t switch the main session’s model per message. The main model still reads every prompt and then hands the job to a helper agent. If the main session runs on opus and the router picks opus, nothing is saved.
  • The cloud timeout overshoots. I set 3 s. The log shows calls of 4.07, 4.31 and 4.41 s. More on why below.
  • “Jev” isn’t what answers. The router asks OpenRouter for typesafe/jev-router. The log says the answers came from three different general-purpose LLMs. What that alias actually is, I don’t know.
  • With the cloud backend, prompts leave the machine. Every prompt goes to OpenRouter and on to whichever vendor serves it. Infra details included. Only the Kev backend keeps them local.
  • Kev hasn’t sized a real prompt yet. The local backend is built and the server runs. But so far the router’s log has only OpenRouter answers, plus two “kev-serve not listening” errors, logged while it was stopped.

It starts as a skill

I didn’t write the router by hand. I wrote a Claude Code skill, model-router, and let Claude build it. (The skill lives here.) The skill is mostly a prompt, written the way I would ask a colleague:

build me a model router that uses Jev. the idea: Jev sizes up every message i send, and small jobs go to a smaller, cheaper AI model instead of the biggest one.

It then lists what I want: four sizes (tiny, everyday, large, hardest), one helper agent per size, each pinned to a model I actually have. When Jev is less than 60% sure, or my message is a short reply like “yes do that”, the main session handles it itself. It needs an on/off switch, a status check with counts and cost, and it has to be OFF by default. And it must never slow me down or block a message.

Two preconditions in the skill turned out to matter more than the recipe:

  • “Find Jev. Do not guess what Jev is.” If the config can’t be found, stop and ask. That one line is why the router logs who actually answered, and that log is where the most interesting finding came from.
  • “Be honest about limits.” Claude Code can’t swap the main model per message, so the skill makes the builder say so up front instead of promising a switch it can’t deliver.

The skill ends with a test: 8 sample messages from tiny to hardest, plus 2 short replies, and a table of what was picked and how sure the sizer was. It also has to end with a reminder that while the router is on, my messages pass through OpenRouter. That reminder is what eventually pushed me towards Kev.

What got built

Claude Code has hooks. UserPromptSubmit fires before a prompt reaches the model, gets the event as JSON on stdin, and can return extra context for the model. That is enough to build a router:

The model router on my workstation, with Kev-4B as the local sizer. Everything inside the pink box sees my prompts when the cloud backend is used.

  1. I type a prompt. The hook pipes the event into router.py, a stdlib-only Python script, with a 5 s hook timeout.
  2. Control words are handled locally: router on, router off, router status and router backend auto|kev|openrouter. The script blocks those prompts, so they never reach any model.
  3. If state/enabled doesn’t exist, the script exits silently. That file is the on/off switch.
  4. Otherwise it asks a sizer to pick one of five labels: tiny, everyday, large, hardest, or contextual for replies that only make sense inside the conversation. Which sizer depends on the backend (next section).
  5. If the answer is a tier and confidence is at least 0.6, the script returns additionalContext: “Jev sized this message as ’large’ (confidence 0.78). Delegate the whole task to the router-large subagent (runs on opus)…”. Anything else routes to “self” and prints nothing.
  6. The main session delegates to router-tiny, -everyday, -large or -hardest, pinned to haiku, sonnet, opus and fable. Each must end its answer with a Model: line, so I can see who did the work.

It fails open. Any exception means no output and exit 0, and the prompt goes through as if the router wasn’t there. A router that can block my work is worse than no router.

And yes: it routed the request to write up its own first session. “large”, 0.78, opus.

Three backends

The router has a backend setting, kept in state/backend per project:

  • openrouter: cloud only. A chat-completions request to typesafe/jev-router, asking for {"size", "confidence"} as JSON.
  • kev: local only. A typed request to Kev on 127.0.0.1:8008. If Kev is stopped, there is no routing, and nothing leaves the machine.
  • auto (the default): Kev if it’s listening, OpenRouter otherwise.

auto is convenient and a bit of a trap. Kev runs on demand, so whenever it’s stopped, auto quietly means “cloud”. For anything private, I use kev or switch the router off. In this blog’s repo it’s set to kev.

Jev in the cloud: what the log says

The first 16 prompts all went through OpenRouter:

routed tocount
tiny (haiku)3
everyday (sonnet)1
large (opus)1
self4
error7

Total cost for sizing: about $0.0032. Latency across the 9 successful calls: 2.99 s on average, from 1.26 to 4.41 s. That’s three seconds before Claude even starts thinking about my prompt. The 7 errors are all ValueErrors, most likely replies without a {…} in them. The first version didn’t log the reply; the current one does, so the next failure will say what came back.

Two things surprised me.

The label lies. I thought I was calling TypeSafe’s Jev, a closed “System One” decision model that answers typed choice and score questions in a single forward pass. The answered_by field in the log says otherwise: openai/gpt-6-luna five times, deepseek/deepseek-v4.1-flash three times, moonshotai/kimi-k3 once. So the router works, but it isn’t calling what I thought it was. A general-purpose LLM, asked nicely to reply with JSON, will sometimes reply with prose instead. That probably explains the errors.

Timeouts aren’t wall-clock. urllib’s timeout=3 applies to each socket operation, not to the whole request. Connecting, sending and waiting for the first byte each get their own 3 s. So a 4.4 s call is still “within” a 3 s timeout. What Claude Code does when it kills a hook at 5 s I haven’t tested. Unknown.

Kev on my GPU

Why send every prompt to a vendor to get a one-word answer, when there’s an RTX 4090 right here?

The candidate is Kev-4B (jaredpalmer/kev-4b, Apache-2.0): a LoRA plus a classification head on Qwen3.5-4B-Base, serving a Jev-shaped POST /v1/systemone. It reports 0.817 against Jev’s 0.857 on Kev’s own Transfer-v4 suite. That’s self-reported, on their benchmark, not mine. I looked at two alternatives. Kev-9B buys +0.005 for about 17 GB of VRAM, so no. Von (395M) is the fallback if Kev-4B turns out to be too heavy.

Kev runs as kev-serve, a local server installed by an Ansible role:

  • it pins the Kev commit, since upstream has no tags or releases
  • it uses Python 3.12 via uv, because my system python3 is 3.14 and Kev wants <3.14
  • it’s a systemd user unit with no [Install] section, so it can’t be enabled by accident: systemctl --user start kev-serve when I want it, stop before I play
  • it binds 127.0.0.1:8008, with no auth unless KEV_API_KEY is set

Measured warm latency: 60–130 ms, against 1.3–4.4 s for OpenRouter. GPU memory was 16.7 GB with the plain kernels and 20.4 GB with the fused ones, at the same latency for one request at a time. So fused stays off. The games need the rest of the card, and that’s also why Kev is on demand rather than always on.

I measured both sizers on latency and reliability:

Kev-4BJev (OpenRouter)
Latency60–130 ms700–1400 ms
Speed edge10–15× faster—
Reliability100%62.5% (API errors)

An important note on accuracy: I initially benchmarked both sizers against a small set of sample prompts, but that methodology turned out to be invalid. The labels I derived (tiny, everyday, large, hardest) had essentially no correlation with actual task complexity (correlation = -0.017). A corrected benchmark using 2,787 real prompts labeled by observed output cost shows Kev-4B at 32.7% accuracy (vs 25% random guessing). Per-tier recall is uneven: 76% for everyday, but 12% for tiny, 14% for large, 0% for hardest. Kev tends toward “everyday”.

Confidence is monotonically informative despite low overall accuracy: 55% accuracy in the top confidence bin (0.5–0.6) versus 32.7% overall. However, Kev’s confidence never reaches the 0.6 gate in practice, so the router has never actually routed a real prompt to a smaller model. Every prompt falls through to the main session. The local path has been a no-op in practice, and measurement is what surfaced this.

Where the real win is: Kev is an order of magnitude faster than the cloud sizer (60–130 ms vs 700–1400 ms), and it knows when it’s unsure. Crucially, all processing stays local when using the Kev backend, which matters for privacy.

Plugging Kev in wasn’t a URL swap

I hoped to change a base URL and be done. No. Kev doesn’t speak chat completions. It speaks the typed /v1/systemone API: you send a state (my prompt) and a set of questions. The router sends exactly one, a choice question called size, with the five labels and their descriptions as criteria. Kev answers with answers.size.choice and a probability for every option.

Two details matter:

  • Kev’s confidence isn’t Jev’s confidence. Kev’s own confidence field is rescaled so that 0 means “uniform over the options”. It doesn’t fit a 60% gate at all. The router uses the probability of the chosen label instead. Same threshold, different number, and without reading the API that would have been a silent bug.
  • A stopped port doesn’t refuse, it hangs. On my workstation (WSL2), connecting to the port while kev-serve is stopped doesn’t fail fast. It waits for the full timeout. With Kev on demand, that would add 1.5 s to every prompt whenever Kev is off. So before calling Kev, the router reads /proc/net/tcp and checks whether anything is listening on 8008. Reading a file is instant; a hanging connect() isn’t.

Where the savings actually are

The math is unkind. The main model reads every prompt, and delegation adds a round trip. Routing pays off only when the main session is expensive and the task is small: opus in the main seat, a rename handed to haiku. For everything else I pay for sizing plus delegation and get roughly what I would have gotten anyway.

Kev changes one part of that. Sizing becomes free and fast, so the overhead is only the delegation. Whether “tiny” is common enough to matter, 16 log lines can’t tell. The router status counter will, eventually.

What’s not done

  • Building a feedback loop. The router has a router feedback command (not yet wired in) that logs which model actually solved the task. Feeding those results back into the sizer’s training data would fix the accuracy problem.
  • Measuring Kev on your workload distribution. The corrected benchmark used 2,787 real prompts from the session archive; the router’s real job is to be accurate on your 60–70 prompts a week over months. That workload may differ.
  • The skill still describes a Jev-based router. It should be rewritten around Kev, with Jev as the optional cloud fallback.
  • Whether Kev should be reachable from other machines is undecided. Without an API key, no.

Lessons

  • Read the log, not the label. “jev-router” was answered by three different LLMs. Log who actually answered.
  • Write “don’t guess” into the skill. The line that forbade guessing what Jev is did more for correctness than any other part of the recipe.
  • A timeout is per socket operation, not per request. Measure wall-clock yourself if it matters.
  • Same threshold, different meaning. Two sizers can both return a “confidence” and mean different things by it. Confidence is calibrated when measured correctly: 55% accuracy in the 0.5–0.6 bin, 34% in the 0.4–0.5 bin.
  • Fail-open is a feature, not a bug. Routing only wins when the router is right. Kev defaults to “self” on uncertain prompts. In practice, Kev never reaches the 0.6 confidence gate, so everything goes through unrouted for now. That is honest measurement, not a bug in the sizer: it knows what it doesn’t know.
  • Routing only pays when the main model is expensive and the task is small. Opus in the main seat, haiku for renamed variables.
  • Local beats remote on privacy and latency (60–180 ms vs 700–1400 ms), at the cost of 14–17 GB of VRAM on a gaming GPU. Hence on demand, and hence backend kev for anything private.