An agent loop makes dozens of small calls that aren’t reasoning at all. Is this ticket about billing. Is this search result relevant. Did the tool call succeed. Should this story be a blog post. Most agent harnesses send every one of those to the same frontier model that writes the code, with a prompt that ends in “return JSON.”

I measured what that costs. The same six-question classification, run 150 times per model: a System One model answered in 163 ms at $53 per million calls. Claude Opus 5 took 2.7 seconds and $11,183. That’s roughly 17x the latency and 200x the price for a yes/no question.

But the savings only survive if the loop around the model owns the thresholds. My data says that part is harder than the model swap.

A System One model answers in types, not text

The name comes from Kahneman’s split between fast, intuitive System 1 and slow, deliberate System 2. TypeSafe’s Jev is the first model sold under that label. It takes a state (text or JSON) and a set of typed questions, and it returns three kinds of answer: a choice from options you list, a score against a rubric you write, and a noul, the probability that a statement is true.

It doesn’t write text. It doesn’t call tools. TypeSafe’s own docs say it plainly: Jev is not a drop-in replacement for the LLM behind Claude Code or Cursor. You keep the LLM for the work that needs one and call Jev for the judgments.

That’s the same argument I made in April about where agent performance comes from: the loop matters more than the weights. A System One model is a new part to put inside the loop.

Code computes, System One judges, the LLM generates

Three layers, and a question belongs to exactly one of them.

LayerOwnsExamples
CodeAnything with an exact answercounts, date math, thresholds, combining scores
System OneBounded judgments with a closed answer setroute this ticket, is this vendor marketing, how durable is this story
LLMGeneration and multi-step reasoningwrite the draft, call tools, plan a fix

TypeSafe publishes where Jev breaks, which is more than most vendors do. The jaggedness page for jev-1.13 lists arithmetic, date comparison, double negatives, oversized state and adversarial text. Every one of those is a boundary in the table above. Push arithmetic down into code. Push anything open-ended up to the LLM.

Worked example: scoring blog topics

The content engine behind this blog pulls about 40 stories a day from TechMeme, Reddit and Firecrawl, and a regex scorer decides which ones are worth writing about. In September I wrote a Jev version of that scorer. One request asks six questions about each story:

substance: {
  type: "noul",
  instructions:
    "Does this contain something concrete to react to — a claim, an outcome, a number, an incident, a decision someone made — as opposed to asking the reader for help or opinions?",
  // criteria: { true: "...", false: "..." }
},

The weights live in code, not in the prompt, so retuning them re-ranks the backlog without spending a token:

const gate = WEIGHTS.substance_floor + (1 - WEIGHTS.substance_floor) * j.substance;
const normalized = merit * gate - WEIGHTS.vendor_fluff * j.vendor_fluff;

The first shadow run, read-only next to the regex scorer, promoted three r/devops help threads to long-form. “Do you really test your backup” scored 2.9 of 3 on beat fit and 0.83 on operator angle. Jev was right about both. A backup question from a practitioner is squarely on beat. It’s also not a blog post, because it asks instead of reports. The judgment was fine. My weighting was wrong, and the fix was a substance question applied as a multiplier rather than another additive term. The Jev scorer lives on a branch and only runs in shadow when I invoke it; the regex scorer is what ships today.

The numbers, and where the savings leak

I ran the six questions against 50 recent stories, three times each, on Jev, Claude Haiku 4.5 (no extended thinking) and Claude Opus 5 at low effort. Same wording, same criteria, Claude held to the same schema with structured outputs. Total spend: $1.93.

Jev 1.13Haiku 4.5Opus 5 (low)
Median latency163 ms1,156 ms2,740 ms
Cost per 1M calls$53$1,610$11,183
Tier changed across its own 3 runs2 of 5020 of 505 of 50

Prices come from TypeSafe’s models page ($0.042 per million input tokens, output free) and Anthropic’s pricing page. The Haiku column surprised me. It’s the default “cheap classifier,” and it was the least consistent model in the test, flipping its own tier on 20 stories out of 50. Prompt caching won’t rescue it here either: this prompt is about 1,300 tokens and Haiku 4.5 won’t cache anything under 4,096. Opus 5 caches from 512 tokens, which brings it to roughly $3,000 per million calls if every call hits the cache. Still about 57x Jev.

Jev agreed with Opus on 274 of the 300 individual answers. It also landed 26 of 50 stories in a different tier.

Both are true because Jev scores the whole set higher (median composite 1.73 against Opus at 1.25) while ranking the stories in roughly the same order. My tier cutoffs sit where the old scorer’s numbers used to cluster. Move to a model with a different baseline and half the backlog changes tier. Haiku and Opus disagreed on tier for 20 of 50 stories too, so this isn’t a Jev defect. Tier cutoffs belong to the model they were tuned on. A model swap is a change to the agent harness, not a config flag.

The obvious fix is confidence routing: let Jev answer, and send low-confidence cases to Opus. Jev returns a confidence with every score and choice. The naive rule, “escalate the story if any answer is below 0.8,” sent 48 of 50 stories to Opus, because a four-level rubric rarely puts 80% on a single level. Durability’s median confidence was 0.65.

TypeSafe warns against carrying one threshold across question types, and the data agrees. A per-type rule (escalate when the choice confidence is below 0.8 or any noul lands between 0.3 and 0.7) sent 24 of 50 stories up. That costs $5,421 per million, 2.1x cheaper than Opus alone, and it catches 17 of the 26 answers where the two models disagree. Real, but nowhere near 200x. On a workload where half the stories are genuinely ambiguous, escalation eats most of the savings.

Where this breaks

Durability was the weakest question. It caused 12 of the 26 disagreements, and on 8 stories Jev split its probability between two levels at least two steps apart. The score it returns is an expected value, so a story that’s either a one-week event or a durable shift comes back as “somewhere in the middle.” TypeSafe’s docs say not to treat that number as a precise magnitude. My composite divides it by three and adds it in, which is exactly what they say not to do.

There’s no answer key in these numbers. Agreement with Opus is not accuracy, and 50 stories from one domain is a small sample. State is also data Jev doesn’t treat as hostile by default, which matters for an agent that reads web pages. And the rate limits on the models page are “adjusting dynamically,” which is a sentence to read twice before routing production traffic through one vendor.

What to change on Monday

Grep your agent harness for model calls whose output is a label, a score or a boolean. Those are System One questions wearing an LLM’s price tag. Shadow them against a System One model for a week, read-only, and log both answers.

Then do the part the benchmark says is hard. Retune your cutoffs on the new model before you switch, set confidence thresholds per question type, and escalate only where the answer can change the decision. The model is the cheap part now. The 26 misfiled stories came from my loop, not from Jev.