← All posts
ai-securityagentsarchitecture

The decider pattern: stop using LLMs for control flow

Here's an expensive mistake most agent pipelines make: asking a frontier LLM to write prose every time the system just needs a branch. "Should this tool call proceed?" does not need a paragraph. It needs a yes or no with a number attached.

A new model category exists for exactly this. They're called decision models — "System One" models, after Kahneman's fast thinking — and the category went from one stealth startup to an industry in about two weeks.

What a decision model is

TypeSafe AI launched the first one, Jev, on September 15, 2026 (alongside a $40M seed led by DCVC). It's named after economist William Stanley Jevons — of Jevons paradox: when intelligence gets cheaper, software uses far more of it. The bet is that judgment is about to get very cheap, and pipelines will use far more of it.

Jev is not an LLM. You send it a state (text or JSON) plus typed questions, and it returns typed answers with calibrated probabilities — no text generation at all. Three primitives: Choice (pick one of up to 255 options, with per-option probabilities), Score (position on an ordered rubric), and a yes/no that returns p(true). All questions in a request evaluate in parallel against the shared state.

The numbers: 70–500ms end to end, $0.042 per million input tokens, output free. Compare that to waking a frontier chat model for every branch decision.

Why chat models are the wrong tool for control flow

Three problems, all structural:

Latency. A decision that needs 100ms should not wait 30 seconds for tokens to stream. Every branch in your agent loop multiplies the wait.

Cost. Frontier inference per decision is cents; decider inference per decision is fractions of a millicent. An eval harness making ten thousand judgments is the difference between a rounding error and a budget line.

Calibration. Chat models are trained to sound equally confident about everything. A decision model returns a number your code can branch on — p(true) = 0.92 means something, and 0.51 means "ask a human." That's the honest difference between prose and probability.

The category is already open

Within two weeks of Jev's launch, OpenAI announced a Decisions API, Cloudflare open-sourced Clef 27B and Clef-flash 9B under Apache 2.0, and AWS Strands Labs shipped a 1.9B-parameter open decider under Apache 2.0 that runs in ~110ms on an RTX 3090 — and on Apple Silicon and plain CPU. The open options mean a security shop can self-host its decider layer at zero marginal cost.

Where deciders fit in security pipelines

This is the pattern we build around:

  • Triage as decisions, not prose. Every finding gets scored — severity, exploitability, priority — as calibrated numbers, not paragraphs. Ten thousand findings, ten thousand cheap judgments, ranked output.
  • Tool-call gating. Before an agent executes anything irreversible, a decider answers one yes/no: does this action match the authorized plan? 100ms, near-zero cost, every single time.
  • Eval-harness judges. "LLM-as-judge" is the wrong tool for scoring evals — you want a judge that returns a number, deterministically, thousands of times. Deciders are built for this.
  • Pipeline routing. Which scan to run next, which finding deserves a human, which alert is noise — all branches, all cheap.

Jevons paradox applies directly: when judgment costs almost nothing, you stop rationing it. Every finding triaged, every tool call gated, every pipeline step scored. Cheap judgment everywhere is what makes agents reliable enough to trust.

If you're building agent pipelines for security work and the control flow is still running through a chat model, that's a conversation worth having — book a scoping call.

Need a pentest, an AI security assessment, or a custom security build?

Human-led testing, production AI builds, and the full loop in between. Book a free 30-minute scoping call.

Book a scoping call