381 tok/s on a 3090: what that number actually means
A post making the rounds (via @Oluwaphilemon1) claims Qwen3.8-27B hits around 381 tokens per second on a single RTX 3090 — 24GB of VRAM, 250 watts. The number is real. But as the post itself notes, there is a very important detail behind it: this isn't normal-chat generation.
Here's what's actually going on, and why it matters more for security work than for chatbots.
Speculative decoding, in one paragraph
Normal autoregressive generation produces one token per forward pass through the model. Speculative decoding cheats: a cheap draft mechanism proposes several tokens at once, and the big model verifies them all in a single pass. Every accepted draft token is a token you got nearly for free. Your speedup equals your acceptance rate. The entire game is making the drafter guess well.
Qwen3.8 ships with native MTP (multi-token prediction) heads, so the draft mechanism is built into the model itself — no separate drafter model eating VRAM. That's the foundation. The 381 number comes from the clever part: prompt-lookup drafting.
The trick: the answer is already in the room
Prompt-lookup drafting doesn't guess from a small model. It guesses from the prompt. When the model is quoting, summarizing, or restructuring a document you pasted — a findings report, a chunk of source code, a log file — the next tokens are often verbatim spans of text already sitting in the context window. The drafter proposes those spans; the verifier accepts them at a very high rate.
381 tok/s happens when the model is quoting a document back at you. It's not a chat benchmark. It's a "the output heavily overlaps the input" benchmark.
An open recipe from Mads Henrichsen (syv.ai) documents the full ladder on a single 3090: 46 tok/s with speculation off, 133 tok/s greedy single-user with the full optimization stack (int8 output head, fp16 recurrent state, int8 activations, a 40,000-entry draft vocabulary counted over the model's own output lifting coverage from 92% to 97.5%, a split-KV kernel), and up to ~1,035 tok/s steady-state decode across 64 concurrent requests. The 381 figure is one rung on that ladder — the prompt-lookup rung.
Why this is a security-workload story
Think about what security agents actually do all day:
- Triage agents read a findings report and restructure it into tickets. The output overlaps the input heavily.
- Code-review agents quote vulnerable functions back with annotations.
- Detection engineers feed logs in and get detection rules out, with IOCs copied verbatim.
- Eval harnesses replay the same scenarios repeatedly — highly predictable output.
These are exactly the workloads where draft acceptance rates are highest. The workloads where speculative decoding looks least impressive — free-form prose, creative writing — are the ones security pipelines rarely run. The benchmark that flatters this technique most is, by coincidence, shaped like our work.
There's a second, quieter point in the recipe worth sitting with: a stranger's RTX 4090 ran the identical stack uncapped at 450 watts and finished single-user inference 1.9% faster. Decode is a bandwidth game, and moving from a 3090 to a 4090 only moved bandwidth 8%. The six-year-old card at 250W is within spitting distance of the flagship because the bottleneck was never compute. For fixed-cost local labs, that means the used 3090 remains the price-performance king — you don't need the new card, you need the new config.
The honest half
None of this is free, and the recipe's own docs are admirably upfront:
- Throughput is task-dependent. Structured, predictable output drafts beautifully. The same setup on free-form prose drops hard — acceptance rate, not hardware, is the ceiling.
- Quality pays a toll. About one point on IFBench-class evaluations across the shipped configs. Grade-school math holds at 95–96.5%. Know which of your workloads can afford that point and which can't.
- Cold starts are real. Around 100 seconds on a 100k-token prompt before the first token. Fine for a long-running service, miserable for interactive use.
- A benchmark can't tell you the output is garbage. High tok/s with low acceptance-correctness is just fast wrongness. Measure acceptance rate and output quality together, or the number is theater.
What we'd actually build on this
For a security team, the shape of the deployment is: a 3090-class box (or two) running Qwen3.8-27B with MTP speculative decoding and prompt-lookup drafting, serving the workloads where output overlaps input — triage, quoting, restructuring, replay. Long-running, warm, behind an API. The workloads that need genuine reasoning get the same model with speculation tuned down and the quality budget spent where it counts.
This is the unglamorous core of the build stage: not the biggest model, not the newest card, but the right decoding strategy for the shape of your workload, measured honestly. The 381 number is real. The useful number is your acceptance rate on your prompts.
If you want a local inference setup designed around your actual workloads — measured, not benchmarketed — book a scoping call. We build these.
Need a pentest, an AI security assessment, or a custom security build?
Human-led testing, production AI builds, and the full loop in between. Book a free 30-minute scoping call.
Book a scoping call