← All posts
ai-securityagentslocal-inference

The $5,000 security agent lab

Three things happened this year that compose into one story. Decision models made machine judgment ~100× cheaper — and open source. Frontier-adjacent model weights went MIT licensed and desktop-sized. And the box to run it all dropped to $4,999.

Put them together and you get something new: an entire agentic security pipeline — triage, tooling, evals — running locally, privately, at fixed cost. No API meter. No data egress. Here's the stack and the math.

The box: NVIDIA DGX Spark

The DGX Spark is a desktop AI system built around NVIDIA's GB10 Grace Blackwell Superchip, with unified memory — CPU and GPU share one pool, so there's no VRAM ceiling to game. It ships with DGX OS plus Ollama, vLLM, llama.cpp, and PyTorch+CUDA preinstalled. It's explicitly positioned for local agents and inference.

The price history tells its own story: launched at $3,999, raised to $4,699 on memory constraints, then up to $6,950 in the memory crisis. But on October 3, 2026, NVIDIA announced a 64GB tier at $4,999, shipping October 23 through Acer, ASUS, Dell, Gigabyte, HP, and MSI — with a Sync Cluster mode that pools two 64GB units to 128GB.

What actually runs on the 128GB box, measured independently:

  • Qwen3.8-27B: ~45 tok/s single-stream — interactive daily driver
  • gpt-oss-120b: ~59 tok/s decode
  • GLM-5.3-Flash at 1–3 bit: fits in 93–128GB — a 320B frontier-adjacent MoE on a desk
  • Dense 70B-class: ~4 tok/s — a batch job, not a chatbot. MoE or bust locally.

The judgment layer: deciders

The agent loop's control flow — triage this finding, gate that tool call, score this eval — runs on decision models, not chat models. Open options under Apache 2.0 run in ~110ms on a single RTX 3090, or on Apple Silicon and CPU. A 1.9B-parameter decider is a rounding error in the weight budget and returns calibrated probabilities your code can branch on.

The worker: MIT frontier weights

GLM-5.3-Flash: 320B total, 18B active, MIT licensed, ungated on Hugging Face. At 1-bit it fits in 93GB — inside the 128GB box's practical budget. It trades blows with Claude Opus 4.8 on agentic coding benchmarks. The model writing your agent's prose and code is frontier-adjacent, and it never phones home.

The economics

A heavy agent workload on frontier APIs — thousands of decisions, long eval runs, continuous triage — is a meter that never stops. At $4,999 for the box, the break-even against serious API spend is measured in weeks, not months. And the cost is fixed: run the eval harness ten thousand times or ten times, the box costs the same.

For security work specifically, the math has a second term: privacy. Client code, findings, and captures stay on hardware you own. No data processing agreements, no retention policies, no third-party risk review for the model vendor — because there is no model vendor in the loop.

What this enables

This is the technical backbone of the build stage: private agent pipelines, private eval harnesses, private detection automation. A pentest firm can run triage agents over client findings without the findings leaving the office. A dev shop can offer AI security builds with a straight answer to "where does our code go?" — nowhere.

The $5,000 lab isn't a compromise anymore. It's the setup the API-reseller shops can't match. If you want one designed around your workloads, book a scoping call — we build these.

Need a pentest, an AI security assessment, or a custom security build?

Human-led testing, production AI builds, and the full loop in between. Book a free 30-minute scoping call.

Book a scoping call