00 / FIELD NOTES
The local revolution is now.
Five reasons running serious AI workflows on your own infrastructure is no longer a 2027 problem, with one signature data point each.
2026-04-29 · 7 min · Optetron
Every quarter, the case for owning your AI stack gets cheaper, faster, and harder to argue against. The five claims below are independent. You don't have to believe all of them. Two are enough.
The price you see is not the price you'll pay
Frontier-API prices are subsidised. Investors are funding your bill so the platform race resolves in someone else's favour. Once the round closes and the moat is built, the price floor is set by whoever has leverage, and by construction that isn't you.
The local crossover happens earlier than people think. At a few tens of millions of tokens per month (the load of a small product team using AI day-to-day), owned hardware beats hosted APIs across the curve.
Cost vs monthly tokens
Self-hosted hardware amortizes; hosted APIs scale linearly with usage.
Data: cclocal snapshot 2026-04-29 · cclocal sha fixture
Open weights are no longer the second best option
The capability gap is closing on the benchmarks people actually care about. Not closed, closing. The point isn't that open-weight models match the strongest frontier model on every task. It's that for the work most teams need done, code edits, agent loops, structured extraction, the cheap-and-local frontier sits well into the "good enough" band.
Capability vs cost — gpqa-diamond
Lower-left = expensive and weak. Lower-right = cheap and weak. Upper-left = expensive and strong. Upper-right = the offer.
Data: cclocal snapshot 2026-04-29 · cclocal sha fixture
The question worth asking is no longer whether the model is strong enough. It's whether the marginal capability is worth the marginal cost, and the marginal loss of control.
Commodity hardware is fast enough
The speed argument used to be that local inference is too slow to feel real. That stopped being true a while ago. A laptop-class GPU runs an 8B-parameter model at conversational throughput. A workstation-class GPU runs a 35B-parameter MoE faster than most users can read.
Throughput (tokens/sec) by model and provider
Local backends on commodity GPUs hit useful throughput; not always the fastest, always under your control.
Data: cclocal snapshot 2026-04-29 · cclocal sha fixture
You don't need an H100. You need a card you can buy, a rack you can plug in, and a workload you understand.
Your data is more exposed than you think
The Cloud Act lets foreign authorities compel access to data held by US-domiciled providers, wherever the servers physically sit. On-prem isn't paranoia. It's the only configuration that survives a subpoena you didn't see coming.
GDPR ambiguity, training-data leakage, terms-of-service drift. You can read your contract today; you cannot read the one your provider will offer in eighteen months. For some industries, sovereignty is already a hard requirement, not a preference. For everyone else, it is becoming one.
The substrate moves under you
APIs deprecate. Models change. Without your consent.
- A model upgrade silently rewrites your prompts.
- A safety filter quietly blocks a category your workflow depends on.
- A pricing tier is restructured the week before your renewal.
- A whole product line is sunset on six months' notice.
You operate a critical workflow on a substrate you do not control. That substrate is a moving target. The only stable surface is the one you own.
The local revolution is here. We help companies put it on their own infrastructure: design, deployment, training, support.