OPTETRON

Founder availability

The founder is currently under an exclusive contract and cannot accept new requests. You can still subscribe to our newsletter or email us directly at contact@optetron.com.

← Blog

00 / FIELD NOTES

The local revolution is now.

Five reasons running serious AI workflows on your own infrastructure is no longer a 2027 problem, with one signature data point each.

2026-04-29 · 7 min · Optetron

Read the manifesto →

Every quarter, the case for owning your AI stack gets cheaper, faster, and harder to argue against. The five claims below are independent. You don't have to believe all of them. Two are enough.

5×
cost crossover at modest team scale

The price you see is not the price you'll pay

Frontier-API prices are subsidised. Investors are funding your bill so the platform race resolves in someone else's favour. Once the round closes and the moat is built, the price floor is set by whoever has leverage, and by construction that isn't you.

The local crossover happens earlier than people think. At a few tens of millions of tokens per month (the load of a small product team using AI day-to-day), owned hardware beats hosted APIs across the curve.

Cost vs monthly tokens

Self-hosted hardware amortizes; hosted APIs scale linearly with usage.

Cost vs monthly tokensSelf-hosted hardware amortizes; hosted APIs scale linearly with usage.1M10M100M1000M$1$10$100$1000$10000Self-hosted vs OpenRouter @ ~36M tok/moAnthropic Sonnet (API)OpenRouter (Qwen3.6-35b)Self-hosted Qwen3.6-35b (A6000)

Data: cclocal snapshot 2026-04-29 · cclocal sha fixture

More on the price economics →

Open weights are no longer the second best option

The capability gap is closing on the benchmarks people actually care about. Not closed, closing. The point isn't that open-weight models match the strongest frontier model on every task. It's that for the work most teams need done, code edits, agent loops, structured extraction, the cheap-and-local frontier sits well into the "good enough" band.

Capability vs cost — gpqa-diamond

Lower-left = expensive and weak. Lower-right = cheap and weak. Upper-left = expensive and strong. Upper-right = the offer.

Capability vs cost — gpqa-diamondLower-left = expensive and weak. Lower-right = cheap and weak. Upper-left = expensive and strong. Upper-right = the offer.$0.1/Mtok$1/Mtok$10/Mtok0255075100qwen3.6-35b-a3bqwen3.6-35b-a3bgemma-4-e4b-itgemma-4-e4b-itclaude-sonnet-4-6

Data: cclocal snapshot 2026-04-29 · cclocal sha fixture

The question worth asking is no longer whether the model is strong enough. It's whether the marginal capability is worth the marginal cost, and the marginal loss of control.

More on capability parity →

Commodity hardware is fast enough

The speed argument used to be that local inference is too slow to feel real. That stopped being true a while ago. A laptop-class GPU runs an 8B-parameter model at conversational throughput. A workstation-class GPU runs a 35B-parameter MoE faster than most users can read.

Throughput (tokens/sec) by model and provider

Local backends on commodity GPUs hit useful throughput; not always the fastest, always under your control.

Throughput (tokens/sec) by model and providerLocal backends on commodity GPUs hit useful throughput; not always the fastest, always under your control.qwen3.6-35b-a3b · llamacpp-local-a600038 t/sqwen3.6-35b-a3b · openrouter95 t/sgemma-4-e4b-it · llamacpp-local-rtx409065 t/sgemma-4-e4b-it · openrouter130 t/sclaude-sonnet-4-6 · anthropic-api120 t/s

Data: cclocal snapshot 2026-04-29 · cclocal sha fixture

You don't need an H100. You need a card you can buy, a rack you can plug in, and a workload you understand.

More on speed and hardware →

Your data is more exposed than you think

The Cloud Act lets foreign authorities compel access to data held by US-domiciled providers, wherever the servers physically sit. On-prem isn't paranoia. It's the only configuration that survives a subpoena you didn't see coming.

GDPR ambiguity, training-data leakage, terms-of-service drift. You can read your contract today; you cannot read the one your provider will offer in eighteen months. For some industries, sovereignty is already a hard requirement, not a preference. For everyone else, it is becoming one.

More on jurisdiction →

The substrate moves under you

APIs deprecate. Models change. Without your consent.

  • A model upgrade silently rewrites your prompts.
  • A safety filter quietly blocks a category your workflow depends on.
  • A pricing tier is restructured the week before your renewal.
  • A whole product line is sunset on six months' notice.

You operate a critical workflow on a substrate you do not control. That substrate is a moving target. The only stable surface is the one you own.

More on volatility →


The local revolution is here. We help companies put it on their own infrastructure: design, deployment, training, support.

Talk to us →