OPTETRON

Founder availability

The founder is currently under an exclusive contract and cannot accept new requests. You can still subscribe to our newsletter or email us directly at contact@optetron.com.

← Blog

01 / FIELD NOTES

Tuned locally, proven on a test set it never saw

We pointed our optimization loop at the 13-way triage classifier inside a caregiver-support health chatbot. Champion accuracy on the untouched test set was roughly 90% overall and 100% on every safety-critical category, running on a 9B model and hardware you can buy.

2026-06-10 · 7 min · Optetron

We build a caregiver-support chatbot with the A2MCL association, for families facing Lewy body dementia. At its heart sits an unglamorous, load-bearing component: a classifier that routes every incoming message into one of 13 categories. Is this an emergency, a symptom question, an administrative request, a caregiver at the end of their rope? Get that wrong and the rest of the pipeline answers the wrong question. In a health context, that is not a cosmetic bug.

The classifier runs on a 9B-parameter model on a workstation-class GPU. No API, no data leaving the machine. This post is about how good a component like that can get, and how you prove it, to yourself and to the association whose users depend on it.

~90%
test-set accuracy, 13-way health triage, local 9B model, test set touched exactly once

The protocol comes before the run

Prompt optimization has a well-earned credibility problem: it is very easy to "improve" a prompt against the same examples you score it on, ship the number, and discover in production that you memorised your eval. So before the optimizer ran, the dataset (1,460 real and synthesized caregiver messages) was split four ways, with strict one-way visibility:

SplitSizeWho sees it
Few-shot anchors13the prompt itself (one example per category)
Train~822the generator, whose next candidates are fed by confusion patterns and failing examples
Validation~405selection only; champions must beat the incumbent here, gated by a McNemar test
Test185nobody, until the run is over. Scored once. Never in the loop.

The generator never sees validation. Selection never sees test. The test set exists for exactly one purpose: to expose the gap between "looks better in the loop" and "is better on data it never touched".

The run

The optimizer, Optetron Loop, explores candidate prompts the way an evolutionary search explores a fitness landscape: an elitist population, recombination between strong parents, every champion confirmed against the incumbent on the validation split before it takes the crown. The generation side ran on a cloud-hosted open model. The classification side, the thing actually being measured, stayed on the local 9B model throughout, because that is what serves users in the chatbot.

On the selection signal, the campaign moved the classifier from 0.929 to 0.954 validation accuracy, a gain of 2.5 points across the search frontier.

Then the run ended, and the test vault was opened once.

What the test set said

Champion prompt, 185 messages it had never influenced in any way: 89.7% overall. The lineage re-score tells the same story from another angle. The seed prompt scored 89.2% on test; the optimized frontier sits at 90–92%.

The average matters less than this row-by-row breakdown:

CategoryTest accuracy
URGENCE (emergency)100%
OFF_TOPIC100%
TROP_VAGUE (too vague; ask, don't guess)100%
APRES_DECES (post-bereavement)100%
ADMIN_A2MCL100%
IDENTIFIER_MALADIE94.7%
REACTION_SYMPTOME / COMPREHENSION_MALADIE92.3%
PRO_SANTE_DEMANDE91.7%
SOUTIEN_AIDANT (caregiver support)83.9%
VIE_QUOTIDIENNE78.6%

Every category where a routing mistake is dangerous (emergencies, bereavement, off-topic deflection, "this is too vague to answer safely") sits at 100% on the held-out set. The residual errors live in the soft boundaries between neighbouring support categories, where two readers would also disagree. On messages tagged hard by annotators, the champion holds 86.4%.

And yes, validation said 0.954, test says 0.897. That ~5-point gap isn't an embarrassment; it is the measurement working. It's the number a vendor quoting only their in-loop score will never show you, and the reason the test set exists at all.

The result that surprised us

As a cross-check, we took the champion prompt (tuned against the 9B classifier) and ran it unchanged on a different local model, a 12B QAT variant. Accuracy fell to 81.6%, with the damage concentrated exactly where it hurts: caregiver-support recall collapsed.

An optimized prompt is not a portable artifact. It is a fit between a prompt and a model, on your distribution. Swap the model (because a better one came out, because your hardware changed) and the tuning has to re-run. That is the strongest argument we know for owning the optimization loop instead of buying a one-off prompt consultation: the loop is the asset, the prompt is just its current output.

What this buys you

You get a health-grade component on commodity hardware: thirteen-way French triage at ~90%, perfect on the safety-critical categories, on a 9B model that runs on a card you can buy, with no tokens leaving the building. The numbers hold up to scrutiny, because train, validation, and test stay separate, champions pass statistical gates, and the test set is touched once. It is the same discipline we'd sign off on for a client audit. And what you own is a loop, not a deliverable. When the model changes, the data drifts, or a new category appears, you re-run the search on your own infrastructure and re-prove it on a fresh test set.

The chatbot this classifier serves is real and demoable; the optimizer that tuned it is the same one we deploy for clients.


We build optimization loops for systems where the numbers have to be defensible, on your data, your hardware, your terms. Tell us where the data can't leave →