← Obsta Labs

The Missing Judge

October 2026 · Obsta Labs

The sentence this essay exists for is in the middle, so here it is first: an AI can do the work, but it must not be the final authority on whether that work was authorised, correct and complete. Everything before it is about why an independent judge matters. Everything after it is about what such a judge looks like when the work is not a chess move or a molecule but a change in somebody's company.

Where the reliability came from

Look at the systems that got reliably good at something, and you find feedback they could not redefine. AlphaZero combined a learned intuition with search, self-play and the outcome of the game; the rules decided who won, and the network learned from results it could not talk its way out of. The systems that now prove theorems propose steps and a proof checker says yes or no, and the proposals improve because the checker never flatters. Coding models became dependable where the tests were strict and the test suite was not theirs to edit.

I am not claiming this is the whole story of progress. Scale, data and compute make models more capable. Feedback does something different: it makes capability reliable in a particular domain, measurable, and usable for something that matters. The word "judge" hides three things, and it helps to separate them. A training judge supplies the feedback a model learns from. A verification judge says whether an output meets a condition set outside the model. An accountability judge says whether an action was permitted, done as claimed, and properly closed. This essay is about the third. The first two are how we get there.

When the judge measures the wrong thing

Drug discovery is the clearest case, because the judge changes halfway through. A 2024 analysis in Drug Discovery Today looked at molecules discovered with AI and found them succeeding in Phase I trials at roughly 80 to 90 percent, well above the historical range, while their Phase II success rate, around 40 percent on a small sample, was comparable to everyone else's. Phase I is mostly about safety, tolerability and how the body handles a compound; it is the part a model of the molecule has the most to say about. Phase II asks whether the drug works in the people it is meant for. Different judge, different verdict.

Then there is torcetrapib, which was not an AI-designed drug and is here for a different reason. It raised HDL cholesterol by about 72 percent on top of a statin, which is exactly what it was designed to do. In December 2006 the trial was stopped: among fifteen thousand patients, more people on the drug had died than people on the statin alone. The judge the designers had chosen, a number in the blood, was consistent, measurable and wrong about the thing that mattered.

That is the lesson I want to carry into the rest of this. A judge must be independent of the model's own claims, and it must also measure the right thing. A perfectly consistent judge of a stand-in measure is how you get a confident failure.

The work problem

A blood number was not a proxy for a patient's survival. A closed ticket is not a proxy for a correct change. With that in mind, put an AI to work. It reads a ticket, changes a system, sends a message, closes a case. What judges it?

Usually a log, and at best an audit trail. A good audit trail already records a great deal: which identity did what, when, under which permission. I do not want to argue with audit trails; they are necessary, and we keep them. But an audit trail records that an action happened. It does not record whether the acceptance conditions for the work were met, what evidence showed that, or who had the authority to say so. It is evidence, not an account of the decision. And the usual arrangement lets the model write the last line of the account: task complete. The model marks its own work done, and there is nothing in the room that can disagree. I wrote in May that the agent must not close its own ticket. This is the argument underneath it.

Tests do not fix this on their own, and torcetrapib is why. An agent can pass every automated check and still have executed a request nobody authorised, changed something outside the approved scope, met the specification while breaking the business rule, or closed work whose consequences have not yet been seen. A test establishes compliance with a condition. Accountable work also has to establish where the condition came from, who had the authority to set it, and what evidence justified the decision to call the work done.

So five things have to be on the record, and they are not the same thing:

Closure can be a person's decision or a deterministic rule, depending on the work. Both are legitimate, and the record has to say which applied. What closure cannot be is the executor's own report.

What the record can and cannot judge

Here is the boundary, stated plainly, because the idea is more credible with it than without. A record of work with authorisation, evidence and closure can show that the work did what it was told, within the scope it was given, and who signed for it. It cannot show that the instruction was wise, fair or good for the business. It supports that judgment by giving the rule or the reviewer something real to judge, and it leaves the judgment where it belongs.

Where I think this is heading

This part is opinion. The models will get sharper, and my bet is that language alone will not make them radically sharper, because the text they learned from carries no consequences. The gains I expect come from models paired with something that pushes back: formal systems with a proof checker, systems that act in the physical world and meet its outcomes, and systems that do work and meet a record with authority over closure. Whether those grow into separate specialists connected by tools or into one shared representation is open; I lean to the second, the way language, vision and space share a skull. Either way the same condition holds: useful intuition grows from contact with consequences. It cannot grow where nothing can push back.

For work, that contact barely exists yet. Common tooling stops at logs and approval steps. Nobody is short of agents that can act. What is short is a record an agent cannot overrule.

What we do about it

I will describe our own practice in terms a reader can check. We run a small company on such a record. Every commit we make names the work order it delivers, enforced by a hook that refuses the commit otherwise. A closure cites the commit and the test result, which prove that a change was recorded and that the specified checks passed; whether that evidence meets the work order's acceptance conditions is decided by the work order's own closure authority, and the identity that executed the work is refused as the identity that closes it.

It refuses us, which is the point, and it refuses us in two different ways. In September our automated promoter closed 43 work orders whose executors had reported them complete, and refused 13, with the reason printed for each: child work orders still open, follow-ups nobody had acknowledged, one commit whose CI run was red. That is evidence refusing closure. A week later a change landed that refuses a hand-off when the actor claiming completion is the same identity that would later verify it. That is authority refusing closure. They are different checks, and a system needs both.

It also catches us. This week we found that our merge script had been dropping the work-order line from the final squashed commit on the main branch, which meant closures were being made by hand against commits the record could not see. That is a work order now, with a number, like every other defect we find in our own machinery, including the ones the AI caused.

As of 8 October 2026 the engine's own project holds 2,410 work orders: 1,119 closed, 786 cancelled (a number inflated by a thirty-day sweep rule we are still correcting), 356 open, 86 landed and awaiting a verifier, and the rest blocked, in progress or pending verification. The cancelled stay in that count. A total that hid them would be its own stand-in measure.

The fleet that builds the record is its first subject. That is the only honest way to sell a judge: submit to it first.

If you run AI in a business and want to see what an auditor's questions look like against such a record, the short version is on hiveram.com/why. If you want to try the record on your own work, there is a trial. Neither will make your model smarter. They will make it possible to say, afterwards, whether it was authorised, whether it did what it claimed, and who said so.

---