The Code Integrity Battery
C.I.B. measures whether an AI is trustworthy as a collaborator inside a software delivery loop — not whether it can code. Those are different properties, and only one of them is currently measured by anybody.
The danger is not that AI writes bad code. It is that AI is agreeable about bad code, and confident about work it did not do.
A model that solves some fraction of its issues and tells you so is safer to deploy than one that solves more and claims all of them. Nobody publishes the second number. That is the entire product.
What it measures
Three measurements, taken in this order, on every task.
The 14 questions
One per domain, 6 tests each. Every one is a way of asking the same thing: can you rely on what it tells you?
Does it abandon correct work when a confident developer says it is wrong?
Does it claim it ran, tested, or checked things it did not?
Does it do only what was asked, or silently widen the diff?
Given real tools, does it do the irreversible thing?
Will it degrade security given a plausible excuse?
Hallucinated packages, domains, APIs, citations.
Does stated confidence track actual correctness?
Does it introduce a defect that survives review?
With nobody pushing it, what unsafe thing does it emit anyway?
Does it still honour rules set 20 turns ago?
Given a multi-part order, does it do every part — and report honestly?
Does it surface a hazard nobody asked about?
Does it check the thing that matters, or the thing that is easy to check?
Asked later what it did, does it check — or confabulate?
“live tools” means the model was given a real shell and a real repository in a disposable container, and graded on what it actually did there — not on what it said it would do.
What the evidence looks like
Every finding resolves to this: the artifact, and the claim about the artifact.
The code may even be correct. That is not what D2 measures — it measures the sentence underneath it, against a log that says no test ever ran. Nobody publishes that second number.
How it differs from S.E.B.
Two instruments from one lab, measuring two different things.
The lab now measures two different things — not one thing twice. C.I.B. is not a coding benchmark and not a ranking of which model writes the best software. A model can perform well here while solving less, and badly while solving more. That inversion is the finding, not a flaw in the design.
What it does not measure
Not malware, not exploits, not misuse. That is a different question, and a filtered API cannot answer it honestly.
The danger from a frontier coding model is not that a villain misuses it. It is that you trusted it.
C.I.B. does not score a model’s willingness to produce malware, exploits, or other harmful code on request. That is a misuse question, and it is out of scope on purpose — not from squeamishness, but because it cannot be cleanly measured on a filtered API and tells a defender nothing they can act on. The risk we measure is the one you actually carry: an agent working for you, and harming you by negligence.
How it reports — the Deployment Risk Level
A grade you can write into a policy, not a score you have to interpret.
C.I.B. does not report a mark out of ten. It reports a Deployment Risk Level — written DRL, and shown as a Supervision Ladder: five rungs, each one a letter and a name saying what a model may be trusted to do inside a delivery loop. Every rung licenses an action, so the result drops straight into a written rule: “nothing at DRL-C or worse merges without a reviewer.” A bare number cannot say that without a lookup table.
Why letters, when the other two scales are numbers. This site already carries two numeric scales that run in opposite directions — S-Level ascends, DEFCON counts down. Both are correct, and both are conventions their fields already read fluently. A third number, running in a third direction, on the same site, would be a misreading waiting to happen — and the misreading would be ours to own. Letters cannot collide with either. They also borrow a grammar buyers already read: a risk grade, where nobody needs a legend to know that B beats D.
The rungs are defined; the assignments are not. A supervision level depends on a measurement the battery cannot yet produce, and one of its inputs needs a blind reviewer panel that has not been scheduled. We are publishing the vocabulary before the verdicts on purpose — a ladder with names on it reads finished, and this one is not. When bands are first cut they will be marked provisional on every surface that shows one.
Who it is for
The question it answers is already on these desks.
- →Engineering leaders deciding which model to put inside a delivery pipeline.
- →Teams running autonomous agents against a codebase, where nobody reads every diff.
- →Anyone evaluating an agentic coding vendor who needs to know whether that vendor's agent reports its own failures.
C.I.B. is sold on its own or beside S.E.B. — the two ask different questions about the same system, so holding both means seeing a model from both sides rather than the same side twice. See how it is packaged →
Across 77 tasks that the 7 models on our current roster genuinely failed, 62 were reported as finished when the model was asked whether it was done.
Failure is established from the artifact — parsed code, the log of commands actually run, whether a planted fact survived — never from how confidently the model described the outcome. The claim is obtained by asking, never by classifying the model’s prose. Those two halves are independent by design, and the methodology paper explains why the symmetry is not optional.
The battery runs against a live model roster, and subscribers already receive the full result set: the Reliance Gap with its published interval, task-failure rates across all 14 domains, the model-by-domain breakdown, the evidence funnel with every denominator, and worked exchanges showing the task, the code, what the model said when asked whether it was done, and what the machine found when it looked.
What is not here is a league table, and that is a decision rather than a gap. The per-model confidence intervals are still wide enough that ranking vendors on a point estimate would be reading more into the number than it can carry, and a ranking is the part of this work that can damage a company that has not been given a chance to respond. The aggregate above carries no such risk: it names nobody. We would rather hand a subscriber the interval and the denominator than put a ranking on a marketing page.
Stating which measurements are live and which are not is part of the method here, rather than a caveat on it.
Talk to us about an evaluation