XLNCXLNC Watch it measure

One ruler. Every mind.

Measurement for AI Transformation

Your score moved. What moved it?

Method published in peer-reviewed venues since 2016. Read the open-access chapter.

Simulation

Set to an error bar of 0.20. Delivered at 0.21 once judge harshness is corrected at each level.

Simulation: 15,000 decisions on judge data from our own runs, seed 20260908. With harshness left in, the same run came in at 1.16. Error bars in logits.

Final read after 139 questions: measured level 12.40, error bar 0.153 logits (one standard error).

AIM measures the judge and locks the questions. A move bigger than its error bar points at the model. Band: a recorded session on a scripted test performer, 139 locked questions, 2026-08-27.

Judge measured. Questions locked.

Judges change under the same name. In one published study, GPT-4's accuracy at telling prime from composite numbers fell from 84% to 51% between March and June 2023 (Chen, Zaharia and Zou, 2024).

  • Model: free to move.
  • Judge: measured.
  • Questions and keys: locked, with a fingerprint.
  • Grading prompt: held constant. Its record is pending.

Your own task mix in production is outside this ruler.

In a simulation on judge telemetry from our own runs, a run set to a standard error of 0.20 came in at 1.16 with judge harshness left in. With it corrected, 0.21.1

Simulation.

789101112131415harsher judgemore lenient judgeAfter correction both readings meet. The ruler does not move.

Some bias can be fixed. Some must go.

A judge that leans the same way each time is corrected, level by level. A judge that rates without a pattern at a level is not used there.

See how, on the Judges page

The Mona Lisa, undistorted: the calibrated measure
True imageReferenceThe calibrated measure
The Mona Lisa, stretched: stands for uniform leniency; correctable
StretchedCorrectableUniform leniency
The Mona Lisa, warped as in a funhouse mirror: stands for hallucination; exclude, cannot be corrected
HallucinationExclude: cannot be correctedFunhouse mirror

Every judge bends the picture. AIM undoes the bends it can measure and stops using a judge where it cannot. Illustration, an analogy. Painting: Leonardo da Vinci, Mona Lisa, public-domain scan (C2RMF) via Wikimedia Commons.

Measured, not averaged.

Each level carries its own error bar. No average sits on top, so a failure at one level stays visible.

Pick any score. It resolves to one question in a locked question set, with its answer key and a fingerprint.

One of our gates failed. A 200-question wave failed the test for one shared dimension. We published it as a finding.

The run behind the 0.21, in the whitepaper

If your domain is not on the list, we say so before the run.

Re-measured, and one outside view.

A judge that drifts without warning moves every score it grades. So we paid for a second run: Qwen3 235B Instruct, DeepSeek V4.1 Flash and TypeSafe (Jev) graded the same calibration items again. No judge moved more than 0.01 logit (DeepSeek +0.005, Qwen3 +0.007, TypeSafe -0.001; standard error about 0.03). Against a margin of plus or minus 0.10 logit set before the run, that drift counts as zero.

Two runs on one ruler: Qwen3 235B Instruct, DeepSeek V4.1 Flash and TypeSafe each sit at zero drift, well inside the plus or minus 0.10 logit band.

Limit: These judges ran at temperature 0, so much of this stability reflects near-deterministic model output; the figures describe LLM judge stability on a fixed item bank at stage-level grain, not the reliability of scores for people. Stability over months is not tested yet.

The re-measurement run, on the Judges page

Headshot of the person quoted.
“I've always stressed the importance of ethics in persuasion, and Dr. Matt Barney's AI assessment tool brings unprecedented scientific rigor to this domain. I am optimistic that his method holds immense promise in proactively preventing the misuse of persuasion techniques, both by people and emerging technologies, and augmenting their long-term use correctly”
Dr. Robert CialdiniNY Times Best Selling Author, Founder Cialdini Institute, Regents Professor Emeritus ASU

The study behind this work: Barney, Wind and Krishna, 2026

Two questions to ask any ruler.

Think of a physical gauge you trust. It is traced to a standard and states its error. Those are the two checks to ask of any ruler.

Has the ruler moved?

The question set is frozen, fingerprinted and dated. Every run checks the fingerprint first.

Re-measurement run: published.

Is the ruler circular?

Each level is fixed before any answer is graded. New questions stay sealed, so no model could have learned them. For new questions, the writer, the attacker and the judges come from separate companies. Each judge is checked against the others.

The judges for levels 7 to 9 are a certified set of 12. Three of them hold credentials level by level; the rest are still being credentialed.

A check against outside human experts: still ahead. See the roster.

Fingerprints checked at the start of every run.2

Know the cheapest judge that can.

Two judges have published ceilings. Today you pick the judge and see how far up the ruler it grades reliably. The run stops when the error bar is tight enough, or at the question limit you set. Hand-off to a stronger judge is in build.

Cheapest judge first. A stronger judge on stand-by, in build. Stops you set: the widest error bar, the most questions, the most spend.3

In build.

cheapest capable judge stronger judge strongest judge stop: widest error barstop: most questionsstop: most spend In build

Open to reviewers. Closed to training.

Engine
Open to reviewers under Apache-2.0, so they can check the method.
Question banks
Closed, so the questions stay out of public training data. Each bank has a sha256 fingerprint.

If you run an eval loop, AIM is one step in it. It replaces nothing you built.

The whole loop, one click down.

No score without its error bar.

That is the standard we hold every number here to. Each of our own measures names its status beside it. Today each is still in calibration.

Free to copy: the rule we use to decide if a judge's score counts. One paragraph. No form. Read the rule

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.

Forward this to the person who signs.

  1. Ceiling rule met. Criterion: a judge's ceiling is the highest level it passes before a break, capped at 13; passes above a break do not count, and the level-14 questions are exploratory and never set a ceiling. Unit: one judgment on one offered question. Source: the reasoning-ceiling probe run of 2026-08-19, checked by our internal quality review.
  2. The simulation used 15,000 judge decisions, dated 2026-09-09. Method and results are in the whitepaper. Whitepaper
  3. Fifteen artifacts sit under sha256 fingerprints, checked at the start of every run.
  4. Two ceilings are published. Both sit at level 11, with a stated condition. Ceiling rule met
  5. Method: Chapter 3 of a book on models and metrology, by Barney and Barney. Edited by Fisher and Pendrill. Published by De Gruyter in 2024, pages 103 to 132. doi.org/10.1515/9783111036496-003