Share this article:

🎧 ▶ Listen: 5-minute briefing
▶ Play audiobook (Google Drive)
Locally synthesized AI audiobook (Qwen3-TTS)

hero

If your retrieval pipeline asks a model whether the passages are good enough to answer, this post is for you. A score above 0.9 does not tell you whether that judge notices a broken line of reasoning.

We built a control that measures exactly that, trained three judges with it at 4B, 9B and 27B, and released them. Their ability to catch a broken chain rose sharply. On contexts with invented names they mostly got worse than before training. We report both.

In plain terms

Think of stepping stones across a stream. To get from the question to the answer you step on two stones. For “Where was the director of this film born?” the first stone is “who directed the film” and the second is “where that person was born”.

A judge should check that every stone is in place. There are two ways to fool it. One is to swap only the second stone for a story about someone else, which breaks the path. The other is to relabel the stones consistently, which leaves the path intact.

A good judge rejects the broken path and lets the relabeled one through. Ordinary benchmarks only test intact against broken. A judge that distrusts every relabeled stone therefore scores close to perfect.

Key-concept summary infographic 1 Infographic generated by NotebookLM from the sources.

A score at the ceiling separates nothing

The usual sufficiency test puts an untouched context next to a broken one. It counts how often the judge scores the untouched one higher. We call that the nominal score.

Most judges we measured sat near or above 0.9 on it. At that level the ranking between systems barely shows. Worse, the number does not say why the judge scored that way. It cannot tell whether the judge saw the broken path or simply reacted to signs of editing.

In other words, the exam is too easy to tell who is good at it.

Three contexts separate the two reactions

For every question, ChainCheck builds three versions of the context.

Context Path Edited
A intact no
B intact (labels swapped consistently) yes
D broken (only the second stone swapped) yes

B and D are both edited, so they differ only in whether the path holds. How much the judge prefers B over D is the chain effect. A and B both have an intact path and differ only in editing. How much the judge prefers A over B is the edit effect.

Chain selectivity is the chain effect minus the size of the edit effect. When it is clearly above zero, the judge is reading the path rather than the edit. We call it clear when the lower end of the 95% interval is above zero.

What we trained

We started from Qwen3.5-4B, Qwen3.5-9B and Qwen3.8-27B. All three judge the same way. They receive the question and passages, answer yes or no, and the score is the log-probability gap between the two.

Training used 7,481 triplets built from the MuSiQue training split. A and B are labeled sufficient and D insufficient. The loss also requires the lower of A and B to sit clearly above D. That tells the model directly not to treat a relabeled path as a broken one.

We trained one LoRA pass and merged it into the base, shipping one full set of weights. On a single B200 the runs took about 46 minutes for 4B, 55 for 9B and two and a half hours for 27B.

The release gate was fixed before any score was seen. On 2Wiki, never used in training, chain selectivity had to be clearly positive for both real and invented names. The MuSiQue test split had to show the same on real names. The 2Wiki nominal score could not drop more than 0.02 below the untrained model of the same size. All three sizes passed.

Results

We scored the untrained and trained models with the same prompt and the same pipeline. Chain selectivity below is on 2Wiki with real entities.

Model Nominal (before → after) Chain selectivity (before → after)
4B 0.932 → 0.983 +0.136 → +0.376
9B 0.898 → 0.993 +0.078 → +0.386
27B 0.905 → 0.997 +0.129 → +0.461

295 real-entity pairs on 2Wiki. 95% intervals from 2,000 resamples: 27B after training [+0.397, +0.485], before [+0.054, +0.200]. Full tables for every size are on the model cards.

2Wiki chain selectivity: up sharply at every size on real entities, down on synthetic entities except for 9B Bars are chain selectivity, lines are 95% intervals. Left: real entities. Right: synthetic entities.

Chain selectivity more than doubled at every size. The untrained 9B had an interval that crossed zero. Before training there was no evidence it read the path at all.

In other words, the untrained judges reacted heavily to editing, and the trained ones catch a broken path far more clearly.

Splitting the two effects shows what changed. On real entities the untrained models had an edit effect of about +0.1 to +0.2. They penalized intact contexts just for being relabeled. After training the chain effect climbed to almost its maximum of 0.5. The judges now separate B from D nearly perfectly.

Where it did not improve

Read this part before you deploy one of these models.

First, the trained models now lean the other way. They score the relabeled context B above the untouched context A. The edit effect turned negative in most cells and reaches −0.5. Chain selectivity subtracts its size, so the numbers above already include the penalty. Still, do not read the score as edit-invariant.

Second, when the labels are invented names the model has never seen, chain selectivity fell below its untrained level in most cells. The one exception is 9B on 2Wiki. At 27B it went from +0.40 to +0.17 on 2Wiki and from +0.42 to +0.05 on the MuSiQue confirmation split. It stayed clearly above zero, which the gate required. The gains, though, came only on real entities.

Third, at 27B the real-entity value on the MuSiQue confirmation split is lower than before training (+0.18 → +0.11). The intervals overlap, so this is not a measured regression. It is not a gain either.

In other words, expect a clear improvement on contexts about real Wikipedia entities. Do not expect one on contexts full of unfamiliar names.

Two things we caught while releasing

The first was missing weights after the merge. Comparing the merged 4B against its base tensor by tensor, 15 were gone. They were the multi-token prediction weights in the Qwen3.5 family. The training code never loads them, so saving silently dropped them. Scores were unaffected, but without them vLLM speculative decoding does not work. We now copy them over from the base unchanged. The release job also stops if any base tensor is missing.

The second was re-scoring the reference judge. We re-scored the paper’s 27B reference judge through this release pipeline. Item rankings matched almost exactly, yet chain selectivity moved a little within its interval. A and B differ by one name and score almost the same. Tiny precision differences flip some A-versus-B comparisons. Compare edit effects only between runs on the same hardware and precision.

What this means for ThakiCloud products

Work agents on Paxis search internal documents and answer from them. The costliest mistake is a fluent answer built on a broken chain of evidence. A judge before generation lets the agent answer, retrieve more, or stop. A judge that reacts to edits blocks a revised internal policy that only renamed a team. It then lets the broken evidence through. These controls catch that difference before release.

On Metis, size is cost. The judge runs once before every answer, so it is called as often as the generator. Even the 4B judge scored higher chain selectivity on 2Wiki than the untrained 27B. Put a small judge on latency-sensitive paths and the 27B where one wrong call is expensive. All three are merged full weights and load without adapter management.

On Maxis, the same recipe can be rerun on a customer’s documents. All it needs are rules for building three contexts: intact, relabeled and broken. No human labels are required.

Try it

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ThakiCloud/ChainCheck-Judge-Qwen3.5-4B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto")

SYSTEM = ("You are a strict evidence auditor for a retrieval-augmented QA system. "
          "Decide whether the given passages, taken together, contain enough information "
          "to fully answer the question. Related-but-insufficient passages do NOT count.")
USER = ("Question:\n{query}\n\nPassages:\n{passages}\n\n"
        "Do the passages together contain sufficient evidence to answer the question? "
        "Answer with a single word: yes or no.")

yes = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"yes", "y"}]
no = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"no", "n"}]

def sufficiency_score(query, passages):
    body = "\n\n".join(f"[{i + 1}] {p[:4000]}" for i, p in enumerate(passages))
    msgs = [{"role": "system", "content": SYSTEM},
            {"role": "user", "content": USER.format(query=query, passages=body)}]
    text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
    ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
    with torch.no_grad():
        lp = torch.log_softmax(model(**ids).logits[0, -1].float(), -1)
    return (torch.logsumexp(lp[yes], -1) - torch.logsumexp(lp[no], -1)).item()  # > 0 leans "sufficient"

The score is a log-odds, not a probability. Pick the threshold on your own validation data. This is not a chat model, so read the yes/no logits as above rather than generating text.

To measure your own judge on the same ruler, use the benchmark kit. Pass one file with a score per item. A numpy-only script reports the nominal score, chain effect, edit effect and chain selectivity with intervals.

python chaincheck_eval.py --data twowiki_replication.jsonl.gz --scores my_scores.jsonl --exclude-cb27b

Infographic generated by NotebookLM from the sources.

The models and kit are released under Apache-2.0, and the kit data keeps the licenses of its source datasets. Each model card carries the full table against its untrained baseline, the limitations above and the training settings. A paper describing the measurement protocol is in preparation.

If your judge already scores above 0.9, a higher score is not what you need next. Check first whether it reads the broken path or the signs of editing.

Measured on one B200, bf16, 2026-10-05. Every number compares against the untrained model of the same size, scored through the same pipeline.

Share this article:

Tags: chaincheck, counterfactual-evaluation, evidence-sufficiency, llm-judge, lora-merge, multi-hop-qa, open-weights, rag

Categories:

Updated: