A strict reviewer is not a good reviewer

Rejecting everything catches every bad draft. A reviewer scorecard needs false passes, false blocks and abstention. We ran the test on our own review gate.

Article highlights

  • Estimated reading time: 8 minutes
  • Published on: October 1, 2026
  • Last updated: October 1, 2026
RA
Published · Updated · 8 min read
Branded cover for the article 'A strict reviewer is not a good reviewer': TrustGrowth wordmark and title with callouts: Strictness is a decision rate, not correctness, False pass, false block, abstention: separate denominators, Grade the outcome and the reason separately.

Article

Tell an AI reviewer to reject everything and it will stop every defective draft. It will also stop every publishable one.

Strictness is a decision rate, not a correctness rate. This piece sets out the figures that would show correctness, and the protocol we will apply to our own review gate before we quote any of them. It reports no measurement.

What we found

Strictness is the share of items a reviewer returns, not evidence that those returns were correct.

  • Our own gate runs in shadow mode beside human decisions. The internal check has not been published, and this piece reports no figure from it.
  • Correctness needs a reference judgment formed without sight of the gate, and it splits into rates that never share a denominator.
  • A reviewer can get an outcome right for a reason the rubric does not support. The audit below grades outcome and reason separately.
  • Abstentions and measurement failures belong in the denominator, or the accuracy figure flatters the reviewer.

What follows is a protocol, not a results report: no figure below is a TrustGrowth measurement. Any team can run it on its own queue.

What do strictness, false pass, false block and abstention mean?

  • Strictness: the share of items a reviewer returns, over every item it was handed. Nothing in that ratio refers to correctness.
  • False pass: an item the gate released although the reference rubric calls it unacceptable. Counted over all unacceptable items.
  • False block: an item the gate returned although the reference rubric calls it acceptable as it stands. Counted over all acceptable items.
  • Abstention: the reviewer declined to decide. Evidence missing, policy silent, or the case escalated without a verdict.
  • Measurement failure: the reviewer produced no usable decision: a malformed response, a timeout or a failed tool call. Not a choice, and counted on its own.

What a false pass hides

A false pass is an item accepted although it fails the rubric that applied to it. A high pass rate conceals these by construction: nothing marks the miss, because the gate produced no finding to check.

A deterministic gate checks what it was told to check. Four defect classes sit outside that list, and the audit will look for each of them in the released set: a title or keyword promise the body never delivers; a protocol described in a method section and then explicitly not run; near-duplicate articles inside one topic cluster; leaked pipeline artefacts. Each can travel through whenever nothing else fires.

We have written before about a fail-open scoring defect in our own publication gate, the same failure family from the other side.

What a false block costs

A false block is a send-back on an item an independent reference calls acceptable as it stands. Blocking hard looks like diligence, and on a dashboard the two look alike.

The audit will test the stated reason on every send-back against the rubric, and will flag reasons that rest on density and format proxies. The kinds a rule-based gate can produce: an overlap check that fires on earlier posts sharing brand vocabulary; a credential pattern that reads a genuine byline as fabrication; a long-form checklist applied to a forty-word social post. Whether ours does any of these is a question the audit answers, not one this piece answers.

A defensible send-back names the defect, quotes the span and states the requirement it violates. One driven by a misread rubric costs a revision nobody needed and teaches writers to serve the proxy.

Abstention is not failure, but it changes the denominator

Abstentions and measurement failures both belong in the denominator of the class they came from. Once they exist, the bad-item block rate stops being one minus the false-pass rate, because some unacceptable items received neither decision. An unverified pass cannot quietly be counted as safe.

Our gate has no abstention state today. "Could not assess this" has nowhere to go. The audit will classify such cases as unresolved, not as a block or a pass.

What will the audit record?

Table 1. Proposed audit protocol, dated 2026-09-06. No cell reports a result, and no TrustGrowth measurement is included.

Element What the audit will record Inputs Every item handed to the reviewer, including items with no verdict. Item type, rubric version, model version if supplied, how the set was built. Decision categories Pass, send-back, abstention, measurement failure. One category per item, taken from the reviewer output alone. Reference-label process Independent labellers grade each item from the text alone, without sight of the gate output, against the same rubric version. Disagreements are recorded, not averaged away. Reason-matching rule A stated reason counts as supported only when it names a defect the reference labeller also named, quotes the span and cites a rubric requirement. Outputs Each rate as numerator over denominator. We will publish the counts with their denominators when the audit has run. Limitations One queue, one window, one rubric family. LLM labels are a proxy, not human truth. Nothing transfers to another task or dataset.

How do you audit a reviewer in ten minutes?

  1. Count every item the reviewer was handed, including any it never gave a verdict for.
  2. Count passes, send-backs, abstentions and measurement failures separately, never folding any into another.
  3. Label each item against a reference rubric applied without sight of the gate output, then mark false passes and false blocks.
  4. Write every rate with its denominator visible, as numerator over denominator.
  5. Recalculate with abstentions excluded, label it a different metric, and keep both.
  6. Read the stated reason on each send-back and ask whether the rubric supports it.
  7. Record the date, the model version if supplied, the rubric version and how the set was built.

Table 2. Metric definitions with placeholders. "Your figure" is a placeholder, not a TrustGrowth measurement.

Metric Numerator Denominator Result What it does not prove Pass rate items released all items your figure that the releases were sound Send-back rate items returned all items your figure that returns were warranted Abstention rate items with no verdict all items your figure that the queue was clean False-pass rate releases judged unacceptable all unacceptable items your figure how much reached readers False-block rate returns judged acceptable all acceptable items your figure each revision's cost

Synthetic illustration, not a measured result. These figures are invented to show the arithmetic; a future TrustGrowth evaluation would replace them with counts and denominators. A reviewer finds 72 of 100 annotated defects while reporting 120 distinct findings, so 72 findings match a real defect and 48 do not. Findings match defects one to one; repeating a valid criticism earns nothing.

Synthetic illustration, not a measured result. Now the item level, again invented. Of 50 items a reference calls acceptable, the reviewer blocks 15. Of 50 it calls unacceptable, it passes 8. Bad items blocked: 42 of 50, which sounds reassuring. Unacceptable items among the releases: 8 of 43, which does not.

The first example is a finding-level metric: does each reported finding match a real defect? The second is an item-level outcome metric: was each item passed or blocked correctly? The two cannot be merged, because their denominators count different things, and an equal split is constructed.

Method and limits

No measurement is reported in this piece. Our gate runs in shadow mode beside human decisions, and the internal check has not been published. Every figure above is a definition, an invented illustration or a placeholder marked "your figure". When the audit has run, we will publish the counts with their denominators, the rubric version and the date.

Blind LLM labels are not human truth. Where LLM judges supply the reference, their labels are a proxy that needs its own check.

Public judge benchmarks do not supply the reference this audit needs; each tests a judge on its own task, not on a gate policy. JudgeBench (Tan et al., 2024) screens correctness judgement on hard response pairs. LLMBar (Zeng et al., 2023) asks whether an evaluator picks the output that followed the instruction against an attractive wrong answer. RewardBench 2 (2025) tests reward ordering, and labels part of its factuality subset by agreement between two model judgements with spot checks.

CriticEval (Lan et al., 2024) separates feedback, comparison, correction and meta-feedback, scoring subjective critiques with GPT-4 against human-refined references. The MT-Bench LLM-as-a-judge study (Zheng et al., 2023) makes the structural point: in an ordinary run the judge is part of the instrument, only its separate human-agreement experiments evaluate the judge. Work on self-preference bias adds that a judge can favour text from its own family. None estimates how often a gate wrongly blocks a publishable draft under a policy.

Do this next

  • Audit one reviewer evaluation end to end before trusting its headline figure.
  • Show every denominator next to its rate, in the same table.
  • Keep abstentions and measurement failures visible as their own rows.
  • Grade the rationale separately from the outcome; a right decision for a wrong reason is a defect.
  • Re-run the audit whenever the rubric, prompt, model or item set changes.

If you run a reviewer over AI drafts, we would like the awkward cases: one miss it let through, or one edit it demanded that you judged unwarranted. Send us the case with the reference judgment.

Last measured

No first-party evaluation is reported here. The benchmark papers are cited at the versions linked above: JudgeBench v2, RewardBench 2 v2, CriticEval v5, the MT-Bench study v4 and the self-preference bias paper v1. The audit date will be published with the counts.

One separate public data point about this site, not about any reviewer: our published proof page records a TrustGrowth score of 43 out of 100 for trustgrowth.ai, labelled Provisional with a 38 to 48 band, as of 2026-09-06, rolling Authority, Visibility, Growth, Corroboration, Technical, Content, E-E-A-T, Performance and GEO signals into one public snapshot built on first-party Google Search Console data. It is a site trust score, not a writing-benchmark score, not a model score, and not evidence about any reviewer.

AI evaluation reviewer evaluation false passes false blocks abstention measurement methods
Share:

Know your site's real SEO score

Free GSC-verified audit, E-E-A-T scoring, and AI-powered content strategy.

Get Started Free

Related Articles