Judge Optimization Lab · final manual review

Every judge prompt, from shipped initial to final candidate

A source-locked review of all 11 rubric bodies: production v2 → v0.3.2 redesign baseline → the 2026-07-14 final candidate panel, with every documented change explained from Child through Expert, plus exact diffs, full prompts, and the open scope-review queue.

Snapshot: 2026-07-15 · final candidates are not the deployed production panel · rubric bodies only; runtime wrappers vary by evaluation unit

11 / 11judges included
5,116 → 3,915prompt lines
2,174 / 3,375lines added / removed
19open scope findings
3promoted records
250change cards explained
0 / 11manually reviewed

Three states are kept separate

1 · Shipped initial

The production rubric body is the v2 prompt, pinned in the Data Lab registry and bundled with a two-line vendoring header. The report verifies the stripped runtime body byte-for-byte against the embedded v2 source before building.

2 · Redesign baseline

v0.3.2 is the shared-chassis state from which the July judge-local manifests were written. It is retained as a separate exact prompt and local diff so the rationale documents line up with the text they actually changed.

3 · Final manual-review candidate

AI Safety remains v0.3.2; Gestalt and Zen are v0.3.4; the other eight are v0.3.3. “Final” here means the end of the documented redesign pass—not deployed production.

4 · Executed prompt caveat

These are rubric bodies. Production inserts each body into pairwise, trajectory, candidate-quality, or classifier-disagreement wrappers; offline benchmarking uses a separate golden-judge wrapper.

Do not read version labels as deployment state. Production still serves the v2/v0.2.1-era rubric bodies. Seven Fable redesigns and Mahayana remain candidates; the July 14 scope pass was recommendations-only and touched no prompt.

Initial → final, judge by judge

Line deltas compare the shipped v2 body with the final candidate. “Open” counts surviving July 14 scope findings naming that judge.

JudgeVersionsCandidate state+ / − linesOpen
v2 · shipped/runtime body → v0.3.2 Promoted · diagnostic evidence +183 / −353 10
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +175 / −328 7
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +222 / −358 12
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +271 / −366 8
v2 · shipped/runtime body → v0.3.4 Promoted · diagnostic evidence +255 / −321 5
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +213 / −325 7
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +167 / −330 7
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +173 / −228 8
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +262 / −335 8
v2 · shipped/runtime body → v0.3.3 Candidate · not promoted +135 / −193 6
v2 · shipped/runtime body → v0.3.4 Promoted · exploratory record +118 / −238 4

What changed before the judge-local July deltas

Why the overhaul existed: v0.2.1 concentrated 83.8% of 1,771 grades in 6–7.5, produced 0/1,771 scores ≥9, fired penalty machinery on 1.0% of grades, and let visible safety effort create a RED>GREEN inversion. The v0.3 chassis made the instrument discriminating again.

Live top band + scenario fit

Scores 9–10 became awardable; the failed complexity gate became fit to this scenario, so proportionate simplicity can be excellent.

Evidence-based failure mechanics

Tier-2 point subtraction became quoted-evidence flags plus a consistency guard; Tier-1 remains a cap. Borderline signals move to observations, not the firing array.

Anti-template + vocabulary neutrality

Portable response templates cannot masquerade as mastery, while plain-language understanding receives the same credit as tradition terminology.

Canonical output + routing-owned relevance

Relevance was removed from each judge; routing determines membership. One JSON contract and one aggregation path replaced bespoke variants.

Ratified decisions 0001–0011

Loading decision trail…

12 confirmed findings + 7 plausible findings still open

These are not implemented changes. They are the concrete queue for this final manual review; each item is also repeated inside every affected judge.

Loading scope findings…

Semantic rationale + exact prompt text

Loading prompt review…

What to settle before calling the panel final

  1. SR-01 + SR-15: ratify one bypassing/positive-reframe construct and acknowledgment bar before validation labeling.
  2. SR-03: settle dependency ownership and tier between AI Safety, Coaching, and Depth Psychology.
  3. SR-04: decide whether safety-theater should multiply as Tier-1 across all 11 judges on YELLOW/ORANGE.
  4. SR-06 + SR-19: resolve the dark-night-dismissal and trauma-blindness name collisions without mixing distinct constructs.
  5. SR-08 + SR-17 + SR-18: decide central homes for validation capture, privacy misrepresentation, and authoritative attainment verification.
  6. SR-09 + SR-13 + SR-16: align referral clocks, grounding protocol ownership, and generic-versus-specialist screening.
  7. Validation probes: run the missing cells for Zen clinical labels, values questions, dual reverse-breathing caps, WM off-domain routing, and acknowledgment-then-reframe.
  8. Promotion discipline: eight candidates still need their sealed-holdout A/Bs; Zen’s spaced test-retest remains due; production must be explicitly promoted and rebaselined afterward.
Notes remain local to this browser until exported.
Final-panel digest 4c8b462535048c1c · authoring prompts commit bb8ff13baab8 · briefing commit 656fe2489e2d · self-contained, no network requests