# Waypoint Deliverables — Full Adjudication

**Date:** 2026-07-16 (post-repair pass). **Adjudicator authority:** deliverable files only — the spec was never edited; spec-targeting findings are routed to SPEC-LEVEL INPUTS (§5) for the report agent.
**Inputs adjudicated:** (a) prior verifier findings (`verifier-findings-raw.md` — marker-bank truncation/absence findings resolved by construction; content findings adjudicated); (b) four fresh verifier reports (2026-07-16); (c) three external reviews (`external/codex-ultra-marker-redteam.md`, `codex-ultra-estimator-review.md`, `codex-ultra-package-review.md`, all read in full); (d) the four deliverables.
**Verdict key:** CONFIRMED (fix applied) · PARTIAL (fix applied in modified form — see DIVERGENCES) · REJECTED (no edit; reason logged) · VERIFIED-NO-ACTION (claim checked out; nothing to fix) · RESOLVED-BY-CONSTRUCTION (pre-repair structural finding) · ROUTED (spec/study-design item → §5).

**Edit operations applied:** literature-dossier.md 19 · marker-bank-v0.md 12 · probe-item-bank-v0.md 13 · estimator-design-memo.md 30 — 74 total across the four deliverables.

---

## 1. Prior verifier findings (verifier-findings-raw.md)

### 1.1 Corpus/digest verifiers (ac945b1dd, a9f130da5, a706f224b, ae344b5ac)
Digests only, no defect findings against deliverables → NOTED, no action.

### 1.2 Verifier a9f591dc9 — literature dossier (was never edited after these findings; all applied now)

| # | Finding | Verdict | Action |
|---|---|---|---|
| 1 | §1.2–1.4 claim [C5] on T+N evidence only (scheme caps 2 families at [C4]); §1.2 calls four *lineages* "four independent families" | CONFIRMED (MAJOR) | All three downgraded to [C4] with inline recount notes; §1.2 reworded ("four traditions + the state/trait literature"); Cahn & Polich and Nisbett & Wilson retagged out of family N (see #7). §1.1 keeps [C5] legitimately (T+P+N) |
| 2 | §5 item 7: Sparby & Sacchet J1–J8 called "top-of-Wheel (51+) cartography" vs canon §8 overlay (jhāna→formless ≈31–50, cessation at 51) | CONFIRMED (MAJOR) | Re-anchored to bands 4–5 (~31–51) with canon §8 cite |
| 3 | §1.5/§1.9 demand off-cushion behavior the spec §5 channels cannot observe above ~41 | CONFIRMED (MAJOR) | Both arrows restated as claim-vs-telemetry consistency + durable linguistic trait-shift, observability gap named explicitly (matches the fix the repaired marker bank already carries in its band-5 preamble) |
| 4 | §1.2 misstates spec §6 ("volitionally reproducible" is not in the spec's HGF rule) | CONFIRMED | Restated as a tradition-derived **recommended amendment** for M1 spec revision, not existing spec content |
| 5 | §4.3/§4.7 "51–60 Low-X cascade" — cascade is 52–56 | CONFIRMED | Heading and §4.7 item 3 renamed 52–56 with boundary note |
| 6 | §3.4 internal inconsistency: expected ~5–15% vs miscalibration outside ~2–15% | PARTIAL | Not unified — lane 4 §7 deliberately uses **both** bands (expectation ~5–15%; wider review-trigger ~2–15%, looser on the low side). Fixed by making the two-band design explicit instead. → DIVERGENCE D1 |
| 7 | Family-tag misassignments: Hanley 2018 tagged (C) twice; Nisbett & Wilson + Cahn & Polich tagged (N) | CONFIRMED | Hanley retagged C→P in §2.1 and §3.2; N&W and C&P moved to "general cognitive science, outside the four families" — both retags reduce apparent cross-family independence, consistent with the C5→C4 downgrades |
| 8 | §4.4 `Lindahl & Britton 2019` bare cite: no flag, no venue, not in must-read list | CONFIRMED | Attributed to lane 4 §7.1 with explicit [UNVERIFIED metadata this pass] flag; verified set (§5.E items 16–18) pointed to. Not deleted: the discriminator itself is real and lane-carried |
| 9 | §1.3 "Nine lineages insist" — only five evidenced | CONFIRMED | Live-checked lane 1 §11.4: it evidences exactly five (Dzogchen, Zen, Neidan, Mahāsi, Christian) for this principle; "nine" is the lane's synthesis banner. Softened to "five lineages directly evidenced… lane 1 claims the pattern across all nine surveyed" |

### 1.3 Verifier acea43b9a — dossier citations (15/15)
All VERIFIED-NO-ACTION. Follow-through applied: `Gamma & Metzinger 2021` "confirm on pull" flags lifted (×2, §2.1 + §5 item 14) since this verifier confirmed it live.

### 1.4 Verifier a21585813 — marker bank

| # | Finding | Verdict | Action |
|---|---|---|---|
| 1 | CRITICAL truncation (ended at M-A4-010) | RESOLVED-BY-CONSTRUCTION | Repair completed the file; confirmed §§2.3–7 all present |
| 2 | M-S4-004 + M-A4-009 keyed to ambiguous Stage N→N+1 gates (canon §9-A) | CONFIRMED | Repair applied the note to M-S4-004 only; **M-A4-009 note completed this pass** (both §9-A readings recorded, Surya adjudicates at M2) |
| 3 | Redmore 1976 PMID 16367343 "range doubt" | REJECTED | PMID live-verified twice by independent verifiers; PubMed legacy retro-assignment explains the high number; bank §7 already records the doubt-and-resolution |
| 4–7 | §9-D day-totals; M-X3-002 anchor gloss; M-A4-009 Mobility; M-S4-005 canon-cite drop | VERIFIED-APPLIED (by repair) | Spot-checked all four in the repaired file; M-S4-005 left a locus-gloss residue → fixed under fresh finding 2.3 |
| 8 | M-X3-001 floor-marker imitability / rule 15 promotion | VERIFIED-APPLIED (by repair) | Rule 15 present; imitability reads "n/a — floor marker" |

### 1.5 Verifier a9912d000 — marker bank citations
Structural finding resolved by construction. Anālayo-ladder gloss and Farias design-gap reframing verified applied at point of use (lines 176/69). Items 4–15 VERIFIED-NO-ACTION. Neidan zhèngyàn UNVERIFIED honesty flag retained.

### 1.6 Verifiers a0dd2968 + ac2605a — probe bank missing; shared-base citations
Missing-file findings RESOLVED-BY-CONSTRUCTION (probe-item-bank-v0.md exists, 52 items). ac2605a content items: Farias reframe applied (rule 4) ✓; preprint caveats for [AL]/[FSNR]/[VOH] carried in bank §7 ✓. VERIFIED-NO-ACTION.

### 1.7 Verifier a056b6af — estimator citations (15/15)
All verified. Two actionable items applied: **Grove et al. 2000** UNVERIFIED flag lifted + DOI 10.1037/1040-3590.12.1.19 added (§15); **Hui–Walter** extension caveat applied — folded into the stronger §10.2 withdrawal (see 3.2).

### 1.8 Verifier aa1bb4f32 — estimator memo

| # | Finding | Verdict | Action |
|---|---|---|---|
| 1 | §10.5 "dark-night at 38" persona vs canon (~52–59) | CONFIRMED (MAJOR) | Renamed **dip-phase at 38**; added a true 52–59 dark-night persona with corroborated boundary history |
| 2 | §9 milestone prior 0.175 is not the day-24 posterior (0.196 post-E2b or 0.115 post-E4) | CONFIRMED (MAJOR) | Recomputed **0.196 → 0.646 → 0.088**, snapshot P(reached)=.09; labeled as the pre-event (post-E2b) state with rationale. Codex independently audited the same number (0.196) → DIVERGENCE D2 records the 0.115 alternative |
| 3 | §7 "~4 weak" = 2.4 < E_min 2.5 | CONFIRMED | "~5 weak (3.0 bits)" with the failing arithmetic shown |
| 4 | §2.1 "bands 1–2 structurally ~0" — grid contains 15–20 (band 2) | CONFIRMED | "band 1 and 8–10 structurally 0; band 2 prior-negligible but in-grid" |
| 5 | Header "before M0 sign-off" conflates milestones; memo self-declares complete without attached reviews | CONFIRMED | Status → **pending-reviews-complete**; M0→M1-close; Codex reviews referenced at their actual paths; Oracle marked not yet attached |
| 6 | §10.6 "never authored, extracted, or judged by same family" vs §11 "acceptable C = B" | CONFIRMED | §10.6 narrowed: never *extracted* by authoring family; M3 judging is deterministic |
| 7 | "every number signed by a human" overclaims (v0 bootstrap is mechanical) | CONFIRMED | Qualified: human-signed from M2; v0 bootstrap machine-derived and traceable |
| 8 | E3/M-A4-008 read as `first_glimpse` conflates canon 39 Non-Dual State with ~41 awakening glimpse (§9-J requires recording the sense) | CONFIRMED | §9 sense note added; snapshot annotated "glimpse-class sense per §9-J"; connects to the bank's new M-PGL-001 milestone dual-class |

---

## 2. Fresh verifier findings

### 2.1 Fresh verifier 1 — marker bank citations (12/12)
All VERIFIED-NO-ACTION (Farias figures exact and self-flagged; Vonk & Visser 2021 dating defensible; Fox 2017 instrument + subscales correct — the "SBS-13" item-count label remains not independently confirmed, noted here, left in place; Redmore, U Paṇḍita, Anālayo, Lindahl 2020, Laukkonen & Slagter all exact).

### 2.2–2.6 Fresh verifier 2 — marker bank structure

| # | Finding | Verdict | Action |
|---|---|---|---|
| 1 | **§1.4 does not exist** — the curriculum→stage bridge is cited 7× (lines 18/36/58/213/230/285/923) as the only path for 8 curriculum markers, but §1 ends at §1.3 | CONFIRMED (CRITICAL) | **§1.4 written** (bank): curriculum-position latent, elicited bridge natural-frequency table, weak-tier-capped v0 bootstrap from canon §10 correspondences, §9-A ambiguity binding, rule-17/exposure inheritance. **Estimator memo aligned** (the memo predated the bank's v0.1 curriculum class): `curriculum` added to both router layer lists + a `c_k` latent bullet in §2.2 referencing the bank §1.4 bridge |
| 2 | Spec §4 names `reverse_breathing | first_glimpse | small_death` as milestone marker_ids; bank classes only small_death as `milestone` — class-based routing would drop two of three from the snapshot | CONFIRMED (MAJOR) | M-S4-004 dual-classed **curriculum + milestone**; M-PGL-001 dual-classed **state + milestone-claim** (first M-PGL-001 event enters `first_glimpse` as claimed; corroboration via M-X5-001/P-049 only). Design intent (no direct band update) preserved. §0 revision log documents it |
| 3 | M-S4-005 locus still says "Theravada-overlay A&P territory" after the Source dropped that cite | CONFIRMED | Locus re-glossed to [AL]'s ñāṇa mapping with the dropped-cite explanation |
| 4 | M-S5-003 cites [CANON] §3 for "49–50 No-thingness rows" — §3 rows 49–50 are Equanimity/Small Death; No-thingness is the §8 overlay | CONFIRMED | Cite corrected to §8 with the row identities named |
| 5 | M-NX-003 foil (Wisdom asserting a nonexistent phenomenon) in tension with spec §11's "probes double as legitimate teaching moves" (finding truncated mid-sentence) | CONFIRMED (as recorded tension) | Tension note added at M-NX-003 (foils are the one probe move the §11 mitigation cannot cover — hence D3 duties/caps/Surya clearance) + added to probe bank §8.3 M2 spec-revision agenda. The underlying ethics question → SPEC-LEVEL INPUTS item 4 |

### 2.7–2.10 Fresh verifier 3 — probe bank citations

| # | Finding | Verdict | Action |
|---|---|---|---|
| 1 | P-009 "genuine equanimity increases contact ([MB] CT-1 via [FSNR])" is an interpretive derivation from one psychopathology caveat, not a paper finding | CONFIRMED (UNSUPPORTED-AS-CITED) | P-009 marked author-derived inference; same qualifier added at the claim's home in the bank (M-E4-005 Source) for consistency |
| 2 | Belmont-duties list misattributed: Belmont's conditions are necessity, minimal undisclosed risk, **debriefing plan**; "independent review" is not Belmont's; debriefing silently dropped | CONFIRMED (attribution overreach) | Disclosure note + §0.3 D3 reworded: duties are **Belmont-adapted, not Belmont's**; Belmont's own three stated; independent review + no-false-teaching named as our additions (latter tracking Belmont's truthful-answers rule); **debrief/dissemination plan added as an open M2 item** |
| 3 | Keys [FSNR]/[NIR]/[AL] used in P-009/P-030/P-032 but absent from §9; integrity statement claims [MB]/[L1]/[L5]/[CXR] only | CONFIRMED (citation-integrity gap) | Three rows added to §9 (resolving through [MB] §7 with preprint caveats); integrity statement amended |
| 4 | P-040 "the ONLY countermeasure with that property" is lane 5's superlative, not Paulhus's claim | CONFIRMED (minor) | Attribution split: paper's confirmed claim vs [L5]'s ranking judgment |
| 5 | Belmont passage, Paulhus, Redmore, Rogers verifications | VERIFIED-NO-ACTION | Paulhus flag also lifted in dossier §5 item 22 on the strength of this pass |

### 2.11–2.14 Fresh verifier 4 — probe bank vs spec

| # | Finding | Verdict | Action |
|---|---|---|---|
| 1 | Support probes P-034/P-036/P-039 permitted **and scored** during YELLOW+/dark-night — spec §5(2)/decision #4 says "never during flagged-vulnerable moments," unqualified; exception never flagged for amendment | CONFIRMED (MAJOR, spec-constraint violation) | Applied **both** halves of the verifier's either/or: (a) while any flag/watch is active, support-probe responses route to **safety/phase feature capture only — never band evidence** (§0.4 rule 1, the three items' caps, §5 preamble: delivery inverts, scoring does not); (b) the carve-out added to §8.3's M2 spec-revision list. → DIVERGENCE D3 |
| 2 | DISCLOSURE DESIGN self-titled "binding recommendation"; "this bank adopts that pattern" — reversing locked spec decision #4 pre-M2; 25 D2 items defined against an unauthorized consent regime | CONFIRMED (MAJOR) | Retitled "recommendation to M2 — requires spec §5.2 amendment; not binding until it lands"; adoption language → "designed to that pattern, contingent on the amendment; until then the canonical spec governs and **D2 is not deliverable**"; §0.3 D2 carries the same contingency; §8.3 records it |
| 3 | P-025 Targets lists M-A5-002 but no key routes to it; key 3 routes to the M-A5-003 family | CONFIRMED | Target corrected to M-A5-003 (M-A5-002 coverage unaffected — P-031 still targets it, §8.2 claims hold) |
| 4 | P-003/P-044 "onboarding" caps contradict §0.4 rule 3's no-exception first-sessions ban (finding truncated: "add an explicit ru…") | CONFIRMED | Explicit rule-3 exception written: onboarding covariate/provenance items produce no band evidence in any direction — personalization/exposure/provenance only |

Fresh verifier 4's clean checks (52 items, tier counts 25/25/2, coverage matrix, canon anchors, §9-A/D/J handling): NOTED, consistent with my own reads.

---

## 3. External reviews

### 3.1 codex-ultra-marker-redteam.md [RT]
Already incorporated by the bank's v0.1 red-team revision. Spot-verified the §0 dispositions against the file: ten-worst dispositions real (M-S4-003 descoped; M-A4-008 no direct band update; M-E3-002/M-S3-002/M-S4-006/M-X3-002/M-A4-010 covariates; M-T4-001 downgraded; M-T4-002 curriculum; M-E3-003 tombstoned), rules 13–17 added, CT-1/2/3/5 criteria adopted, devotion/destabilization operationalized (M-PIT-001/002), misassigned markers re-classed, missing-list #1–11 closed by completion, #12–13 carried as declared limits (§6 items 7–8). VERIFIED-INCORPORATED — no further action. RT's strongest structural demand (split latent families; only trait/integration evidence updates stage) is implemented in weakened form (class system + weak-tier bridge) — pre-existing divergence, recorded at D6.

### 3.2 codex-ultra-estimator-review.md + codex-ultra-package-review.md — deliverable-level items
(The two reviews substantially overlap; the package file also contains duplicated body text and pasted terminal/skill artifacts around lines 96–181 — external inputs, outside deliverable authority, left unedited; noted for the report agent.)

**CONFIRMED and applied to the memo** (details in memo §14.1 "Accepted and applied"):
- Silence/stationary contradiction: ν+/ν− = 6 ⇒ silence drifts mass up-band toward the stationary distribution and sharpens — the opposite of the claimed "diffuses, uncertainty grows." Fixed: drift zeroed during unobserved gaps (symmetric gap rates), forward drift only over observed practice; boundary behavior specified; residual +0.75 pts/month observed-window drift recorded as an M3 decision point (§4.1).
- "No teleporting" softened to suppressed-not-impossible (§4.1).
- Clamp breaks the redistribution equation at boundaries → truncate-and-renormalize (§4.5).
- Archetype "neutrality" was false as mechanized (one-sided offset shift moves overall through coupling) → archetype deltas recentered to sum to zero across pillars; M3 swap-invariance check; review's stronger drop-archetypes-entirely position recorded as the fallback (§5).
- θ_o/δ identifiability ridge + static-δ vs per-pillar-diffusion inconsistency: recorded as named design risks with candidate resolutions (Σδ=0 / info-fraction reporting / derive-overall-from-pillars), cross-referenced to dossier §4.7-5's aggregation-rule question for Surya — not silently re-architected (§4.5). → DIVERGENCE D5
- BKT elicitation question elicited the posterior P(m|claim), not the likelihoods → re-worded to the two likelihood-direction questions (§4.6).
- Fractional matrix powers of the weekly phase matrix need not stay stochastic → continuous-time phase generator (§4.2).
- Possible double-tempering → confidence shrinkage declared compile-time-only (§3.3).
- "Evidence bits" ≠ absorbed information (toy: ~2.6 capacity-bits vs ~0.1–0.2 realized KL-bits); guard contradictions counted toward reportability → semantics renamed potential-capacity, guards excluded from the sufficiency sum, toy audit recomputed (2.64 this window; overall ≈ 9.7) (§7, §9).
- Toy-example audit: milestone prior 0.196 (concordant with verifier), bare-claim counterfactual 0.49 → **0.54**, starved-pillar "keeps the set wider" claim corrected (flat message = multiplicative identity), operation-ordering (kernel/evidence interleave) and band-kernel-vs-CTMC (0.002 vs ~10⁻⁹) fidelity note added, machine-readable M3 fixture committed (§9).
- Hui–Walter "embryonic design" withdrawn as an identification claim (§10.2).
- Coverage counts descriptive-only; pooled events not independent; practitioner as resampling unit (§10.3).
- Intervals labeled **model-conditional** until external calibration (§2.1).
- Consolidation estimate-blindness sharpened: posterior-aware span-flagging admits events only through estimate-blind re-extraction; residual selection-bias risk recorded (§4.3).
- Dossier: [C5] grades explicitly declared non-systematic-review expert grades (preamble caveat).

**Recorded as risks with mitigations, re-architecture deferred to M3 decision points** (memo §14.1 middle block; each kept-design choice is a judgment divergence, D4): opportunity-model absence (the deepest critique — positive-only events confound stage with volume/expressiveness/probe-policy/extractor recall; mitigations: ω-normalization, rule-13 episode clustering, probe categorical keys; count/point-process model is the M3 redesign candidate); episode-level dependence beyond retells; KL caps/run-overrides as data-dependent tempering vs contamination-mixture; phase-gating self-sealing vs joint update (F3 monitor + no-gating ablation); cross-layer evidence reuse (joint-factor ablation); Beta claim-reliability non-conjugacy (fixed conservative likelihood + `disputed` fallback); elicited-parameter uncertainty (correlated-perturbation ensembles at M3); M3-tuning-as-calibration-to-generator renamed and scoped.

**ROUTED to §5 (spec/study design):** evaluation circularity/freeze; N=5–10 statistics; single-rater circularity/"gold"/ceiling/smearing; covert-probe ethics + hidden-profile governance + DPIA + team power asymmetry; posterior/interval semantics as a spec output-object question; 1–100 point estimate removal; stage-blind safety/teacher-flag triggers; performative lock-in; construct honesty ("Surya-aligned teaching-profile estimate"); radical-v1-simplification recommendation.

---

## 4. DIVERGENCES (both positions recorded; none silently resolved)

- **D1 — §3.4 rate bands.** Verifier: unify ~5–15% vs ~2–15% (internal inconsistency). Adjudicator: lane 4 §7 deliberately distinguishes an expectation band (~5–15%) from a wider review-trigger band (~2–15%, looser low side); unifying would misstate the source. Fix applied = clarification naming both bands as deliberate. If Surya prefers one number at M2, unify then.
- **D2 — toy milestone prior state.** Verifier offered 0.196 (post-E2b) or 0.115 (post-E4 under strict §4.2 ordering). Chose **0.196**: the pre-event state avoids E4 both moving the stage prior and re-entering the milestone chain it primes (the cross-layer double-count codex flagged), and codex's independent audit reproduced 0.196. The 0.115 strict-ordering alternative is recorded here; if M3 fixes strict ordering with a joint factor, renumber.
- **D3 — support probes.** Verifier offered either unscore-during-flags or flag-for-amendment. Applied **both**: scoring restricted to safety/phase features while flagged (spec compliance in the interim — band evidence from flagged moments would violate §5(2) today) *and* the carve-out put on the M2 amendment agenda (care-delivery during flags is genuinely indicated; the spec should authorize it explicitly).
- **D4 — codex re-architecture demands vs shipped design.** Codex: remove KL caps, adaptive reliability, archetype effects, stage dynamics, phase gating from v1; adopt opportunity-aware episode-level observation model; or downgrade the product claim. Adjudicator: kept the design with (i) honest relabeling applied everywhere codex's semantics critique bites (model-conditional intervals, capacity-bits, descriptive coverage), (ii) each disputed mechanism given a named diagnostic, ablation, and fallback at M3, (iii) codex's coarse static ordinal model adopted as the **pre-registered ablation baseline** the rich model must beat prospectively. Rationale: for an internal-only, advisory, abstention-heavy instrument at N≈5–10, an auditable deterministic scorecard with honest labels is a legitimate v1; wholesale re-architecture is a design decision above an adjudication pass. Codex's position preserved verbatim in `external/` and summarized in memo §14.1.
- **D5 — identifiability ridge fix.** Codex: impose Σδ_k=0 globally (overall = pillar center) or derive overall from pillars. Adjudicator: applied zero-sum **to archetype deltas only** (repairing the specific false neutrality claim); the global constraint changes what "overall" means — that is the pillar→overall aggregation question the dossier §4.7-5 already routes to Surya (M2). Ridge + all three candidate resolutions recorded in memo §4.5.
- **D6 — [RT] latent-family split strength (pre-existing).** RT: only validated trait/integration evidence + repeated within-person change should update stage. Bank v0.1: class system + weak-tier curriculum bridge + state-first recurrence rules — a weakened implementation. Recorded in bank §0.6 and §6 declared limits; construct-validity controls and independent criterion validation remain calibration-protocol work (spec §7), out of bank scope.
- **D7 — §1.4 restoration path.** Fresh verifier 2 offered restore-in-bank or repoint-to-memo. The memo defined no bridge either (it predates the bank's v0.1 curriculum class), so restored §1.4 in the bank **and** aligned the memo (curriculum router layer + c_k latent) — repointing alone would have moved the dangling reference, not resolved it.

---

## 5. SPEC-LEVEL INPUTS (for the report agent — external-review findings that target the spec/study design; NOT deliverable defects; deliverables were not contorted to dodge them)

1. **Evaluation circularity — freeze the dossier before the interview.** (Codex estimator §5; package fixes 6–8.) Spec §7 ingests the gold interview as evidence while the estimator runs afterward → target leakage; post-M4 table revision would train on the validation set. **Second opinion: AGREE, adopt fully.** Cheap and pure upside: hash-freeze dossier + estimator forecast pre-interview; interview transcript enters only a separately labeled post-evaluation condition; freeze all elicited tables before labels; any post-M4 revision is v2, validated only on future windows/people. Fold into spec §7 at M1 revision.
2. **N=5–10 cannot support the claimed validation statistics.** (Coverage 8/10 has a 95% CI of ~44–97%; pooled pillar×round events are not independent; drop 3-bin reliability; κ inappropriate at this N/range.) **Second opinion: AGREE on the statistics — they are arithmetic, not opinion.** M4 should be re-scoped in the spec as a feasibility/reliability case series: descriptive coverage counts, per-case RPS + signed band error + complete case displays, practitioner as the resampling unit. Mild disagreement on absolutism: weighted κ can still be *shown* as a descriptive alongside case-level displays, never as an inferential claim (memo §10.3 now labels accordingly).
3. **Single-rater circularity; "gold" naming; the "information ceiling."** (Surya authors construct, parameters, probes, and criterion → agreement measures fidelity-to-Surya; interview-vs-dossier delta is a within-rater mode discrepancy, not a ceiling — regularization can even beat it.) **Second opinion: AGREE in substance.** Rename "gold label" → "Surya reference rating" (costless honesty; spec §7 already logs single-rater as a threat and the house divergence convention fits). Adopt: repeated blinded ratings after washout, duplicated anchor vignettes, Surya-elicited probability vectors instead of post-hoc smearing, second qualified rater on a subset when one exists. On the ceiling: **PARTIAL** — not a mathematical ceiling, but the delta remains the right headline *diagnostic* of what report-only data supports for this rater; keep the measurement, drop the "ceiling" language.
4. **Covert-probe disclosure ethics (+ hidden-profile governance).** (Assessment-aware item-blind consent; Belmont conditions; notice/opt-out/correction/deletion; GDPR Art. 35 DPIA; team power asymmetry; foils vs spec §11's "probes double as teaching moves.") **Second opinion: AGREE that assessment-aware, item-blind is the defensible floor** — "invisible as a program" inside a trust relationship is not defensible incomplete disclosure, and the probe bank is now explicitly contingent on the §5.2 amendment. Team-cohort v1 under spec §8's informed consent approximates it; alpha requires the enrollment language + opt-out. DPIA: agree it is strongly indicated before any alpha processing (religious/mental-health-adjacent profiling); do it once, early. Team protections (no manager access, independent stewardship): agree, write into §8. **Foils specifically: my own view is stricter than the deliverable's** — P-040 is the one item family that cannot double as a teaching move and asserts a nonexistent phenomenon to a student; if Surya or independent review balks at the D3 duties (now including a debrief plan), drop foils entirely — the bank survives on paraphrase pairs + perturbations, and Paulhus-style credibility weighting is a nice-to-have, not load-bearing.
5. **Are the intervals a real posterior?** (Generalized-Bayes tempered scorecard; fixed elicited parameters; nominal 80% has no established coverage.) **Second opinion: AGREE with the diagnosis; DISAGREE that v1 must be a coherent joint Bayesian model before it is useful.** Applied the honest-labeling half (model-conditional intervals, capacity-bits, descriptive coverage) to the memo; the spec's §4 output object should adopt the same labeling ("credible_interval" → model-conditional until M5 evidence). The right test is codex's own: pre-register the coarse static ordinal model as the baseline; if the rich machinery cannot beat it on locked prospective prediction, adopt the simple model. Spec §7/§12 should name that ablation gate explicitly.
6. **Additional spec-targeting items, briefly:** (a) *Remove the 1–100 point estimate* — PARTIAL AGREE: keep `point` in the schema as a display derivative of the posterior mean; make band+interval canonical (dossier §3.1 and the memo already treat sub-band precision as unsupported). (b) *Stage-blind safety / Tier-A5* — AGREE stage context must never lower an acute threshold or downgrade a flag (deliverables already implement flag-on-features + never-adjudicate; make "stage never downgrades safety" an explicit spec §8 line); PARTIAL on removing the posterior-mass-≥51 trigger: keep it as a supportive-triage signal (highest-support-need territory is the spec's stated rationale) while outreach criteria remain impairment/uncontrollability/persistence/suicidality/user-request. (c) *Performative lock-in (thermometer→thermostat)* — AGREE it is under-priced in spec §11: adaptation can manufacture the trajectory it later cites; v1's offline/advisory posture is the de facto shadow mode — name it, and add randomized audit probes + periodic no-profile counterfactual checks before runtime integration. (d) *Construct honesty* — AGREE: until independent raters exist, the validated construct is "a Surya-aligned teaching-profile estimate"; spec §6's constructs-on-trial stance already concedes this — say it in §7's metrics framing. (e) *Package-review file hygiene* — the external file contains duplicated review text and pasted terminal/skill-spec artifacts (≈ lines 96–181); harmless as input, but the report agent should quote from the deduplicated sections.

---

*Adjudication complete. Deliverables edited in place; spec untouched; external reviews untouched. Provenance markers ("adjudicator fix / verifier fix / external-review fix, 2026-07-16") are embedded at every edit site; the marker bank's §0.7 and the memo's §14.1 carry per-file summaries.*
