The codex-ultra package review answered "would a top psychometrician sign off on this?" with a flat No — "an unusually thoughtful, auditable expert-system prototype, not a validated developmental-stage instrument," and "not ethically deployable as written." The adjudicated response is not to contest that verdict but to shrink the claim to what it licenses: v1 is instrument-development research — internal-only, advisory, abstention-heavy, honestly labeled — with the validity work (independent raters, real coverage, prospective skill) staged at M4–M6. A fourth review — a stress-test of the adjudicated fixes themselves — then set the surviving path precisely: "No, not as a deployment-validation study. Yes, as a strict shadow-mode feasibility study — with one amendment: insert the blinded stage-free actionability gate before M2 and make passing it necessary to continue." Details in External Verdicts; the concrete spec changes in SPEC-REVISIONS.md.
Band-level (10-point) is the defensible resolution; sub-band point claims from language are not supported — with two real exceptions: discrete, telemetry-corroborated milestone events (reverse breathing, first glimpse, small death) and the per-pillar profile, which is where the finest genuine resolution lives. Below ~41 finer resolution exists but comes from telemetry + gate completion, not from parsing language; from ~41 up even band-level gets noisy and abstention is the default. The 56-point internal grid stays (it can hold bimodal "genuine 48 vs well-read 30" posteriors), but what is reported is band + interval. [C3–C4]
A ≈55–75 minute battery: NADA-T + NADA-S (top pick — closest instrument to the Wheel's non-dual axis), MODTAS, Hood M-scale (targeted 51+ ceiling slot only), MPE-92M, a STAGES-style sentence-completion protocol (the differentiator — a human-scorable instantiation of Waypoint's own method), and FFMQ with its DIF caveat baked in; optionally one self-transcendence scale. All once, calibration-only, scored as band-conditional convergent anchors — never the estimator, never a product surface. Rejected: MAAS (floors), MEDEQ-as-stage, PNSE/Finders (the cautionary artifact), Hawkins LOC (the negative control), neural markers in v1. [C4] — full table below.
The traditions score the structure of a debrief, never doctrinal content, and defeat scripting with unscripted fresh-angle follow-ups (Zen sassho: 20–100 checking questions per kōan precisely because the "answers" are semi-public). Five moves transfer directly: structure-not-content scoring; fresh-angle follow-up after any advanced claim; reproducibility-at-will + persistence before the slow latent moves; show-don't-tell; and under-claiming calibration (genuine attainers under-report; readers over-report). These became the 52-item probe bank, built on lane 1's 20 verification-question patterns. [C4]
A two-tier design over Britton-schema events (phenomenology, valence, duration, impairment, practice-link as separate fields): Tier A at monthly consolidation is conjunctive (cluster AND persistence AND impact) keyed to the empirically supported persistence-risk signature — dysregulated arousal, not sadness; Tier B expedited is disjunctive on the consensus intervention triggers (uncontrollability, loss of critical attitude, impairment, suicidality). History-neutral thresholds, probe-elicited difficulty down-weighted, expected Tier-A rate ~5–15% (review band ~2–15%). Binding rules: the flag nowcasts a present state, never predicts risk (event-prediction PPV ≤ 0.01 at these base rates), and Waypoint records features but never asserts the dark-night-vs-depression differential — a human adjudicates. [C4]
The Wheel's deep structure is strongly confirmed (position-vs-phase is Sufi maqām/ḥāl verbatim; per-pillar unevenness; non-linear intensifiable top; small-death milestone placement). Three disagreements go to Surya at M2: the Hawkins LOC numbers carry no epistemic weight (refuted method; the column is even non-monotonic in band 3) — keep as provenance at most; 100-point resolution overclaims — sub-band labels are phase/milestone vocabulary, not estimator targets; and the 52–56 "Low-X cascade" is a co-occurring, recurring phase cluster, not a fixed ordered sequence (scoring the order would reward map-readers). Plus one label collision worth renaming: point 43 "Dissociation" vs the clinical risk signature. Full list: dossier §4.7.
"No. A top psychometrician would call this an unusually thoughtful, auditable expert-system prototype — not a validated developmental-stage instrument. Its engineering is ahead of its measurement science."— Verdict section
"As written, it is not ethically deployable." … "State exactly what each score is claimed to mean and what decision it supports. Then show that independent trained raters, using a frozen operational rubric, can distinguish the proposed bands and pillars reliably."— Ethical soundness / What a top psychometrician would demand first
What it demands: the central problem is circularity — Surya defines the construct, approves the probes, supplies the likelihoods, and produces the "gold" labels, so agreement measures fidelity-to-Surya, not validity of the Wheel. Its single highest-value change: replace M4 with a hard, transparently consented shadow-mode validation phase (no stage output influencing anything, frozen scoring manual, repeated blinded + independent ratings, simple baselines first). Until independent raters exist, the honest construct name is "a Surya-aligned teaching-profile estimate." Also: assessment-aware consent, GDPR Art. 35 DPIA, stage never lowering a safety threshold, and the unpriced failure mode performative lock-in — the thermometer becoming the thermostat.
Adjudicator's verdict: the diagnosis is accepted nearly wholesale — 9 of its 10 spec-targeting findings adjudicated AGREE or PARTIAL and folded into SPEC-REVISIONS.md (freeze, rename, case-series stats, disclosure, DPIA, stage-blind safety, lock-in controls, construct honesty). The one substantive disagreement: v1 need not be a coherent joint Bayesian model before it is useful — instead codex's own simple model becomes the pre-registered baseline the rich design must beat prospectively (divergence D4/R16). File hygiene note: the review file contains duplicated text + pasted terminal artifacts (~lines 96–181); quote only from its top section.
"The design has good engineering bones — typed evidence, deterministic arithmetic, provenance, abstention, proper-score ambitions, and estimate-blind extraction — but it is not yet a coherent Bayesian state-space model. It is an elicited, tempered scorecard combining several Bayesian-looking components with posterior-dependent heuristics."— Overall assessment
Its three highest-value changes: (1) eliminate evaluation circularity (freeze before the interview, redefine M4 as a feasibility case series); (2) replace the heuristic updater with one coherent observation model (opportunity-aware emissions, episode clustering, joint phase handling, one likelihood per datum); (3) radically simplify v1 for N=5–10 (coarse static ordinal states, no dynamics/archetypes/adaptive reliability). It also caught real math errors — the silence kernel that would have advanced practitioners up-band during gaps, a boundary-breaking clamp, a false archetype-neutrality claim, wrong-direction BKT elicitation, and "evidence bits" overstating realized information ~20×.
Adjudicator's verdict: ~15 findings CONFIRMED and applied to the memo (silence-drift zeroed in gaps, truncate-and-renormalize, zero-sum archetype deltas + swap-invariance check, likelihood-direction elicitation, capacity-bits semantics, model-conditional intervals, toy-example recomputation, Hui–Walter claim withdrawn). The deepest structural critiques (opportunity model, phase self-sealing, KL caps, cross-layer reuse) are recorded as named M3 decision points with diagnostics, ablations, and fallbacks rather than silently adopted or dismissed — divergence D4 records both positions; change (1) went to the spec (R6/R9), change (3) became the M5 ablation gate (R16).
"This bank is not yet construct-valid as a stage estimator. It currently mixes at least six different things: meditation-path development, verbal/narrative sophistication, personality and cultural style, psychotherapy or coaching exposure, compliance with LIFE's curriculum, and transient meditation states."— Bottom line (reviewing the truncated v0)
What it demanded: split the bank into separate latent families and let only validated trait/integration evidence update stage; its "ten worst markers" list led with the Neidan verification signs (suggestion-prone, map-readable) and the non-dual glimpse (highly imitable, also produced by DPDR/psychedelics/panic).
Adjudicator's verdict: VERIFIED-INCORPORATED. Bank v0.1 was rebuilt around it before adjudication — all ten worst dispositioned (1 descoped, 1 deleted, 6 reclassified covariate/curriculum, 2 downgraded), a Class/Contamination field on every marker, five new global rules (episode clustering, capacity norming, down-faking, honesty-texture caution, alternative-route provenance), and its CT-1/2/3/5 discriminator criteria adopted verbatim. Its strongest structural demand is implemented in deliberately weakened form (evidence classes + weak-tier curriculum bridge, not fully split latents) — recorded as standing divergence D6, not hidden.
"The four fixes improve honesty more than validity. … The deeper category error is that Waypoint is tested as a predictor or Surya-imitation instrument but intended to operate as a controller of teaching and safety decisions."— Overall assessment
"No, not as a deployment-validation study. Yes, as a strict shadow-mode feasibility study — with one amendment: insert the blinded stage-free actionability gate before M2 and make passing it necessary to continue."— Closing verdict
Fix-by-fix: the hash-freeze is a partial resolution (a hash proves bytes didn't change, not independent assembly — freeze the whole evaluation bundle and embargo M4 dossiers from M2 elicitation); the feasibility case series is sound only if M5 loses deployment authority (strongest earnable claim: "merits a preregistered shadow-mode alpha study"); "reference ratings" is mostly relabeling without the 2×2 rating design (two dossier + two interview ratings, randomized, washout — the delta only means something against within-mode baselines); and disclosure must be use-aware, with the DPIA before the team M4 and P-040 dropped outright. Its sharpest new content is the closed loop nobody had modeled: the estimator counts its own decisions as confirmation (practice-intensity drift, adaptive probing, full-history smoothing erasing treatment effects) — hence evidence origin tags, a measurement_environment_version, contemporaneous frozen forecasts for any policy claim, and safety as an executable, property-tested invariant (stage-blind primary lane, severity = max of lanes, ≥51 flags shadow-only, alerts lead with observations, never "unbinding-adjacent").
Disposition: landed post-adjudication, so it is folded straight into the spec proposals — R17 (the blinded actionability gate as new milestone M1.5, pass required to reach M2), R18 (origin tags + measurement-environment versioning; controller actions never count as competence), R19 (frozen contemporaneous forecasts; dropout as a policy outcome), and amendments to R6, R8, R12, R13, R15, R16. Its cheapest, highest-value demand: the M1.5 gate costs ~two short Surya sessions and answers whether stage information improves teaching decisions at all — before 4–6 sessions are spent eliciting marker rows.
| Territory | Items |
|---|---|
| Entry & floor, bands 3→4 (21–35) | 4 |
| Band 4 by pillar (31–40) — the core app band | 16 |
| Band 4→5 boundary & band 5 (41–50) | 12 |
| Bands 6–7 (51+) — recognition & support | 3 |
| Phase & confusion-pair discriminators | 4 |
| Anti-contamination (foils, paraphrase, perturbation, provenance) | 5 |
| Behavior–claim consistency (telemetry crosses) | 8 |
3 designated support probes (P-034, P-036, P-039) stay deliverable during flags — as care, scored for safety/phase only.
The review caught that the original rates would have advanced practitioners up-band during silence and eventually sharpened the posterior. Corrected: drift is zeroed during unobserved gaps (silence now genuinely widens uncertainty); forward drift applies only over observed practice, and a >6-point monthly shift trips a mandatory-rationale audit.
Per-pillar reporting requires ≥2.5 capacity-bits of trailing evidence, an 80% interval ≤2 bands, and no unresolved conflict — otherwise insufficient or conflicted (a distinct state: contradictions are never averaged into a midpoint). Guard events can never make a pillar more reportable.
The Bronlet 2025 house pattern (decomposed structural rubrics, temperature 0, median-of-N) hits weighted κ≈0.78 vs expert raters on the closest published task — that, not perfection, is the v1 extractor target vs Surya. Model bumps are recalibration events gated by golden-transcript regressions (lane 7: no formal LLM measurement-invariance framework exists; ours would be first).
| Instrument | What it anchors | Why it earns the slot |
|---|---|---|
| NADA-T + NADA-S | The Wheel's central non-dual axis (trait + state) | Top pick — closest validated instrument to the construct; its self-transcendence/bliss split mirrors the corpus's epistemic-vs-affective distinction (α .81–.94) |
| STAGES-style sentence completion | Language-structure stage scoring | The differentiator: a second, human-scorable instantiation of Waypoint's own method — the cleanest place to measure LLM-vs-trained-scorer agreement (κ 0.7–0.8 bar) |
| MODTAS | Absorption capacity | Discriminates the absorption↔dullness pair (CT-3) |
| MPE-92M | Pure-awareness / minimal phenomenal experience profile | Validated top-of-awareness-pillar phenomenology |
| Hood M-scale | Introvertive mysticism | Targeted 51+ ceiling slot only — near-useless in 22–45, but the only validated vocabulary reaching territory the estimator must still recognize |
| FFMQ (DIF caveat baked in) | Dispositional mindfulness | Kept because its item-level DIF across experience levels is itself informative at calibration |
| Optional: one self-transcendence trait scale | Trait-level transcendence | Cheap convergent redundancy if session time allows |
Explicitly rejected: MAAS (floors above basic mindfulness) · MEDEQ as a stage measure (session-depth scale — mined for probe wording only) · Jeffery Martin PNSE/Finders (the Goodhart cautionary tale, not an instrument) · Hawkins LOC calibration (the negative control: single unvalidated rater, false precision, non-reproducible procedure — the anti-pattern Waypoint's whole validation protocol repudiates) · neural cessation markers in v1 (single-digit-adept, internally contradictory).
| ID | What stays contested | Resolution path |
|---|---|---|
| D1 | Flag-rate bands: ~5–15% expectation vs ~2–15% review-trigger — deliberate two-band design, not an inconsistency | Surya may unify at M2 |
| D2 | Toy milestone prior 0.196 (pre-event state) vs 0.115 (strict operation ordering) | Renumber if M3 adopts joint factors |
| D3 | Support probes during flags: unscore vs amend the spec — both applied | M2 amendment (R4) |
| D4 | Codex's re-architecture demands vs the shipped v1 design — kept with honest labels, named diagnostics, and codex's simple model as the pre-registered baseline | M3 decision points; M5 ablation gate (R16) |
| D5 | Identifiability ridge fix: zero-sum applied to archetype deltas only; the global constraint is the pillar→overall aggregation question | Surya at M2 (dossier §4.7-5) |
| D6 | Red-team's full latent-family split vs the bank's class-system implementation | Declared limit; construct-validity work lives in spec §7 |
| D7 | §1.4 restoration path: restore-in-bank AND align memo (not either/or) | Done; noted for provenance |
| ID | Spec § | Proposed change (one line) |
|---|---|---|
| R1 | §2 | "Surya gold labels" → "Surya reference ratings" everywhere — agreement measures fidelity-to-Surya, not truth |
| R2 | §2 | "Information ceiling" → within-rater mode discrepancy; the delta stays the headline diagnostic |
| R3 | §4 | Intervals labeled model-conditional until M5 coverage evidence; point demoted to display derivative; conflicted + per-pillar sufficiency; param_version; P(reached) |
| R4 | §5 | Fully-covert probes → assessment-aware, item-blind disclosure at enrollment + the support-probe carve-out (delivery as care, safety/phase scoring only) |
| R5 | §6 | Peak states move the slow latent only when volitionally reproducible + persistent (the traditions' one missing criterion) |
| R6 [ST] | §7 | Freeze the entire evaluation bundle before the interview (dossier-construction rules, probe policy, event log, code/tables, analysis plan + forecast); interview transcript quarantined; M4 dossiers embargoed from M2 elicitation; post-M4 revisions are v2, validated on genuinely new practitioners |
| R7 | §7 | Headline-number wording updated at the protocol site (delta, not ceiling) |
| R8 [ST] | §7 | Single-rater mitigations: the 2×2 rating design (two dossier + two interview ratings, randomized, washout; delta reported against within-mode baselines), anchor vignettes (drift-detection only), probability vectors elicited at rating time, second rater on a subset when one exists — until then the claim is "replication of Surya's reference ratings" |
| R9 | §7 | Metrics reframed as a feasibility case series honest at N=5–10: per-case RPS + signed error + complete displays; coverage descriptive with exact binomials; κ descriptive-only; practitioner = resampling unit |
| R10 | §7 | Construct honesty: what M4–M5 validates is a "Surya-aligned teaching-profile estimate," scoped to map-exposed cohort members in the observed range |
| R11 | §7 | Synthetic QA labeled scenario tests + arithmetic checks (SBC-style) — never coverage evidence |
| R12 [ST] | §8 | Stage never downgrades safety as an executable invariant: primary classifier stage-blind, stage-aware lane separate, severity = max of the two, never suppress/downgrade/delay a flag or swap clinical language for contemplative framing, property-tested under wrong-stage injections; ≥51 flag shadow-only; outreach on raw features; alerts lead with observations, never "unbinding-adjacent" |
| R13 [ST] | §8 | Use-aware consent (profile inference + its uses disclosed) and DPIA before the team M4 study, not just alpha; power-asymmetry protections (independent stewardship, no manager access); notice / opt-out / correction / deletion; passive vs disclosed-probe evidence analyzed separately |
| R14 | §11 | New risk row: performative lock-in (thermometer→thermostat) — v1 named as de facto shadow mode; randomized audit probes + no-profile counterfactuals before runtime |
| R15 [ST] | §11 | The foil family (P-040) is dropped outright — it can prime reports, contaminate the evidence stream, and damage trust; paraphrase pairs + perturbations carry the anti-contamination load |
| R16 [ST] | §12 | M5 keeps the pre-registered simplicity ablation and loses deployment authority: strongest earnable claim is "merits a preregistered shadow-mode alpha study"; runtime adaptation, stage-aware safety, and flag activation are decided only by the later estimate-visibility policy trial |
| R17 [ST] | §12 | New milestone M1.5 — blinded stage-free actionability gate, pass required to continue: on ~12–20 frozen cases, Wisdom recommendations under no-estimate / reference / wrong-estimate views, ranked blind by Surya; redundant latent, brittle policy, or unsafe escalation each stop the program cheaply |
| R18 [ST] | §4–6 | Evidence origin tags (spontaneous / standardized_anchor / policy_elicited / post_teaching / post_outreach) + measurement_environment_version (Wisdom prompt/model, curriculum, probe scheduler, safety classifier, flag policy); assigned curriculum and controller-driven cadence never count as competence; outreach recorded as an intervention |
| R19 [ST] | §7 | Policy/causal evaluation scores contemporaneous frozen forecasts, never smoothed history (smoothing can rewrite an intervention as pre-existing state); dropout tracked as a policy outcome and possible harm, not missing data |
Deliberately not spec edits: the Wheel-model items Surya owns at M2 (Hawkins numbers, 52–56 as phase cluster, the point-43 "Dissociation" rename, pillar→overall aggregation, the §9-A stage-numbering ambiguity that displaces the reverse-breathing gate by a full band) and the estimator re-architecture candidates parked as M3 decision points (D4).
Main research pass: a ~30-agent run (corpus digests, 7 lanes, 4 deliverable authors, finder→verifier adversarial checks on every empirical claim). Two agents died mid-stream and a session limit stalled the pass — a same-day repair pass completed the marker bank (§§2.3–7 had truncated at M-A4-010) and the missing probe bank, then re-ran four fresh verifiers over everything. The adjudication pass applied 74 edits across the four deliverables, with provenance markers at every edit site. This report itself was delayed: two earlier report attempts died on transient API 529 errors before writing anything; this is the third attempt. A fourth external review — the adjudication stress-test — landed after the first build of this report and is folded into the verdicts section and SPEC-REVISIONS (R17–R19 + six amendments). Still outstanding: the Oracle (GPT Pro) review that spec §10 requires attached to the estimator memo — the memo carries status pending-reviews-complete, so M1 does not close until it lands (plus sign-off on the spec revisions).