SOME · Waypoint · M1 Research Phase

Waypoint Research Report

Estimating a practitioner's position on the Wheel of Life from conversation, probes, and telemetry — what the research phase found, what the external reviewers said, and what should change in the spec before anything is built.
Date 2026-07-16 Spec waypoint-practitioner-stage-assessment.md (draft v0.1, untouched) Deliverables 4 of 4, adjudicated Owner Fionn · Teacher anchor Surya

Where Waypoint stands after the research phase

All four deliverables exist, were adversarially verified, externally reviewed at ultra effort, and adjudicated line-by-line. The construct survives; the claims around it get smaller and more honest. No published computational estimator of contemplative stage exists anywhere — Waypoint is first, which means the calibration protocol carries the full evidential load.
4 + 4
deliverables + external reviews
91 / 52
marker entries / probe items
74
adjudication edits applied
19
proposed spec revisions
Headline external verdict: NO

The codex-ultra package review answered "would a top psychometrician sign off on this?" with a flat No — "an unusually thoughtful, auditable expert-system prototype, not a validated developmental-stage instrument," and "not ethically deployable as written." The adjudicated response is not to contest that verdict but to shrink the claim to what it licenses: v1 is instrument-development research — internal-only, advisory, abstention-heavy, honestly labeled — with the validity work (independent raters, real coverage, prospective skill) staged at M4–M6. A fourth review — a stress-test of the adjudicated fixes themselves — then set the surviving path precisely: "No, not as a deployment-validation study. Yes, as a strict shadow-mode feasibility study — with one amendment: insert the blinded stage-free actionability gate before M2 and make passing it necessary to continue." Details in External Verdicts; the concrete spec changes in SPEC-REVISIONS.md.

Headline answers to the five questions the research had to settle

Each is answered in full, with sources, in the literature dossier §3. Confidence grades are the dossier's own strict scheme ([C4] = two independent source families or one family + strong replication).

1 · What granularity does evidence actually support? band-level

Band-level (10-point) is the defensible resolution; sub-band point claims from language are not supported — with two real exceptions: discrete, telemetry-corroborated milestone events (reverse breathing, first glimpse, small death) and the per-pillar profile, which is where the finest genuine resolution lives. Below ~41 finer resolution exists but comes from telemetry + gate completion, not from parsing language; from ~41 up even band-level gets noisy and abstention is the default. The 56-point internal grid stays (it can hold bimodal "genuine 48 vs well-read 30" posteriors), but what is reported is band + interval. [C3–C4]

2 · Which instruments earn calibration-battery slots? 6 + 1 optional

A ≈55–75 minute battery: NADA-T + NADA-S (top pick — closest instrument to the Wheel's non-dual axis), MODTAS, Hood M-scale (targeted 51+ ceiling slot only), MPE-92M, a STAGES-style sentence-completion protocol (the differentiator — a human-scorable instantiation of Waypoint's own method), and FFMQ with its DIF caveat baked in; optionally one self-transcendence scale. All once, calibration-only, scored as band-conditional convergent anchors — never the estimator, never a product surface. Rejected: MAAS (floors), MEDEQ-as-stage, PNSE/Finders (the cautionary artifact), Hawkins LOC (the negative control), neural markers in v1. [C4] — full table below.

3 · What do ñāṇa interviews and kōan checking teach probe design? 5 transfers

The traditions score the structure of a debrief, never doctrinal content, and defeat scripting with unscripted fresh-angle follow-ups (Zen sassho: 20–100 checking questions per kōan precisely because the "answers" are semi-public). Five moves transfer directly: structure-not-content scoring; fresh-angle follow-up after any advanced claim; reproducibility-at-will + persistence before the slow latent moves; show-don't-tell; and under-claiming calibration (genuine attainers under-report; readers over-report). These became the 52-item probe bank, built on lane 1's 20 verification-question patterns. [C4]

4 · Difficulty markers → flag thresholds? two-tier nowcast flag

A two-tier design over Britton-schema events (phenomenology, valence, duration, impairment, practice-link as separate fields): Tier A at monthly consolidation is conjunctive (cluster AND persistence AND impact) keyed to the empirically supported persistence-risk signature — dysregulated arousal, not sadness; Tier B expedited is disjunctive on the consensus intervention triggers (uncontrollability, loss of critical attitude, impairment, suicidality). History-neutral thresholds, probe-elicited difficulty down-weighted, expected Tier-A rate ~5–15% (review band ~2–15%). Binding rules: the flag nowcasts a present state, never predicts risk (event-prediction PPV ≤ 0.01 at these base rates), and Waypoint records features but never asserts the dark-night-vs-depression differential — a human adjudicates. [C4]

5 · Where does the Wheel disagree with the strongest external work? 3 disagreements, for Surya

The Wheel's deep structure is strongly confirmed (position-vs-phase is Sufi maqām/ḥāl verbatim; per-pillar unevenness; non-linear intensifiable top; small-death milestone placement). Three disagreements go to Surya at M2: the Hawkins LOC numbers carry no epistemic weight (refuted method; the column is even non-monotonic in band 3) — keep as provenance at most; 100-point resolution overclaims — sub-band labels are phase/milestone vocabulary, not estimator targets; and the 52–56 "Low-X cascade" is a co-occurring, recurring phase cluster, not a fixed ordered sequence (scoring the order would reward map-readers). Plus one label collision worth renaming: point 43 "Dissociation" vs the clinical risk signature. Full list: dossier §4.7.

What the external reviewers actually said — and what happened to every finding

Four independent reviews (all gpt-5.6-sol at reasoning effort ultra, 2026-07-16): three pre-adjudication, then a fourth stress-testing the adjudicated fixes themselves. Quoted honestly; full texts in external/; dispositions in ADJUDICATION.md and SPEC-REVISIONS.md. The Oracle (GPT Pro) review the spec also requires is still outstanding — M1 does not close until it is attached.
Package review — verdict: NO

codex-ultra-package-review.md — "would a top psychometrician sign off?"

"No. A top psychometrician would call this an unusually thoughtful, auditable expert-system prototype — not a validated developmental-stage instrument. Its engineering is ahead of its measurement science."— Verdict section
"As written, it is not ethically deployable." … "State exactly what each score is claimed to mean and what decision it supports. Then show that independent trained raters, using a frozen operational rubric, can distinguish the proposed bands and pillars reliably."— Ethical soundness / What a top psychometrician would demand first

What it demands: the central problem is circularity — Surya defines the construct, approves the probes, supplies the likelihoods, and produces the "gold" labels, so agreement measures fidelity-to-Surya, not validity of the Wheel. Its single highest-value change: replace M4 with a hard, transparently consented shadow-mode validation phase (no stage output influencing anything, frozen scoring manual, repeated blinded + independent ratings, simple baselines first). Until independent raters exist, the honest construct name is "a Surya-aligned teaching-profile estimate." Also: assessment-aware consent, GDPR Art. 35 DPIA, stage never lowering a safety threshold, and the unpriced failure mode performative lock-in — the thermometer becoming the thermostat.

Adjudicator's verdict: the diagnosis is accepted nearly wholesale — 9 of its 10 spec-targeting findings adjudicated AGREE or PARTIAL and folded into SPEC-REVISIONS.md (freeze, rename, case-series stats, disclosure, DPIA, stage-blind safety, lock-in controls, construct honesty). The one substantive disagreement: v1 need not be a coherent joint Bayesian model before it is useful — instead codex's own simple model becomes the pre-registered baseline the rich design must beat prospectively (divergence D4/R16). File hygiene note: the review file contains duplicated text + pasted terminal artifacts (~lines 96–181); quote only from its top section.

Estimator review — needs revision

codex-ultra-estimator-review.md — the statistical deep-dive

"The design has good engineering bones — typed evidence, deterministic arithmetic, provenance, abstention, proper-score ambitions, and estimate-blind extraction — but it is not yet a coherent Bayesian state-space model. It is an elicited, tempered scorecard combining several Bayesian-looking components with posterior-dependent heuristics."— Overall assessment

Its three highest-value changes: (1) eliminate evaluation circularity (freeze before the interview, redefine M4 as a feasibility case series); (2) replace the heuristic updater with one coherent observation model (opportunity-aware emissions, episode clustering, joint phase handling, one likelihood per datum); (3) radically simplify v1 for N=5–10 (coarse static ordinal states, no dynamics/archetypes/adaptive reliability). It also caught real math errors — the silence kernel that would have advanced practitioners up-band during gaps, a boundary-breaking clamp, a false archetype-neutrality claim, wrong-direction BKT elicitation, and "evidence bits" overstating realized information ~20×.

Adjudicator's verdict: ~15 findings CONFIRMED and applied to the memo (silence-drift zeroed in gaps, truncate-and-renormalize, zero-sum archetype deltas + swap-invariance check, likelihood-direction elicitation, capacity-bits semantics, model-conditional intervals, toy-example recomputation, Hui–Walter claim withdrawn). The deepest structural critiques (opportunity model, phase self-sealing, KL caps, cross-layer reuse) are recorded as named M3 decision points with diagnostics, ablations, and fallbacks rather than silently adopted or dismissed — divergence D4 records both positions; change (1) went to the spec (R6/R9), change (3) became the M5 ablation gate (R16).

Marker red-team — incorporated

codex-ultra-marker-redteam.md — the construct-validity attack on the bank

"This bank is not yet construct-valid as a stage estimator. It currently mixes at least six different things: meditation-path development, verbal/narrative sophistication, personality and cultural style, psychotherapy or coaching exposure, compliance with LIFE's curriculum, and transient meditation states."— Bottom line (reviewing the truncated v0)

What it demanded: split the bank into separate latent families and let only validated trait/integration evidence update stage; its "ten worst markers" list led with the Neidan verification signs (suggestion-prone, map-readable) and the non-dual glimpse (highly imitable, also produced by DPDR/psychedelics/panic).

Adjudicator's verdict: VERIFIED-INCORPORATED. Bank v0.1 was rebuilt around it before adjudication — all ten worst dispositioned (1 descoped, 1 deleted, 6 reclassified covariate/curriculum, 2 downgraded), a Class/Contamination field on every marker, five new global rules (episode clustering, capacity norming, down-faking, honesty-texture caution, alternative-route provenance), and its CT-1/2/3/5 discriminator criteria adopted verbatim. Its strongest structural demand is implemented in deliberately weakened form (evidence classes + weak-tier curriculum bridge, not fully split latents) — recorded as standing divergence D6, not hidden.

Adjudication stress-test — conditional yes

gpt56-sol-ultra-adjudication-stress-test.md — the final layer: do the adjudicated fixes hold?

"The four fixes improve honesty more than validity. … The deeper category error is that Waypoint is tested as a predictor or Surya-imitation instrument but intended to operate as a controller of teaching and safety decisions."— Overall assessment
"No, not as a deployment-validation study. Yes, as a strict shadow-mode feasibility study — with one amendment: insert the blinded stage-free actionability gate before M2 and make passing it necessary to continue."— Closing verdict

Fix-by-fix: the hash-freeze is a partial resolution (a hash proves bytes didn't change, not independent assembly — freeze the whole evaluation bundle and embargo M4 dossiers from M2 elicitation); the feasibility case series is sound only if M5 loses deployment authority (strongest earnable claim: "merits a preregistered shadow-mode alpha study"); "reference ratings" is mostly relabeling without the 2×2 rating design (two dossier + two interview ratings, randomized, washout — the delta only means something against within-mode baselines); and disclosure must be use-aware, with the DPIA before the team M4 and P-040 dropped outright. Its sharpest new content is the closed loop nobody had modeled: the estimator counts its own decisions as confirmation (practice-intensity drift, adaptive probing, full-history smoothing erasing treatment effects) — hence evidence origin tags, a measurement_environment_version, contemporaneous frozen forecasts for any policy claim, and safety as an executable, property-tested invariant (stage-blind primary lane, severity = max of lanes, ≥51 flags shadow-only, alerts lead with observations, never "unbinding-adjacent").

Disposition: landed post-adjudication, so it is folded straight into the spec proposals — R17 (the blinded actionability gate as new milestone M1.5, pass required to reach M2), R18 (origin tags + measurement-environment versioning; controller actions never count as competence), R19 (frozen contemporaneous forecasts; dropout as a policy outcome), and amendments to R6, R8, R12, R13, R15, R16. Its cheapest, highest-value demand: the M1.5 gate costs ~two short Surya sessions and answers whether stage information improves teaching decisions at all — before 4–6 sessions are spent eliciting marker rows.

Coverage: stage-class markers by band × pillar

What the extractor can actually score toward a band, after the red-team reclassifications. Empty cells are declared gaps, not oversights: bands 6–7 have no per-pillar evidence by design (no source supports pillar attribution above 51 — live-teacher territory), and band-3 placement is mostly the absence of band-4 structure. Curriculum-class markers (+c) reach the band posterior only through the §1.4 bridge, capped at the weak tier.
Sensations Emotions Thoughts Awareness Cross-pillar Band 7 · 61–70 Enlightenment Declared gap — no per-pillar evidence above 51; live-teacher territory (canon §10) Declared gap — no per-pillar evidence above 51 Declared gap — no per-pillar evidence above 51 Declared gap — no per-pillar evidence above 51 M-X7-001 stabilization claims · M-X7-002 integration signature · M-X7-003 practice-motivation return 3 recognition only Band 6 · 51–60 Unbinding Declared gap — highest support need; recognition + safety only Declared gap Declared gap Declared gap M-X6-001 polarity switch (51) · M-X6-002 Low-X loss cascade (52–56, phase-cluster) · M-X6-003 non-dual loop (57) · M-X6-004 deconstruction texture (58–59) · M-X6-005 luminosity–emptiness (60) 5 band-level + safety Band 5 · 41–50 Awakening M-S5-001 whole-field flux (trait) · M-S5-002 volitional absorption, five masteries · M-S5-003 formless-sphere ordering M-E5-001 compassion surge (47) · M-E5-002 detachment with function (48) · M-E5-003 equanimity with empathy-cost (49) M-T5-001 lived anicca/anattā · M-T5-002 construct-freedom + re-authoring · M-T5-003 emptiness ladder 3–4 M-A5-001 reproducible non-dual access · M-A5-002 self-model ladder position M-X5-001 glimpse trait-ization · M-X5-002 recognition-reframe (42) · M-X5-003 early-glimpse suffering (41) · M-A5-003 small-death claim (50 boundary, milestone) 3 3 3 2 4 incl. small-death Band 4 · 31–40 core app band Stage: M-S4-001 somatic vocabulary · M-S4-005 vibration onset (state-first). Curriculum: M-S4-002 breath mechanics · M-S4-004 reverse breathing (+milestone). Descoped: M-S4-003 Neidan signs Stage: M-E4-001 live-catching · M-E4-002 somatic anchoring · M-E4-003 self-acceptance · M-E4-005 staying-with · M-E4-006 warmth (counter-indicator). Curriculum: M-E4-004 compassion targets Stage: M-T4-001 thought-as-event (weak) · M-T4-006 rumination half-life. Curriculum: M-T4-002 beliefs-as-constructs · M-T4-003 mortality contemplation · M-T4-004 values/mission · M-T4-005 visualization M-A4-001 meta-awareness · M-A4-002 noticing latency · M-A4-003 three-phase structure · M-A4-004 effortless attention† · M-A4-005 open field · M-A4-006 attention-on-attention · M-A4-007 absorption+clarity† · M-A4-008 non-dual glimpse† (state/milestone). Curriculum: M-A4-009 cascade fluency Declared gap — M-A4-009 (curriculum) and M-A4-010 (covariate) reclassified out; band-4 integration evidence lives in M-E4-005 / M-T4-006 2 +2 curriculum 5 +1 curriculum 2 +4 curriculum 8 3 state-first · +1 curr. Band 3 · 21–30 Intelligence · pre-path M-S3-001 somatic opacity — conditional (capacity-gated, rule 14). M-S3-002 reclassified covariate M-E3-001 reified emotion. M-E3-002 covariate · M-E3-003 tombstoned M-T3-001 thought–fact equivalence · M-T3-002 hindsight-only rumination (28–33 transition) M-A3-001 unnoticed mind-wandering · M-A3-002 attention as external force M-X3-001 content-collapsed reports (floor; fires only on invited reports). M-X3-002/003 covariates 1 conditional 1 2 2 1 floor marker
Stage-class markers per cell: 1 2 3 4–5 6+ declared gap
Hover any cell for its marker ids. Counts are stage-class only; "+c" markers are curriculum-class (band reach via the §1.4 bridge, weak-tier cap); † state-first markers update phase first and reach the band only via recurrence/trait forms. Source: marker-bank-v0.md §6 coverage matrix.
The bank in one breath: 91 entries — 63 band-locus markers (12 band 3 · 28 band 4 · 15 band 5 · 5 band 6 · 3 band 7) + 13 phase markers (glimpse / plateau / dip / integration / intensification) + 15 cross-band guards (M-NX: contamination, implausibility, look-alike redirects). Of the 63: 8 curriculum-class, 6 reclassified covariates (score nothing), 1 descoped (Neidan signs), 1 deleted with tombstone. Every marker carries a Class/Contamination field naming which of the five red-team confounds could mimic it, 17 global scoring rules bind the extractor, and the expanded confusion table (CT-1…CT-8) makes the five spec look-alike pairs — plus cessation near-neighbors, DPDR, and practice-meaninglessness-at-three-positions — consulted-not-scored. Guard logic and the confusion table are crown-jewel internal (narrower distribution than the Wheel itself).

52 conversational items, all DRAFT until Surya's M2 gate

Calibrated moves Wisdom weaves into natural conversation — never agreement-scored, never a discoverable correct answer; the extractor scores the shape of the reply against the marker bank. Delivery: ≤1 scored probe per session, ≤2 per week, suspended during flags and dark-night watches (support probes excepted, safety/phase capture only), never verbatim twice, never exam-like.
D1 — a good teacher's move anyway; assessment is secondary use D2 — exists mainly to elicit scorable structure D3 — verification with mild incomplete disclosure (2: foils + paraphrase pairs)
Disclosure tiers under the assessment-aware, item-blind design. Source: probe-item-bank-v0.md §8.1.

Composition by territory

TerritoryItems
Entry & floor, bands 3→4 (21–35)4
Band 4 by pillar (31–40) — the core app band16
Band 4→5 boundary & band 5 (41–50)12
Bands 6–7 (51+) — recognition & support3
Phase & confusion-pair discriminators4
Anti-contamination (foils, paraphrase, perturbation, provenance)5
Behavior–claim consistency (telemetry crosses)8

3 designated support probes (P-034, P-036, P-039) stay deliverable during flags — as care, scored for safety/phase only.

The strings attached (all deliberate)

  • The 25 D2 items are not deliverable until the spec §5.2 assessment-aware amendment lands (R4) — a deliverable cannot overrule its spec, and the bank says so on page one.
  • The foil family (P-040) is the ethical edge case: it asserts a nonexistent phenomenon, so it cannot double as a teaching move. It runs only under the D3 duties (necessity, minimal undisclosed risk, independent review, no false teaching, debrief plan) — and is dropped entirely if Surya or independent review balks (R15).
  • Deliberate non-coverage: no probe for the descoped Neidan signs (probing would manufacture the phenomena), none for passive-texture markers (unpromptedness is their evidence), none for covariates, none for band 6–7 per-pillar cells.
  • Every item is a family, not a string — paraphrase-rotated, perturbation follow-ups in follow-up space; nothing dies on disclosure except D3 timing.

Estimator at a glance: LLM proposes, the mechanical layer integrates

A discrete-grid Bayesian filter per pillar (grid 15–70), coupled through the overall wheel latent, expert-elicited emissions from the marker bank, a continuous-time anti-teleporting kernel, a fast phase layer that absorbs surprise before stage does, BKT milestone side-chains, claim-gated consistency machinery, and abstention as a first-class output. Zero fitted parameters at v1; the canonical posterior is a pure deterministic function of (event log, versioned tables, priors) — every disagreement traces to an inspectable row.
LLM PROPOSES MECHANICAL LAYER INTEGRATES — OWNS ALL ARITHMETIC Evidence dossier transcripts · probe log · telemetry (offline export, v1) Extraction LLM, family A typed marker events estimate-blind Event router deterministic · dedup keys phase | milestone | pillar | curriculum | guard | safety Bayesian layer Δt kernel (no teleporting) phase gate → pillar updates → coupling → milestones (BKT) → consistency r̂ · surprise caps Snapshot band + pillar posteriors phase · milestones cited rationale abstains below sufficiency Teacher-flag path (two-tier) reads raw events — never phase-gated, never waits for stage sufficiency Monthly consolidation (canonical) LLM merge/reconcile pass → full-history re-smoothing nowcast = filter · consolidation = smoother Emissions: every event compiles to a likelihood column over the grid from Surya-elicited natural-frequency rows (set, read, and veto every number). Absence is never evidence · repeats merge to one episode · claims are gated until corroborated · silence widens the posterior (post-review fix).
v1 runs offline over exported dossiers; runtime integration is gated on M5. Full design: estimator-design-memo.md; model-class survey in lane 3.

Anti-teleporting, fixed

The review caught that the original rates would have advanced practitioners up-band during silence and eventually sharpened the posterior. Corrected: drift is zeroed during unobserved gaps (silence now genuinely widens uncertainty); forward drift applies only over observed practice, and a >6-point monthly shift trips a mandatory-rationale audit.

Abstention is the competence claim

Per-pillar reporting requires ≥2.5 capacity-bits of trailing evidence, an 80% interval ≤2 bands, and no unresolved conflict — otherwise insufficient or conflicted (a distinct state: contradictions are never averaged into a midpoint). Guard events can never make a pillar more reportable.

κ ≈ 0.7–0.8 is the realistic bar

The Bronlet 2025 house pattern (decomposed structural rubrics, temperature 0, median-of-N) hits weighted κ≈0.78 vs expert raters on the closest published task — that, not perfection, is the v1 extractor target vs Surya. Model bumps are recalibration events gated by golden-transcript regressions (lane 7: no formal LLM measurement-invariance framework exists; ours would be first).

The one-time instrument battery — ≈55–75 minutes, calibration-only

Convergent-validity anchors for the M4 study, scored as band-conditional DIF-aware evidence against Surya's reference ratings. Never a product surface, never the estimator spine — self-report content scales are structurally unidentifiable between genuine development and a shifted response frame (the FFMQ DIF finding is the empirical backbone of the "insight vocabulary vs lived insight" confusion pair).
InstrumentWhat it anchorsWhy it earns the slot
NADA-T + NADA-SThe Wheel's central non-dual axis (trait + state)Top pick — closest validated instrument to the construct; its self-transcendence/bliss split mirrors the corpus's epistemic-vs-affective distinction (α .81–.94)
STAGES-style sentence completionLanguage-structure stage scoringThe differentiator: a second, human-scorable instantiation of Waypoint's own method — the cleanest place to measure LLM-vs-trained-scorer agreement (κ 0.7–0.8 bar)
MODTASAbsorption capacityDiscriminates the absorption↔dullness pair (CT-3)
MPE-92MPure-awareness / minimal phenomenal experience profileValidated top-of-awareness-pillar phenomenology
Hood M-scaleIntrovertive mysticismTargeted 51+ ceiling slot only — near-useless in 22–45, but the only validated vocabulary reaching territory the estimator must still recognize
FFMQ (DIF caveat baked in)Dispositional mindfulnessKept because its item-level DIF across experience levels is itself informative at calibration
Optional: one self-transcendence trait scaleTrait-level transcendenceCheap convergent redundancy if session time allows

Explicitly rejected: MAAS (floors above basic mindfulness) · MEDEQ as a stage measure (session-depth scale — mined for probe wording only) · Jeffery Martin PNSE/Finders (the Goodhart cautionary tale, not an instrument) · Hawkins LOC calibration (the negative control: single unvalidated rater, false precision, non-reproducible procedure — the anti-pattern Waypoint's whole validation protocol repudiates) · neural cessation markers in v1 (single-digit-adept, internally contradictory).

Every finding dispositioned — 74 edits, 7 recorded divergences, nothing silently resolved

One adjudication pass (2026-07-16, post-repair) processed the prior verifier findings, four fresh verifier reports, and all three external reviews against the four deliverables. Verdict vocabulary: CONFIRMED (fix applied) · PARTIAL (applied modified — divergence logged) · REJECTED (reason logged) · VERIFIED-NO-ACTION · ROUTED (spec/study-design → SPEC-REVISIONS). The spec itself was never touched. Full log: ADJUDICATION.md; raw findings: verifier-findings-raw.md.
74
edit operations applied
19 · 12 · 13 · 30
dossier · markers · probes · memo
7
divergences recorded (D1–D7)
10
findings routed to the spec

The fixes that mattered most

The seven standing divergences (both positions kept)

IDWhat stays contestedResolution path
D1Flag-rate bands: ~5–15% expectation vs ~2–15% review-trigger — deliberate two-band design, not an inconsistencySurya may unify at M2
D2Toy milestone prior 0.196 (pre-event state) vs 0.115 (strict operation ordering)Renumber if M3 adopts joint factors
D3Support probes during flags: unscore vs amend the spec — both appliedM2 amendment (R4)
D4Codex's re-architecture demands vs the shipped v1 design — kept with honest labels, named diagnostics, and codex's simple model as the pre-registered baselineM3 decision points; M5 ablation gate (R16)
D5Identifiability ridge fix: zero-sum applied to archetype deltas only; the global constraint is the pillar→overall aggregation questionSurya at M2 (dossier §4.7-5)
D6Red-team's full latent-family split vs the bank's class-system implementationDeclared limit; construct-validity work lives in spec §7
D7§1.4 restoration path: restore-in-bank AND align memo (not either/or)Done; noted for provenance

19 proposed revisions — quoted, replaced, and justified in SPEC-REVISIONS.md

Proposed only; the spec is untouched. R1–R16 cover all 10 adjudicated spec-level inputs plus the deltas the deliverables themselves requested; R17–R19 come from the post-adjudication stress-test, which also amended R6, R8, R12, R13, R15, and R16 (marked [ST] below). The pattern across all 19: keep the design, shrink the claims — every revision removes an overclaim, adds a protection, or converts a disputed design choice into a falsifiable gate.
IDSpec §Proposed change (one line)
R1§2"Surya gold labels" → "Surya reference ratings" everywhere — agreement measures fidelity-to-Surya, not truth
R2§2"Information ceiling" → within-rater mode discrepancy; the delta stays the headline diagnostic
R3§4Intervals labeled model-conditional until M5 coverage evidence; point demoted to display derivative; conflicted + per-pillar sufficiency; param_version; P(reached)
R4§5Fully-covert probes → assessment-aware, item-blind disclosure at enrollment + the support-probe carve-out (delivery as care, safety/phase scoring only)
R5§6Peak states move the slow latent only when volitionally reproducible + persistent (the traditions' one missing criterion)
R6 [ST]§7Freeze the entire evaluation bundle before the interview (dossier-construction rules, probe policy, event log, code/tables, analysis plan + forecast); interview transcript quarantined; M4 dossiers embargoed from M2 elicitation; post-M4 revisions are v2, validated on genuinely new practitioners
R7§7Headline-number wording updated at the protocol site (delta, not ceiling)
R8 [ST]§7Single-rater mitigations: the 2×2 rating design (two dossier + two interview ratings, randomized, washout; delta reported against within-mode baselines), anchor vignettes (drift-detection only), probability vectors elicited at rating time, second rater on a subset when one exists — until then the claim is "replication of Surya's reference ratings"
R9§7Metrics reframed as a feasibility case series honest at N=5–10: per-case RPS + signed error + complete displays; coverage descriptive with exact binomials; κ descriptive-only; practitioner = resampling unit
R10§7Construct honesty: what M4–M5 validates is a "Surya-aligned teaching-profile estimate," scoped to map-exposed cohort members in the observed range
R11§7Synthetic QA labeled scenario tests + arithmetic checks (SBC-style) — never coverage evidence
R12 [ST]§8Stage never downgrades safety as an executable invariant: primary classifier stage-blind, stage-aware lane separate, severity = max of the two, never suppress/downgrade/delay a flag or swap clinical language for contemplative framing, property-tested under wrong-stage injections; ≥51 flag shadow-only; outreach on raw features; alerts lead with observations, never "unbinding-adjacent"
R13 [ST]§8Use-aware consent (profile inference + its uses disclosed) and DPIA before the team M4 study, not just alpha; power-asymmetry protections (independent stewardship, no manager access); notice / opt-out / correction / deletion; passive vs disclosed-probe evidence analyzed separately
R14§11New risk row: performative lock-in (thermometer→thermostat) — v1 named as de facto shadow mode; randomized audit probes + no-profile counterfactuals before runtime
R15 [ST]§11The foil family (P-040) is dropped outright — it can prime reports, contaminate the evidence stream, and damage trust; paraphrase pairs + perturbations carry the anti-contamination load
R16 [ST]§12M5 keeps the pre-registered simplicity ablation and loses deployment authority: strongest earnable claim is "merits a preregistered shadow-mode alpha study"; runtime adaptation, stage-aware safety, and flag activation are decided only by the later estimate-visibility policy trial
R17 [ST]§12New milestone M1.5 — blinded stage-free actionability gate, pass required to continue: on ~12–20 frozen cases, Wisdom recommendations under no-estimate / reference / wrong-estimate views, ranked blind by Surya; redundant latent, brittle policy, or unsafe escalation each stop the program cheaply
R18 [ST]§4–6Evidence origin tags (spontaneous / standardized_anchor / policy_elicited / post_teaching / post_outreach) + measurement_environment_version (Wisdom prompt/model, curriculum, probe scheduler, safety classifier, flag policy); assigned curriculum and controller-driven cadence never count as competence; outreach recorded as an intervention
R19 [ST]§7Policy/causal evaluation scores contemporaneous frozen forecasts, never smoothed history (smoothing can rewrite an intervention as pre-existing state); dropout tracked as a policy outcome and possible harm, not missing data

Deliberately not spec edits: the Wheel-model items Surya owns at M2 (Hawkins numbers, 52–56 as phase cluster, the point-43 "Dissociation" rename, pillar→overall aggregation, the §9-A stage-numbering ambiguity that displaces the reverse-breathing gate by a full band) and the estimator re-architecture candidates parked as M3 decision points (D4).

How this was produced, and where everything lives

Run provenance

Main research pass: a ~30-agent run (corpus digests, 7 lanes, 4 deliverable authors, finder→verifier adversarial checks on every empirical claim). Two agents died mid-stream and a session limit stalled the pass — a same-day repair pass completed the marker bank (§§2.3–7 had truncated at M-A4-010) and the missing probe bank, then re-ran four fresh verifiers over everything. The adjudication pass applied 74 edits across the four deliverables, with provenance markers at every edit site. This report itself was delayed: two earlier report attempts died on transient API 529 errors before writing anything; this is the third attempt. A fourth external review — the adjudication stress-test — landed after the first build of this report and is folded into the verdicts section and SPEC-REVISIONS (R17–R19 + six amendments). Still outstanding: the Oracle (GPT Pro) review that spec §10 requires attached to the estimator memo — the memo carries status pending-reviews-complete, so M1 does not close until it lands (plus sign-off on the spec revisions).

Deliverables (post-adjudication)

External reviews (codex gpt-5.6-sol, effort ultra)

Research lanes

Corpus notes (seed papers + canon)

Library & inputs