# Waypoint Spec — Proposed Revisions (M1 close)

**Status:** PROPOSED ONLY — the spec (`../waypoint-practitioner-stage-assessment.md`) has **not** been edited. Each item quotes the current text verbatim, proposes a replacement, and gives a one-line rationale, for Fionn (and Surya where flagged) to accept, modify, or reject at M1 close.
**Date:** 2026-07-16.
**Sources:** `deliverables/ADJUDICATION.md` §5 SPEC-LEVEL INPUTS (all 10 routed items covered); estimator memo §14 spec deltas; probe bank §8.3 asks; literature dossier §1.2 / §3.1 / §4.7; the four external reviews in `external/` — three pre-adjudication plus **[ST]** `external/gpt56-sol-ultra-adjudication-stress-test.md` (GPT-5.6 Sol, effort ultra), a post-adjudication stress-test of the fixes themselves. [ST]'s demands are folded in as **R17–R19** and as amendments to **R6 / R8 / R12 / R13 / R15 / R16** (each tagged [ST] in its provenance).
**Convention:** revisions are ordered by spec section. "Provenance" names where the finding came from and the adjudicator's verdict on it.

---

## §2 — Decisions table

### R1 — Rename "gold labels" to "Surya reference ratings" (decision #3)

**Current (§2, decision #3):**
> | 3 | Ground truth | Triangulated calibration | Surya gold labels + one-time instrument battery + longitudinal behavior; divergences recorded, not overridden |

**Proposed:**
> | 3 | Ground truth | Triangulated calibration | **Surya reference ratings** (structured interview + blinded dossier-only rating) + one-time instrument battery + longitudinal behavior; divergences recorded, not overridden |

**Rationale:** Surya authors the construct, the parameters, and the criterion — agreement measures fidelity-to-Surya, not independent truth; "gold" overclaims and the rename is costless honesty. Apply the same rename everywhere "gold label" appears (§7, §12 M4).
**Provenance:** codex package review (Construct validity); ADJUDICATION §5.3 — second opinion AGREE.

### R2 — "Information ceiling" becomes "within-rater mode discrepancy" (decision #7)

**Current (§2, decision #7):**
> | 7 | Gold-label protocol | Surya structured interview AND blinded dossier-only rating per subject | The interview-vs-dossier delta measures the information ceiling of any data-only estimator |

**Proposed:**
> | 7 | Reference-rating protocol | Surya structured interview AND blinded dossier-only rating per subject | The interview-vs-dossier delta is the **headline v1 diagnostic** of what report-only data supports for this rater — a within-rater mode discrepancy, not a mathematical ceiling |

**Rationale:** both ratings share one rater, one ontology, and one memory, so the delta mixes information access with intra-rater noise and mode effects — an estimator can even beat the dossier rating through regularization; keep the measurement, drop the "ceiling" claim (also in §6's LLM bullet and §7 step 3 — see R7).
**Provenance:** codex estimator review §5; ADJUDICATION §5.3 — PARTIAL (keep delta as headline diagnostic, drop ceiling language).

---

## §4 — Output object

### R3 — Honest interval labeling, point demoted to display, richer `data_sufficiency`, provenance fields

**Current (§4 snapshot schema, three lines):**
> `overall: { band_posterior: p[10 bands], point: 1–100, credible_interval },`
> `milestones: [ { marker_id, e.g. reverse_breathing | first_glimpse | small_death, status: claimed|corroborated, evidence_refs[], confidence } ],`
> `data_sufficiency: ok | insufficient   // below threshold → abstain, never guess`

**Proposed:**
> `overall: { band_posterior: p[10 bands],          // canonical claim`
> `           point: 1–100,                          // display derivative of the posterior mean — never a standalone claim`
> `           credible_interval },                   // model-conditional (heuristic) until M5 coverage evidence; labeled so wherever surfaced`
> `milestones: [ { marker_id, e.g. reverse_breathing | first_glimpse | small_death,`
> `                status: claimed|corroborated, P_reached, evidence_refs[], confidence } ],`
> `data_sufficiency: ok | insufficient | conflicted  // per-pillar AND overall; conflicted ≠ insufficient — contradictions are never averaged into a midpoint`
> `param_version: { marker_bank_hash, prior_table_version, extractor_model_id },  // reproducibility contract`
> `credible_set: reserved                            // M6+ conformal sets (needs ≳20–30 labeled dossiers)`

**Rationale:** the nominal 80% interval is conditional on hundreds of elicited numbers being right and has no demonstrated coverage until M5 — label it so; band+interval is the canonical resolution (sub-band point claims are unsupported by language evidence, dossier §3.1); `conflicted`, per-pillar sufficiency, `P(reached)`, and `param_version` are the estimator memo's §14 deltas 1–5 folded in.
**Provenance:** ADJUDICATION §5.5 + §5.6a (AGREE/PARTIAL — keep `point` as display derivative); estimator memo §14.

---

## §5 — Evidence channels (probes)

### R4 — Assessment-aware, item-blind disclosure replaces fully-covert delivery; support-probe carve-out

**Current (§2, decision #4):**
> | 4 | Evidence channels | Passive conversation + active conversational probes + practice telemetry | No standalone questionnaire in-product; probes are woven invisibly into Wisdom conversation |

**Current (§5, channel 2):**
> Delivery constraints: invisible (never framed as assessment), frequency-capped, never during flagged-vulnerable moments, never repeated verbatim, logged with item-id for scoring.

**Proposed (decision #4):**
> | 4 | Evidence channels | Passive conversation + active conversational probes + practice telemetry | No standalone questionnaire in-product; probes are **assessment-aware, item-blind**: disclosed as a program at enrollment, never exam-like in conversation |

**Proposed (§5, channel 2):**
> Delivery constraints — **assessment-aware, item-blind**: at enrollment, practitioners are told in plain language that (a) their conversations and practice data are used to estimate where they are on the path so teaching can adapt, (b) some of Wisdom's conversational questions are deliberately designed for that purpose, (c) who sees the result (Wisdom adaptation + teacher triage; never shown in-app, never a public score), and (d) they can opt out without service penalty. Item identities, timing, and scoring keys stay hidden (item-blind); the existence and purpose of assessment does not. In conversation, probes are never exam-like — the surviving sense of "invisible." Frequency-capped; never repeated verbatim; logged with item-id for scoring. **During flagged-vulnerable moments and active dip/dark-night watches, all probing is suspended except the designated support probes, whose function is care and triage; their responses route to safety/phase feature capture only — never band evidence.**

**Rationale:** "invisible as a program" inside a trust-based teaching relationship is indefensible incomplete disclosure (Belmont conditions); assessment-aware item-blind is the defensible floor, and the probe bank's 25-item D2 tier is explicitly undeliverable until this amendment lands; the support-probe carve-out authorizes the care-delivery the unqualified "never during flagged-vulnerable moments" forbids, while keeping the band-scoring ban absolute (adjudication divergence D3 applied both halves).
**Provenance:** codex package review (Ethical soundness); ADJUDICATION §5.4 — AGREE; probe bank DISCLOSURE DESIGN note + §8.3.

---

## §6 — Estimator architecture

### R5 — Volitional-reproducibility criterion added to the error-routing rule

**Current (§6, substrate bullet):**
> Prediction errors spend themselves on fast state first and revise slow latents only through persistent runs of error (hierarchical-Gaussian-filter style) — a practitioner does not teleport from 30 to 55 in a week.

**Proposed:**
> Prediction errors spend themselves on fast state first and revise slow latents only through persistent runs of error (hierarchical-Gaussian-filter style) — a practitioner does not teleport from 30 to 55 in a week. **Peak-state evidence advances the slow latent only when volitionally reproducible and persistent (Sufi maqām/ḥāl; Pa-Auk five masteries); otherwise it annotates phase.**

**Rationale:** the one addition the traditions insist on that §6 lacks — state ≠ station is stated almost verbatim by four traditions and the state/trait literature ([C4]); a verifier confirmed the criterion is currently *not* in the spec despite the dossier initially citing it as if it were.
**Provenance:** literature dossier §1.2 (recommended amendment, verifier fix); lane 1 §7.

---

## §7 — Calibration & validation protocol

### R6 — Hash-freeze the dossier, forecast, and tables BEFORE the interview; quarantine the interview transcript

**Current (§7, per-subject step 2):**
> 2. **Surya structured interview** → gold label: overall band + pillar ratings + milestone judgments + rationale. Interview protocol is a research-phase deliverable (derived from wheel markers; recorded and itself ingested as evidence).

**Proposed (new step between 1 and 2, plus step-2 rewording) — [ST]-amended to cover the whole evaluation bundle:**
> 1b. **Freeze before the interview — the entire evaluation bundle, not just the dossier:** hash-freeze, before the interview happens: the raw-data cutoff and dossier-construction rules; the probe-selection policy and every delivered probe; the extracted + merged event log; prompt, extractor-model, code, and parameter-table versions; missing-data and exclusion rules; and the estimator's forecast together with the analysis plan and pass/fail criteria. The interview transcript enters only a separately-labeled **post-evaluation condition** — never the evaluated dossier. **M4 dossiers are embargoed from M2 elicitation** (Surya must not calibrate marker rows while looking at the cohort dossiers he will later rate). Any post-M4 revision of tables or bank is **v2**, validated on **genuinely new practitioners** — not merely future windows of the same recognizable people. Stated honestly: on a five-person team, "blinded dossier rating" means masked to estimator output, not to identity.
> 2. **Surya structured interview** → Surya reference rating: overall band + pillar ratings + milestone judgments + rationale. Interview protocol is a research-phase deliverable (derived from wheel markers; recorded — the transcript joins the post-evaluation condition only, per 1b).

**Rationale:** as written, §7 ingests the criterion interview as estimator evidence and lets post-M4 revisions train on the validation set — and a hash proves bytes didn't change, not that the frozen object was assembled independently (posterior-aware dossier construction and subject knowledge entering M2 elicitation are leakage paths a dossier-only freeze misses).
**Provenance:** codex estimator review §5 + "three changes" #1; package review fixes 6–8; ADJUDICATION §5.1 — AGREE, adopt fully; [ST] fix-(a) verdict ("partial resolution" → bundle freeze + M2 embargo + new-people rule).

### R7 — Headline number rewording (delta, not ceiling)

**Current (§7, step 3):**
> 3. **Surya blinded dossier-only rating** of the same subject (order counterbalanced, interview and dossier ratings separated in time). The **interview-vs-dossier delta is the headline v1 number**: the information ceiling of any data-only estimator.

**Proposed:**
> 3. **Surya blinded dossier-only rating** of the same subject (order counterbalanced, interview and dossier ratings separated in time). The **interview-vs-dossier delta is the headline v1 diagnostic**: a within-rater mode discrepancy that bounds what report-only evidence supports *for this rater* — not a mathematical ceiling.

**Rationale:** same finding as R2, applied at the protocol site (and to §6's "interview-vs-dossier ceiling is the headline v1 number" clause and §12 M5's "information ceiling" wording).
**Provenance:** ADJUDICATION §5.3 — PARTIAL.

### R8 — Single-rater mitigations: the 2×2 rating design, anchor vignettes, elicited probability vectors, second rater on subset

**Current (§7, known validity threats):**
> **Known validity threats, tracked from day one:** N is tiny; single-rater anchor (Surya) with no inter-rater check in v1; …

**Proposed (add to the per-subject protocol, and extend the threats line) — [ST]-amended to the 2×2 design:**
> Per-rater protocol additions: (i) **the 2×2 rating design** — per subject, **two dossier ratings and two interview ratings**, randomized in order and separated by washout; report **dossier–dossier disagreement, interview–interview disagreement, and the interview–dossier delta relative to those two within-mode baselines** (an unanchored delta conflates channel information with memory, order, interview performance, subject change, and ordinary rater noise); (ii) **duplicated anchor vignettes** hidden in rating batches — for intra-rater drift detection only (vignettes authored from the same Wheel standardize Surya's scale use without testing external correspondence; never cite them as validity evidence); (iii) ratings elicited as **probability vectors over bands at rating time** (replacing post-hoc adjacent-band smearing, which constructs a blur that flatters diffuse forecasts); (iv) a **second qualified rater on a subset** the moment one exists — until then, the valid claim is strictly **"replication of Surya's reference ratings."** Threats line: append "single-rater mitigations per protocol (2×2 repeated ratings, anchors, elicited vectors); residual single-rater bias is not statistically identifiable from these data and is carried as a standing limit."

**Rationale:** with one rater, only repeated measures make rater noise visible at all, and the interview-vs-dossier delta is uninterpretable without within-mode baselines to compare it against; Park et al. 2024 (lane 7) shows normalizing against the rater's own consistency is what makes such numbers meaningful.
**Provenance:** codex estimator review §5; ADJUDICATION §5.3 — AGREE (adopt list); lane 7 §2; [ST] fix-(c) verdict (2×2 design; anchor-vignette lock-in caveat; "replication of Surya's reference ratings").

### R9 — Metrics reframed as a feasibility case series honest at N=5–10

**Current (§7, metrics):**
> **Metrics:** estimator-vs-interview agreement (band-level accuracy, rank correlation, pillar-level agreement), uncertainty calibration (do 80% intervals contain gold 80% of the time), dossier-vs-interview ceiling, longitudinal movement plausibility. Human-vs-estimator divergences are recorded as **open questions with both positions** (house label-divergence convention), not silently overridden.

**Proposed:**
> **Metrics (M4 is a feasibility/reliability case series at N=5–10; the practitioner — not pillars, events, or rounds — is the resampling unit):** per-case Ranked Probability Score, signed band error, and complete per-practitioner case displays; band hit / adjacent-band hit as the legible pair; **coverage counts reported descriptively with exact binomial intervals** (8/10 coverage spans ~44–97% — never claimed as established calibration); weighted κ may be **shown as a descriptive** alongside the case displays, never as an inferential claim; no 3-bin reliability or ECE analyses at this N; interview-vs-dossier delta (per R7); longitudinal movement plausibility; **bits saved vs the tenure base-rate null and the persistence null at H1 horizons** as the continuous, no-reference-rating channel. Human-vs-estimator divergences are recorded as **open questions with both positions** (house label-divergence convention), not silently overridden.

**Rationale:** the current metric claims are arithmetic impossibilities at this N — pooled pillar×round events share practitioner, dossier, rater, and extractor and are not independent calibration cases; the statistics here are arithmetic, not opinion.
**Provenance:** codex estimator review §6 + package review; ADJUDICATION §5.2 — AGREE (with the κ-as-descriptive softening); estimator memo §10.3 already implements this framing.

### R10 — Construct honesty: scope what M4–M5 validates

**Current (§7, metrics paragraph — addition, no deletion):**
> *(append)*

**Proposed (append to §7 metrics):**
> Until a second qualified rater exists, what M4–M5 validates is a **Surya-aligned teaching-profile estimate** — fidelity to Surya's reading of the Wheel, not independent construct validity of the Wheel itself (§6's constructs-on-trial stance already concedes this; this line says it where the metrics are read). Every M5 conclusion is scoped to **map-exposed cohort members in the observed band range** — it validates nothing about naive-column scoring, high bands, rare milestones, or high-stage safety flags.

**Rationale:** names the honest construct so M5 numbers cannot be over-read; the scoping sentence is the external review's identifiability point made operational.
**Provenance:** codex package review (Construct validity / "Surya-aligned teaching-profile estimate"); ADJUDICATION §5.6d — AGREE; codex estimator review §2.

### R11 — Synthetic QA wording: scenario tests + arithmetic checks, never coverage evidence

**Current (§7, engineering QA):**
> **Engineering QA before humans (default, strike if unwanted):** synthetic-dossier discrimination test — personas authored at known wheel positions from the marker bank; the estimator must separate and rank them. QA of the instrument only; never calibration data.

**Proposed:**
> **Engineering QA before humans (default, strike if unwanted):** synthetic-dossier discrimination test — personas authored at known wheel positions from the marker bank (by a third model family); the estimator must separate and rank them, and confident wrong calls on the confusion-pair personas fail the gate. QA of the instrument only; never calibration data. **Synthetic coverage checks test the integrator's arithmetic under the bank's own assumptions (simulation-based-calibration style); they demonstrate nothing about real-world coverage — hand-authored personas are scenario tests, not calibration samples.**

**Rationale:** prevents M3's synthetic 80%-coverage pass from being read as calibration evidence; aligns the spec with the memo §10.5(iii) wording the external review forced.
**Provenance:** codex estimator review §6 (Talts et al. SBC); estimator memo §10.5.

---

## §8 — Safety & ethics

### R12 — Stage never downgrades safety — as an executable, tested invariant

**Current (§8, stage-aware safety bullet):**
> **Stage-aware safety.** Waypoint context (band + phase) feeds the constitutional classifier's context so thresholds are stage-aware (it already knows practice stage; this sharpens it — dark-night-adjacent phenomenology around Unbinding reads differently from the same words at 25%).

**Proposed (replace the bullet) — [ST]-amended from a policy sentence to an executable invariant:**
> **Stage-aware safety — one-directional by construction.** (i) The **primary safety classifier runs stage-blind**; (ii) any stage-aware support lane runs **separately**; (iii) **operational severity is the maximum (union) of the two lanes**; (iv) adding stage context may never suppress a flag, lower severity, delay response, or replace clinical-language guidance with a contemplative explanation — "reads differently at Unbinding" is diagnostic overshadowing when it points down; (v) the invariant is **property-tested under deliberately wrong stage and phase injections** before any flag activates. The posterior-mass-≥51 teacher flag is **shadow-only in v1**; human outreach is triggered only by raw features — impairment, uncontrollability, persistence, severe sleep loss, suicidality, rapid functional change, or user request — and teacher-facing alerts **lead with those observed features in the practitioner's own words, never with "unbinding-adjacent" framing** (automation/framing bias). Before teacher flags ever activate, the flag path carries a named operational duty: ownership, acknowledgement/response times, capacity limits, escalation and closure rules, and audit of missed or late responses — a flag without a care path creates an expectation of care the system cannot meet. Outreach is recorded as an **intervention** (see R18) and never updates stage.

**Rationale:** the raw ungated marker path already prevents phase-gating suppression, but nothing prevented the classifier or a teacher from *interpreting* danger as "advanced territory" — only lane separation, max-severity composition, and wrong-stage property tests make "stage never downgrades safety" executable rather than aspirational.
**Provenance:** codex package review (Teacher flags and safety); ADJUDICATION §5.6b — AGREE, PARTIAL on the ≥51 trigger; [ST] §4 (executable invariant, shadow-only ≥51 flag, raw-feature outreach, observation-first alerts, operational duty).

### R13 — Use-aware consent; DPIA before the team M4 study; power-asymmetry and subject protections

**Current (§8, consent & governance bullet):**
> **Consent & governance.** Team cohort: informed consent — dossiers expose inner life to Fionn + Surya; a team member can decline or exit with dossier deletion. Alpha users: explicit enrollment consent before any Waypoint processing. KB/data-governance rules apply to any derived content.

**Proposed — [ST]-amended (DPIA moved before team M4; consent upgraded to use-aware):**
> **Consent & governance.** Consent is **use-aware as well as assessment-aware**: participants are told the system infers a stage-like spiritual/psychological profile and may use it for practice selection, pacing, language register, triage, and human outreach — item-blindness protects test integrity, never consequential uses. A formal **DPIA (GDPR Art. 35) is completed before the team M4 study, not merely before external alpha** — Waypoint infers religious/philosophical and mental-health-adjacent states, and the workplace power asymmetry exists now. Team cohort: informed consent — dossiers expose inner life to Fionn + Surya; declining is penalty-free; exit comes with dossier deletion; **plus power-asymmetry protections: independent data stewardship, no manager access to individual profiles, no employment consequences from participation or estimates.** Alpha users: explicit enrollment consent carrying the §5 disclosure language before any Waypoint processing. Standing subject protections: notice, opt-out without service penalty, correction/annotation of disputed factual evidence, human review, access and deletion of derived records. **Disclosure changes the measurement condition:** marker/probe behavior calibrated under "invisible assessment" is not assumed valid under assessment-aware participation — passive and disclosed-probe evidence are analyzed separately (via the R18 origin tags). KB/data-governance rules apply to any derived content.

**Rationale:** hidden profiling of inner life without these protections is not deployable; a DPIA is a review process, not a safeguard, so it must precede the first real processing (the team study), and consent that discloses assessment but not its uses is still incomplete.
**Provenance:** codex package review (Governance); ADJUDICATION §5.4 — AGREE; [ST] fix-(d) verdict (use-aware consent; DPIA before team M4; measurement-condition change).

---

## §11 — Risks

### R14 — New risk row: performative lock-in (thermometer becomes thermostat)

**Current (§11 — addition, no deletion):**
> *(no current row; nearest neighbors are the "Probe reactivity" and "Goodhart" rows)*

**Proposed (add row):**
> | Performative lock-in (adaptation manufactures the trajectory it later cites: low estimate → simpler teaching → fewer chances to show more; high estimate → advanced framing → confirming evidence) | v1's offline/advisory posture is **named as de facto shadow mode**; before runtime integration: randomized audit probes, periodic no-profile counterfactual sessions, and adaptation evaluated as its own intervention — gated at M5 (see R16) |

**Rationale:** more serious than ordinary probe reactivity and currently unpriced — hiddenness makes the feedback loop harder to catch, and it can corrupt the longitudinal record Waypoint later validates against.
**Provenance:** codex package review (Failure mode not priced in); ADJUDICATION §5.6c — AGREE.

### R15 — Probe-reactivity row: drop the foil family

**Current (§11, probe-reactivity row):**
> | Probe reactivity (measurement changes practice) | Frequency caps; probes double as legitimate teaching moves |

**Proposed — [ST]-amended from conditional-drop to drop-now:**
> | Probe reactivity (measurement changes practice) | Frequency caps; probes double as legitimate teaching moves — **except the foil family (P-040 / M-NX-003), which cannot, and is dropped**: a foil that asserts a nonexistent phenomenon to a student can prime later reports, contaminate the evidence stream, and damage trust, and it is not load-bearing enough to justify deception, debriefing complexity, or construct-owner discretion. Paraphrase pairs + perturbations carry the anti-contamination load; Paulhus-style credibility weighting is forfeited as a nice-to-have |

**Rationale:** the one probe move the "teaching moves" mitigation cannot cover; the adjudicator's position was already stricter than the deliverable's (drop if review balks), and the stress-test removes the condition — the cost/benefit never favors deception here.
**Provenance:** fresh verifier 2 finding 5 (M-NX-003 tension); ADJUDICATION §5.4 (foils paragraph); [ST] fix-(d) ("P-040 should be removed now").

---

## §12 — Milestones

### R16 — M5 gate: feasibility report + simplicity ablation, and NO deployment authority

**Current (§12, M5):**
> - **M5** — calibration report: agreement, uncertainty calibration, information ceiling. Go/no-go on runtime integration + teacher-flag activation.

**Proposed — [ST]-amended (M5 loses deployment authority):**
> - **M5** — feasibility report: case-series agreement + descriptive coverage (per §7/R9), the 2×2 delta analysis (per R7/R8), and the **pre-registered simplicity ablation: the full estimator must beat a coarse static ordinal scorecard (no stage dynamics, no archetype effects, no adaptive claim reliability) on locked prospective H1 prediction — if it cannot, v2 adopts the simple model.** **M5 carries no deployment authority.** The strongest conclusion a passing M5 can support is: *"the pipeline is executable, auditable, produces interpretable disagreements, and merits a preregistered shadow-mode alpha study."* Runtime adaptation, stage-aware safety decisions, and teacher-flag activation are decided only by that later prospective study — a preregistered **estimate-visibility policy trial** (Wisdom randomized to receive identical proximal facts with vs without the Waypoint posterior; primary safety lane stage-blind and identical in both arms; outcomes not selected by Waypoint: practitioner-rated fit, goal progress, functioning/distress, adverse events, blinded teacher correction) — under the R14 lock-in controls.

**Rationale:** a feasibility case series cannot validate coverage, transportability, or a safety-relevant flag — softening the statistics while leaving the consequential go/no-go unchanged would be relabeling, not rescoping; the ablation gate stays (the adjudication's D4 position), and deployment questions move to the one experiment that can answer them.
**Provenance:** codex estimator review §7 + "three changes" #3; ADJUDICATION §5.5 and divergence D4; [ST] fix-(b) verdict + §2 (prescriptive validity under intervention; estimate-visibility trial design).

---

## Post-adjudication additions from the stress-test (R17–R19)

*These three arrived with `external/gpt56-sol-ultra-adjudication-stress-test.md` [ST] after the R1–R16 set was drafted; they are new spec matter, not amendments.*

### R17 — New pre-M2 milestone: the blinded stage-free actionability gate (pass required to continue)

**Current (§12 — addition between M1 and M2; no deletion):**
> *(no current milestone; M2 is currently "Surya calibrates marker bank + probe items; interview protocol agreed")*

**Proposed (insert as M1.5):**
> - **M1.5 — blinded actionability gate (pass required to proceed to M2).** On ~12–20 frozen real or independently authored cases, generate Wisdom teaching recommendations under three randomized views: (1) raw case facts with no Wheel estimate; (2) the same facts plus the Surya reference distribution; (3) the same facts plus a plausible-but-wrong or ±1-band-shifted distribution. Surya ranks the recommendations for teaching fit and safety **blind to condition**, with disguised subsets repeated after washout; stop/coarsen rules pre-registered. Three cheap answers before any calibration effort is spent: if adding stage does not improve recommendations, the latent is redundant (defer the Wheel latent, keep phase/pillars/safety features); if a one-band perturbation causes major curriculum changes, the policy is too brittle for current uncertainty (coarsen); if a wrong estimate produces unsafe escalation, runtime use is disqualified regardless of agreement statistics. Cost: a prompt harness plus ~two short Surya sessions — against the 4–6 sessions M2 budgets for eliciting ~60 marker rows.

**Rationale:** the whole program assumes stage information *improves teaching actions* over the proximal facts Wisdom already has (prescriptive validity), and nothing in M2–M5 tests that premise — this gate checks it before the expensive elicitation, and a fail is a cheap, dignified exit that keeps the useful components.
**Provenance:** [ST] §3 (single highest-expected-value change) + §5 closing verdict ("insert the blinded stage-free actionability gate before M2 and make passing it necessary to continue").

### R18 — Evidence origin tags + `measurement_environment_version`; controller actions never count as competence

**Current (§4 snapshot schema / §5 evidence channels / §6 evidence log — addition, no deletion):**
> *(current schema carries `param_version` per R3; events carry no origin field; §6's evidence log does not distinguish policy-elicited from spontaneous evidence)*

**Proposed:**
> Every evidence event carries an **origin tag**: `spontaneous | standardized_anchor | policy_elicited | post_teaching | post_outreach`. The probe scheduler maintains **fixed policy-independent anchor items plus logged random exploration** (entropy-driven selection alone is exploitation without falsification — the posterior decides what becomes observable, then scores itself on it). Snapshots carry a **`measurement_environment_version`** alongside `param_version`, covering: Wisdom prompt/model version, curriculum version, probe-scheduler policy, safety-classifier version, flag policy, and outreach history — evidence generated under materially different regimes is never silently pooled as exchangeable. Two hard rules: **assigned or unlocked curriculum, and controller-influenced practice cadence, never count as competence** (otherwise the §6 practice-intensity multiplier lets a high estimate raise practice, which mechanically advances the next prior — the estimator counting its own decisions as confirmation); and **outreach is recorded as an intervention** — it changes subsequent evidence and never updates stage or enters a later reference-rating replicate without explicit separation.

**Rationale:** at runtime Waypoint is a controller, not a passive observer — the evidence model must record which observations its own actions produced, or the closed loop (estimate → teaching → practice/reporting → estimate) manufactures its own confirmation and the false-high/false-low attractors go undetected.
**Provenance:** [ST] §4 (controller-produced evidence; adaptive-probing MNAR; co-adaptation/measurement-environment drift; outreach-as-intervention).

### R19 — Policy and causal evaluation use contemporaneous frozen forecasts, never smoothed history

**Current (§7 metrics / §6 consolidation — addition, no deletion):**
> *(§6's monthly consolidation re-smooths the full history; §7 does not distinguish smoothed states from frozen forecasts)*

**Proposed (add to §7):**
> Any policy or causal evaluation — H1 skill claims, adaptation-benefit analyses, the M5 ablation, and the eventual estimate-visibility trial — scores **contemporaneous forecasts frozen at the time they were made**, never retrospectively smoothed states: full-history re-smoothing can rewrite an intervention's genuine effect as pre-existing latent state ("they were already high at month one"), making any teaching policy look prescient. Smoothed trajectories remain the canonical *descriptive* record; frozen forecasts are the only *evaluative* record. **Dropout is tracked as a policy outcome and possible harm** — under a false-low estimate, disengagement-then-silence reads as abstention when it may be evidence the teaching policy failed; churn is reported alongside the accuracy metrics, not filed as missing data.

**Rationale:** the smoother's job (best hindsight estimate) and the evaluator's job (was the forecast/policy good at the time) are different questions on the same log — conflating them erases treatment effects and hides policy harms.
**Provenance:** [ST] §4 (full-history smoothing can erase treatment effects; false-low attractor/dropout).

---

## Not proposed as spec edits (routed elsewhere)

- **Wheel/canon items — Surya owns these at M2** (dossier §4.7 consolidated list): Hawkins LOC numbers (keep as provenance, drop as ordinal check, decide on retention); 52–56 "Low-X cascade" re-framed as a co-occurring phase cluster, not a fixed sequence; the point-43 "Dissociation" label collision with the clinical risk signature (possible rename); the pillar→overall aggregation rule; the canon §9-A stage-numbering ambiguity that displaces the reverse-breathing gate by a full band; devotional intensity (46) confirmed phase-not-position. These change the Wheel model, not the spec.
- **Estimator re-architecture items (opportunity model, joint phase update, contamination mixtures, joint (θ_o, δ) posterior)** — kept as named M3 decision points with diagnostics, ablations, and fallbacks in the estimator memo §14.1; adjudication divergence D4 records both positions. They become spec matter only if M3/M5 flips them. [ST]'s closed-loop observation model — conditioning evidence on teaching/probe-policy/outreach, P(E | θ, φ, x, teaching, policy) — joins this M3 list; R18's origin tags are the data-collection prerequisite that makes it fittable later.
- **External package-review file hygiene** — `external/codex-ultra-package-review.md` contains duplicated review text and pasted terminal artifacts (~lines 96–181); quote only from its deduplicated top section (lines 1–95). Report-agent note, not a spec change.
- **Oracle review attachment** — spec §10 already requires it; it is simply outstanding (see report provenance). Status item, not a spec edit.

*End of proposed revisions — 19 items: R1–R16 covering all 10 ADJUDICATION §5 SPEC-LEVEL INPUTS plus the deliverable-proposed schema/protocol deltas (memo §14, probe bank §8.3, dossier §1.2), then R17–R19 from the post-adjudication stress-test [ST], which also amended R6, R8, R12, R13, R15, and R16 in place. [ST]'s closing verdict binds the set together: "No, not as a deployment-validation study. Yes, as a strict shadow-mode feasibility study — with one amendment: insert the blinded stage-free actionability gate before M2 and make passing it necessary to continue."*


