# GPT-5.6 Sol (codex exec, effort ultra) — adjudication stress-test — 2026-07-16

## Overall assessment: needs revision

The four fixes improve honesty more than validity. Fix (a) closes a real leakage path; (b) is valid only if M5 loses deployment authority; (c) mostly renames single-rater circularity; and (d) remains an unimplemented governance proposal. The deeper category error is that Waypoint is tested as a predictor or Surya-imitation instrument but intended to operate as a controller of teaching and safety decisions.

## 1. Fix-by-fix verdicts

### (a) Hash-freeze dossier and forecast before the Surya interview

**Verdict: partial resolution. It closes direct target leakage, not evaluation circularity as a whole.**

A hash establishes that bytes did not change; it does not establish that the frozen object was assembled independently. The dossier can still be selectively constructed, summarized, deduplicated, probe-enriched, or filtered using beliefs about the practitioner. This matters because consolidation includes posterior-aware span selection, despite estimate-blind re-extraction (*estimator-design-memo.md* §4.3).

The freeze must therefore cover the entire evaluation bundle:

- Raw-data cutoff and dossier-construction rules.
- Probe-selection policy and every delivered probe.
- Extracted and merged event log.
- Prompt, extractor model, code and parameter-table versions.
- Missing-data and exclusion rules.
- Forecast, analysis plan and pass/fail criteria.

There is also upstream leakage. Surya sets or vetoes marker likelihoods, dynamics and thresholds at M2 while already knowing this tiny team cohort (*estimator-design-memo.md* §§3, 13). Freezing the final forecast does not undo subject-specific knowledge that entered parameter elicitation. M4 dossiers should be embargoed from M2 elicitation, and any post-M4 retuning must be tested on genuinely new people—not merely future windows from the same recognizable practitioners.

Finally, interview/dossier counterbalancing does not eliminate memory carryover. “Blinded dossier rating” likely means masked to estimator output, not blinded to identity: a five-person team’s prose and practice history may be obvious.

So: **the fix resolves the immediate interview-transcript leak, but only a preregistered, fully frozen evaluation pipeline resolves the broader look-ahead problem** (*ADJUDICATION.md* §5.1; spec §7).

### (b) Feasibility case series; practitioner as resampling unit; κ descriptive-only

**Verdict: sound rescoping only if deployment claims are removed. Otherwise it is relabeling.**

Treating the practitioner as the unit is correct. Complete case displays, signed band errors, abstention decisions and descriptive coverage counts are appropriate. κ and practitioner bootstrap results can appear in an appendix, but with 5–10 narrow-range, map-exposed cases they are too unstable to guide a product decision (*estimator-design-memo.md* §10.3).

The unresolved contradiction is milestone authority. The canonical spec still makes M5 a “go/no-go on runtime integration + teacher-flag activation” and asks for agreement and uncertainty calibration (*waypoint-practitioner-stage-assessment.md* §§7, 12). A feasibility case series cannot validate nominal posterior coverage, transportability to app users, or a safety-relevant flag.

The strongest positive conclusion M4/M5 can support is:

> The pipeline is executable, auditable, produces interpretable disagreements, and merits a preregistered shadow-mode alpha study.

It cannot support “calibrated enough to steer Wisdom” or “safe enough to activate teacher flags.” Unless M5 is rewritten accordingly, the statistical claims have been softened while the consequential decision remains unchanged.

### (c) “Surya reference ratings,” repeated ratings and anchor vignettes

**Verdict: mostly relabeling, with a useful intra-rater reliability improvement.**

“Surya reference rating” is the honest term. Repeated ratings can quantify Surya’s own repeatability, and disguised duplicate anchors can detect scale drift. Neither establishes construct validity.

Surya still:

- Defines and teaches the construct.
- Calibrates markers, probes, likelihoods and dynamics.
- Conducts the criterion interview.
- Rates the dossier.
- Interprets disagreements.

Agreement therefore measures fidelity to an operationalization of Surya’s model, not independent correctness (*waypoint-practitioner-stage-assessment.md* §§2, 7; *estimator-design-memo.md* §§3, 10.2, 13).

Anchor vignettes authored from the same Wheel can actually strengthen common-method lock-in: they standardize Surya’s scale use without testing whether the scale corresponds to anything external.

The interview–dossier delta also remains uninterpretable unless compared with within-mode disagreement. It conflates information-channel differences with memory, order, interview performance, subject change and ordinary rater noise. The minimum useful design is two dossier ratings and two interview ratings, randomized in order and separated by washout. Report:

- Dossier–dossier disagreement.
- Interview–interview disagreement.
- Interview–dossier disagreement relative to those two baselines.

Until an independent rater or prospective external outcome exists, the valid claim is strictly **“replication of Surya’s reference ratings.”**

### (d) Assessment-aware, item-blind disclosure; DPIA; possible removal of P-040

**Verdict: ethically partial if implemented, but currently unresolved.**

The canonical spec still requires invisible probes (*waypoint-practitioner-stage-assessment.md* §2 decision 4 and §5.2), and *ADJUDICATION.md* says the repaired disclosure regime is contingent on a future spec amendment (§2.11, §5.4). Thus the fix is not yet operative.

Assessment-aware consent must also be **use-aware**. Participants need to know that the system infers a stage-like spiritual/psychological profile and may use it for practice selection, pacing, language, triage and human outreach. Item blindness can protect test integrity; it cannot conceal consequential uses.

A DPIA is a review process, not a safeguard. It should occur before team M4, not merely before external alpha. The team cohort introduces workplace power asymmetry now: refusal must be penalty-free, managers should have no access, and stewardship should be independent.

Disclosure also changes the measurement condition. Marker and probe behavior calibrated under “invisible assessment” cannot simply be assumed valid under assessment-aware participation; passive and disclosed-probe evidence should be analyzed separately.

P-040 should be removed now. A foil that asserts a nonexistent phenomenon can prime later reports, contaminate the evidence stream and damage trust. It is not load-bearing enough to justify deception, debriefing complexity or construct-owner discretion.

## 2. The biggest unexamined assumption

The missing assumption is **prescriptive validity under intervention**:

> Conditional on the raw facts Wisdom already has, adding a Waypoint posterior improves the choice of teaching action—and the posterior remains valid after those actions alter what the practitioner practices, experiences and reports.

Waypoint may possess prognostic validity while having no prescriptive value. A stage can predict future language or gate progression without moderating the effect of any teaching choice. Formally, the required claim is not merely that \(\hat\theta\) predicts \(Y\), but that:

\[
E[Y(A(X,\hat\theta))] > E[Y(A(X))]
\]

where \(X\) contains the proximal facts Wisdom already observes.

The current study tests agreement with Surya and prediction of probe responses, milestone corroborations and gate progression (*waypoint-practitioner-stage-assessment.md* §7; *estimator-design-memo.md* §10.1). Those H1 targets are especially weak because Wisdom or the posterior-driven probe scheduler may select them. Nothing tests whether stage information improves practitioner benefit, fit or safety over direct use of current goals, phenomenology, pillars, phase and risk features.

The exposing experiment is a preregistered **estimate-visibility policy trial**:

- Compute and freeze Waypoint for everyone.
- Randomize eligible low-risk participants or blocks to Wisdom receiving either the raw proximal profile alone or the identical information plus Waypoint.
- Keep the primary safety lane stage-blind and identical across arms.
- Freeze and log model, curriculum and policy versions.
- Use outcomes not selected by Waypoint: practitioner-rated fit/helpfulness, goal progress, functioning/distress, adverse events and blinded teacher correction.
- Record every action and its assignment probability.

The team cohort can test mechanics only. Any benefit or safety claim requires a larger prospective alpha study.

## 3. Single highest-expected-value change before M2–M4

**Insert a blinded actionability gate before Surya spends M2 calibrating the marker bank.**

On roughly 12–20 frozen real or independently authored cases, generate Wisdom recommendations under three randomized views:

1. Raw case facts with no Wheel estimate.
2. The same facts plus the Surya reference distribution.
3. The same facts plus a plausible but wrong or ±1-band-shifted distribution.

Have Surya rank the recommendations for teaching fit and safety without knowing the condition; repeat disguised subsets after washout. Pre-register a stop/coarsen rule.

This answers three cheap questions:

- If adding stage does not improve recommendations, the latent is redundant.
- If a one-band perturbation causes major curriculum changes, the policy is too brittle for current uncertainty.
- If a wrong estimate produces unsafe escalation, runtime use is disqualified regardless of agreement statistics.

This is not proof of live benefit; it is a premise check. Its expected value is high because it likely costs a prompt harness and about two short Surya sessions. By comparison, the memo budgets 4–6 Surya sessions merely to elicit approximately 60 marker rows, before M3 implementation and M4 interviews (*estimator-design-memo.md* §3.4). Synthetic QA explicitly cannot test the assumptions (§10.5).

## 4. Missed closed-loop and safety failures

*ADJUDICATION.md* §5.6(c) does mention “thermometer → thermostat.” What both layers failed to do is represent the loop in the data-generating process:

\[
\hat\theta_t \rightarrow \text{Wisdom action}_t
\rightarrow \text{practice/exposure/reporting}_{t+1}
\rightarrow \hat\theta_{t+1}
\]

The estimator models evidence largely as \(P(E\mid\theta,\phi,x)\), when runtime requires at least \(P(E\mid\theta,\phi,x,\text{teaching},\text{probe policy},\text{outreach})\).

### Controller-produced evidence

Wisdom will choose pacing, practices, language and probes using the estimate (spec §8), while curriculum position, practice cadence, gate progression and future probe responses are treated as evidence or validation targets (spec §5; memo §§2.2, 4.1, 10.1).

That lets the estimator count its own decisions as confirmation. The most direct ratchet is memo §4.1: observed practice carries positive forward drift even without discriminating evidence, multiplied by practice intensity. If Wisdom increases practice cadence because of a high estimate, the previous estimate mechanically advances the next prior.

Assigned or unlocked curriculum must never count as competence. Controller-influenced engagement and practice cadence must not validate the controller.

### False-high and false-low attractors

A false high can produce advanced vocabulary and intensive/nondual practices. The practitioner then mirrors the language or experiences destabilization, producing markers that confirm the high estimate. Phase gating may classify the consequences as “dip” or “intensification” and mute downward stage evidence (memo §4.4).

A false low produces beginner content and fewer opportunities to display advanced capacities. The practitioner disengages, reports less or churns; silence then becomes abstention rather than evidence that the teaching policy was wrong. This is a hidden soft-gating system even if no content is formally prohibited.

Dropout must therefore be treated as a policy outcome and possible harm—not merely missing evidence or weak discipline.

### Adaptive probing creates missing-not-at-random evidence

The consolidation selects probes by expected posterior entropy reduction (memo §4.3). The posterior determines what becomes observable, and the resulting response updates the same posterior. This is exploitation without guaranteed falsification.

A system can appear increasingly predictive by selecting easy, confirmatory probes. Fixed policy-independent anchor items and logged random exploration are needed. Every evidence event should carry an origin such as `spontaneous`, `standardized_anchor`, `policy_elicited`, `post_teaching` or `post_outreach`.

### Full-history smoothing can erase treatment effects

Monthly full-history re-smoothing revises the trajectory from \(t_0\) using later evidence (memo §4.3). If Wisdom teaches vocabulary or causes genuine development at month three, the smoother may retrospectively infer that the practitioner was already high at month one. The intervention is rewritten as latent-state evidence, making the policy appear prescient.

Contemporaneous frozen forecasts—not smoothed historical states—must be used for policy evaluation and causal claims.

### Co-adaptation and measurement-environment drift

The practitioner learns Wisdom’s vocabulary and implicit curriculum; Wisdom changes register; the extractor changes version; and users learn what kinds of answers produce which teaching. Hidden scores do not stay behaviorally hidden when the treatment reveals them.

`param_version` covers the marker bank, priors and extractor, but not the Wisdom prompt/model, curriculum, probe scheduler, safety classifier, flag policy or outreach history (memo §§4.3, 14). These should form a `measurement_environment_version`. Evidence generated under materially different regimes should not be silently pooled as exchangeable.

### Safety can fail through diagnostic overshadowing

The spec explicitly says identical words may “read differently” depending on stage (spec §8). That invites spiritualization of depression, mania, psychosis, trauma or dissociation. These are not mutually exclusive with contemplative difficulty.

“Stage never downgrades safety” must become an executable invariant:

- Run a primary stage-blind safety classifier.
- Run any stage-aware support lane separately.
- Operational severity is the union or maximum of the two.
- Adding stage context may never suppress a flag, lower severity, delay response or replace clinical-language guidance with a contemplative explanation.
- Property-test this under deliberately wrong stage and phase injections.

The raw ungated marker path prevents phase-gating suppression; it does not prevent the safety classifier or human teacher from interpreting danger as “advanced territory.”

A posterior-mass-\(\ge 51\) flag should remain shadow-only. Human outreach should require direct raw features such as impairment, uncontrollability, persistence, severe sleep loss, suicidality, rapid functional change or a user request. The teacher-facing alert should describe those observations, not lead with “unbinding-adjacent,” which creates automation and framing bias.

### “Human in the loop” creates an operational duty

Monthly consolidation is not an acute-risk mechanism. Before teacher flags activate, the design needs named ownership, acknowledgement and response times, capacity limits, escalation and closure rules, and audit of missed or late responses. Otherwise the flag creates an expectation of care without a reliable care path.

Outreach itself changes subsequent evidence and must be recorded as an intervention. It must not update stage or enter a later reference-rating reliability replicate without explicit separation.

## 5. Closing verdict

**No, not as a deployment-validation study. Yes, as a strict shadow-mode feasibility study—with one amendment: insert the blinded stage-free actionability gate before M2 and make passing it necessary to continue.**

Even a pass would authorize only the M2–M4 shadow case series and a later prospective alpha trial. It would not justify runtime adaptation, stage-aware safety decisions or teacher-flag activation. If the gate fails, retain the useful proximal components—phase, pillars, raw safety features and current teaching needs—and defer the overall Wheel latent.
