# External Review: Waypoint Practitioner Stage Assessment Package (Codex GPT, reasoning effort: ultra, date: 2026-07-16)

## Verdict

No. A top psychometrician would call this an unusually thoughtful, auditable expert-system prototype—not a validated developmental-stage instrument.

Its engineering is ahead of its measurement science. Abstention, versioning, cited evidence, confusion pairs, phase/stage separation, and keeping arithmetic outside the LLM are all strong. But none resolves the central problem: Surya defines the construct, approves the probes, supplies the likelihoods, and produces the “gold” labels. Agreement can establish fidelity to Surya, not independent validity of the Wheel.

The [AERA/APA/NCME Standards](https://www.testingstandards.net/uploads/7/6/6/4/76643089/standards_2014edition.pdf) require each intended score interpretation and use to be validated separately. Waypoint currently conflates three:

- A real developmental stage.
- Surya’s teaching judgment.
- Prediction of future conversation, telemetry, and curriculum behavior.

Those are not interchangeable.

## Construct validity

The design has content-generation work, not yet construct validity.

The major blockers are:

- **Circular criterion:** Surya is construct author, parameter elicitor, probe approver, and sole criterion rater. “Gold label” should be renamed “Surya reference rating.”
- **Untested dimensionality:** The overall latent plus four pillar offsets assumes a general developmental factor, locally independent pillars, and stable offsets. The same pillar profile can support multiple overall/offset combinations; the chosen prior largely locates the overall score.
- **Reflective-versus-formative ambiguity:** The pillars may constitute overall development rather than reflect a single underlying cause. Until that is resolved empirically, an overall score is not justified.
- **Unsupported metric structure:** An ordinal Wheel does not license equal one-point distances, posterior means such as 36.4, or fixed one-point transition rates. Remove the 1–100 point estimate from v1.
- **Measurement invariance is assumed:** Marker probabilities will vary with language, education, culture, tradition, neurotype, therapy vocabulary, verbosity, disclosure style, distress, and relationship with Wisdom. Map exposure is only one source of DIF.
- **Transfer claims are too strong:** LLM reliability on STAGES scoring supports the feasibility of automated rubric scoring. It does not validate transfer to contemplative development in open-ended adaptive conversation.
- **Cross-tradition convergence is overstated:** Traditional texts, psychometric precedents, clinical safety research, and predictive-processing theory provide different kinds of evidence. Their thematic resemblance does not establish one common latent continuum.

If independent raters cannot be recruited, the honest construct is: **“a Surya-aligned teaching-profile estimate.”** That may still be useful, but it is not an objective stage measure or a validation of the Wheel.

## Scientific defensibility

As written, the output is not a calibrated Bayesian posterior. It is a transparent, tempered evidence score expressed in Bayesian language.

Key reasons:

- Expert-elicited likelihoods are fixed as though known. Sensitivity sweeps do not propagate parameter uncertainty or detect wrong likelihood shapes.
- Marker observation is missing-not-at-random. What appears depends on prompts, topic opportunity, transcript volume, probe selection, disclosure, and extractor behavior. Ignoring absences does not make observed presences unbiased.
- Semantic deduplication cannot remove dependence among multiple markers arising from the same episode, vocabulary style, or probe sequence. Multiplication will overstate certainty.
- Phase gating, KL caps, window budgets, and persistent-run overrides are useful heuristics but do not constitute coherent joint Bayesian updating. Nominal “80% credible intervals” therefore have no established 80% coverage.
- A person’s stage may change slowly, but the system’s belief should sometimes change sharply when strong evidence reveals the prior was wrong. The design conflates slow ontic movement with slow epistemic correction.
- The “evidence bits” sufficiency score is not realized information gain; the toy example counts 1.32 bits for evidence that changes the posterior by only 0.11 bits.
- N≈5–10 can support feasibility and debugging only. It cannot establish calibration, dimensionality, measurement invariance, interval coverage, or meaningful agreement statistics. Pillars and months are not independent replacement subjects.
- The interview–dossier difference is not an “information ceiling.” It mixes information access, rater unreliability, memory, order, interpersonal effects, and modality.
- Synthetic dossiers are useful regression tests, but personas constructed from the marker bank cannot validate the construct or probability calibration.
- High-band predictions and 51+ flags would be extrapolations beyond the expected calibration cohort.

The literature dossier is a useful hypothesis source, but its `[C5]` labels should not be presented as evidence grades without a systematic search, risk-of-bias assessment, and explicit separation of empirical findings from cross-domain analogy.

## Ethical soundness

As written, it is not ethically deployable.

**Covert probes:** “Invisible” and “never framed as assessment” is incomplete disclosure inside a trust-based teaching relationship. Broad Waypoint consent is only adequate if users are explicitly told that some conversational questions are deliberately designed to assess practice development, who receives the resulting profile, and how it affects their experience. The defensible principle is **assessment-aware, item-blind**: scoring keys can remain hidden; the assessment purpose cannot. Incomplete-disclosure research ordinarily requires independent review, necessity, minimal undisclosed risk, and debriefing where appropriate. [The Belmont Report states those conditions plainly.](https://www.hhs.gov/ohrp/regulations-and-policy/belmont-report/read-the-belmont-report/)

**Hidden estimates:** It can be reasonable not to push an identity-laden stage number at practitioners. It is not reasonable to conceal that profiling exists while it changes teaching or human attention. Minimum protections include notice, opt-out without service penalty, correction of factual evidence, annotation of disagreement, human review, access/deletion of derived records, and the option not to see the raw stage label.

**Teacher flags and safety:** Stage must never lower an acute-safety threshold, cancel a flag, or explain away impairment as “dark night” or “Unbinding.” That creates diagnostic overshadowing. Safety detection should remain stage-blind; stage may only add context after human review. Outreach should be triggered by present impairment, uncontrollability, persistence, suicidality, or user request—not merely posterior mass above 51.

**Governance:** These dossiers infer religious/philosophical beliefs and mental-health-adjacent states. In the EU, that strongly suggests a formal DPIA and stringent profiling safeguards before processing, not merely reliance on existing “KB rules.” [GDPR Article 35](https://eur-lex.europa.eu/eli/reg/2016/679/art_35/par_3/oj) is directly relevant. Team-member participation also needs protection against employer/teacher power: independent recruitment and data stewardship, no employment consequences, and no manager access to individual profiles.

## Biggest unexamined assumption

That there is one stable, ordered latent stage whose manifestations are invariant across people and contexts—and that Surya’s perception, conversational structure, future behavior, teaching relevance, and safety relevance all identify that same latent variable.

If that assumption fails, the system converts verbal sophistication, cultural fit, disclosure style, curriculum exposure, and agreement with Surya’s ontology into confident-looking “development.”

## Failure mode not priced in

**Performative lock-in: the thermometer becomes the thermostat.**

A low estimate produces simpler language, practices, and probes, reducing opportunities to demonstrate advanced capacities. A high estimate produces advanced framing and practice exposure that generates confirming evidence. Teacher attention creates further observable difficulty. A mistaken telemetry contradiction lowers the hidden claim-reliability score, causing later genuine reports to count less.

The system can therefore manufacture the trajectory it later cites as longitudinal validation. Hiddenness makes correction less likely. This is more serious than ordinary probe reactivity.

## Single highest-value change

Replace M4 with a **hard, transparently consented shadow-mode validation phase**:

- No stage output influences Wisdom, practice selection, safety thresholds, probe selection, or teacher outreach.
- Use a fixed, counterbalanced core probe panel; randomize any exploratory probes and log assignment probabilities.
- Freeze the scoring manual and intended score interpretations.
- Obtain repeated blinded ratings from Surya and independent qualified raters.
- Use a diverse, held-out prospective cohort and outcomes not incorporated into scoring.
- Compare against simple baselines before testing the complex dynamic model.

Only after this succeeds should adaptation itself be evaluated as a separate intervention.

## What a top psychometrician would demand first

> “State exactly what each score is claimed to mean and what decision it supports. Then show that independent trained raters, using a frozen operational rubric, can distinguish the proposed bands and pillars reliably.”

Until that happens, Waypoint is promising instrument-development research—not instrument-building done right, and not ready for teaching adaptation or safety use.
                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         [35m[3mcodex[0m[0m
One important limitation is already clear: the design is not a single Bayesian model as written. It mixes a generative filter with data-dependent power likelihoods, KL clipping, threshold-triggered reweighting, and post-hoc deduplication; those choices can be useful robustness heuristics, but their intervals cannot be interpreted as Bayesian credible intervals without external calibration. I’m separating “repairable heuristic” from “breaks the claimed probability semantics” in the final review.
[35m[3mcodex[0m[0m
The toy example’s multiplication and normalization are mostly arithmetically correct, but it exposes two concrete inconsistencies: the milestone prior is 0.175 only if the discounted repeat is ignored (the live nowcast included it, which gives 0.196), and the toy band kernel allows a two-band monthly jump with probability 0.002 while the stated point-level CTMC rates imply roughly \(3.5\times10^{-9}\) for the minimum required forward jumps. The example is therefore not an implementation-faithful test fixture yet.
[35m[3mexec[0m[0m
[1m/bin/zsh -lc "sed -n '1,320p' '/Users/fionnenglish/.codex/plugins/cache/openai-curated-remote/data-analytics/0.2.8-13ceeea1f599/skills/build-report/specifications/technical-report.md'"[0m in /Users/fionnenglish/.claude-2/jobs/14f60959/tmp
[32m succeeded in 974ms:[0m
# Technical Report Specification

Use this specification when the user explicitly wants a technical audience or when the report is primarily about data methodology, measurement design,
statistical analysis, experimentation, modeling, or other audit-heavy work.

Shared artifact, visualization, layout, citation, and sharing rules live in
$build-report. This file only captures what is different for the technical shape.

## What This Shape Optimizes For

- Lead with the main result, then show the definitions, evidence, methods, and limitations needed to trust it.
- Preserve technical precision. Simplify wording only when meaning is unchanged.
- Define metrics, cohorts, filters, baselines, and units before leaning on them in argument.
- Keep assumptions, uncertainty, and failure modes close to the findings they affect.
- Distinguish descriptive, diagnostic, inferential, predictive, and causal claims.
- Integrate implications into the technical summary, section-level interpretation, and next-step framing rather than isolating them in a standalone section.

## Required Structure

Default to this order:

1. Title
2. Technical summary
3. Key findings with visual evidence
4. Scope, data, and metric definitions
5. Methodology
6. Limitations, uncertainty, and robustness checks
7. Recommended next steps
8. Further questions

These entries define section roles, not literal heading text. Use headings that fit the actual result and method story instead of copying the template labels verbatim.

If the work is model-heavy, experimental, or inferential, add a dedicated section for model specification, experimental design, or validation details instead of hiding them inside footnotes.

## Drafting Rules

- The technical summary should stand on its own and state the main result directly.
- Start sections with the result, not with methodology setup.
- Use section headings that reflect the actual result or method question rather than copying the template labels verbatim when a more specific heading would improve readability.
- For major findings, make the heading the substantive result or driver, not a generic topic label. Prefer headings like `Digital Natives and Startups drove most of the segment expansion` over labels like `Customer and segment drivers`.
- Define uncommon or overloaded terms on first use.
- Make definitions, cohorts, denominators, and comparison baselines explicit before the reader needs them.
- If a claim depends on modeling choices, thresholds, priors, sampling rules,
  or feature engineering, say so explicitly.
- Include negative results, counterexamples, sensitivity, or failure modes when they change interpretation.
- Use top-line metrics only when each one adds analytical value.
- If the work builds a model, keep fitting, diagnostics, and interpretation in the companion notebook unless a reusable support module is clearly necessary.
- Recommendations or next steps should follow from the evidence, not from generic best practice.

## Section Pattern

For each major section:

1. Result headline
2. Precise interpretation
3. Visual evidence
4. Evidence note with definition or sample context
5. Method or assumption note when needed
6. Limitation, implication, or open question

Use the heading to state the result the section proves. Avoid generic topic headers such as `Segment drivers`, `Customer analysis`, or `Performance trends`.

Do not separate results from the definitions or assumptions needed to trust them. Do not open with methods text that delays the result.

## Technical QA Addendum

- A first-time technical reader can audit the logic without guessing what was measured or how.
- The technical summary and body tell the same story at different depth.
- Definitions, cohorts, denominators, and baselines are explicit.
- Each major section clearly distinguishes result, evidence, and limitation.
- Implications appear in the summary, relevant sections, or next steps rather than as a standalone block.
- Uncertainty is quantified when possible and described precisely when not.
- If the work includes forecasting, experiments, causal claims, or modeling,
  the report states what the result does and does not establish.
- Recommendations and implications remain clearly signposted and evidence-backed.

[1mcollab:[0m Wait
[1mcollab:[0m Wait
[35m[3mcodex[0m[0m
# Adversarial statistical review

## Overall assessment: needs revision

The design is promising as an auditable deterministic evidence scorecard. It is not, as written, one coherent Bayesian model, and its “80% credible intervals” do not currently have defensible probability semantics.

The central problem is not any elicited default. It is the combination of:

- positive-only selected observations;
- conditional-independence assumptions;
- posterior-dependent tempering and KL clipping;
- sequential phase gating;
- circular milestone/reliability updates;
- fixed, uncertainty-free expert parameters.

Determinism makes the calculation reproducible. It does not make the resulting distribution calibrated.

At \(N=5\)–10, M4 can test implementation, gross discrimination, agreement with Surya, and prospective prediction for these individuals. It cannot validate interval coverage, archetype effects, transition rates, individual claim reliability, or generalization.

## Unsound or incoherent update rules

### 1. The elicitation unit and update unit do not match

The elicited quantity is:

\[
f_m(g)=P(\text{marker fires at least once in a typical month}\mid\theta=g).
\]

The integrator applies this monthly probability to individual timestamped events, applies it again at exponent 0.25 to repeats, and ignores non-firings. That is not the likelihood of the elicited experiment.

If markers are genuinely assessed once per window, the likelihood for assessable markers should be:

\[
L_w(g)=\prod_m f_m(g)^{y_m}\{1-f_m(g)\}^{1-y_m}.
\]

Markers without a real opportunity should be missing, not zero.

If markers are event arrivals, elicit rates and use an exposure-aware Poisson/negative-binomial or marked-point-process likelihood. If conditioning on the number of detected events, marker identity must be normalized by total event intensity; multiplying raw \(f_m(g)\) omits that denominator.

Concrete fix:

- Define `opportunity_id`, `episode_id`, exposure, and eligible marker families.
- Either update once per fixed window, or replace monthly probabilities with event rates.
- Until then, call the output a generalized-Bayes evidence distribution, not a calibrated posterior.

### 2. Deduplication does not establish independence

The merge pass handles repeated descriptions of one episode. It does not handle dependence among:

- several markers fired from the same experience;
- markers sharing vocabulary or writing style;
- several extractor decisions from one quote;
- probe responses selected because of the current posterior;
- telemetry and guard markers generated by the same underlying event.

Ten correlated one-bit markers remain ten fictitious bits after simple retell deduplication.

Concrete fix:

- Make the episode or opportunity the evidence unit.
- Permit at most one most-specific factor per evidence family per episode in v1.
- Later use joint bundle likelihoods, testlet/random effects, or a cluster-level composite-likelihood weight.
- Run leave-one-episode-out sensitivity, not only row-value perturbations.

### 3. KL caps and persistent-run overrides are data-dependent pseudo-likelihoods

Choosing a tempering exponent after observing the event, specifically to keep posterior KL below 1.5 bits, means the likelihood depends on the prior and realized outcome. Splitting one transcript into two sessions, changing event order, or changing bundle boundaries can change the answer.

The persistent-run rule is worse: later observations retroactively increase the weight of earlier observations selected because they pointed in the same direction. A wrong, narrow prior also clips precisely the evidence needed to escape it.

Concrete fix:

Use a predeclared contamination mixture:

\[
L_u(\theta)=(1-\epsilon)L_{\text{signal},u}(\theta)+\epsilon L_{\text{noise},u},
\]

or a latent outlier/regime process. Persistence should matter because an explicit transient state cannot plausibly explain a long run, not because a threshold releases a larger cap.

If power likelihoods remain, their weights must be fixed before seeing the outcome and assigned once per independent evidence cluster.

### 4. Phase gating is not joint Bayesian error routing

The system first infers phase from the session and then uses that inferred phase to choose an exponent for stage evidence from the same session. This is plug-in, data-dependent gating—not marginalization.

It also creates a self-sealing failure: downward evidence can induce “dip,” which then mutes that same downward evidence.

A coherent update is:

\[
p(\theta_t,\phi_t\mid e_t)\propto
P(e_t\mid\theta_t,\phi_t)
\sum_{\theta_{t-1},\phi_{t-1}}
P(\theta_t,\phi_t\mid\theta_{t-1},\phi_{t-1})
p(\theta_{t-1},\phi_{t-1}).
\]

Concrete fix:

- Define phase-specific joint emissions and update phase and stage simultaneously.
- Require at least some phase evidence independent of the stage evidence being moderated.
- Prefer a transient session-effect or hidden semi-Markov nuisance state over named phase gates until phase is empirically distinguishable.

The toy glimpse does not demonstrate Bayesian surprise routing. Stage remains unchanged because the marker was manually routed away from stage.

### 5. The CTMC does not deliver the stated silence behavior

With constant interior rates, the usual reflecting-boundary birth–death chain has:

\[
\frac{\pi(g+1)}{\pi(g)}=\frac{\nu_+}{\nu_-}=6.
\]

Therefore long unobserved periods eventually concentrate mass near 70. If silence changes only the forward rate to \(0.03\times0.3=0.009\), the ratio is still 1.8, so it still drifts upward. A finite ergodic chain does not diffuse forever; it eventually approaches and sharpens around its stationary distribution.

Other problems:

- \(\exp(Q\Delta t)\) gives positive probability to every reachable distance for every \(\Delta t>0\). It discourages teleporting; it does not make it impossible.
- The net monthly drift at default intensity is \((0.03-0.005)\times30=0.75\) points.
- “Variance ×2” is undefined without saying how both rates change; changing rates generally changes drift too.
- Practice intensity is also stage-related telemetry, so using it both as evidence and as a transition driver risks double counting and endogeneity.

Concrete fix:

- Specify boundary conditions.
- Construct a reversible generator around a broad desired stationary distribution using detailed balance.
- Use near-zero drift during unobserved intervals.
- Separate drift and diffusion parameters.
- Consider a sticky semi-Markov model with state-specific dwell times instead of a time-homogeneous birth–death chain.

### 6. The longitudinal pillar model contradicts the static-offset model

The stated model has dynamic \(\theta_o(t)\), static \(\delta_k\), and \(\theta_k(t)=\theta_o(t)+\delta_k\). Under that model, all pillar movement is induced by the common state and the posterior over each persistent offset must be retained.

NOWCAST instead diffuses pillar marginals separately and then recouples them using the original offset prior. This either loses the learned offset posterior or implicitly redraws offsets each window.

There is also a location-identifiability ridge:

\[
\theta_o'=\theta_o+c,\qquad \delta_k'=\delta_k-c
\]

produces identical pillar likelihoods. Cross-pillar markers and the offset prior break the ridge, meaning overall location is substantially prior-defined.

Concrete fix:

- Impose \(\sum_k\delta_k=0\), making overall location definitional.
- Maintain the actual joint posterior over persistent offsets.
- Avoid `clamp`; restrict valid support and renormalize. At boundaries, several offsets map to one clamped pillar state, so the memo’s redistribution equation is incorrect.
- Simpler v1: estimate four pillar states and define overall explicitly from posterior samples of their mean or another declared functional.

### 7. Milestones, guards, and reliability reuse the same evidence

The design derives a milestone prior from stage, updates the milestone from a claim and telemetry, and then feeds the milestone posterior back into stage. Unless implemented as one joint factor, that counts the stage prior and evidence twice.

E4’s telemetry contradiction is used to:

1. move stage through a guard marker;
2. move milestone probability;
3. increment the reliability Beta;
4. potentially trigger conflict status.

The one-sentence dedup rule does not prevent this cross-layer reuse.

Correct structure:

\[
\theta_t\rightarrow M_t\rightarrow
\{\text{claim},\text{telemetry}\},
\]

with:

\[
P(e\mid\theta_t)=
\sum_M P(e\mid M,r_p)P(M\mid\theta_t).
\]

A guard record may summarize the result, but should not be another likelihood factor.

### 8. The Beta reliability update is not conjugate to the actual data

Beta–Bernoulli updating assumes claim truth is observed without error and verification is ignorable. Here corroboration and contradiction are noisy, selectively obtained, and sometimes produced using the same telemetry as the stage update.

Using \(E[r_p]\) as an exponent has no corresponding claim model. A plausible model is instead:

\[
L_{\text{claim}}(g,r)
=rL_{\text{genuine}}(g)+(1-r)L_{\text{contaminated}}(g),
\]

integrated over uncertainty in \(r\).

One disputed event changing every future claim across every domain is also excessive cross-domain generalization.

Concrete fix:

- Do not estimate individual reliability in v1.
- Use fixed conservative channel-specific error rates.
- Later model latent claim truth, corroborator sensitivity/specificity, and domain-specific reliability jointly.
- Do not apply an independent guard likelihood to evidence already used to adjudicate claim truth.

### 9. “Evidence bits” are not information absorbed

Summing each marker’s maximum adjacent-band log-LR is neither realized information gain nor an effective sample size. A marker receives full credit even when:

- its likelihood is flat where the posterior currently lies;
- it contradicts other evidence;
- it is redundant;
- it is a guard that pushes away from a claim.

In the toy, unresolved E4 helps sensations cross the reporting threshold. Contradiction can therefore make the system more reportable.

Concrete fix:

Use:

- a minimum count of independent evidence clusters;
- multiple sessions and preferably multiple channels;
- leave-one-cluster-out stability;
- robustness across parameter/dependence scenarios;
- posterior expected decision loss for the intended downstream action.

If the current quantity is retained, rename it `nominal_evidence_score`, not bits.

## Identifiability at \(N=5\)–10

| Component | What is not identifiable |
|---|---|
| 56-point stage | Five elicited anchors and band-level labels cannot distinguish 35.8 from 36.4. Integer precision is interpolation plus prior. |
| Overall vs offsets | Pillar evidence identifies sums \(\theta_o+\delta_k\), not their decomposition. |
| Stage vs phase | A sequence can be explained by stage, phase, extractor error, emission strength, or transition rates. |
| Marker prevalence | True prevalence, conversational opportunity, extractor sensitivity, and extractor specificity are confounded. |
| Archetype effects | Path, individual, stage, exposure, reporting style, offset shift, and marker DIF are nearly collinear. |
| Exposure effects | The team cohort is exposed everywhere, so the naive/exposed contrast is unobserved. |
| Claim reliability | A handful of selectively verified claims cannot identify a person-specific reliability parameter. |
| High stages | With subjects concentrated around 22–45, behavior in 51+ is entirely prior/expert extrapolation. |

At this N, repeated sessions help estimate within-person predictive consistency, but they do not turn five practitioners into dozens of independent calibration subjects.

Reasonable things M4 can learn:

- whether the integrator implements the specified calculation;
- whether adversarial cases produce obvious failures;
- Surya’s test–retest consistency;
- descriptive agreement with Surya;
- whether frozen predictions beat nulls prospectively for these specific people;
- which dossiers and assumptions dominate the conclusion.

It cannot learn:

- nominal 80% coverage;
- population/archetype priors or DIF;
- \(Q\), practice-intensity effects, or phase dynamics;
- marker curves and extractor error simultaneously;
- individual reliability;
- generalization to alpha users or high-stage practitioners.

## Where partial pooling will fail

This is structured hierarchical shrinkage, but not data-learned partial pooling: population, archetype, offset, emission, and transition parameters are fixed expert tables. Their uncertainty is absent from the posterior.

Specific failures:

- Path × exposure cells will contain zero or one subject.
- The same path information changes offset priors and likelihood rows, so the two effects are inseparable.
- A breathwork offset shift changes the inferred overall stage from identical sensations evidence, despite the claim that path does not affect overall position.
- Total pooling of emissions assumes measurement invariance across terse, verbose, humble, enthusiastic, native, and map-literate practitioners.
- `KL(individual || archetype prior)` grows with evidence volume and posterior sharpness; it is not individuation accuracy.
- A wrong-archetype swap tests sensitivity to tables selected by archetype. Separation can occur by construction.
- With \(\Lambda_k\equiv1\), a starved pillar sends a constant message and has no effect on the overall posterior. It cannot “keep the overall set wider,” as the toy claims.

Concrete v1 correction:

- Use one wide common prior.
- Remove archetype effects from stage inference; retain path as metadata or for probe selection.
- Constrain offsets to sum to zero, or remove the overall latent.
- Treat passive-marker invariance as an explicit sensitivity dimension.
- Do not activate archetype/DIF estimation until compared cells have genuine overlap across stage—roughly 10 practitioners per cell is only a bare planning floor, not sufficient validation.

## Miscalibration risks

### Parameter uncertainty is almost entirely omitted

The interval conditions on exact marker curves, \(Q\), offset spread, phase gates, extractor behavior, dependence assumptions, and reliability weights. Those are probably the dominant uncertainties.

`firm|rough|guess` shrinkage and independent ±20/40/60% perturbations do not represent:

- wrong marker direction;
- shared expert optimism;
- correlated row errors;
- extractor false positives;
- uncertain conversational opportunity;
- misspecified dynamics.

Elicit distributions on logit-scale anchors and propagate joint parameter draws. Report a robust envelope or union of intervals across plausible models. Until external validation, label intervals `model-conditional support intervals`.

### Synthetic QA cannot calibrate human probabilities

Synthetic dossiers can test signs, routing, arithmetic, and obvious adversarial failures. They cannot validate real marker prevalence, dependence, extractor errors, base rates, or human coverage.

The memo says synthetic QA is “never calibration” but assigns M3 the job of tuning caps, confidence temperatures, evidence thresholds, and abstention rules. That is calibration to the synthetic generator.

Split M3 into:

- simulation-based calibration from the exact stated generative model, testing code only;
- held-out adversarial transcripts, testing extraction/routing only;
- no tuning of human probability or abstention claims.

### Coverage cannot be estimated at M4

If 8 of 10 independent cases fall inside an 80% interval, the exact 95% interval for actual coverage is approximately 44%–97%. For 4 of 5 it is approximately 28%–99.5%.

Pillars, rounds, and events do not create 40–80 independent observations because they share practitioner, dossier, rater, extractor, and parameters. Reliability bins, paired bootstraps, and risk–coverage curves will look more informative than they are.

Approximately 62 independent calibration units are needed merely to estimate 80% coverage within ±10 percentage points under ideal binomial assumptions; clustering requires more. That still would not identify the full hierarchy.

### Validation targets are partly endogenous

Probe selection depends on the posterior, and the probe may change subsequent behavior. Milestone corroboration is also scheduled in response to claims. Score only forecasts frozen before outcomes, under a logged evaluation policy.

Use fixed or randomized sentinel probes, include nonresponse and no-event windows, and use target-specific nulls. A stage prior is not a null forecast for a probe category or an event time.

H2 “movement plausibility” is a diagnostic, not a scored outcome. H2 needs a later independent rating or a prespecified observable trajectory.

Also add a direct no-stage predictor of future observables. Beating base-rate and persistence nulls does not show that the stage latent is useful if a recent-history model predicts equally well.

## What breaks under one expert rater

Surya defines the construct, supplies the prior/emission tables, approves their implications, and supplies the reference label. Agreement therefore partly measures successful encoding of Surya’s own rubric—incorporation bias—not independent validity.

Additional problems:

- Systematic rater bias cannot be separated from true stage.
- Interview and dossier ratings by the same person have correlated construct and threshold errors.
- Their difference is not a clean “information ceiling”; it also contains modality effects, memory, familiarity, and occasion noise.
- The proposed Hui–Walter analogy is inapplicable: the readings are dependent, ordinal, applied to one narrow population, and far too few.
- Mechanical adjacent-band smearing assumes symmetric, known, unbiased rater error and rewards diffuse forecasts.
- If M4 residuals cause table revision, the revised model has been trained on M4 and cannot be validated on it.
- A different LLM family is not a substitute for a second qualified domain rater.

Concrete fixes:

1. Rename the target `agreement with Surya’s dossier judgment`, not gold or truth.
2. Have Surya provide an explicit probability vector over bands and pillars.
3. Have every dossier rated twice after washout, in randomized blinded order, with duplicated anchor vignettes.
4. Obtain a second qualified rater on at least a subset.
5. Model rater thresholds/confusion only after repeated or multi-rater data exist.
6. Freeze extractor, tables, priors, and thresholds before M4.
7. Any post-M4 revision creates v2, evaluated only on future windows or people.
8. Explicitly exclude the gold interview transcript and instrument results from the estimator’s evaluation dossier.
9. Compare estimator disagreement with Surya against Surya’s own test–retest disagreement.

Drop weighted \(\kappa\) at this sample size and range. Report per-case RPS against Surya’s elicited distribution, signed band error, absolute band error, and all individual paired differences.

## Toy-example audit

Most displayed multiplication and normalization is correct. The example is nevertheless not an implementation-faithful fixture:

- After E2b, the milestone prior is approximately 0.196, not 0.175. The reported value ignores E2b.
- The bare-claim counterfactual gives band-5 mass approximately 0.538 from the actual pre-E4 nowcast, not 0.49.
- The toy band kernel assigns probability 0.002 to b3→b5 in a month. Starting from the top of b3 requires at least 11 forward point-jumps; with forward mean 0.9/month, that probability is approximately \(3.5\times10^{-9}\) before accounting for backward jumps.
- The toy multiplies all monthly evidence and then applies one kernel. The stated algorithm requires transitions between event timestamps; those operations do not commute.
- A flat starved-pillar message has exactly zero effect on the overall posterior.

Ship a machine-readable fixture containing every event timestamp, exact production-grid \(Q\), phase generator, priors, likelihood rows, and expected intermediate posterior.

## Better model classes

For \(N=5\)–10, the best replacement is simpler, not more elaborate.

### Immediate v1: coarse probe-anchored ordinal model

- Use 3–4 coarse states: 22–30, 31–40, 41–50, and 51+ manual-review territory.
- Estimate a static state over a 60–90-day window.
- Use standardized probes with explicit opportunities and full categorical outcomes as the measurement backbone.
- Add telemetry through explicit exposure/censoring models.
- Treat passive markers as episode-capped, low-weight composite evidence.
- Use one common wide prior.
- Omit archetype stage effects, personalized reliability, dynamic offsets, and phase gating.
- Integrate elicitation and extractor uncertainty into the output.

This model says less but can be audited and prospectively scored.

### Later: robust ordinal state-space model

A coherent longitudinal extension could use:

\[
z_{p,t}=z_{p,t-1}+\mu_p\Delta t+\eta_t,
\]

\[
u_{p,s}=\rho u_{p,s-1}+\xi_s,
\]

\[
z_{pk,t}=z_{p,t}+d_{pk},\qquad \sum_k d_{pk}=0.
\]

Here \(z\) is slow stage, \(u\) is a transient session effect with Student-\(t\), mixture, or semi-Markov innovations, and \(d_k\) are identified pillar offsets. Observations use joint opportunity-aware emissions, including extractor error. Persistent observations move \(z\) because the transient component cannot explain a long run—not because a KL threshold is released.

Model milestones as noisy observations of an interval-censored attainment time, not as a BKT posterior reinjected into stage.

Earn complexity prospectively:

1. target-specific base rate;
2. persistence/recent-history predictor;
3. episode-capped unit-weighted score;
4. probe-anchored coarse ordinal model;
5. add telemetry/exposure;
6. add dynamics;
7. add passive markers;
8. add phase;
9. add pillar coupling and archetypes.

A component should enter only if it improves held-out prequential score for most practitioners and does not worsen worst-case sensitivity or abstention behavior.

## The three changes with highest expected value

1. **Replace the positive-only marker multiplication and adaptive gates/caps with one opportunity-aware, episode-level observation model.** Update phase, stage, milestones, telemetry, and claim truth jointly; otherwise explicitly downgrade the product claim to a generalized-Bayes scorecard with non-nominal intervals.

2. **Collapse v1 to a coarse, mostly static, probe-anchored ordinal model with a common prior and propagated elicitation uncertainty.** Remove archetype stage effects, individualized reliability, dynamic pillar offsets, and phase gating until each earns its place through prospective ablation.

3. **Redesign M4 as a frozen feasibility and prospective-prediction study, not calibration.** Use repeated probabilistic Surya ratings, a second-rater subset, fixed/randomized sentinel outcomes, practitioner-level analysis, and future people or windows for every post-revision validation.
[2mtokens used[0m
70,206
# Adversarial statistical review

## Overall assessment: needs revision

The design is promising as an auditable deterministic evidence scorecard. It is not, as written, one coherent Bayesian model, and its “80% credible intervals” do not currently have defensible probability semantics.

The central problem is not any elicited default. It is the combination of:

- positive-only selected observations;
- conditional-independence assumptions;
- posterior-dependent tempering and KL clipping;
- sequential phase gating;
- circular milestone/reliability updates;
- fixed, uncertainty-free expert parameters.

Determinism makes the calculation reproducible. It does not make the resulting distribution calibrated.

At \(N=5\)–10, M4 can test implementation, gross discrimination, agreement with Surya, and prospective prediction for these individuals. It cannot validate interval coverage, archetype effects, transition rates, individual claim reliability, or generalization.

## Unsound or incoherent update rules

### 1. The elicitation unit and update unit do not match

The elicited quantity is:

\[
f_m(g)=P(\text{marker fires at least once in a typical month}\mid\theta=g).
\]

The integrator applies this monthly probability to individual timestamped events, applies it again at exponent 0.25 to repeats, and ignores non-firings. That is not the likelihood of the elicited experiment.

If markers are genuinely assessed once per window, the likelihood for assessable markers should be:

\[
L_w(g)=\prod_m f_m(g)^{y_m}\{1-f_m(g)\}^{1-y_m}.
\]

Markers without a real opportunity should be missing, not zero.

If markers are event arrivals, elicit rates and use an exposure-aware Poisson/negative-binomial or marked-point-process likelihood. If conditioning on the number of detected events, marker identity must be normalized by total event intensity; multiplying raw \(f_m(g)\) omits that denominator.

Concrete fix:

- Define `opportunity_id`, `episode_id`, exposure, and eligible marker families.
- Either update once per fixed window, or replace monthly probabilities with event rates.
- Until then, call the output a generalized-Bayes evidence distribution, not a calibrated posterior.

### 2. Deduplication does not establish independence

The merge pass handles repeated descriptions of one episode. It does not handle dependence among:

- several markers fired from the same experience;
- markers sharing vocabulary or writing style;
- several extractor decisions from one quote;
- probe responses selected because of the current posterior;
- telemetry and guard markers generated by the same underlying event.

Ten correlated one-bit markers remain ten fictitious bits after simple retell deduplication.

Concrete fix:

- Make the episode or opportunity the evidence unit.
- Permit at most one most-specific factor per evidence family per episode in v1.
- Later use joint bundle likelihoods, testlet/random effects, or a cluster-level composite-likelihood weight.
- Run leave-one-episode-out sensitivity, not only row-value perturbations.

### 3. KL caps and persistent-run overrides are data-dependent pseudo-likelihoods

Choosing a tempering exponent after observing the event, specifically to keep posterior KL below 1.5 bits, means the likelihood depends on the prior and realized outcome. Splitting one transcript into two sessions, changing event order, or changing bundle boundaries can change the answer.

The persistent-run rule is worse: later observations retroactively increase the weight of earlier observations selected because they pointed in the same direction. A wrong, narrow prior also clips precisely the evidence needed to escape it.

Concrete fix:

Use a predeclared contamination mixture:

\[
L_u(\theta)=(1-\epsilon)L_{\text{signal},u}(\theta)+\epsilon L_{\text{noise},u},
\]

or a latent outlier/regime process. Persistence should matter because an explicit transient state cannot plausibly explain a long run, not because a threshold releases a larger cap.

If power likelihoods remain, their weights must be fixed before seeing the outcome and assigned once per independent evidence cluster.

### 4. Phase gating is not joint Bayesian error routing

The system first infers phase from the session and then uses that inferred phase to choose an exponent for stage evidence from the same session. This is plug-in, data-dependent gating—not marginalization.

It also creates a self-sealing failure: downward evidence can induce “dip,” which then mutes that same downward evidence.

A coherent update is:

\[
p(\theta_t,\phi_t\mid e_t)\propto
P(e_t\mid\theta_t,\phi_t)
\sum_{\theta_{t-1},\phi_{t-1}}
P(\theta_t,\phi_t\mid\theta_{t-1},\phi_{t-1})
p(\theta_{t-1},\phi_{t-1}).
\]

Concrete fix:

- Define phase-specific joint emissions and update phase and stage simultaneously.
- Require at least some phase evidence independent of the stage evidence being moderated.
- Prefer a transient session-effect or hidden semi-Markov nuisance state over named phase gates until phase is empirically distinguishable.

The toy glimpse does not demonstrate Bayesian surprise routing. Stage remains unchanged because the marker was manually routed away from stage.

### 5. The CTMC does not deliver the stated silence behavior

With constant interior rates, the usual reflecting-boundary birth–death chain has:

\[
\frac{\pi(g+1)}{\pi(g)}=\frac{\nu_+}{\nu_-}=6.
\]

Therefore long unobserved periods eventually concentrate mass near 70. If silence changes only the forward rate to \(0.03\times0.3=0.009\), the ratio is still 1.8, so it still drifts upward. A finite ergodic chain does not diffuse forever; it eventually approaches and sharpens around its stationary distribution.

Other problems:

- \(\exp(Q\Delta t)\) gives positive probability to every reachable distance for every \(\Delta t>0\). It discourages teleporting; it does not make it impossible.
- The net monthly drift at default intensity is \((0.03-0.005)\times30=0.75\) points.
- “Variance ×2” is undefined without saying how both rates change; changing rates generally changes drift too.
- Practice intensity is also stage-related telemetry, so using it both as evidence and as a transition driver risks double counting and endogeneity.

Concrete fix:

- Specify boundary conditions.
- Construct a reversible generator around a broad desired stationary distribution using detailed balance.
- Use near-zero drift during unobserved intervals.
- Separate drift and diffusion parameters.
- Consider a sticky semi-Markov model with state-specific dwell times instead of a time-homogeneous birth–death chain.

### 6. The longitudinal pillar model contradicts the static-offset model

The stated model has dynamic \(\theta_o(t)\), static \(\delta_k\), and \(\theta_k(t)=\theta_o(t)+\delta_k\). Under that model, all pillar movement is induced by the common state and the posterior over each persistent offset must be retained.

NOWCAST instead diffuses pillar marginals separately and then recouples them using the original offset prior. This either loses the learned offset posterior or implicitly redraws offsets each window.

There is also a location-identifiability ridge:

\[
\theta_o'=\theta_o+c,\qquad \delta_k'=\delta_k-c
\]

produces identical pillar likelihoods. Cross-pillar markers and the offset prior break the ridge, meaning overall location is substantially prior-defined.

Concrete fix:

- Impose \(\sum_k\delta_k=0\), making overall location definitional.
- Maintain the actual joint posterior over persistent offsets.
- Avoid `clamp`; restrict valid support and renormalize. At boundaries, several offsets map to one clamped pillar state, so the memo’s redistribution equation is incorrect.
- Simpler v1: estimate four pillar states and define overall explicitly from posterior samples of their mean or another declared functional.

### 7. Milestones, guards, and reliability reuse the same evidence

The design derives a milestone prior from stage, updates the milestone from a claim and telemetry, and then feeds the milestone posterior back into stage. Unless implemented as one joint factor, that counts the stage prior and evidence twice.

E4’s telemetry contradiction is used to:

1. move stage through a guard marker;
2. move milestone probability;
3. increment the reliability Beta;
4. potentially trigger conflict status.

The one-sentence dedup rule does not prevent this cross-layer reuse.

Correct structure:

\[
\theta_t\rightarrow M_t\rightarrow
\{\text{claim},\text{telemetry}\},
\]

with:

\[
P(e\mid\theta_t)=
\sum_M P(e\mid M,r_p)P(M\mid\theta_t).
\]

A guard record may summarize the result, but should not be another likelihood factor.

### 8. The Beta reliability update is not conjugate to the actual data

Beta–Bernoulli updating assumes claim truth is observed without error and verification is ignorable. Here corroboration and contradiction are noisy, selectively obtained, and sometimes produced using the same telemetry as the stage update.

Using \(E[r_p]\) as an exponent has no corresponding claim model. A plausible model is instead:

\[
L_{\text{claim}}(g,r)
=rL_{\text{genuine}}(g)+(1-r)L_{\text{contaminated}}(g),
\]

integrated over uncertainty in \(r\).

One disputed event changing every future claim across every domain is also excessive cross-domain generalization.

Concrete fix:

- Do not estimate individual reliability in v1.
- Use fixed conservative channel-specific error rates.
- Later model latent claim truth, corroborator sensitivity/specificity, and domain-specific reliability jointly.
- Do not apply an independent guard likelihood to evidence already used to adjudicate claim truth.

### 9. “Evidence bits” are not information absorbed

Summing each marker’s maximum adjacent-band log-LR is neither realized information gain nor an effective sample size. A marker receives full credit even when:

- its likelihood is flat where the posterior currently lies;
- it contradicts other evidence;
- it is redundant;
- it is a guard that pushes away from a claim.

In the toy, unresolved E4 helps sensations cross the reporting threshold. Contradiction can therefore make the system more reportable.

Concrete fix:

Use:

- a minimum count of independent evidence clusters;
- multiple sessions and preferably multiple channels;
- leave-one-cluster-out stability;
- robustness across parameter/dependence scenarios;
- posterior expected decision loss for the intended downstream action.

If the current quantity is retained, rename it `nominal_evidence_score`, not bits.

## Identifiability at \(N=5\)–10

| Component | What is not identifiable |
|---|---|
| 56-point stage | Five elicited anchors and band-level labels cannot distinguish 35.8 from 36.4. Integer precision is interpolation plus prior. |
| Overall vs offsets | Pillar evidence identifies sums \(\theta_o+\delta_k\), not their decomposition. |
| Stage vs phase | A sequence can be explained by stage, phase, extractor error, emission strength, or transition rates. |
| Marker prevalence | True prevalence, conversational opportunity, extractor sensitivity, and extractor specificity are confounded. |
| Archetype effects | Path, individual, stage, exposure, reporting style, offset shift, and marker DIF are nearly collinear. |
| Exposure effects | The team cohort is exposed everywhere, so the naive/exposed contrast is unobserved. |
| Claim reliability | A handful of selectively verified claims cannot identify a person-specific reliability parameter. |
| High stages | With subjects concentrated around 22–45, behavior in 51+ is entirely prior/expert extrapolation. |

At this N, repeated sessions help estimate within-person predictive consistency, but they do not turn five practitioners into dozens of independent calibration subjects.

Reasonable things M4 can learn:

- whether the integrator implements the specified calculation;
- whether adversarial cases produce obvious failures;
- Surya’s test–retest consistency;
- descriptive agreement with Surya;
- whether frozen predictions beat nulls prospectively for these specific people;
- which dossiers and assumptions dominate the conclusion.

It cannot learn:

- nominal 80% coverage;
- population/archetype priors or DIF;
- \(Q\), practice-intensity effects, or phase dynamics;
- marker curves and extractor error simultaneously;
- individual reliability;
- generalization to alpha users or high-stage practitioners.

## Where partial pooling will fail

This is structured hierarchical shrinkage, but not data-learned partial pooling: population, archetype, offset, emission, and transition parameters are fixed expert tables. Their uncertainty is absent from the posterior.

Specific failures:

- Path × exposure cells will contain zero or one subject.
- The same path information changes offset priors and likelihood rows, so the two effects are inseparable.
- A breathwork offset shift changes the inferred overall stage from identical sensations evidence, despite the claim that path does not affect overall position.
- Total pooling of emissions assumes measurement invariance across terse, verbose, humble, enthusiastic, native, and map-literate practitioners.
- `KL(individual || archetype prior)` grows with evidence volume and posterior sharpness; it is not individuation accuracy.
- A wrong-archetype swap tests sensitivity to tables selected by archetype. Separation can occur by construction.
- With \(\Lambda_k\equiv1\), a starved pillar sends a constant message and has no effect on the overall posterior. It cannot “keep the overall set wider,” as the toy claims.

Concrete v1 correction:

- Use one wide common prior.
- Remove archetype effects from stage inference; retain path as metadata or for probe selection.
- Constrain offsets to sum to zero, or remove the overall latent.
- Treat passive-marker invariance as an explicit sensitivity dimension.
- Do not activate archetype/DIF estimation until compared cells have genuine overlap across stage—roughly 10 practitioners per cell is only a bare planning floor, not sufficient validation.

## Miscalibration risks

### Parameter uncertainty is almost entirely omitted

The interval conditions on exact marker curves, \(Q\), offset spread, phase gates, extractor behavior, dependence assumptions, and reliability weights. Those are probably the dominant uncertainties.

`firm|rough|guess` shrinkage and independent ±20/40/60% perturbations do not represent:

- wrong marker direction;
- shared expert optimism;
- correlated row errors;
- extractor false positives;
- uncertain conversational opportunity;
- misspecified dynamics.

Elicit distributions on logit-scale anchors and propagate joint parameter draws. Report a robust envelope or union of intervals across plausible models. Until external validation, label intervals `model-conditional support intervals`.

### Synthetic QA cannot calibrate human probabilities

Synthetic dossiers can test signs, routing, arithmetic, and obvious adversarial failures. They cannot validate real marker prevalence, dependence, extractor errors, base rates, or human coverage.

The memo says synthetic QA is “never calibration” but assigns M3 the job of tuning caps, confidence temperatures, evidence thresholds, and abstention rules. That is calibration to the synthetic generator.

Split M3 into:

- simulation-based calibration from the exact stated generative model, testing code only;
- held-out adversarial transcripts, testing extraction/routing only;
- no tuning of human probability or abstention claims.

### Coverage cannot be estimated at M4

If 8 of 10 independent cases fall inside an 80% interval, the exact 95% interval for actual coverage is approximately 44%–97%. For 4 of 5 it is approximately 28%–99.5%.

Pillars, rounds, and events do not create 40–80 independent observations because they share practitioner, dossier, rater, extractor, and parameters. Reliability bins, paired bootstraps, and risk–coverage curves will look more informative than they are.

Approximately 62 independent calibration units are needed merely to estimate 80% coverage within ±10 percentage points under ideal binomial assumptions; clustering requires more. That still would not identify the full hierarchy.

### Validation targets are partly endogenous

Probe selection depends on the posterior, and the probe may change subsequent behavior. Milestone corroboration is also scheduled in response to claims. Score only forecasts frozen before outcomes, under a logged evaluation policy.

Use fixed or randomized sentinel probes, include nonresponse and no-event windows, and use target-specific nulls. A stage prior is not a null forecast for a probe category or an event time.

H2 “movement plausibility” is a diagnostic, not a scored outcome. H2 needs a later independent rating or a prespecified observable trajectory.

Also add a direct no-stage predictor of future observables. Beating base-rate and persistence nulls does not show that the stage latent is useful if a recent-history model predicts equally well.

## What breaks under one expert rater

Surya defines the construct, supplies the prior/emission tables, approves their implications, and supplies the reference label. Agreement therefore partly measures successful encoding of Surya’s own rubric—incorporation bias—not independent validity.

Additional problems:

- Systematic rater bias cannot be separated from true stage.
- Interview and dossier ratings by the same person have correlated construct and threshold errors.
- Their difference is not a clean “information ceiling”; it also contains modality effects, memory, familiarity, and occasion noise.
- The proposed Hui–Walter analogy is inapplicable: the readings are dependent, ordinal, applied to one narrow population, and far too few.
- Mechanical adjacent-band smearing assumes symmetric, known, unbiased rater error and rewards diffuse forecasts.
- If M4 residuals cause table revision, the revised model has been trained on M4 and cannot be validated on it.
- A different LLM family is not a substitute for a second qualified domain rater.

Concrete fixes:

1. Rename the target `agreement with Surya’s dossier judgment`, not gold or truth.
2. Have Surya provide an explicit probability vector over bands and pillars.
3. Have every dossier rated twice after washout, in randomized blinded order, with duplicated anchor vignettes.
4. Obtain a second qualified rater on at least a subset.
5. Model rater thresholds/confusion only after repeated or multi-rater data exist.
6. Freeze extractor, tables, priors, and thresholds before M4.
7. Any post-M4 revision creates v2, evaluated only on future windows or people.
8. Explicitly exclude the gold interview transcript and instrument results from the estimator’s evaluation dossier.
9. Compare estimator disagreement with Surya against Surya’s own test–retest disagreement.

Drop weighted \(\kappa\) at this sample size and range. Report per-case RPS against Surya’s elicited distribution, signed band error, absolute band error, and all individual paired differences.

## Toy-example audit

Most displayed multiplication and normalization is correct. The example is nevertheless not an implementation-faithful fixture:

- After E2b, the milestone prior is approximately 0.196, not 0.175. The reported value ignores E2b.
- The bare-claim counterfactual gives band-5 mass approximately 0.538 from the actual pre-E4 nowcast, not 0.49.
- The toy band kernel assigns probability 0.002 to b3→b5 in a month. Starting from the top of b3 requires at least 11 forward point-jumps; with forward mean 0.9/month, that probability is approximately \(3.5\times10^{-9}\) before accounting for backward jumps.
- The toy multiplies all monthly evidence and then applies one kernel. The stated algorithm requires transitions between event timestamps; those operations do not commute.
- A flat starved-pillar message has exactly zero effect on the overall posterior.

Ship a machine-readable fixture containing every event timestamp, exact production-grid \(Q\), phase generator, priors, likelihood rows, and expected intermediate posterior.

## Better model classes

For \(N=5\)–10, the best replacement is simpler, not more elaborate.

### Immediate v1: coarse probe-anchored ordinal model

- Use 3–4 coarse states: 22–30, 31–40, 41–50, and 51+ manual-review territory.
- Estimate a static state over a 60–90-day window.
- Use standardized probes with explicit opportunities and full categorical outcomes as the measurement backbone.
- Add telemetry through explicit exposure/censoring models.
- Treat passive markers as episode-capped, low-weight composite evidence.
- Use one common wide prior.
- Omit archetype stage effects, personalized reliability, dynamic offsets, and phase gating.
- Integrate elicitation and extractor uncertainty into the output.

This model says less but can be audited and prospectively scored.

### Later: robust ordinal state-space model

A coherent longitudinal extension could use:

\[
z_{p,t}=z_{p,t-1}+\mu_p\Delta t+\eta_t,
\]

\[
u_{p,s}=\rho u_{p,s-1}+\xi_s,
\]

\[
z_{pk,t}=z_{p,t}+d_{pk},\qquad \sum_k d_{pk}=0.
\]

Here \(z\) is slow stage, \(u\) is a transient session effect with Student-\(t\), mixture, or semi-Markov innovations, and \(d_k\) are identified pillar offsets. Observations use joint opportunity-aware emissions, including extractor error. Persistent observations move \(z\) because the transient component cannot explain a long run—not because a KL threshold is released.

Model milestones as noisy observations of an interval-censored attainment time, not as a BKT posterior reinjected into stage.

Earn complexity prospectively:

1. target-specific base rate;
2. persistence/recent-history predictor;
3. episode-capped unit-weighted score;
4. probe-anchored coarse ordinal model;
5. add telemetry/exposure;
6. add dynamics;
7. add passive markers;
8. add phase;
9. add pillar coupling and archetypes.

A component should enter only if it improves held-out prequential score for most practitioners and does not worsen worst-case sensitivity or abstention behavior.

## The three changes with highest expected value

1. **Replace the positive-only marker multiplication and adaptive gates/caps with one opportunity-aware, episode-level observation model.** Update phase, stage, milestones, telemetry, and claim truth jointly; otherwise explicitly downgrade the product claim to a generalized-Bayes scorecard with non-nominal intervals.

2. **Collapse v1 to a coarse, mostly static, probe-anchored ordinal model with a common prior and propagated elicitation uncertainty.** Remove archetype stage effects, individualized reliability, dynamic pillar offsets, and phase gating until each earns its place through prospective ablation.

3. **Redesign M4 as a frozen feasibility and prospective-prediction study, not calibration.** Use repeated probabilistic Surya ratings, a second-rater subset, fixed/randomized sentinel outcomes, practitioner-level analysis, and future people or windows for every post-revision validation.
