# Lane 3 — Statistical Machinery for the Waypoint Estimator

**Status:** research-phase dossier v1 — 2026-07-16
**Feeds:** estimator design memo (deliverable 4), marker bank v0 representation (deliverable 2), probe bank scoring keys (deliverable 3), M4/M5 metrics.
**Inputs read:** `waypoint-practitioner-stage-assessment.md` (spec), `predictive-profile-harness-v3-design.md` (harness), `waypoint-research/inputs/fable-max-bayesian-profiling-eval.md` (binding principles), corpus digests 1–7.
**Citation policy:** every empirical/methodological work cited below was live-verified (Crossref / publisher / ACL Anthology / JMLR / NeurIPS proceedings / arXiv) on 2026-07-16 unless explicitly marked UNVERIFIED-BY-LIVE-CHECK (canonical works whose indexes didn't resolve; bibliographic details from training knowledge, flagged for the finder→verifier pass).

---

## 0. Operating point and binding constraints

The v1 estimator runs at an unusual corner of model space, and most of the literature surveyed below was built for a different corner. Fix the operating point first:

- **N = 5–10 practitioners** (team cohort), stage-range-narrow, map-contaminated.
- **Sparse, irregular longitudinal evidence:** conversation-derived marker events, occasional probes (frequency-capped), telemetry. Days-to-weeks between informative observations; gaps are informative in themselves.
- **Expert-set likelihoods initially** (research phase + Surya); refit only when labeled data exists (M6, ~2027).
- **Ordinal latent target:** 10 bands / 1–100 points, humans at 22–62, app users mostly 22–45; per-pillar profile (sensations/emotions/thoughts/awareness) + phase annotation.
- **Abstention is mandatory:** `data_sufficiency: insufficient` below threshold, never guess.
- **Offline pipeline** over exported dossiers; nowcast per session + monthly consolidation.
- **Single-rater gold anchor** (Surya), with interview + blinded-dossier double rating.

Binding principles carried in from the Fable-max review and spec §6 (these are constraints on the machinery, not preferences):

1. **LLM proposes, mechanical layer integrates.** The extractor emits typed marker events; a deterministic integrator owns the arithmetic and the scorecard. Never let the judge add up its own scorecard (Meehl's division of labour; see the review's Grove et al. discussion).
2. **Partial pooling:** population prior → practice-arc archetype → individual.
3. **Score every prediction at its own horizon against a base-rate null.** Never reward durability itself.
4. **Family separation:** marker-extractor and any validation labeler on different model families.
5. **Hierarchical error routing:** prediction errors spend themselves on fast state (phase) first; slow latents (stage) revise only under persistent runs of error.

The decisive consequence of the operating point: **at N=5–10, nothing with data-fitted structural parameters is estimable.** The v1 statistical question is not "which model fits best" but "which fixed-parameter scoring machine is auditable, calibratable in principle, and abstains honestly." Every model class below is judged against that.

---

## 1. Problem formalization

Per practitioner *p*:

- **Slow latent (stage):** θ_p(t) on an ordered grid. Recommended internal grid: **integer points 15–70** (bounded slightly beyond the practical human range 22–62 so mass isn't artificially clipped), with the 10-band posterior derived by summing grid mass. A 56-point grid costs nothing computationally, avoids band-edge artifacts, and answers the spec's granularity question empirically: report band + point + credible interval, and let observed CI widths tell you whether sub-band resolution in 22–45 is real (§6, Q1).
- **Pillar latents:** θ_p^{sens}, θ_p^{emo}, θ_p^{tho}, θ_p^{awa} on the same grid, coupled (§3.1, rank-1 coupling through the overall latent). The overall posterior is derived, not independently estimated.
- **Fast latent (phase):** φ_p(t) ∈ {glimpse, plateau, dip, integration, intensification, …} — a switching regime, not a position change.
- **Milestone latents:** m_j ∈ {not-reached, reached} per milestone (reverse_breathing, first_glimpse, small_death, …), with monotone reach-once semantics plus corroboration state.
- **Evidence stream:** marker events e_i = (marker_id, polarity, confidence, evidence_ref, t_i); probe responses (item_id, ordinal category, t); telemetry features (aggregated per window).

Everything the estimator consumes — marker, probe response, telemetry feature, milestone corroboration — reduces to the same mathematical object: **an emission likelihood column L(e | θ = g) over the grid** (possibly conditioned on phase and covariates). That unification is what makes the machinery thin.

---

## 2. Survey of candidate machinery

Each family: what it is → what it offers Waypoint → fit at the v1 operating point → risks → data requirements.

### 2.1 Item Response Theory (Rasch, Graded Response)

**What it is.** Latent-trait measurement: P(response category | θ, item parameters). Rasch model (Rasch 1960: *Probabilistic Models for Some Intelligence and Attainment Tests* — UNVERIFIED-BY-LIVE-CHECK for edition details; canonical) for dichotomous items with difficulty only; Samejima's graded response model (Samejima 1968/1969, *Estimation of Latent Ability Using a Response Pattern of Graded Scores*, Psychometrika Monograph / ETS RB, DOI 10.1007/bf03372160) for ordered polytomous responses with per-item discrimination *a* and ordered thresholds *b_k*; partial-credit (Masters 1982, Psychometrika, DOI 10.1007/bf02296272) and generalized partial credit (Muraki 1992, APM, DOI 10.1177/014662169201600206) as alternatives. Standard texts: Embretson & Reise 2000, *Item Response Theory for Psychologists* (UNVERIFIED-BY-LIVE-CHECK; canonical); De Ayala 2009, *The Theory and Practice of Item Response Theory* (UNVERIFIED-BY-LIVE-CHECK; canonical).

**What it offers.** The load-bearing insight is the **separation of item calibration from person scoring**. Operational CAT programs estimate a test-taker's θ from a handful of responses because item parameters were fixed *in advance*. Waypoint is in exactly that regime, with one substitution: item parameters come from **expert elicitation** instead of a 500-person calibration sample. A probe item = a GRM item whose thresholds live on the Wheel scale ("practitioners below ~35 typically answer in category 1–2; 35–45 in category 3; …"). Each scored probe response then contributes a likelihood column over the grid — the same object as a marker.

**Fit at v1.** As a **scoring grammar for the probe bank: strong.** As an estimation procedure: dead on arrival. Fitting even Rasch item parameters wants ~100+ persons (Linacre 1994, "Sample Size and Item Calibration Stability," *Rasch Measurement Transactions* 7(4):328 — UNVERIFIED-BY-LIVE-CHECK, newsletter piece, standard reference); GRM conventionally 500+ (De Ayala 2009). N=5–10 estimates nothing.

**Risks.** (a) Unidimensionality: the construct is explicitly 4-pillar; score probes against the pillar they target, not a common θ; full multidimensional IRT (Reckase) is data-hungry — avoid. (b) Expert-set thresholds will be miscalibrated somewhere; run the sensitivity analysis of §5.3 and re-elicit the flagged items. (c) Local independence violations: multiple probe responses inside one conversation are correlated; cap per-conversation evidence (§6, T2).

**Data to graduate.** Empirical item analysis (observed category × gold band tables) becomes possible per-item once the alpha cohort produces ≳50–100 labeled probe administrations; full GRM refit is a 2028+ prospect at best.

### 2.2 Computerized Adaptive Testing (probe scheduling)

**What it is.** Sequentially choose the next item to maximize information about θ. Bayesian roots: Owen 1975 (JASA, DOI 10.1080/01621459.1975.10479871). Item selection by Fisher information is standard but misleading early in a test when θ̂ is bad; Chang & Ying 1996 ("A Global Information Approach to Computerized Adaptive Testing," APM, DOI 10.1177/014662169602000303) argue for KL/global information integrated over the plausible θ range; van der Linden 1998 ("Bayesian Item Selection Criteria for Adaptive Testing," Psychometrika, DOI 10.1007/bf02294775) formalizes fully Bayesian criteria (e.g. minimum expected posterior variance). Exposure control: Sympson & Hetter 1985 (Proc. 27th Military Testing Association meeting — UNVERIFIED-BY-LIVE-CHECK; canonical) and the randomesque method of Kingsbury & Zara 1989 ("Procedures for Selecting Items for Computerized Adaptive Tests," *Applied Measurement in Education*, DOI 10.1207/s15324818ame0204_6).

**What it offers.** Waypoint is *permanently* in the early-test regime CAT theory worries about: the posterior never gets sharp, so **Bayesian/KL selection is the right rule, not Fisher-at-the-point-estimate**. Concretely: for each candidate probe, compute expected posterior entropy reduction under the current grid posterior (exact, cheap — sum over grid × response categories); rank; then filter through the delivery constraints. The delivery constraints in spec §5.2 (frequency caps, never verbatim, never in flagged-vulnerable moments) are exactly CAT's exposure-control and content-constraint machinery wearing product clothes — Sympson-Hetter caps ≈ frequency caps; randomesque selection (pick randomly among top-k informative items) is directly reusable to stop Wisdom hammering the single most informative probe.

**Fit at v1.** Strong, as a **ranking policy** the consolidation pass emits ("highest-value probes for this practitioner next month: …"). Not an autonomous loop in v1 — Wisdom delivers opportunistically; the scheduler just orders the menu.

**Risks.** Probe reactivity is outside CAT's frame (test items don't change the examinee; teaching probes do) — the frequency cap is the mitigation, and it means the scheduler is starved: with ~2 probes/week, expected-information ranking matters more, not less. Also beware selection-induced dependence: probes chosen *because* the posterior is ambiguous between bands X/Y will oversample the confusion region — fine for inference (likelihoods are conditioned on θ, selection on the posterior is ignorable), but it biases naive item-quality statistics later; log the posterior at selection time.

### 2.3 Bayesian Knowledge Tracing and learner modeling

**What it is.** BKT (Corbett & Anderson 1995, "Knowledge Tracing: Modeling the Acquisition of Procedural Knowledge," *UMUAI*, DOI 10.1007/bf01099821): per-skill 2-state HMM (unlearned→learned) with parameters P(L0), P(T), guess *g*, slip *s*. Individualized variants: Yudelson, Koedinger & Gordon 2013 ("Individualized Bayesian Knowledge Tracing Models," AIED 2013, LNCS 7926, DOI 10.1007/978-3-642-39112-5_18) — student-specific learn rates beat student-specific priors. Deep Knowledge Tracing (Piech et al. 2015, NeurIPS — UNVERIFIED-BY-LIVE-CHECK for page numbers; canonical) replaces the HMM with an RNN.

**What it offers.** Not the stage estimator — stage is not binary mastery. But **exactly the right shape for the milestone layer.** Each milestone (reverse_breathing, first_glimpse, small_death, nirodha-samāpatti-class events) is a 2-state latent with:
- **guess ≡ claimed-but-not-real** — vocabulary contamination, look-alike states (the cessation-vs-7th-jhana near-neighbors from the corpus), enthusiasm inflation;
- **slip ≡ real-but-not-reported** — spiritual-humility deflation, and the corpus's specific warning that cessations are systematically *under*-claimed (noticed only as a "reset," Laukkonen et al. 2023 digest).

The spec's `claimed | corroborated` milestone status is then two read-thresholds on one posterior: *claimed* = a claim-type marker fired; *corroborated* = P(reached) crosses a threshold after behavioral/after-effect evidence (post-cessation trait residue, gate telemetry) arrives. One warning: classical BKT assumes no forgetting; meditation regresses. Give milestone latents a small reversion rate where the domain says the capacity is trainable-and-losable (e.g. absorption depth), and none where it is ratchet-like (insight milestones per the canon).

**Fit at v1.** Strong, with all five parameters per milestone expert-set (they are natural-frequency questions: "of 100 practitioners who claim X without the gating scaffold, how many actually had it?"). Milestone posteriors then feed the stage filter as **soft ordering constraints** — the "no-teleporting" ladder (e.g. a 7-day-cessation claim without 8-jhana + ethical-transformation scaffolding defaults to look-alike/abstain, per the nirodha paper's gating ladder) becomes: P(milestone_j reached | θ) curves that are near-zero below the milestone's canonical band.

**Risks.** Milestone-marker circularity: don't let the same transcript sentence count once as a milestone claim and again as a stage marker (§6, T2). DKT and any fitted learner model: data-starved and opaque at this N — rejected for v1 and probably v2.

### 2.4 Latent-state models over sparse irregular longitudinal evidence

**What it is.** Discrete HMMs (Rabiner 1989, Proc. IEEE, DOI 10.1109/5.18626): forward filtering for online state estimates; forward-backward smoothing for retrospective ones. For irregular observation times, the continuous-time formulation: a generator matrix Q with transition kernel P(Δt) = exp(QΔt) — multi-state panel models (Jackson 2011, "Multi-State Models for Panel Data: The msm Package for R," *J. Statistical Software* 38(8), DOI 10.18637/jss.v038.i08) and CT-HMMs at scale (Liu, Li, Li, Song & Rehg 2015, "Efficient Learning of Continuous-Time Hidden Markov Models for Disease Progression," NeurIPS 2015 — verified in proceedings). Latent transition analysis (Collins & Lanza 2010, *Latent Class and Latent Transition Analysis*, Wiley — UNVERIFIED-BY-LIVE-CHECK; canonical) is the psychometric cousin. Continuous alternatives: linear-Gaussian state-space/DLM (Durbin & Koopman, *Time Series Analysis by State Space Methods* — UNVERIFIED-BY-LIVE-CHECK; canonical).

**What it offers — this is the v1 spine.** Three structural gifts:

1. **The spec's cadence is literally filter/smoother.** Rolling nowcast = forward filter (cheap, per-session); monthly consolidation = smoothing pass over the accumulated evidence (re-reads everything end-to-end, reconciles, writes canonical snapshot). The architecture decision in spec §6 falls out of Rabiner's two algorithms.
2. **Continuous-time transition kernel handles sparse irregularity honestly.** Structure Q as a birth–death chain on the ordered grid (nearest-neighbor moves only, slow rates, mild forward drift, nonzero regression rate). Then a practitioner who goes silent for six weeks gets exp(Q·6wk) applied — the posterior *diffuses*, uncertainty grows with silence, and long-gap practitioners fall below the abstention threshold automatically. No hand-built time-decay heuristic needed: the no-teleporting prior and the time-decay rule in spec §6 are both just Q.
3. **Multimodal posteriors for free.** The confusion pairs create *legitimately bimodal* states of knowledge — "either a genuine ~48 or a well-read ~30 speaking the vocabulary." A discrete grid posterior represents that bimodality exactly; any Gaussian filter (Kalman/DLM, and the literal HGF below) structurally cannot. This is the decisive argument for discrete-grid over continuous-state machinery, beyond mere convenience.

**Fit at v1.** Recommended core (§3, rank 1) — with **every structural parameter expert-set**: Q's rates from the canon's timeline material (stage dwell times: 2.1=120d, 2.2=90d, etc., plus Surya's judgment on plateau/regression rates), emissions from the marker bank. Exact inference by matrix arithmetic on a 56-point grid; the whole integrator is a few hundred lines of NumPy, fully auditable.

**Risks.** (a) First-order Markov assumption is wrong in detail (trajectory shape carries information — e.g. dark-night sequences); the phase variable (§2.5) absorbs the worst of it, and the consolidation LLM pass narrates what the Markov chain can't see. (b) Expert-set Q wrong → over/under-diffusion; sensitivity-check dwell times ±50% (§5.3). (c) **Fitting** transition matrices (LTA, standard Baum-Welch) needs hundreds of trajectories — rejected for v1; Q stays elicited until M6+.

### 2.5 Hierarchical Gaussian Filter

**What it is.** Mathys et al. 2011 ("A Bayesian Foundation for Individual Learning under Uncertainty," *Frontiers in Human Neuroscience*, DOI 10.3389/fnhum.2011.00039) and Mathys et al. 2014 ("Uncertainty in Perception and the Hierarchical Gaussian Filter," *Frontiers in Human Neuroscience*, DOI 10.3389/fnhum.2014.00825): a variational filter over stacked Gaussian random walks where each level's step size (volatility) is governed by the level above, yielding precision-weighted prediction-error updates with level-specific learning rates. Implementation: TAPAS toolbox (Frässle et al. 2021, *Frontiers in Psychiatry*, DOI 10.3389/fpsyt.2021.680811).

**What it offers.** The **principle** the spec already committed to: errors spend themselves on fast levels first; slow beliefs revise only under persistent error runs; volatility estimates gate learning rates. It is also the aesthetically satisfying choice (the estimator models the practitioner the way predictive processing says the practitioner models the world — the Laukkonen corpus's own frame).

**Fit at v1.** **Adopt the principle, not the filter.** Literal HGF assumes dense trial-by-trial input streams, Gaussian (unimodal) latents, and per-subject inversion with tuned priors — all three fail here (sparse irregular markers; bimodal confusion posteriors; N=5–10). The discrete cousin that keeps the principle: **phase-conditional dynamics** — when the phase layer detects a volatile regime (dip/intensification), (a) emission reliabilities are down-weighted (a dark-night week's despair-flavored markers shouldn't drag stage down — the spec's "a dark-night week is not regression"), and (b) the transition kernel's variance is temporarily raised (genuine reconfiguration is *more* likely near milestones — the punctuated-transition observation in the Vohryzek digest). That is HGF's volatility-gating, discretized.

**Risks.** Phase misclassification becomes an error *amplifier* (wrongly calling "dip" mutes real regression evidence). Mitigate: phase posteriors, not hard labels; and the safety duty never routes through the mute — teacher-flag logic reads raw markers, not phase-discounted ones.

**Data to graduate.** A literal continuous state-space/HGF layer becomes worth testing when practitioners have dense multi-year histories (alpha cohort, 2028+); it may never beat the grid filter and is not on the critical path.

### 2.6 Hierarchical / partial-pooling Bayes at very small N

**What it is.** Multilevel models shrinking unit estimates toward group means: Gelman & Hill 2007 (*Data Analysis Using Regression and Multilevel/Hierarchical Models* — UNVERIFIED-BY-LIVE-CHECK; canonical); priors for group-level variances when groups are few: Gelman 2006 (*Bayesian Analysis*, DOI 10.1214/06-ba117a — half-Cauchy family instead of inverse-gamma). Tooling: Stan (Carpenter et al. 2017, JSS — UNVERIFIED-BY-LIVE-CHECK; canonical), brms (Bürkner 2017, JSS 80(1), DOI 10.18637/jss.v080.i01).

**What it offers.** The binding three-level structure: population prior → archetype → individual. At v1 this enters as **structured elicited priors, not fitted random effects**:
- *Population level:* cold-start prior = wide distribution over 22–45 shaped by tenure base rates. Keep it shallow — the corpus is emphatic that expertise ≠ hours and progression is non-linear/inverted-U (many-to-(n)one digest), so tenure buys little sharpening.
- *Archetype level:* practice-path archetypes (breathwork-led vs love-led vs clarity-led vs awareness-led; map-exposed vs naive) shift pillar priors and marker-contamination handling, elicited from Surya as small deltas on the population prior.
- *Individual level:* the filter posterior. **Individuation = measured drift of the individual posterior away from its archetype prior** — which imports the harness's `N_wrong_profile_matched` control directly: run each dossier against its own prior and against a matched wrong archetype; if estimates don't separate, Waypoint is stereotyping, not measuring (§5.2).

**Fit at v1.** Necessary frame; nearly free to implement (it's just how priors are built). Actually *fitting* hierarchical variance components on 5–10 practitioners is at the ragged edge even with Gelman-2006-style priors — do it only in the M6 refit, and only with informative priors centered on the elicited values.

**Risks.** The team cohort is stage-range-narrow: pooling toward a narrow population can homogenize away the very differences being measured. Mitigation: report the individuation contrast (above) as a first-class metric, and keep the population prior honest about its provenance (team ≠ app population).

### 2.7 Ordinal structure and abstention

**Ordinal models.** The band scale is ordered; both model and metrics must respect adjacency. Cumulative-link (ordered logit/probit) models are the canonical regression form (Agresti 2010, *Analysis of Ordinal Categorical Data*, 2nd ed., Wiley — UNVERIFIED-BY-LIVE-CHECK; canonical); in Waypoint they appear implicitly — GRM *is* a cumulative-link model, and the grid filter's output is already a full ordinal distribution. No separate ordinal regression layer is needed at v1; the requirement bites in **metrics** (§5.1: ranked probability score, quadratic-weighted κ) and in **elicitation** (monotone/unimodal likelihood shapes, §4).

**Abstention.** Two literatures:
- *Reject option:* Chow 1970 ("On Optimum Recognition Error and Reject Tradeoff," *IEEE Trans. Information Theory*, DOI 10.1109/tit.1970.1054406) — optimal rejection thresholds posterior confidence; error–reject curves.
- *Selective prediction:* El-Yaniv & Wiener 2010 ("On the Foundations of Noise-free Selective Classification," *JMLR* 11:1605–1641) — the **risk–coverage curve** as the object to report: accuracy among non-abstained cases as a function of coverage. This should be an M5 headline artifact: Waypoint's value proposition is being *right when it speaks*, and the risk–coverage curve is the honest picture of that trade.

**Recommended abstention rule (v1)** — abstain (`data_sufficiency: insufficient`) when ANY of:
1. **Evidence mass:** total absorbed evidence weight < E_min bits, where each event contributes min(|log-LR|, cap) summed over the window (thresholds set in the design memo; the cap is §6 T2's anti-double-counting cap);
2. **Posterior width:** the 80% credible interval spans > 2 bands (equivalently, band-posterior entropy above threshold);
3. **Conflict:** the claims-vs-behavior consistency check fires above threshold — and this must surface as a *distinct* value (`conflicted`, not `insufficient`): averaging a contradiction into a midpoint estimate is the one behavior worse than silence.
Per-pillar abstention is independent (a practitioner can be estimable on awareness and unestimable on emotions).

**Conformal prediction** (Angelopoulos & Bates 2021, "A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification," arXiv:2107.07511): distribution-free prediction *sets* with coverage guarantees — the principled upgrade to threshold abstention, and ordinal adaptive prediction sets fit the band scale nicely. But guarantees need a calibration set; at fewer than ~20–30 labeled dossiers the coverage statement is vacuous. Park for M6+; design the snapshot schema so a credible *set* (not just interval) can be emitted later.

### 2.8 Uncertainty calibration and proper scoring

**Proper scoring rules.** Gneiting & Raftery 2007 ("Strictly Proper Scoring Rules, Prediction, and Estimation," *JASA* 102:359–378 — verified via DTIC preprints; JASA version canonical) is the frame; Brier 1950 (UNVERIFIED-BY-LIVE-CHECK; canonical) and log score the workhorses. For **ordinal** targets the right headline is the **Ranked Probability Score** (Epstein 1969, "A Scoring System for Probability Forecasts of Ranked Categories," *J. Applied Meteorology* 8(6):985–987, DOI 10.1175/1520-0450(1969)008<0985:ASSFPF>2.0.CO;2; note by Murphy 1971, *J. Applied Meteorology* 10:155–156, DOI 10.1175/1520-0450(1971)010<0155:anotrp>2.0.co;2): it penalizes a band-4 call less than a band-9 call when truth is band 5. Log-loss stays as the secondary (it is the bits-saved currency and punishes confident wrongness hardest).

**Nulls (binding: score at horizon vs base rate).** Two nulls, both cheap:
1. **Base-rate null:** the tenure-conditioned population prior — what you'd say knowing only "practitioner, 8 months in."
2. **Persistence null:** last consolidated snapshot carried forward — meaningful from the second consolidation on; this is the null that keeps the monthly pass honest ("did re-reading everything actually add information?").
Report **bits saved per snapshot vs each null**, per practitioner, with paired comparisons only (at this N, aggregate means are noise; the harness's paired-bootstrap discipline applies).

**Calibration at tiny N.** ECE and reliability diagrams (Naeini, Cooper & Hauskrecht 2015, "Obtaining Well Calibrated Probabilities Using Bayesian Binning," AAAI, DOI 10.1609/aaai.v29i1.9602; Guo et al. 2017, "On Calibration of Modern Neural Networks," ICML, PMLR 70:1321–1330) are meaningless with 5–10 gold labels — bins go empty. Honest v1 reporting:
- **Coverage counts, not curves:** "the 80% CI contained Surya's band in k of N subjects," raw counts with a binomial interval;
- **Pool prediction *events*:** pillar × subject × consolidation round gives ~40–80 scoreable events from a 10-person cohort over a few months — enough for a coarse 3-bin reliability check;
- **Synthetic-dossier QA measures machinery calibration** at arbitrary n (does the integrator's 80% mean 80% *when the marker bank's assumptions hold*) — necessary, and never sufficient, since it can't test the assumptions themselves (spec §7's QA-only stance is exactly right).

**Extractor-confidence caveat.** LLM-emitted confidences are verbal and miscalibrated; treat extractor confidence as an *ordinal feature* mapped through an elicited weighting (e.g. high/med/low → likelihood tempering exponents), never as a probability (harness §11.3's rule, applied to the extractor).

### 2.9 Label noise and single-rater models

**What exists.** Dawid & Skene 1979 ("Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm," *Applied Statistics* 28(1):20–28, DOI 10.2307/2346806): per-rater confusion matrices + latent true labels, from *multiple raters per item*. Successors: MACE (Hovy, Berg-Kirkpatrick, Vaswani & Hovy 2013, "Learning Whom to Trust with MACE," NAACL-HLT 2013:1120–1130), Raykar et al. 2010 ("Learning From Crowds," *JMLR* 11:1297–1322), the item-level Bayesian treatment of Passonneau & Carpenter 2014 ("The Benefits of a Model of Annotation," *TACL* 2, DOI 10.1162/tacl_a_00185), and the comparison in Paun et al. 2018 ("Comparing Bayesian Models of Annotation," *TACL* 6, DOI 10.1162/tacl_a_00040).

**The v1 problem: one rater.** With a single anchor rater, per-rater error rates are unidentifiable from labels alone. What transfers:
1. **Model-the-rater mindset.** Surya's label is an instrument reading, not truth. Represent it with an elicited confusion structure — minimally "adjacent-band tolerance" (probability the gold band is off-by-one), elicited from Surya about himself per band region (self-assessed difficulty is higher in 22–35 where phenomenology is thin). The calibration metrics of §5 then score against a *smeared* gold, not a delta function.
2. **The interview-vs-dossier double rating is two instruments, not two raters.** Same rater, different information channels: the delta estimates the *channel's* information ceiling (the spec's headline number), not rater noise. Keep both labels forever; record divergences with both positions (house convention).
3. **Hui & Walter 1980** ("Estimating the Error Rates of Diagnostic Tests," *Biometrics* 36:167–171, DOI 10.2307/2530508): with ≥2 imperfect tests on the same subjects (and some structure), error rates are estimable *without any gold standard*. Waypoint's triangulation — interview rating, blinded dossier rating, instrument battery, estimator — is a 4-test Hui-Walter design in embryo. Underpowered at N=5–10, but **log all four per subject now**, in analysis-ready per-subject rows, and the latent-class analysis becomes runnable the day the alpha cohort lands. This is the statistical formalization of spec decision #3 ("triangulated calibration") and the correct answer to the single-rater validity threat: the fix is more *instruments*, not more raters, until a second qualified rater exists.
4. **Family separation** (binding): the validation labeler shares no model family with the extractor; nothing in this section works if labels and evidence share failure modes.

### 2.10 Measurement invariance across practitioners

**What it is.** A measure is invariant when the same latent level produces the same response distribution regardless of group membership (Meredith 1993, "Measurement Invariance, Factor Analysis and Factorial Invariance," *Psychometrika* 58:525–543, DOI 10.1007/bf02294825; Millsap 2011, *Statistical Approaches to Measurement Invariance*, Routledge — UNVERIFIED-BY-LIVE-CHECK; canonical). Item-level violations = differential item functioning (Holland & Wainer 1993, *Differential Item Functioning*, Erlbaum — UNVERIFIED-BY-LIVE-CHECK; canonical).

**Where it bites Waypoint.** Four concrete DIF exposures:
- **Map exposure** (the big one): a dereification phrase is strong evidence from a naive practitioner and weak evidence from someone who has read the Wheel — the identical marker has different likelihoods by group. This is reading-ahead contamination *stated as DIF*, which makes the mitigation precise: **condition marker likelihoods on an exposure covariate** (two columns, naive/exposed) wherever Surya judges the marker learnable; the beautiful-loop digest says exactly which family that is (insight vocabulary).
- **Tradition/vocabulary background:** a Zen-reading practitioner and a Neidan-trained one describe the same territory in non-overlapping registers; marker bank needs register-diverse indicators per construct (lane 1/2 handoff).
- **Verbal fluency / language:** articulate ≠ advanced; structure-scoring over content-scoring (Cook-Greuter precedent, lane 2) is the design-level counter.
- **Pillar path:** breathwork-path practitioners emit somatic markers at higher base rates for path reasons, not stage reasons — pillar-conditional base rates in the elicitation (§4).

**Fit at v1.** No statistical invariance *testing* is possible at N=5–10. Handle by design: contamination-susceptibility tags + conditional columns in the marker bank now; formal DIF analysis is an M6 deliverable once alpha-cohort N supports it. The team cohort being 100% map-exposed is actually a small mercy: v1 calibration only exercises the "exposed" column, and says nothing about naive users — flag that explicitly in the M5 report.

---

## 3. Ranked model classes for v1

| # | Model class | v1 role | Params fit from data at v1 | Data to graduate | Main failure mode |
|---|-------------|---------|---------------------------|------------------|-------------------|
| 1 | **Grid-state Bayes filter** (ordinal HMM, continuous-time sticky kernel, expert-set emissions) | Core integrator: nowcast = filter, consolidation = smoother | **None** | Labeled trajectories → refit Q, emissions (M6) | Miscalibrated elicited likelihoods; correlated-evidence overconfidence |
| 2 | **Milestone BKT layer** (per-milestone 2-state latent, guess/slip expert-set, optional reversion) | Milestone posteriors; claimed/corroborated; ordering constraints into #1 | None | ~50+ labeled milestone claims | Claim/marker double counting; wrong ratchet assumptions |
| 3 | **GRM-structured probe scoring + Bayesian (KL/expected-entropy) probe scheduler with exposure caps** | Probe bank scoring keys; per-practitioner probe rankings from consolidation | None | 50–100 labeled probe administrations per item region | Expert thresholds wrong; local dependence within conversations |
| 4 | **Three-level partial-pooling prior** (population → archetype → individual), elicited | Cold start; archetype deltas; individuation contrast | None (fit at M6 with informative priors) | Alpha cohort (~2027) | Homogenization on a narrow cohort; stereotype lift |
| 5 | **Phase-switching emission regimes** (HGF-inspired volatility gating, discretized) | Phase layer: down-weight emissions + widen kernel in volatile regimes | None (phase classified by LLM, gates expert-set) | Dense longitudinal histories | Phase misclassification muting true regression |
| — | *Rejected for v1:* fitted GRM/2PL/MIRT; fitted HMM/LTA transitions; DKT/learned embeddings; literal HGF inversion; conformal sets; Dawid-Skene multi-rater; Kalman/DLM continuous state | — | — | see §2 | data-starved, unidentifiable, or structurally unimodal |

**The composite recommendation** (what the estimator design memo should specify): one grid filter per pillar plus a coupling to the overall latent; milestones as side-chains feeding ordering likelihoods; probes and markers entering identically as likelihood columns; phase gating both emissions and kernel; priors from the pooled hierarchy; abstention per §2.7; scoring per §5. Total v1 statistical surface: **zero fitted parameters, one afternoon of NumPy, every number in it signed by a human.** That last property is the point — v1 is instrument-building, and every disagreement with Surya's gold labels must be traceable to a specific elicited number someone can inspect and revise.

Pillar coupling, concretely: model overall θ as the latent driver and pillar offsets δ_k as slowly-varying deviations with an elicited cross-pillar spread (Surya's judgment of how uneven development typically is, e.g. "±1 band typical, 2+ bands notable"). Pillar evidence updates its pillar posterior directly and the overall posterior through the coupling; the spec's "2.4-deep in awareness, untrained in emotional integration" practitioner is then representable without letting one loud pillar drag the whole profile.

## 4. Representing marker likelihoods for a non-statistician teacher

**Requirement:** Surya must be able to set, read, and veto every number; the mechanical layer must be able to consume them exactly; the format must expose miscalibration to inspection. Log-likelihood ratios, logits, and Gaussians all fail the first clause. The evidence-backed answer is **natural frequencies** (Gigerenzer & Hoffrage 1995, "How to Improve Bayesian Reasoning Without Instruction: Frequency Formats," *Psychological Review* 102(4):684–704, DOI 10.1037/0033-295x.102.4.684 — frequency formats roughly double correct Bayesian reasoning in experts and novices alike), elicited under a SHELF-style protocol (Gosling 2018, "SHELF: The Sheffield Elicitation Framework," in *Elicitation*, Springer ISOR 261, DOI 10.1007/978-3-319-65052-4_4; background: O'Hagan et al. 2006, *Uncertain Judgements: Eliciting Experts' Probabilities*, Wiley — UNVERIFIED-BY-LIVE-CHECK; canonical; Cooke 1991, *Experts in Uncertainty*, OUP — UNVERIFIED-BY-LIVE-CHECK; canonical).

**The elicitation unit.** One marker = one question Surya answers a handful of times:

> "Out of 100 practitioners sitting in [band range], how many would show *this* in a typical month of conversations with Wisdom?"

- Anchors at **five coarse regions** (22–30, 31–40, 41–50, 51–60, 61+), not ten bands and never 56 grid points — the machine interpolates.
- Surya picks a **shape template** first, then fills anchors: `ramp-up` (rises with stage), `ramp-down`, `bump` (peaks in a region — most phase-linked and transition markers), `gate` (near-zero until a threshold band — milestone-linked). Templates enforce the monotone/unimodal structure ordinal evidence should have and cut the elicitation to shape + 2–3 numbers.
- **Confidence tag per row:** `firm / rough / guess` → mechanical shrinkage of the row toward flat (a `guess` marker can never move the posterior more than ~0.5 bits; a `firm` one up to the global cap). This is Dirichlet-style tempering the teacher never needs to see.
- **Contamination column where flagged:** markers tagged reading-learnable get two rows (naive / map-exposed); pillar-path-sensitive markers get base-rate adjustments per path.

**The review affordance (as important as the format).** Elicitation research is unanimous that experts are better at criticizing implications than at stating parameters — so after each batch, show Surya the numbers *inverted*:
1. **Implication cards:** "A practitioner your table puts at 35 shows markers A, B, C this month → the estimate moves to 41 ± one band. Feel right?"
2. **Confusion-pair audits:** for each of the five spec confusion pairs, a side-by-side of what the current table does to the look-alike (e.g. "a dissociating practitioner showing flat-affect markers would currently read as equanimity at 43 — here's the discriminating marker whose absence should block that"). Discrimination failures found here are missing markers, not bad numbers — route back to the marker bank.
3. **Synthetic-vignette replay:** run the elicited bank over 5–10 short vignettes at known positions; show the posterior per vignette; sign-off happens on outcomes, not on tables.

**Storage format:** versioned YAML in the marker bank, one block per marker, provenance mandatory:

```yaml
marker: dereification_language          # "anger arose and passed" vs "I am angry"
pillar: thoughts
polarity: positive
shape: ramp-up
freq_per_100_month:                     # naive practitioners
  22-30: 4
  31-40: 15
  41-50: 45
  51-60: 75
  61+: 85
freq_per_100_month_map_exposed:         # reading-learnable → second row
  22-30: 20
  31-40: 30
  41-50: 55
  51-60: 80
  61+: 88
confidence: rough                       # firm | rough | guess → shrinkage
contamination: high                     # lane-5 tag; forces the second row
source: "Laukkonen 2025 beautiful-loop §dereification; Surya elicitation 2026-08-xx"
set_by: surya
reviewed: pending
max_bits: 1.5                           # derived, displayed, capped
```

The integrator normalizes rows into emission columns and derives log-LRs mechanically; `max_bits` is *displayed* so reviewers see each marker's maximum pull. Changes are diffs in git; every disagreement between estimator and gold label at M4 traces back to specific rows.

**Elicitation protocol sketch (M2 input):** batches of 10–15 markers per session (fatigue is the known SHELF failure mode); definition → shape → anchors → immediate implication cards → tag → next. Budget ~4–6 sessions for a v0 bank of ~60 markers.

## 5. Scoring and validation specifics (feeds M4/M5)

### 5.1 Metric suite
- **Primary: Ranked Probability Score** over the 10-band posterior vs gold band (Epstein 1969) — ordinal-aware, proper.
- **Secondary: log-loss** (bits currency) and Brier; **band hit / adjacent-band hit** as the legible headline pair.
- **Pillar agreement:** quadratic-weighted κ vs Surya's pillar ratings (Cohen 1968, "Weighted Kappa," *Psychological Bulletin* 70(4):213–220, DOI 10.1037/h0026256).
- **Calibration:** 80% CI coverage counts with binomial intervals; 3-bin reliability over pooled prediction events (§2.8).
- **Skill:** bits saved vs tenure-base-rate null and vs persistence null, per practitioner, paired.
- **Abstention:** risk–coverage curve (El-Yaniv & Wiener 2010) + abstention rate by reason (insufficient / conflicted).
- **Milestones:** precision/recall of `corroborated` calls vs Surya's milestone judgments, with the under-claiming asymmetry reported (missed real cessations vs endorsed look-alikes are different products; count separately).

### 5.2 Controls (harness imports)
- **Matched-wrong-dossier:** estimator on practitioner A's evidence scored against B's gold (matched on tenure/archetype). If real dossiers don't beat swapped ones, Waypoint is measuring the archetype, not the person.
- **Shuffled-time control:** same markers, permuted timestamps — tests whether the temporal machinery (kernel, phase gating) earns anything beyond a bag-of-markers naive-Bayes readout.
- **Family separation** enforced end-to-end: extractor family ≠ validation-labeler family; synthetic personas (below) authored by a third family.

### 5.3 Synthetic-dossier QA (M3 gate) and sensitivity analysis
- Personas authored at known positions **by a different model family from the extractor**, including the adversarial set the confusion table implies: vocabulary contaminator at 28 talking like 48; humble under-claimer at 45; dark-night at 38 vs clinical-depression-flavored at 25; dissociative flatness vs equanimity; absorption vs dullness. Pass criterion: correct ranking + abstain-or-separate on each adversarial pair. This QAs the instrument, never calibrates it.
- **Sensitivity analysis is v1's substitute for model fitting:** perturb every elicited frequency within its confidence-tag band, re-run all dossiers, report which conclusions flip. Markers whose perturbation flips a band assignment are the priority re-elicitation list for M2's second pass. (Cheap: the whole pipeline is a deterministic matrix computation.)

## 6. Traps, and answers to the spec's open questions

**T1 — Don't reward durability.** Score each snapshot at its own horizon vs the nulls; a stage estimate that "stays right" mostly restates that stage is slow. (Binding; Fable-max review §2.)

**T2 — Correlated evidence is the silent killer.** Naive Bayes over marker events assumes independence; a practitioner who mentions dissolving self-boundaries five times in one enthusiastic session is one observation, not five. Mitigations, all three: (a) per-marker, per-window firing caps (a marker counts once per consolidation window at full weight, repeats at steep discount); (b) per-session total-evidence cap in bits; (c) the consolidation LLM pass explicitly merges repeated evidence of the same underlying event before the smoother runs — this is a *statistical* duty of the monthly pass, not just narrative hygiene. Milestone claims and their supporting markers are deduplicated against each other (§2.3).

**T3 — The extractor never touches the arithmetic.** Marker emission is the LLM's job; likelihoods, caps, posteriors, and scores live in the deterministic layer. Extractor confidence enters only through the elicited tempering map (§2.8).

**T4 — Tenure priors stay weak.** Expertise ≠ hours; inverted-U trajectories (corpus digest 1). The tenure-conditioned prior should be visibly wide; if the estimator's skill comes mostly from the prior, bits-saved vs the base-rate null will expose it — that's what the null is for.

**Spec open question 1 (granularity):** run the grid at 1-point resolution internally, report band + point + CI, and let M4's observed CI widths answer empirically. Expectation from this lane: bands are defensible; sub-band point estimates in 22–45 will carry ±4–6 point CIs at best given evidence sparsity — report points, but never let a consumer read them as finer than the CI.

**Spec open question 2 (instrument battery slots):** lane 2's call on content; this lane's constraint: prefer instruments with published ordinal scoring and any measurement-invariance evidence, because battery scores enter the Hui-Walter triangulation (§2.9) as a third imperfect test — an instrument that can't be treated as a noisy ordinal reading of the same latent wastes its slot.

**Spec open question 4 (flag thresholds):** the statistical form should be a posterior-probability trigger on (band ≥ Unbinding-adjacent) ∧ (phase ∈ {dip, intensification}), thresholded for high recall at tolerable false-flag rate — a rare-event detection problem where PR-thinking, not accuracy, applies (harness §9.3). Never phase-discounted (§2.5): the safety path reads raw evidence.

---

## References (verified unless marked)

- Agresti, A. (2010). *Analysis of Ordinal Categorical Data*, 2nd ed. Wiley. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Angelopoulos, A. N. & Bates, S. (2021). A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. arXiv:2107.07511.
- Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. *Monthly Weather Review* 78:1–3. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Bürkner, P.-C. (2017). brms: An R Package for Bayesian Multilevel Models Using Stan. *J. Statistical Software* 80(1). DOI 10.18637/jss.v080.i01.
- Carpenter, B. et al. (2017). Stan: A Probabilistic Programming Language. *J. Statistical Software* 76(1). UNVERIFIED-BY-LIVE-CHECK (canonical).
- Chang, H.-H. & Ying, Z. (1996). A Global Information Approach to Computerized Adaptive Testing. *Applied Psychological Measurement* 20(3):213–229. DOI 10.1177/014662169602000303.
- Chow, C. K. (1970). On Optimum Recognition Error and Reject Tradeoff. *IEEE Trans. Information Theory* 16(1):41–46. DOI 10.1109/tit.1970.1054406.
- Cohen, J. (1968). Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. *Psychological Bulletin* 70(4):213–220. DOI 10.1037/h0026256.
- Collins, L. M. & Lanza, S. T. (2010). *Latent Class and Latent Transition Analysis*. Wiley. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Cooke, R. M. (1991). *Experts in Uncertainty: Opinion and Subjective Probability in Science*. Oxford University Press. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Corbett, A. T. & Anderson, J. R. (1995). Knowledge Tracing: Modeling the Acquisition of Procedural Knowledge. *User Modelling and User-Adapted Interaction* 4:253–278. DOI 10.1007/bf01099821.
- Dawid, A. P. & Skene, A. M. (1979). Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm. *Applied Statistics* 28(1):20–28. DOI 10.2307/2346806.
- De Ayala, R. J. (2009). *The Theory and Practice of Item Response Theory*. Guilford. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Durbin, J. & Koopman, S. J. (2012). *Time Series Analysis by State Space Methods*, 2nd ed. OUP. UNVERIFIED-BY-LIVE-CHECK (canonical).
- El-Yaniv, R. & Wiener, Y. (2010). On the Foundations of Noise-free Selective Classification. *JMLR* 11:1605–1641.
- Embretson, S. E. & Reise, S. P. (2000). *Item Response Theory for Psychologists*. Erlbaum. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Epstein, E. S. (1969). A Scoring System for Probability Forecasts of Ranked Categories. *J. Applied Meteorology* 8(6):985–987. DOI 10.1175/1520-0450(1969)008<0985:ASSFPF>2.0.CO;2.
- Frässle, S. et al. (2021). TAPAS: An Open-Source Software Package for Translational Neuromodeling and Computational Psychiatry. *Frontiers in Psychiatry* 12:680811. DOI 10.3389/fpsyt.2021.680811.
- Gelman, A. (2006). Prior distributions for variance parameters in hierarchical models. *Bayesian Analysis* 1(3):515–534. DOI 10.1214/06-ba117a.
- Gelman, A. & Hill, J. (2007). *Data Analysis Using Regression and Multilevel/Hierarchical Models*. CUP. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Gigerenzer, G. & Hoffrage, U. (1995). How to Improve Bayesian Reasoning Without Instruction: Frequency Formats. *Psychological Review* 102(4):684–704. DOI 10.1037/0033-295x.102.4.684.
- Gneiting, T. & Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. *JASA* 102(477):359–378.
- Gosling, J. P. (2018). SHELF: The Sheffield Elicitation Framework. In *Elicitation* (Springer ISOR 261). DOI 10.1007/978-3-319-65052-4_4.
- Guo, C., Pleiss, G., Sun, Y. & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. ICML, PMLR 70:1321–1330.
- Holland, P. W. & Wainer, H., eds. (1993). *Differential Item Functioning*. Erlbaum. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Hovy, D., Berg-Kirkpatrick, T., Vaswani, A. & Hovy, E. (2013). Learning Whom to Trust with MACE. NAACL-HLT 2013:1120–1130.
- Hui, S. L. & Walter, S. D. (1980). Estimating the Error Rates of Diagnostic Tests. *Biometrics* 36:167–171. DOI 10.2307/2530508.
- Jackson, C. H. (2011). Multi-State Models for Panel Data: The msm Package for R. *J. Statistical Software* 38(8). DOI 10.18637/jss.v038.i08.
- Kingsbury, G. G. & Zara, A. R. (1989). Procedures for Selecting Items for Computerized Adaptive Tests. *Applied Measurement in Education* 2(4):359–375. DOI 10.1207/s15324818ame0204_6.
- Linacre, J. M. (1994). Sample Size and Item Calibration Stability. *Rasch Measurement Transactions* 7(4):328. UNVERIFIED-BY-LIVE-CHECK (standard reference, newsletter).
- Liu, Y.-Y., Li, S., Li, F., Song, L. & Rehg, J. M. (2015). Efficient Learning of Continuous-Time Hidden Markov Models for Disease Progression. NeurIPS 2015 (verified in proceedings index).
- Masters, G. N. (1982). A Rasch Model for Partial Credit Scoring. *Psychometrika* 47:149–174. DOI 10.1007/bf02296272.
- Mathys, C. et al. (2011). A Bayesian Foundation for Individual Learning under Uncertainty. *Frontiers in Human Neuroscience* 5:39. DOI 10.3389/fnhum.2011.00039.
- Mathys, C. D. et al. (2014). Uncertainty in Perception and the Hierarchical Gaussian Filter. *Frontiers in Human Neuroscience* 8:825. DOI 10.3389/fnhum.2014.00825.
- Meredith, W. (1993). Measurement Invariance, Factor Analysis and Factorial Invariance. *Psychometrika* 58:525–543. DOI 10.1007/bf02294825.
- Millsap, R. E. (2011). *Statistical Approaches to Measurement Invariance*. Routledge. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Muraki, E. (1992). A Generalized Partial Credit Model: Application of an EM Algorithm. *Applied Psychological Measurement* 16(2):159–176. DOI 10.1177/014662169201600206.
- Murphy, A. H. (1971). A Note on the Ranked Probability Score. *J. Applied Meteorology* 10:155–156. DOI 10.1175/1520-0450(1971)010<0155:anotrp>2.0.co;2.
- Naeini, M. P., Cooper, G. F. & Hauskrecht, M. (2015). Obtaining Well Calibrated Probabilities Using Bayesian Binning. AAAI. DOI 10.1609/aaai.v29i1.9602.
- O'Hagan, A. et al. (2006). *Uncertain Judgements: Eliciting Experts' Probabilities*. Wiley. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Owen, R. J. (1975). A Bayesian Sequential Procedure for Quantal Response in the Context of Adaptive Mental Testing. *JASA* 70:351–356. DOI 10.1080/01621459.1975.10479871.
- Passonneau, R. J. & Carpenter, B. (2014). The Benefits of a Model of Annotation. *TACL* 2:311–326. DOI 10.1162/tacl_a_00185.
- Paun, S., Carpenter, B., Chamberlain, J., Hovy, D., Kruschwitz, U. & Poesio, M. (2018). Comparing Bayesian Models of Annotation. *TACL* 6:571–585. DOI 10.1162/tacl_a_00040.
- Piech, C. et al. (2015). Deep Knowledge Tracing. NeurIPS 2015. UNVERIFIED-BY-LIVE-CHECK for pages (canonical).
- Rabiner, L. R. (1989). A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition. *Proceedings of the IEEE* 77(2):257–286. DOI 10.1109/5.18626.
- Rasch, G. (1960). *Probabilistic Models for Some Intelligence and Attainment Tests*. Danish Institute for Educational Research. UNVERIFIED-BY-LIVE-CHECK (canonical).
- Raykar, V. C. et al. (2010). Learning From Crowds. *JMLR* 11:1297–1322.
- Samejima, F. (1968/1969). Estimation of Latent Ability Using a Response Pattern of Graded Scores. *Psychometrika Monograph Supplement 17* / ETS RB. DOI 10.1007/bf03372160.
- Sympson, J. B. & Hetter, R. D. (1985). Controlling item-exposure rates in computerized adaptive testing. Proc. 27th Annual Meeting, Military Testing Association. UNVERIFIED-BY-LIVE-CHECK (canonical, gray literature).
- van der Linden, W. J. (1998). Bayesian Item Selection Criteria for Adaptive Testing. *Psychometrika* 63:201–216. DOI 10.1007/bf02294775.
- Wang, X., Berger, J. O. & Burdick, D. S. (2013). Bayesian analysis of dynamic item response models in educational testing. *Annals of Applied Statistics* 7(1):126–153. DOI 10.1214/12-aoas608. (Dynamic-IRT existence proof; data-hungry, not for v1.)
- Yudelson, M. V., Koedinger, K. R. & Gordon, G. J. (2013). Individualized Bayesian Knowledge Tracing Models. AIED 2013, LNCS 7926:171–180. DOI 10.1007/978-3-642-39112-5_18.
