# Waypoint — Literature Dossier (research deliverable 1 of 4)

**Status:** research-phase deliverable, v1 — compiled 2026-07-16.
**Question it answers:** *How is stage on the contemplative path actually judged — by the traditions that invented the maps, by contemporary meditation science, and by the psychometrics of scoring humans from language?* Filtered for what Waypoint can use; every claim source-cited.
**Inputs synthesized:** spec `/Users/fionnenglish/LIFE/specs/waypoint-practitioner-stage-assessment.md`; corpus notes (`corpus/internal-canon.md` + 6 seed-paper notes); lane dossiers 1 (traditions), 2 (psychometrics), 4 (safety), 6 (citation map) read in full; lane 3 (statistics) and lane 5 (Goodhart) via their delivered digests (full detail in `lanes/3-statistical-machinery.md`, `lanes/5-goodhart-contamination.md`).
**Reader:** technically fluent, has read the spec.

**Verification convention (carried, not laundered):** `[VERIFIED]` = citation metadata confirmed live in an upstream lane this pass. `[UNVERIFIED]` = real-and-standard but not re-confirmed, or safeguard-flagged — treat as a lead. Speculation is labelled inline. Where a lane flagged a numeric as unre-extractable, that flag is preserved here.

**Confidence scheme for §1** (deliberately strict about what counts as *independent*). These are structured expert confidence grades over the corpus this research pass assembled — not systematic-review evidence grades (no exhaustive search or formal risk-of-bias assessment behind them; external-review caveat, 2026-07-16):
- **[C5] Very high** — independent convergence across ≥3 source families *and* a stated mechanism.
- **[C4] High** — 2 independent families, or 1 family + strong replication.
- **[C3] Moderate** — theory-strong but empirically thin, or single-family.
- **[C2] Low** — suggestive, contested, or small-N.

The four source families are treated as independent: **(T)** contemplative traditions (9 lineages, lane 1); **(P)** developmental/mindfulness psychometrics (Loevinger→STAGES; FFMQ/NADA validity work, lane 2); **(C)** clinical/adverse-events science (Britton/Lindahl/Goldberg, lane 4); **(N)** the Laukkonen predictive-processing network (6 seed papers, lane 6). **Binding caveat (lane 6):** family **N** is *one author-network* — Laukkonen, Slagter, Sacchet, Lutz, Metzinger, Friston, Sparby recur as authors *and* as each other's authorities. N-only agreement is internally consistent, not corroboration, and is capped at **[C3]**.

---

## 1. CONVERGENT FINDINGS — what independently marks progression

Ranked by confidence. These are the load-bearing design facts; the recurring pattern is that all four families independently reinvented the same handful of moves.

### 1.1 Score the *structure* of a report, never the doctrinal content — **[C5]**
The single most-replicated finding in the whole corpus. Every durable tradition scores *how* an experience is reported and *what the practitioner does next*, never whether the words are doctrinally correct, precisely because the words are the faked layer (T: Mahāsi three-phase report; Zen *jakugo* capping-phrase; Dzogchen grasped-vs-released — lane 1 §§1,3,4). The ego-development lineage independently formalized this as scoring language *structure* against a rubric (P: Loevinger WUSCT → O'Fallon STAGES three-dimension rule, `O'Fallon et al. 2020, Heliyon 6(3):e03472` [VERIFIED]). Cognitive science supplies the mechanism: felt insight / noetic certainty is *inducible and can make false content feel true* (N: `Grimmer, Laukkonen et al. 2022, Psychon. Bull. Rev. 29(3):954–970` [VERIFIED]; `Laukkonen et al. 2020, "dark side of Eureka," Cognition 196:104122`; plus general cognitive science outside the four defined families: `Nisbett & Wilson 1977, Psychol. Rev. 84(3):231–259` — retagged from N, verifier fix), so conviction is not evidence. And the decisive asymmetry: on sentence-completion, faking *up* a structural level mostly fails while faking *down* is easy (P: Redmore 1976, via lane 5 digest) — structure-scoring is not just principled, it is empirically hard to game upward. **→ Waypoint's core bet (LLM extracts structure, deterministic layer aggregates) is the best-evidenced choice in the dossier.**

### 1.2 State ≠ station: a peak experience advances the slow latent only if reproducible-at-will + persistent — **[C4]**
The spec's position-vs-phase split (§3) is stated almost verbatim by four traditions and echoed by the meditation-science state/trait literature — two families under the §1 scheme, hence [C4] (downgraded from [C5], verifier fix 2026-07-16: lineages within family T are not independent families). Sufism separates *maqām* (station: acquired, stable, volitionally re-enterable) from *ḥāl* (state: given, transient, cannot be summoned) — this *is* the spec's split (T: al-Qushayrī *al-Risāla*; al-Hujwīrī *Kashf al-Maḥjūb*, lane 1 §7). Dzogchen separates *nyams* (transient meditation experiences) from realization (unchanging), the diagnostic being grasped-vs-released not which experience arose (T: lane 1 §4). Patañjali warns *ānanda* (bliss) is explicitly not terminal (T: Yoga Sūtras 1.17–18). Theravāda treats fetter-removal as permanent trait change vs a temporary altered state (T/N: `Anālayo 2021, "The Four Levels of Awakening," Mindfulness 12(4):831–840`). Meditation neuroscience supplies the same cut as state-vs-trait (general meditation-neuroscience, outside the four defined families: `Cahn & Polich 2006, Psychol. Bull. 132(2):180–211`; N: f-SNR's explicit state/trait/depth separation, `Laukkonen 2026, Clear Mind, arXiv:2606.29698`). **→ Strong convergent backing for the spec §6 error-routing rule (slow latents revise only under persistent evidence) — with one addition the spec does not yet contain: reproducibility-at-will. The volitional-reproducibility criterion is a tradition-derived recommended amendment (fold into the spec revision at M1 close), not existing spec §6 content (verifier fix).**

### 1.3 An information ceiling on report-only assessment is real and large — **[C4]**
The spec's headline v1 number (interview-vs-dossier delta, §7.3) is asserted by the traditions and conceded by the science — T + N, two families, hence [C4] (downgraded from [C5], verifier fix). Five lineages are directly evidenced insisting some part of attainment is visible only to a *live* teacher: pointing-out (Dzogchen, live), *dokusan/sanzen* (Zen, live), *huǒhòu* fire-timing (Neidan, oral-only), ñāṇa placement (Mahāsi, teacher across successive interviews), spiritual direction (Christian) — lane 1 §11.4; lane 1's synthesis claims the pattern recurs across all nine lineages it surveyed, but only these five are evidenced for this specific principle (verifier fix: softened from "nine"). The PP network concedes the same from the other side: deep non-conceptual states "may not [permit] any useful self-reports whatsoever," motivating no-report paradigms (N: `Laukkonen & Slagter 2021, NBR 128:199–217`; `Tsuchiya et al. 2015, TiCS 19(12):757–770`). The internal canon states exactly where the ceiling begins to bind: **below ~41 the canon's transition evidence is behavioral/telemetric (countable gates, demonstrable skills); from ~41 up it is phenomenological; from 51 up a teacher-in-the-loop is asserted to be structurally required** (canon §10). **→ Treat a data-only estimator as structurally handicapped above ~41, quantify the gap rather than assume it away, and expect the delta to widen with band.**

### 1.4 The map contaminates the report — **[C4]**
T + N evidence — two families, hence [C4] (downgraded from [C5], verifier fix; the mechanism is strong but the third family is our own cohort observation, not independent literature). Every tradition with a *public* map (ñāṇas, kōan answers, nyams, bhūmis) has a named scripting problem and a purpose-built countermeasure (T: honesty injunction; unscripted *sassho*; the teacher's live read — lane 1 §11.5). The PP corpus is itself the contamination vector, and says so: all six seed papers supply rich, quotable vocabulary ("dereification," "luminosity," "non-dual," "cessation," "clarity") that a map-reading practitioner reproduces without the attainment (N: cautions in every corpus note; `Metzinger 2024, The Elephant and the Blind` is 500+ reports' worth of exactly this vocabulary). The spec's own cohort is fully map-contaminated (team has read the Wheel, §7). This is the "insight vocabulary vs lived insight" confusion pair (§3). **→ Never a probe with a discoverable correct answer; upweight spontaneous pre-exposure usage, discount post-exposure (lane 5 exposure-ledger recommendation); score structure over vocabulary.**

### 1.5 Above ~Awakening, trait/off-cushion behavior discriminates; peak-states do not — **[C4]**
Mahāyāna is explicit: meditative equipoise is *identical* across all ten bhūmis — only the *post-meditation* qualities (the paramitas) differ (T: lane 1 §5). Mahāsi confirms fruition by trait change, not by the cessation report alone; Zen's tenth ox-herding picture is marketplace *return* (T). The cessation literature converges: the reliable residue of a genuine `small_death` is a *durable* shift toward de-reified, low-grasping, present-centered, low-self-referential language and affect — not the event description (N: `Laukkonen et al. 2023, Prog. Brain Res. 280:61–87`, after-effects §1e; `Agrawal & Laukkonen 2024` fetter model). **→ Above ~41 weight durable trait residue heavily — but the observability gap must be named (verifier fix): spec §5's channels are conversation, probes, and telemetry, so above ~41 "off-cushion behavior" is not directly observable — it operationalizes as (a) claim-vs-telemetry consistency and (b) durable linguistic trait-shift across months of unrelated conversation, not observed conduct. A vivid peak report with no such residue is a state, not a stage.**

### 1.6 The path is non-linear — U-shaped, spiral, plateau-punctuated — **[C4]**
"Advanced" is not "permanently rarefied." Zen's marketplace return and Batchelor's spiral (T), the canon's own plateau–peak ladder and twice-recurring practice-meaninglessness phase (canon §7), the f-SNR U-curve (initial destabilization before a new stable basin; N: `Laukkonen 2026`), the inverted-U in startle habituation (moderate experience > experts on some markers; N: `Antonova et al. 2015, PLoS One 10(5):e0123512`), and the dukkha-ñāṇa cycling (C/T: Lindahl) all say the same thing. **→ The phase model is not optional. A dark-night week or a plateau is not a regression; the aggregator must phase-gate any monotone "higher band = higher clarity" prior or it will misread transitions.**

### 1.7 Destabilization has characteristic signatures distinct from progress — **[C4]**
Traditions and clinical science converge on a differential. Neidan names a miscultivation syndrome (uncontrolled energy to the head, insomnia, nervous-system strain) requiring teacher correction (T). John of the Cross's three signs are a centuries-tested dark-night-vs-melancholy differential (T: *Dark Night* I.9). The clinical literature localizes the persistence-risk signature precisely: **dysregulated arousal (hyperarousal + dissociation), not sadness** (C: `Britton et al. 2021, Clin. Psychol. Sci. 9(6):1185–1204` [VERIFIED]), and gives eleven teacher-used criteria whose reliable members are uncontrollability, loss of critical attitude, sustained impairment, and suicidality-regardless-of-framework (C: `Lindahl et al. 2020, Front. Psychol. 11:1905` [VERIFIED]). **→ Safety flags key on specific signatures (valence/duration/impairment/practice-link as separate fields), never on "difficulty" per se; and Waypoint records features, never asserts the differential (§3.4 below).**

### 1.8 Development is uneven across pillars — **[C3→C4]**
Vidyādhara levels make mind-ripening and body/energy transformation *explicitly separable axes* (the matured vidyādhara has ripened the mind but not purified the body); Teresa's mansions have "many rooms" and "not everyone follows exactly the same path" (T: lane 1 §§5,9). f-SNR is channel-specific — different practices tune different declutter/amplify channels (N: `Laukkonen 2026` Table 1), and the phenomenological matrix is multi-axis by construction (N: `Lutz, Jha, Dunne & Saron 2015, Am. Psychol. 70(7):632–658` [VERIFIED]). Confidence is [C3] as a *tested* empirical claim (no study measures pillar-orthogonality directly) but [C4] as a design principle multiple families endorse. **→ The per-pillar profile (§3) is well-motivated; a single scalar hides exactly what Wisdom needs. But note the corollary problem in §4: deriving one overall band from four pillars needs a defensible aggregation rule.**

### 1.9 Reification/dereification is the strongest single *linguistic* axis — with the loudest caveat — **[C3]**
The clearest transcript-level signal the corpus offers: "anger arose and passed" (ownerless, impermanent, constructed) vs "I *am* an angry person" (naïve-realist reification), graded from OM (lowers) to ND (removes) (N: `Laukkonen & Slagter 2021`; `Agrawal & Laukkonen 2024` emptiness ladder gross→subtle). It maps onto STAGES' object-sophistication dimension (P), giving cross-family support for the *axis*. But it is largely N-family for the meditation-specific gradient, and it is **the most contamination-prone marker in the bank** — dereified phrasing is exactly what map-readers learn to produce (§1.4). **→ Use as a high-value axis, but only when corroborated by structure + the behavioral channels Waypoint actually sees (telemetry consistency, durable trait-shift — the §1.5 observability caveat); never score the phrase alone.**

### 1.10 Absorption vs dullness is resolvable by clarity-of-knowing — **[C3]**
Same low-content report, opposite depth: genuine absorption/MPE is high epistemic depth + gathered hyper-precision; dullness is low epistemic depth (N: `Metzinger 2020, Phil. Mind Sci. 1(I)`; beautiful-loop §1d). Patañjali's object-subtlety gradient refines it (T). The discriminating prediction — genuine absorption *sharpens* deviance-sensitivity while dullness blunts it — has thin but real empirical support (N: `Mago et al. 2025, arXiv:2511.20990` jhāna sharpens MMN; the MMN literature is "mixed," `Fucci et al. 2022`). **→ Probe clarity-of-knowing / deviance-sensitivity ("would you have caught something unexpected?") before scoring any low-content state as advanced.**

---

## 2. MEASUREMENT APPROACHES THAT SURVIVE SCRUTINY — and those that do not

### 2.1 Survive — build on these

**Ego-development sentence-completion structure-scoring (Loevinger → Cook-Greuter → O'Fallon STAGES).** The load-bearing precedent and the one mature discipline that does *exactly* Waypoint's job: assign a human an ordinal developmental stage by scoring free-text language structure against a rubric, with trained-rater reliability and a *deterministic protocol-level aggregation rule* (the ogive — you do not average item stages; the protocol stage is the highest stage at which a threshold cumulative proportion of responses sits). This maps one-to-one onto "LLM marker extraction → thin Bayesian aggregation" (spec §6). STAGES' three-dimension generative rubric — perspective (individual↔collective), agency (passive↔reciprocal), object-sophistication (concrete→subtle→MetAware) — turns stage-scoring from memorized exemplars into a rule an LLM can apply (P: `O'Fallon et al. 2020, Heliyon 6(3):e03472` [VERIFIED]; WUSCT manuals `Loevinger & Wessler 1970`, `Hy & Loevinger 1996` [UNVERIFIED metadata, foundational]; Cook-Greuter post-conventional extension [UNVERIFIED metadata]). **Why it survives:** 50 years of reliability, structure-over-content by design (§1.1), and it degrades *gracefully* upward — you extend the stems and rubric without re-norming a whole instrument (lane 2 §5).

**LLM structure-extraction at near-human-rater agreement.** The direct empirical green light: an LLM replicates STAGES scoring at **weighted κ ≈ 0.779 (95% CI 0.672–0.887) per single sentence, 0.705 for 10-sentence aggregates**, cross-checked across GPT/Claude/LLaMA, with ~5 items a floor and ~10 optimal for stability (SD 0.138 at 10) (P: `Bronlet 2025, Front. Psychol. 16:1488102, DOI 10.3389/fpsyg.2025.1488102` [VERIFIED]). **Why it survives:** it is the exact leg Waypoint asks an LLM to do, it hands over a data-sufficiency floor for the abstention threshold, and its GPT-vs-Claude-vs-LLaMA cross-check is a ready template for the spec's different-model-family rule (§6d).

**Structured phenomenological debrief scored on granularity + latency.** The Mahāsi three-phase report (what occurred / how you noted it / what happened to it) scored on metacognitive granularity and noticing-latency, not emotional content (T: `Sayadaw U Paṇḍita, In This Very Life, 1992`); microphenomenology to recover what a subject cannot spontaneously narrate (N: `Petitmengin et al. 2017/2019`); neurophenomenological button-press labelling time-locked to data (N: `Yang et al. 2024, Cereb. Cortex bhad408`). **Why it survives:** it is a probe *scoring key*, hard to fake, and it is how three independent families actually verify.

**Unscripted fresh-angle follow-up (the anti-faking instrument).** Zen *sassho* — 20–100 checking questions per major kōan, from angles the published map doesn't cover, that only lived experience can improvise (T: `Hori, Zen Sand, 2003`). **Why it survives:** it is the traditions' purpose-built defence against exactly the scripting problem the spec's public Wheel creates (§1.4). Directly transferable, and it degrades an attacker's advantage precisely because the correct answer is not discoverable.

**Behavior-over-claim / telemetry consistency.** Objective markers the subject cannot easily fake — startle habituation, attentional-blink reduction (N: `Antonova et al. 2015`; `Slagter et al. 2007`); expertise makes internal states more *decodable* (65.6% experts vs chance in novices, N: `Guidotti et al. 2023, Brain Topography 36(3):409–418`; `Weng et al. 2020`). Waypoint's channel-analogue: practice telemetry, claim-vs-behavior consistency penalties (spec §6), and the synthetic-persona discrimination QA (higher bands should be *easier* to classify with less variance — the behavioral echo of MVPA separability). **Why it survives:** it is the spec's anti-Goodhart spine and the one channel immune-ish to self-presentation.

**Nowcast a present state; never predict a future event.** Framing the teacher flag as "this practitioner has reported X-cluster for Y weeks with Z impairment" (base rate ~10%) rather than "this practitioner will deteriorate" (C: `Belsher et al. 2019, JAMA Psychiatry 76(6):642–651` [VERIFIED]: event-prediction PPV ≤0.01 at low base rates). **Why it survives:** it is the single design choice that makes the flag statistically defensible, and every flag is human-verifiable from cited evidence.

**Convergent-validity instrument battery — used *once*, calibration-only (never the estimator spine).** The shortlist that survives contact with bands 4–6: **NADA-T + NADA-S** (best fit to the wheel's central non-dual axis; α .81–.94; P: `Hanley, Nakamura & Garland 2018, Psychological Assessment 30(12):1625–1639` [VERIFIED; retagged C→P, verifier fix — this is psychometrics, not clinical/adverse-events work]); **MODTAS** (absorption capacity, discriminates absorption↔dullness; `Tellegen & Atkinson 1974` [VERIFIED]); **Hood M-scale** (the only validated vocabulary reaching 51+ introvertive mysticism — a targeted ceiling slot, `Hood 1975` [VERIFIED]); **MPE-92M** (pure-awareness profile; `Gamma & Metzinger 2021, PLoS One 16(7):e0253694` [VERIFIED — confirmed live 2026-07-16, flag lifted]); a **short WUSCT/STAGES-style sentence-completion protocol** (the differentiator — a second, human-scorable instantiation of Waypoint's own method); optionally one self-transcendence trait scale and **FFMQ with the DIF caveat baked in**. Total ≈ 55–75 min (lane 2 §7). **Why they survive — with a leash:** each is a validated *convergent anchor*, scored as band-conditional DIF-aware evidence, never as a metric with a stable zero.

### 2.2 Do not survive — reject, with reasons

**Hawkins' Map of Consciousness calibration — the negative control.** Surya credits Hawkins as the numeric-scale *inspiration* (canon §1), and every wheel point carries a Hawkins-LOC-extension value (LOC 0 → 1,000,000, S4). But Hawkins assigned those numbers by **applied-kinesiology muscle-testing**, which has no demonstrated reliability or validity, fails double-blind testing, and is classed as pseudoscience (T: `Hawkins, Power vs. Force, 1995`; critiques `PMC2000870`, `PMC1847521` [VERIFIED]). It is the worked example of the three anti-patterns Waypoint's validation protocol exists to repudiate: single unvalidated rater with no inter-rater check, false-precision point-claims with no error bars, and a non-reproducible private-felt-sense procedure. **Reject the method; the spec's distributions + credible intervals + triangulated ground truth + abstention are a point-by-point antidote.** (Consequence for the Wheel itself: see §4.1.)

**Self-report content scales as the estimator spine.** FFMQ shows differential item functioning across meditators vs non-meditators — the same item does not carry the same latent meaning across experience levels (P: `Van Dam et al. 2009, Personality and Individual Differences 47(5):516–521, DOI 10.1016/j.paid.2009.05.005` [VERIFIED]; per-facet breakdown [UNVERIFIED]). TMS measurement invariance across proficiency is the open question its own literature is probing (P: `Ireland et al. 2019, J. Clin. Psychol., DOI 10.1002/jclp.22709` [VERIFIED]). "Mind the Hype" generalizes it (P/N: `Van Dam et al. 2017/2018, Perspect. Psychol. Sci. 13:36–61` [VERIFIED]). **Reject as spine:** a raw-score increase is unidentifiable between genuine development and a shifted response frame. This DIF *is* the empirical backbone of the "insight vocabulary vs lived insight" pair, and the reason the spine must be structure-extraction over conversation.

**Jeffery Martin's PNSE / "Fundamental Wellbeing" four-cluster continuum.** A four-"location" map of awakening from self-report, sold via a ~$3k course with no control groups and no independent replication; the continuum may be a clustering artifact (P: lane 2 §3, `[VERIFIED that the work + claim exist; peer-reviewed status weak]`). **Reject as a scoring scheme:** it is *precisely* the Goodhart/artifact trap Waypoint must avoid — its failure mode is Waypoint's cautionary tale. At most, informal marker candidates to be independently validated.

**Neural markers of cessation/depth, in v1.** The cessation EEG evidence is single-digit-adept and internally contradictory — alpha-desynchronization (`Chowdhury et al. 2023`) vs gamma-synchronization (`Berkovich-Ohana 2017`) point opposite ways, and the authors say so (N: cautions in the nirodha + emptiness notes). **Reject for v1 likelihood weights:** out of Waypoint's evidence channels anyway (language + telemetry), and not yet established. Keep as method exemplar (intensive single-subject sampling legitimizes tiny-N longitudinal design), not as signal.

**Hours/tenure as a depth proxy.** Expertise is "notoriously difficult to pin down," hours "does not necessarily indicate how advanced someone is," and FA→OM→ND progression is "not guaranteed" (N: `Laukkonen & Slagter 2021`; `Van Dam et al. 2018`). **Reject strong tenure priors:** the spec is right to keep base-rate-by-tenure priors weak and never let session counts drive the estimate.

**Conviction / noetic certainty as evidence.** Aha-feelings are inducible, misattributable, and make false content feel true; detailed warnings only *reduce*, not prevent, false insights (N: false-insight cluster, `Grimmer et al. 2022/2023`; `McGovern et al. 2024`). **Reject conviction intensity as a positive marker:** high-intensity mystical/noetic language should trigger a *reification check*, not band promotion.

**Any single-scalar band claim.** Teresa: growth is "gradual and imperceptible"; f-SNR: "signal and noise are tradition- and goal-relative," a global scalar under-determines which pillar advanced; contentless reports are "neither truly contentless nor identical" (`Woods et al. 2022`). **Reject the bare scalar** in favour of the per-pillar profile + band posterior + credible interval the spec already mandates.

**MAAS and MEDEQ-as-a-stage-measure.** MAAS floors above basic mindfulness (skip); MEDEQ is a *session-depth* (fast-state) instrument whose top level collapses everything transpersonal into one bucket — mine its wording for probes, do not use it as a stage measure (lane 2 §§1,3).

---

## 3. ANSWERS TO THE SPEC'S OPEN QUESTIONS (§10)

Each open question quoted verbatim, then answered from the evidence.

### 3.1 "What granularity does evidence actually support — band-level, or finer within 22–45?"
**Answer: band-level (10-point) is the defensible working resolution; sub-band point resolution is *not* supported by language evidence, with two exceptions — discrete milestone events, and the per-pillar profile.** Confidence [C3–C4].

The strongest scoring precedent resolves to *one ordinal stage per protocol* from ≥5–10 items (P: `Bronlet 2025` κ≈0.78 single-item, SD 0.138 at 10; the ogive yields one Total Protocol Rating). That is roughly *one band* of resolution, not one scale-point. Decodability studies are expert-vs-novice coarse (N: `Guidotti 2023`). And the canon itself changes measurement modality by region: **below ~41 evidence is countable and behavioral** (7 full sessions to close Stage 1; 14/10/5/7-session gates per pillar; reverse-breathing demonstrable) — here finer resolution *is* available, but it comes from **telemetry + gate completion**, not from parsing language. **From ~41 up evidence is phenomenological and the ceiling binds** (§1.3) — here even band-level gets noisy and abstention should be the default (spec `data_sufficiency: insufficient`).

Practical recommendation: (a) report a **band-level posterior** as the canonical overall output; (b) within 22–45, allow finer *milestone*-level resolution only where it is **event-anchored and telemetry-corroborated** (`reverse_breathing`, `first_glimpse`, `small_death` are crisp events with entry/exit — lane 3's milestone-BKT layer and the cessation literature's "crisp events as temporal anchors" both support this); (c) locate the finest genuine resolution in the **per-pillar profile**, because the four pillars are separately gated and separately evidenced (a practitioner can be 2.4-deep in awareness with untrained emotion — §1.8); (d) never emit a sub-band single-scale-point claim (e.g. "53 not 54") from language alone — that is the Hawkins false-precision failure (§2.2). Lane 3's discrete 56-point grid (15–70) is the right *internal* representation because it can hold legitimately **bimodal** posteriors ("genuine 48 vs well-read 30") that a Gaussian filter structurally cannot; but the *reported* granularity out of it is band + interval, not point.

### 3.2 "Which established instruments (if any) are worth the calibration battery slots?"
**Answer: a ≈55–75 min battery of NADA-T + NADA-S, MODTAS, Hood M-scale (ceiling only), MPE-92M, a short STAGES-style sentence-completion protocol, and FFMQ-with-DIF-caveat; optionally one self-transcendence trait scale. All calibration-only, scored as band-conditional DIF-aware evidence, never a product surface and never the estimator.** Confidence [C4] on the shortlist.

Rationale per slot is in §2.1 and lane 2 §7. The three non-obvious calls: (1) **NADA is the top pick** — closest existing instrument to the wheel's central non-dual axis, and its self-transcendence/bliss split even mirrors the corpus's epistemic-vs-affective distinction (P: `Hanley et al. 2018` [VERIFIED; retagged C→P]). (2) **The sentence-completion protocol is the differentiator, not a nicety** — it is a second, human-scorable instantiation of Waypoint's own method on the same subject, and the cleanest place to measure "does the LLM extractor agree with a trained structure-scorer" (target κ 0.7–0.8, the Bronlet/STAGES bar). (3) **Hood M-scale earns only a targeted 51+ ceiling slot** — near-useless for 22–45, but the only validated vocabulary that reaches introvertive-mysticism phenomenology the estimator must still recognize. **Explicitly excluded:** MAAS (floors), MEDEQ-as-stage (state scale — mine for probe wording), PNSE/Finders (not an instrument, §2.2). Additional cataloguing tools worth a read, not a slot: `Schmidt & Berkemeyer 2018, Altered States Database` (the shopping catalog); `Desbordes et al. 2015` equanimity measure. **Every battery score is triangulation for Surya's gold label, not a competing ground truth** (the DIF caveat, §2.2).

### 3.3 "How do ñāna-interview and koan-checking traditions structure *verification questions* — and what transfers to probe design?"
**Answer: they score the structure of a debrief and defeat scripting with unscripted fresh-angle follow-ups and show-don't-tell demonstration; five moves transfer directly.** Confidence [C4] (well-documented tradition practice); lane 1 §12 delivers 20 transferable patterns for the probe bank.

The mechanics: **Mahāsi ñāṇa interview** never asks "what stage are you in?" — it runs a tightly structured debrief (what occurred → how you noted it, incl. noticing-latency → what happened to it) and infers stage from the *structure* of the report across successive interviews, scoring metacognitive granularity and temporal resolution (T: `U Paṇḍita 1992`; Panditārāma demands reports be "short and to the point," best-sitting-only). **Zen kōan checking** tests understanding that must be *demonstrated not described*: *sassho* checking-questions (20–100 per major kōan) come from fresh unscripted angles specifically to defeat the semi-public "correct answers"; *jakugo* requires an indirect expressive capping-phrase (proof-of-solution, not the solution) (T: `Hori 2003`). **Sufism** supplies the reproducibility test (summonable-at-will *maqām* vs unreproducible *ḥāl*); **Pa-Auk** the five masteries (pre-resolve a duration, sustain, exit cleanly).

The five transfers to Waypoint's invisible probes: **(1) score structure not content** (the three-phase episode probe); **(2) unscripted fresh-angle follow-up after any advanced claim** — no probe with a discoverable correct answer (the core anti-faking instrument, defeats the map-contamination in §1.4); **(3) reproducibility-at-will + off-cushion persistence** to decide whether evidence moves the slow latent (§1.2); **(4) show-don't-tell** — invite an image or concrete moment, score aptness/embodiment over doctrinal correctness; **(5) under-claiming calibration** — genuine deep attainers under-report, readers over-report, so weight *spontaneous, structurally-rich, hedged* reports above *confident, vocabulary-heavy, unprompted* ones (corrects the spec's bidirectional self-presentation threat, §7). Delivery constraints (invisible, frequency-capped, never during flagged-vulnerable moments, never verbatim-repeated, rotated so no probe accrues a discoverable answer) are in the spec §5 and reinforced by lane 5's disclosure-robustness rule (keep detection logic internal-only; knowing the detection strategy degrades detection more than knowing the symptoms).

### 3.4 "What is the defensible mapping between 'difficult territory' markers and flag thresholds?"
**Answer: a two-tier nowcast flag over a Britton-schema marker event, with a conjunctive persistence core, history-neutral thresholds, an expected ~5–15% Tier-A rate, and an absolute rule that Waypoint records features and never asserts the differential.** Confidence [C4] — this is the best-evidenced lane; full design in lane 4 §7.

The mapping: difficulty markers carry `{phenomenology_category (VCE 59×7), valence, duration, impairment, practice_link, elicitation}` as **separate fields** (collapsing them to "negative: yes/no" reproduces the measurement failure Britton 2021 was written to fix). **Tier A (monthly-consolidation, persistent-difficulty):** dysregulated-arousal cluster (the empirical persistence-risk signature — hyperarousal + dissociation, `Britton et al. 2021` [VERIFIED]); dissociative cluster with negative valence/impairment; trauma re-experiencing; sustained dark-night phenomenology (≥3 weeks, or any duration with impairment); Unbinding-adjacent territory (posterior mass ≥0.3 on 51+ — flags on *stage*, highest support need); post-intensification destabilization. **Tier B (expedited):** the `Lindahl et al. 2020` [VERIFIED] consensus intervention triggers — loss of critical attitude, uncontrollability, escalating impairment, suicidality (classifier authority regardless of framework). **Threshold logic:** Tier A is **conjunctive** (cluster AND cross-session persistence AND ≥1 of valence/impairment/uncontrollability); Tier B **disjunctive** (sensitivity dominates); probe-elicited difficulty carries *lower* weight than spontaneous (the `Farias et al. 2020` [VERIFIED] 3.7%-vs-33.2% elicitation gap); **history-neutral** (psychiatric history neither raises nor lowers thresholds — `Lindahl 2020` circularity + `Farias` AEs-without-history); **rate-monitored** (sustained Tier-A outside ~2–15% triggers threshold review — a deliberately *wider* review band than the ~5–15% expectation band, looser on the low side; both bands are lane 4 §7's design, clarified here after a verifier read them as inconsistent — `Goldberg et al. 2022` [VERIFIED] base rates). **The binding constraint (Belsher):** the flag asserts an observed *present* state, never future risk — no risk scores ever. **And the ontological rule:** identical phenomenology is appraised as progress by some and pathology by others (`Lindahl et al. 2017` [VERIFIED]), so the flag never adjudicates "genuine dark night vs depression" — it records features + support-need and routes to a human (LIFE Team identity, never Wisdom). Flag text names observed patterns in the practitioner's own words, never nosology (`FDA General Wellness` boundary, lane 4 §6a).

### 3.5 "Where does the Wheel's structure disagree with the strongest empirical work, and does anything in the model need Surya's revisiting?"
**Answer: the Wheel agrees with the external work on its deep structure (banded, position-vs-phase, per-pillar, non-linear top) and disagrees on three points worth Surya's attention — the Hawkins numbers, sub-band ordinal precision, and two specific labels.** Full treatment in §4; the items flagged for Surya are consolidated in §4.7.

---

## 4. WHERE THE WHEEL AGREES AND DISAGREES WITH THE STRONGEST EXTERNAL WORK

The research-validation duty. Disagreements are findings, not embarrassments; several are genuine strengths of the Wheel, and the honest ones are flagged for M2.

### 4.1 DISAGREE (substantive): the Hawkins LOC inheritance carries no epistemic weight
The Wheel's design DNA and its `hawkins_level` column both come from Hawkins (canon §1), whose calibration method is refuted pseudoscience (§2.2). The tell is internal: the LOC column is **non-monotonic in band 3** (point 24 "Cooperation" = LOC 100–299, overlapping its three predecessors — canon §9-H), a data-entry artifact that would be impossible if the numbers indexed anything real. **Finding:** the Hawkins numbers are decorative provenance, not an ordinal scale. **Recommendation:** never use the LOC column as an independent ordinal check; and Surya should decide whether keeping the Hawkins numbers at all is worth the association with a discredited method, given the whole validation protocol reads as a repudiation of it. This is the sharpest disagreement in the dossier — but it is a disagreement with the Wheel's *borrowed scaffolding*, not its phenomenology.

### 4.2 DISAGREE (calibration): 100-point resolution overclaims what evidence supports
The Wheel names 100 distinct scale-points (52 Low Pain, 53 Low Pleasure, 54 Low Meaning…). The strongest external work resolves to ~band precision (§3.1): Teresa's "gradual and imperceptible," the ego-development ogive's one-stage-per-protocol, Bronlet's ~one-stage LLM resolution, f-SNR's non-strict ordering. **Finding:** the *fine* granularity is a useful phenomenological *narrative* (Surya narrates the path in nearly scale-point order — canon §3) but not a measurable ordinal scale at single-point resolution. **Recommendation (already spec-aligned):** report band + interval; treat the sub-band point-labels as milestone/phase vocabulary, not as estimator targets. No Wheel change needed — this is a measurement-scope clarification.

### 4.3 DISAGREE (specific ordering): the 52–56 "Low-X cascade" as a fixed public sequence
The canon presents Low Pain → Pleasure → Meaning → Empathy → Knowledge as an ordered scale-point sequence (canon §3, §7). The dukkha-ñāṇa literature it corresponds to says these textures **cycle and vary** within days-to-weeks and are not traversed in a fixed public order (C/T: `Lindahl et al. 2017`; Shinzen's days-vs-months distinction, lane 4 §4b; the canon *itself* elsewhere calls dark-night phases cycling — canon §7). **Finding:** the phenomenology is real and well-attested; the fixed monotone ordering is not empirically supported and, if scored literally, would let a well-read practitioner script the sequence (§1.4). **Recommendation:** treat 52–56 as a *phase cluster* whose members co-occur and recur, not as five ordered checkpoints; score the cluster, not the order.

### 4.4 TENSION (label collision worth Surya's eye): "Dissociation" at 43
The Wheel labels point 43 "Dissociation" as a *normal recognition phase* ("I'm these, I'm not that" — canon §6, §9-M). The clinical literature makes dissociation half of the **persistence-risk signature** (`Britton et al. 2021`) and a Tier-A flag trigger (lane 4 §7.1). The canon already flags the collision (§9-M: a marker bank "must not treat the wheel label as license to read dissociative symptoms as progress"). **Finding:** a genuine, named tension — the same word denotes a progress-phase in the Wheel and a risk-signature in the safety lane. **Recommendation:** keep the marker bank's dissociation-as-difficulty scoring *independent* of the wheel label; and Surya may wish to rename point 43 (e.g. to a recognition/differentiation term) to remove the collision at the source. The discriminator the corpus offers: positive-valence self-boundary change *without* distress/impairment is stage evidence; negative-valence, uncontrollable, or impairing detachment is a flag (the Lindahl & Britton discriminator, carried via lane 4 §7.1; the standalone `Lindahl & Britton 2019` citation is [UNVERIFIED metadata this pass] — the verified members of the set are items 16–18 in §5.E).

### 4.5 AGREE (strongly): position-vs-phase, per-pillar, non-linear top, milestone events
Where the Wheel is at its best, the external work backs it emphatically. **Position vs phase** is Sufi *maqām/ḥāl* verbatim (§1.2). **Per-pillar unevenness** is the vidyādhara mind-vs-body separable axes and f-SNR channel-specificity (§1.8). **The non-linear, intensifiable top** — the canon's "enlightenment is not an endpoint… it's a function that can intensify, intensify" and its plateau–peak ladder (canon §6) — agrees strikingly with the PP/f-SNR "depth is continuous" framing and `Sparby & Sacchet 2025`'s advanced-meditation/human-development map. **The small-death milestone at 50/51** aligns with stream-entry-as-first-cessation (`Anālayo 2021`; `Agrawal & Laukkonen 2024`), and the canon's deliberate **exclusion of the "big death"** (full nirodha samāpatti) as rare-and-unnecessary is *consistent* with the science, which treats NS as a real but exceedingly rare attainment gated behind eight-jhāna mastery + non-returner-class transformation (`Laukkonen et al. 2023`). These are not lucky coincidences — they are a domain expert and four literatures converging.

### 4.6 AGREE (with a caution the Wheel should absorb): devotional/bliss intensity at ~46
The Wheel places devotional/mystical intensity ("Ascension," ~46) as a positioned phase (canon §6–7). Patañjali, Dzogchen, and the false-insight literature all warn bliss/noetic-certainty is *not* terminal and is inducible/false (§1.2, §1.9). **Finding:** near-agreement — the canon already frames 46 as a *phase*, not an automatic position gain. **Recommendation:** ensure the marker bank scores devotional intensity as phase (reproducible + persistent to count toward position), and treats high-conviction mystical language as a reification-check trigger, not a promotion signal.

### 4.7 Consolidated for Surya (M2)
1. **Hawkins numbers** — keep as provenance, drop as ordinal check; decide whether to retain them at all (§4.1).
2. **Sub-band precision** — confirm band+interval is the reporting resolution; sub-band points are milestone/phase vocabulary (§4.2).
3. **52–56 Low-X cascade** — re-frame as a co-occurring/recurring phase cluster, not a fixed ordered sequence (§4.3; 51 is Unbinding onset and 57–60 are Non-Dual Loop/Deconstruction/Void/Luminosity — the cascade proper is 52–56, verifier fix).
4. **"Dissociation" (43) label** — resolve the collision with the clinical risk-signature, possibly by renaming (§4.4).
5. **Overall-band aggregation from four pillars** — the per-pillar profile is well-supported, but the rule for deriving *one* overall band from four unevenly-developed pillars is undecided and must not be a naïve max or mean (flag for the estimator design memo; §1.8).
6. **Devotional intensity (46) / bliss** — confirm phase-not-position scoring (§4.6).

---

## 5. ANNOTATED MUST-READ LIST (the 15–25 works a future builder must know)

Ranked within groups by decision-relevance. Verification flags carried from upstream lanes.

### A. The scoring method (read first — this is what Waypoint *is*)
1. **O'Fallon, Polissar, Neradilek & Murray (2020).** *The validation of a new scoring method for assessing ego development based on three dimensions of language.* Heliyon 6(3):e03472. DOI 10.1016/j.heliyon.2020.e03472. **[VERIFIED; inter-rater κ / concordance numerics UNVERIFIED — publisher 403 upstream.]** — The three-dimension generative rubric (perspective / agency / object-sophistication). The single most transferable scoring discipline in the dossier.
2. **Bronlet (2025).** *Leveraging on large language model to classify sentences: STAGES scoring for sentence completion.* Frontiers in Psychology 16:1488102. DOI 10.3389/fpsyg.2025.1488102. **[VERIFIED.]** — The direct empirical green light: LLM structure-scoring at κ≈0.78, ≥5–10 item floor, GPT/Claude/LLaMA cross-check. The extractor's template and QA bar.
3. **Loevinger & Wessler (1970)** / **Hy & Loevinger (1996).** WUSCT manuals. **[UNVERIFIED metadata; foundational.]** — The ogive protocol-aggregation rule Waypoint's deterministic layer should copy literally (do not average item stages). Cook-Greuter's post-conventional extension **[UNVERIFIED metadata]** carries the ceiling upward into band-4–6 territory.

### B. How the traditions verify (the probe-design source)
4. **Sayadaw U Paṇḍita (1992).** *In This Very Life.* Wisdom Publications (full text aimwell.org, bps.lk). — The Mahāsi three-phase report: the probe scoring key for the awareness pillar (granularity + noticing-latency).
5. **Hori, Victor Sōgen (2003).** *Zen Sand: The Book of Capping Phrases for Kōan Practice.* U. Hawai'i Press. — *Sassho* (unscripted fresh-angle follow-up) and *jakugo* (show-don't-tell): the tradition's purpose-built anti-scripting instruments.
6. **al-Qushayrī, *al-Risāla*** + **al-Hujwīrī, *Kashf al-Maḥjūb*.** **[verified via secondary comparative sources.]** — *Maqām* (station/position) vs *ḥāl* (state/phase), articulated in the 11th c. with the precision the spec's §3 split needs.
7. **Sparby & Sacchet (2024).** *Toward a Unified Account of Advanced Concentrative Absorption Meditation: Classification of Jhāna.* Mindfulness 15(6):1375–1394. DOI 10.1007/s12671-024-02367-w. **[VERIFIED.]** — The teacher-anchored J1–J8 ladder: bands 4–5 milestone cartography on the Wheel (the canon's Theravada overlay places jhāna-factor → formless territory at ~31–50, with only cessation at 51 — canon §8; re-anchored from "top-of-Wheel 51+", verifier fix) that Waypoint must recognize but will rarely see at depth.
8. **Sparby & Sacchet (2025).** *Toward a Unified Model of Advanced Meditation, Human Development, Meditation Maps, and Transtradition Metaphors.* Mindfulness 16(9):2472–2482. DOI 10.1007/s12671-025-02632-6. **[VERIFIED — confirmed live 2026-07-16; flag lifted.]** — The paper in the corpus closest to Waypoint's own "map humans onto a developmental scale" problem.

### C. The construct spine (the scientific twin)
9. **Laukkonen & Slagter (2021).** *From many to (n)one.* Neurosci. Biobehav. Rev. 128:199–217. DOI 10.1016/j.neubiorev.2021.06.021. **[VERIFIED.]** — The FA→OM→ND / temporal-depth continuum the spec adopts as the Wheel's scientific twin; source of the dereification / non-grasping / self-ladder markers.
10. **Lutz, Jha, Dunne & Saron (2015).** *Investigating the phenomenological matrix of mindfulness-related practices.* American Psychologist 70(7):632–658. DOI 10.1037/a0039585. **[VERIFIED.]** — The multi-axis dimensional decomposition; the marker-bank rubric skeleton (per-pillar axes over a scalar). Highest recurrence in the corpus.
11. **Dahl, Lutz & Davidson (2015).** *Reconstructing and deconstructing the self.* Trends Cogn. Sci. 19(9):515–523. DOI 10.1016/j.tics.2015.07.001. **[VERIFIED.]** — Attentional / constructive / deconstructive practice families → the practice-to-pillar typing scheme (breathwork≈attentional, love≈constructive, clarity/awareness≈deconstructive).
12. **Laukkonen et al. (2023).** *Cessations of consciousness in meditation: nirodha samāpatti.* Prog. Brain Res. 280:61–87. DOI 10.1016/bs.pbr.2022.12.007. **[VERIFIED; neural-signature claims preliminary/n≈1 — do not weight.]** — The small-death/big-death definitions, the necessary-and-sufficient cessation rubric (absence + no-retrospective-content + clarity-after), and the after-effect trait residue that is the real corroborating evidence.

### D. Instruments for the calibration battery
13. **Hanley, Nakamura & Garland (2018).** *The Nondual Awareness Dimensional Assessment (NADA).* Psychological Assessment 30(12):1625–1639. DOI 10.1037/pas0000615. **[VERIFIED.]** — Best fit to the wheel's central non-dual axis; trait (NADA-T) + state (NADA-S) in one family. Top battery pick.
14. **Gamma & Metzinger (2021).** *The minimal phenomenal experience questionnaire (MPE-92M).* PLoS One 16(7):e0253694. DOI 10.1371/journal.pone.0253694. **[VERIFIED — confirmed live 2026-07-16; flag lifted.]** — Validated pure-awareness profile for the top awareness band.
15. **Metzinger (2020).** *Minimal phenomenal experience: meditation, tonic alertness, "pure" consciousness.* Phil. Mind Sci. 1(I):1–44. DOI 10.33735/phimisci.2020.I.46. **[VERIFIED.]** — Defines MPE and the MPE-vs-dullness distinction (the absorption-vs-dullness confusion pair's conceptual anchor).

### E. The safety / teacher-flag core (read as a set)
16. **Lindahl, Fisher, Cooper, Rosen & Britton (2017).** *The Varieties of Contemplative Experience.* PLoS ONE 12(5):e0176239. DOI 10.1371/journal.pone.0176239. **[VERIFIED.]** — The 59×7 difficulty taxonomy + 26 influencing factors: the controlled vocabulary for phase annotation and the flag-context features. (Criterion-sampled — severity, not prevalence.)
17. **Lindahl, Cooper, Fisher, Kirmayer & Britton (2020).** *Progress or Pathology?* Front. Psychol. 11:1905. DOI 10.3389/fpsyg.2020.01905. **[VERIFIED.]** — The eleven teacher-used criteria; the reliable intervention triggers (uncontrollability, loss of critical attitude, impairment, suicidality) and the unreliable ones (psychiatric history, retrospective benefit). Reframe: "what support does this person need?" not "spiritual or pathological?"
18. **Britton, Lindahl, Cooper, Canby & Palitsky (2021).** *Defining and Measuring Meditation-Related Adverse Effects.* Clin. Psychol. Sci. 9(6):1185–1204. DOI 10.1177/2167702621996340. **[VERIFIED.]** — The persistence-risk signature is dysregulated arousal (not sadness); the separate-fields marker schema (valence/duration/impairment/attribution).
19. **Goldberg, Lam, Britton & Davidson (2022).** *Prevalence of meditation-related adverse effects in a population-based sample.* Psychotherapy Research 32(3):291–305. DOI 10.1080/10503307.2021.1933646. **[VERIFIED.]** — App-user-like base rates (~10% lasting ≥1mo; 1.2% lasting impairment) → the ~5–15% expected Tier-A flag-rate calibration.
20. **Belsher et al. (2019).** *Prediction Models for Suicide Attempts and Deaths.* JAMA Psychiatry 76(6):642–651. DOI 10.1001/jamapsychiatry.2019.0174. **[VERIFIED.]** — Event-prediction PPV ≤0.01 at low base rates → the flag must nowcast a present state, never predict a future one. The design choice that makes it statistically defensible.

### F. Anti-Goodhart / contamination (why structure-scoring, not conviction)
21. **Grimmer, Laukkonen, Tangen & von Hippel (2022).** *Eliciting false insights with semantic priming.* Psychon. Bull. Rev. 29(3):954–970. DOI 10.3758/s13423-021-02049-x. **[VERIFIED.]** — with companions **Laukkonen et al. (2020)** "dark side of Eureka" (Cognition 196:104122) and **McGovern et al. (2024)** (Commun. Psychol. 2(1):69). Aha-feelings are inducible and make false content feel true; warnings only reduce, not prevent (Grimmer et al. 2023). The footing for structure-over-conviction.
22. **Paulhus, Harms, Bruce & Lysy (2003).** *The over-claiming technique.* JPSP 84(4):890–904, PMID 12703655. **[VERIFIED — confirmed live 2026-07-16 incl. the warned/fake-good robustness claim; flag lifted.]** — Overclaiming foils that uniquely survive *both* warning and fake-good instructions → a per-practitioner credibility weight that works on a fully map-aware cohort. (Foil list needs Surya clearance against the tradition overlays.)
23. **Van Dam et al. (2018).** *Mind the Hype.* Perspect. Psychol. Sci. 13:36–61. DOI 10.1177/1745691617709589. **[VERIFIED.]** — The field's own skeptical checklist (controls, preregistration, effect sizes, expertise≠hours) and the DIF backbone; the bar the spec's construct-validity claims (§11) must clear. Companion: **Van Dam et al. (2009)** FFMQ DIF, Personality and Individual Differences 47(5):516–521, DOI 10.1016/j.paid.2009.05.005 **[VERIFIED].**
24. **Nisbett & Wilson (1977).** *Telling more than we can know.* Psychol. Rev. 84(3):231–259. — Long retrospective introspective reports are unreliable → structure + behavior over content-of-claim.

### G. The negative control (what not to do)
25. **Hawkins, David R. (1995).** *Power vs. Force.* — with critiques **PMC2000870**, **PMC1847521** (manual-muscle-testing / applied-kinesiology reliability reviews) **[VERIFIED].** — The worked example of the numeric-consciousness-scale ambition undisciplined by measurement rigour: single unvalidated rater, false-precision points, non-reproducible calibration. Cite as the anti-pattern the entire validation protocol repudiates (§4.1).

**Deliberately not on the list (with reason):** the whole-brain-modeling machinery (Deco/Hopf, criticality, parcellation) is off-scope for a language-based estimator; Jeffery Martin PNSE is the cautionary artifact, not a source; the raw seed-paper reference lists are exhaustively catalogued in lane 6 and the corpus notes. The highest-value *external* acquisitions still outstanding are the out-of-network staging instruments the seed corpus never cites — the ego-development lineage above (items 1–3) fills that gap.

---

## Appendix — one-line provenance for the four families
- **Traditions (T):** lane 1 (`lanes/1-traditions-verification.md`) — 9 lineages + Hawkins negative control + 20 probe patterns.
- **Psychometrics (P):** lane 2 (`lanes/2-psychometrics-scoring.md`) — ego-development scoring lineage + instrument verdicts + battery.
- **Clinical/safety (C):** lane 4 (`lanes/4-safety-adverse-events.md`) — base rates, differential, two-tier flag design.
- **PP network (N):** lane 6 (`lanes/6-citation-map.md`) + 6 corpus notes — the Laukkonen network's own best-of (one network; not independent corroboration).
- **Statistics / Goodhart** (referenced via digests): lanes 3 and 5 — grid-state Bayes filter + milestone-BKT + natural-frequency marker likelihoods (lane 3); structure-over-content + overclaiming foils + exposure ledger (lane 5).
