# Lane 2 — Modern Psychometrics & Scoring Methodology

**For:** Waypoint (SOME) practitioner-stage assessment. Spec: `/Users/fionnenglish/LIFE/specs/waypoint-practitioner-stage-assessment.md`.
**Scope:** existing instruments for meditation depth / stage / awakening, their validity under advanced-practitioner phenomenology, and — the load-bearing part for Waypoint — the **ego-development sentence-completion scoring lineage** (Loevinger → Cook-Greuter → O'Fallon STAGES) as the closest methodological precedent for scoring humans from *language structure* rather than content.
**Reader:** has read spec §§1–7, 10. Written 2026-07-16.
**Verification convention:** every citation carries author/year/venue + DOI where confirmed via Crossref/PubMed/publisher this pass. `[VERIFIED]` = citation metadata confirmed live this pass. `[UNVERIFIED]` = plausibly real but not confirmed in this pass (treat as a lead, not evidence). Speculation is labeled inline.

---

## 0. TL;DR verdicts (decision-relevant)

1. **No existing instrument measures the Waypoint construct** (position on a 1–100 non-dual developmental wheel with per-pillar resolution). Every scale below measures either a *trait* (dispositional mindfulness / absorption / self-transcendence), a *state* (session-level depth / nonduality / mysticism), or *difficulty*. None gives a stage on a validated ordinal ladder that reaches bands 4–6. Use them **once, calibration-only, for convergent validity** exactly as §7 already scopes — never as the estimator.
2. **The one true methodological precedent is ego-development sentence-completion scoring** (Loevinger WUSCT → Cook-Greuter MAP → O'Fallon STAGES). It is the only mature tradition of *assigning a human an ordinal developmental stage by scoring the structure of their free-text language against a rubric, with trained-rater reliability and a deterministic protocol-level aggregation rule (ogive)*. This maps almost one-to-one onto Waypoint's "LLM marker extraction → thin Bayesian aggregation" (spec §6). Adopt its **structure-over-content** discipline, its **item/stem** design, and its **ogive-style protocol aggregation**; it is the strongest external validation that Waypoint's core bet is sound.
3. **STAGES scoring has a published LLM replication** (Bronlet 2025, Frontiers in Psychology) reaching **weighted κ ≈ 0.78** (single sentence) vs. human experts — direct evidence that an LLM can do the marker-extraction leg of Waypoint at near-human-rater agreement, and a ready template for the extractor prompt and its QA.
4. **The single biggest psychometric threat is real and named in the literature: differential item functioning / response-shift in advanced practitioners** (Van Dam 2009; Van Dam et al. 2017 "Mind the Hype"; Grabovac; the TMS invariance work). The same words mean different things at 25% vs. 55% on the wheel. This is the empirical backbone of the spec's "insight vocabulary vs. lived insight" confusion pair (§3) and the reason self-report scales cannot be the spine of Waypoint.
5. **Recommended one-time calibration battery** (§7): NADA-T + NADA-S (nonduality), MODTAS (absorption trait, discriminates absorption↔dullness confusion), FFMQ (dispositional mindfulness, *with DIF caveat baked in*), Hood M-scale (mysticism ceiling for the rare 51+ subject), a Cloninger TCI Self-Transcendence subscale OR ASPIRES (self-transcendence trait), and — the differentiator — a **short WUSCT/STAGES-style sentence-completion protocol** scored for structure. Total ≈ 55–75 min. Rationale + per-instrument verdict in §7 below.

---

## 1. Dispositional-mindfulness scales (FFMQ, MAAS, TMS) — and why they break at the top

### FFMQ — Five Facet Mindfulness Questionnaire
- **Instrument:** 39-item, five facets (observing, describing, acting with awareness, non-judging, non-reactivity). Baer et al. 2006/2008.
- **The finding that binds Waypoint — differential item functioning:** Van Dam, Earleywine & Danoff-Burg (2009), *Differential item function across meditators and non-meditators on the Five Facet Mindfulness Questionnaire*, **Personality and Individual Differences 47(5):516–521, DOI 10.1016/j.paid.2009.05.005** `[VERIFIED]`. Items function differently across meditators vs. non-meditators — i.e., a given FFMQ item does **not** carry the same latent meaning across experience levels, so raw score comparisons across the range are not psychometrically defensible. (The *observing* facet is the one most widely reported as behaving inconsistently across groups — positively valenced in trained samples, ambiguous in untrained ones — but I could not re-confirm the exact per-facet result from the abstract this pass; treat the "observe facet" specifics as `[UNVERIFIED]` while the DIF finding itself is `[VERIFIED]`.)
- **Broader critique:** Van Dam, van Vugt, Vago, … Meyer (2017), *Mind the Hype: A Critical Evaluation and Prescriptive Agenda for Research on Mindfulness and Meditation*, **Perspectives on Psychological Science 12(6), DOI 10.1177/1745691617709589** `[VERIFIED]` — canonical statement that self-report mindfulness scales suffer construct proliferation, semantic drift, and weak validity; scoring self-report agreement across expertise levels is exactly what it warns against.
- **Verdict for Waypoint:** **Convergent-validity anchor only, and a cautionary one.** Useful in the calibration battery to locate a subject in the low-to-mid range relative to norms, but its DIF is itself a *demonstration* of the spec's core problem (advanced practitioners re-interpret the items). Do **not** use FFMQ deltas as ground truth for movement. Its main gift is negative: it proves why Waypoint must score *structure of free response*, not agreement with fixed items.

### MAAS — Mindful Attention Awareness Scale
- **Instrument:** Brown & Ryan (2003), 15 items, single factor (present-moment attention/awareness, reverse-scored inattention). `[UNVERIFIED citation metadata this pass — very well established]`.
- **Verdict:** Weakest of the three for Waypoint. Unidimensional, measures ordinary attentional lapse, floor/ceiling issues in practitioners, no traction above basic mindfulness. **Skip** for the battery; nothing it measures reaches bands 4–6.

### TMS — Toronto Mindfulness Scale (state)
- **Instrument:** Lau, Bishop, Segal et al. (2006), state measure with two factors (Curiosity, Decentering). `[UNVERIFIED citation metadata this pass]`.
- **Directly relevant validity work:** Ireland, Day & Clough (2019), *Exploring scale validity and measurement invariance of the Toronto Mindfulness Scale across levels of meditation experience and proficiency*, **Journal of Clinical Psychology, DOI 10.1002/jclp.22709** `[VERIFIED]`. This is the measurement-invariance test Waypoint needs to cite: it explicitly asks whether TMS means the same thing across proficiency levels — the empirical form of the response-shift worry.
- **Verdict:** The *Decentering* factor is conceptually the closest self-report handle on the dereification marker the seed corpus (Laukkonen) treats as the strongest linguistic stage signal. But it is a **state** scale and invariance across proficiency is the open question its own literature is probing. **Optional battery inclusion** if a decentering anchor is wanted; not core.

---

## 2. Absorption, mysticism, nonduality, self-transcendence (state/trait content scales)

### MODTAS / Tellegen Absorption Scale — absorption trait
- **Original:** Tellegen & Atkinson (1974), Tellegen Absorption Scale (TAS), 34 items, **DOI 10.1037/t14465-000 (PsycTESTS)** `[VERIFIED]`. **Modified (MODTAS):** Jamieson (2005) `[UNVERIFIED citation metadata this pass; the Modified 34-item Likert version is standard]`.
- **Why it earns a battery slot:** the spec names **absorption vs. dullness** as a confusion pair (§3). The seed corpus (Vohryzek 2025; Laukkonen f-SNR) resolves it as *high epistemic depth + gathered precision (absorption) vs. low clarity-of-knowing (dullness)* — genuine absorption *sharpens* deviance-sensitivity. MODTAS is the validated trait handle on absorption capacity and correlates with jhāna access; it gives a convergent anchor for "can this person actually absorb" independent of self-narrated depth.
- **Verdict:** **Include in the battery.** Cheap, well-validated, discriminates a confusion pair Waypoint must handle.

### Hood Mysticism Scale (M-scale)
- **Instrument:** Hood (1975), *Mysticism Scale — Research Form D*, 32 items, **DOI 10.1037/t04627-000 (PsycTESTS)** `[VERIFIED]`; short version Spilka, Hood & Gorsuch (1985), **DOI 10.1037/t04629-000** `[VERIFIED]`. Built on Stace's common-core phenomenology; the standard structure is a two/three-factor solution (introvertive mysticism, extrovertive mysticism, interpretation) — factor structure and cross-cultural invariance are actively debated in Hood's own later work `[factor-count detail UNVERIFIED this pass]`.
- **Verdict for Waypoint:** **Include specifically as the ceiling instrument for rare 51+ (Unbinding/Enlightenment) subjects.** Its introvertive-mysticism items (ego loss, timelessness, unity, ineffability) are the only validated self-report vocabulary that reaches the top-of-wheel phenomenology the estimator must still recognize (spec §3). Near-useless for discriminating 22–45 (the app population), so it is a *targeted* slot, not a workhorse. Note the same reification caveat: M-scale endorsement is learnable vocabulary.

### NADA — Nondual Awareness Dimensional Assessment ★ strongest topical fit
- **Instrument & citation:** Hanley, Nakamura & Garland (2018), *The Nondual Awareness Dimensional Assessment (NADA): New Tools to Assess Nondual Traits and States of Consciousness Occurring Within and Beyond the Context of Meditation*, **Psychological Assessment 30(12):1625–1639, DOI 10.1037/pas0000615** `[VERIFIED]`.
- **Structure (verified from PMC full text this pass):**
  - **NADA-T (trait), 13 items, two dimensions:** *Self-transcendence* (9 items; e.g. "I experienced all notion of self and identity dissolve away", "my mind expanded into space") and *Bliss* (4 items; e.g. "I have experienced an all-embracing love"). Composite reliability: total **.93**, self-transcendence **.94**, bliss **.81**. Both load on a second-order nondual-awareness factor (bifactor ESEM across N=338/221/166).
  - **NADA-S (state), 3 items,** present-tense adaptation for immediate post-practice nondual shift. Reliability .67–.88 by form/timepoint.
  - **Known-groups:** meditators > non-meditators on all subscales (NADA-T total 38.49 vs 31.76, p<.001; N=528 PCA sample); scores rise with practice frequency/duration.
- **Verdict:** **Top pick for the battery.** It is the closest existing instrument to the wheel's central axis (non-dual depth). The self-transcendence/bliss split even mirrors the seed corpus's distinction between *epistemic* de-reification and *affective* by-products. Use **both** NADA-T (trait, slow-latent anchor) and NADA-S (state, phase anchor). Caveat: it is a self-report *content* scale — subject to the same vocabulary contamination — so it convergently validates, it does not adjudicate.

### Self-transcendence: Cloninger TCI-ST and ASPIRES
- **Cloninger TCI Self-Transcendence:** Cloninger, Przybeck, Svrakic & Wetzel (1994), Temperament and Character Inventory, **DOI 10.1037/t03902-000 (PsycTESTS)** `[VERIFIED]`. Self-Transcendence is one of three character dimensions (self-forgetfulness, transpersonal identification, spiritual acceptance). Well-normed, but a *personality-trait* frame — measures spiritual disposition, not attainment; known to conflate with magical thinking at high scores.
- **ASPIRES:** Piedmont & Toscano, *Assessment of Spirituality and Religious Sentiments (ASPIRES) Scale*, **DOI 10.1007/978-3-319-24612-3_87** `[VERIFIED]`. Spiritual Transcendence (connectedness, universality, prayer fulfillment) + Religious Sentiments. Cross-culturally validated as a personality-level facet.
- **Verdict:** **Pick at most one, optional.** Either gives a convergent self-transcendence-trait anchor, but both measure dispositional spirituality rather than wheel position and will not move with practice on Waypoint timescales. Prefer **NADA-T over these** if a slot must be cut. If included, TCI-ST for the normed-personality tie-in; ASPIRES if you want the connectedness/universality facet specifically.

---

## 3. Meditation-depth and awakening instruments

### Piron MEDEQ / MEDI — Meditation Depth Questionnaire / Index
- **Citation:** Piron, H. (2001), *The Meditation Depth Index (MEDI) and the Meditation Depth Questionnaire (MEDEQ)*, Journal for Meditation and Meditation Research 1:69–92 `[UNVERIFIED exact page metadata]`; reference-work summary Piron (2022/2025), *Meditation Depth Questionnaire (MEDEQ) and Meditation Depth Index (MEDI)*, in **Handbook of Assessment in Mindfulness Research, DOI 10.1007/978-3-030-77644-2_41-1** (and 2025 edition **10.1007/978-3-031-47219-0_41**) `[VERIFIED reference-work citation]`.
- **Structure (widely reported; not fully re-verified from source this pass — `[UNVERIFIED]`):** 30 items sorting session depth into **five ascending levels — (1) Hindrances, (2) Relaxation, (3) Concentration / Personal Self, (4) Essential/Transpersonal Qualities, (5) Non-Duality / Transpersonal Self.** Empirically applied in EEG work (e.g. alpha/theta inversely related to MEDI depth, Neuroscience of Consciousness 2021) `[VERIFIED that such a paper exists]`.
- **Verdict:** **Conceptually the closest published "depth ladder" to the wheel** and worth reading for item design of the *within-session* depth probes — its five-level ascent parallels the Wheel's own within-band progression. But it is a **session-depth** instrument (fast state), not a developmental-stage instrument (slow latent), and its top level collapses everything transpersonal into one bucket, so it cannot resolve bands 4→5→6. **Do not add to the battery as a stage measure;** mine it as prior art for state-depth probe wording (feeds the probe item bank, lane not-this-one).

### Sacchet lab — Advanced Concentrative Absorption Meditation / Jhāna (ACAM-J)
- **This is the most methodologically aligned live research program** for top-of-wheel phenomenology + neurophenomenology. Verified outputs this pass:
  - Chowdhury, van Lutterveld, Laukkonen, Slagter, Ingram & Sacchet (2023/2024), *Investigation of advanced concentrative absorption meditation (ACAM-J) …* — intensively-sampled single-expert case study; **bioRxiv → Cerebral Cortex 35(4):bhaf079, DOI via PubMed 40215476** (7T fMRI, functional-connectivity-gradient reorganization across jhānas) `[VERIFIED]`.
  - Potash, … Sacchet (2024/2025), *Integrated phenomenology and brain connectivity demonstrate changes in nonlinear processing in jhana advanced meditation*, **PMC11642738 / bioRxiv 2024.11.29.626048** `[VERIFIED]`; and a geometric-eigenmode analysis follow-up (ResearchGate 389586506) `[VERIFIED exists]`.
- **Method to copy (matches spec §6 and the seed-corpus digests):** time-locked first-person state labels (button-press) anchored to a **published cartography (J1–J8)** rather than ad hoc labels; distributions and stability over point estimates; effect sizes not p-values. The ACAM-J phenomenology instruments are the model for how Waypoint should structure *milestone* markers for the rare 51+ subject.
- **Verdict:** **Not a battery instrument** (it is an N=1–handful adept protocol, not a scalable questionnaire), but the **template for milestone-marker phenomenology** and a source of validated jhāna descriptors for the marker bank. Directly complements seed-corpus digests 5 (Vohryzek) and 6.

### Jeffery Martin — PNSE / "Fundamental Wellbeing" / Finders Course clusters
- **Claim:** a cross-tradition classification of "Persistent Non-Symbolic Experience" into four "locations" on a single continuum (Martin, *Clusters of Individual Experiences form a Continuum of Persistent Non-Symbolic Experiences in Adults*; conference/self-published venues, Center for the Study of Non-Symbolic Consciousness) `[VERIFIED that the work + continuum claim exist; peer-reviewed status weak]`.
- **Skeptical read (why the spec says "treat skeptically"):** (1) the "locations"/continuum construct may be a **clustering artifact** of self-report data with no independent structure; (2) headline attainment-rate claims come from a **paid course (~$3k) without control groups or independent replication**; (3) most output is via the author's own platforms/books/talks rather than mainstream peer review; (4) commercial incentive to over-claim. Multiple contemplative-community critiques (including from associates) call the model "deeply flawed."
- **Verdict:** **Do not use PNSE as a scoring scheme or ground-truth ladder.** It is worth exactly one thing to Waypoint: a reminder that a four-cluster "map of awakening" from self-report is *precisely* the Goodhart/artifact trap Waypoint must avoid — its failure mode is Waypoint's cautionary tale. If Surya finds the "locations" descriptions phenomenologically useful, treat them as informal marker candidates to be independently validated, never as a rubric.

### Lindahl & Britton — Varieties of Contemplative Experience (VCE)
- **Citation:** Lindahl, Fisher, Cooper, Rosen & Britton (2017), *The varieties of contemplative experience: A mixed-methods study of meditation-related challenges in Western Buddhists*, **PLoS ONE 12(5):e0176239, DOI 10.1371/journal.pone.0176239** `[VERIFIED]`.
- **Structure:** **59 meditation-related experiences across 7 domains** (cognitive, perceptual, affective, somatic, conative, sense of self, social) + **26 influencing factors across 4 domains** (practitioner-level, practice-level, relationships, health behaviors).
- **Verdict:** **The single best empirical taxonomy for the marker bank's difficulty/phase axis and the safety duty (spec §8, lane 4).** Its 7-domain × 59-experience grid is directly reusable as a controlled vocabulary for phase annotation and for the dark-night-vs-depression confusion pair. Its 26 influencing factors are exactly the covariates the aggregation layer should condition on. Not a scoring instrument — a phenomenological codebook. Ingest it.

---

## 4. Ego-development sentence-completion scoring — the load-bearing precedent

This is the methodological heart of the lane: the only mature discipline that assigns a human an **ordinal developmental stage by scoring the structure of their free-text language against a rubric**, with trained-rater reliability and a deterministic protocol-level aggregation rule. It is Waypoint's closest external ancestor.

### 4.1 The lineage
- **Loevinger WUSCT** — Washington University Sentence Completion Test. 36 sentence stems (e.g. "When people are helpless…"); each completion scored to an ego-development stage (E2–E9 / Impulsive → Integrated) against a manual; a **cumulative "ogive" rule** deterministically converts the distribution of item-level ratings into one Total Protocol Rating (TPR). Manuals: Loevinger & Wessler (1970); Hy & Loevinger (1996). `[UNVERIFIED exact citation metadata this pass; foundational and well established]`.
- **Cook-Greuter MAP** — extended the ceiling upward (Construct-aware, Unitive stages) and refined the scoring for post-conventional/transpersonal levels — i.e. exactly the band-4–6 territory generic ego-development scoring is thin on. Cook-Greuter (1999/2000 dissertation & *Journal of Adult Development* work) `[UNVERIFIED exact metadata; well established]`. Certification-based rater training; the MAP is the commercial descendant.
- **O'Fallon STAGES** — re-derives the whole ladder from **three orthogonal dimensions of language**, which is the structural insight most transferable to Waypoint:
  - **Citation `[VERIFIED]`:** O'Fallon, Polissar, Neradilek & Murray (2020), *The validation of a new scoring method for assessing ego development based on three dimensions of language*, **Heliyon 6(3):e03472, DOI 10.1016/j.heliyon.2020.e03472**.
  - **The three dimensions:** (1) **individual ↔ collective** (perspective), (2) **passive/receptive ↔ active/reciprocal** (agency), (3) **object sophistication: concrete → subtle → MetAware** (what kind of object the person operates on). Every stage = a repeating pattern across these three as they cycle. This turns stage-scoring from memorized exemplars into a **generative rule** — which is exactly what you want an LLM extractor to apply.
  - STAGES scores are reported to agree with the Cook-Greuter/Loevinger protocol (concurrent validity) `[VERIFIED as claimed in the validation paper]`; exact concordance correlation and inter-rater κ from the Heliyon paper could not be re-extracted this pass (publisher 403) — `[UNVERIFIED numeric]`.

### 4.2 Reliability & protocol rules (what "good" looks like)
- **Rater competence bar:** STAGES certifies scorers when scoring accuracy is **≥ ~85%** against a gold key `[VERIFIED from the LLM-replication paper's description]`. WUSCT/Cook-Greuter historically report high trained-rater inter-rater reliability (commonly cited in the .80s–.90s range) `[UNVERIFIED specific figures this pass — do not cite a number without pulling the manual]`.
- **Total-protocol rule (the piece Waypoint should copy literally):** you do **not** average item stages. The ogive rule assigns the protocol stage as the highest stage at which a threshold *cumulative proportion* of responses sits — an explicit, deterministic aggregation from many item-level judgments to one protocol-level stage. **This is the exact shape of Waypoint's "thin Bayesian aggregation" (spec §6):** the LLM scores each response's structure (the item-level judgment), and a deterministic layer combines them into a posterior — never the LLM adding up its own scorecard. Ego-development scoring is 50 years of evidence that this division of labor works.

### 4.3 The published LLM replication ★ direct evidence for Waypoint's estimator
- **Citation `[VERIFIED]`:** Bronlet, X. (2025), *Leveraging on large language model to classify sentences: a case study applying STAGES scoring methodology for sentence completion test on ego development*, **Frontiers in Psychology 16:1488102, DOI 10.3389/fpsyg.2025.1488102** (Integral Transpersonal Institute, Milan; PMC11839766).
- **Results (verified from full text this pass):**
  - **Weighted Cohen's κ (LLM vs. expert): 0.779 (95% CI 0.672–0.887) per single sentence; 0.705 (95% CI 0.686–0.725) for aggregated 10-sentence averages.** "Substantial" agreement.
  - **Models tested:** GPT-3.5-turbo, GPT-4, GPT-4-turbo, GPT-4o, Claude Sonnet, Claude Opus, LLaMA-3-70B. GPT-4o chosen as workhorse (context-window reasons); GPT-3.5-turbo lowest variance but weaker.
  - **Aggregation finding directly useful to Waypoint:** a minimum of **~5 sentences** is acceptable and **~10 sentences** gives an optimal stability/variance tradeoff (SD 0.138 at 10). This is a concrete floor for how many scored responses Waypoint needs before a per-pillar posterior is trustworthy — feeds the **abstention / data-sufficiency threshold** in the output object.
  - **Limitation:** test set = 58 sentences (generalizability caveat, author-stated).
- **Takeaways for Waypoint:**
  1. An LLM does the *structure-scoring* leg at near-trained-rater agreement (κ≈0.78) — **empirical green light for the marker-extractor** (spec §6 extraction).
  2. Keep the **extractor and any validation labeler on different model families** (spec §6(d)): this study literally cross-checked GPT vs Claude vs LLaMA — reuse that as the correlated-error control.
  3. Report **κ and CIs**, aggregate over ≥5–10 scored items before committing a stage, and treat single-item scores as high-variance — all directly portable to the synthetic-dossier QA (§7) and calibration metrics.

### 4.4 What structure-over-content scoring *buys* Waypoint
- **Immunity to vocabulary contamination (spec §3, lane 5).** Content-scoring rewards saying "anger arose and passed"; structure-scoring asks whether the *whole response is organized* the way that stage organizes experience (object sophistication, agency, perspective). A practitioner who has read the map can borrow the phrase but not (easily) fake the structural signature across many unrehearsed responses. This is the single strongest countermeasure to reading-ahead contamination and it is battle-tested.
- **A generative rubric that rides model upgrades (spec's "fat skill, thin layer").** STAGES' three-dimensional rule is compact enough to put in a prompt and general enough that a better model applies it better — versus a frozen keyword list.
- **A legitimacy precedent for the whole enterprise.** "Score a human's developmental stage from free language against a rubric, with reliability and error bars" is not novel or fringe — it is a 50-year psychometric tradition with a 2025 LLM replication. Waypoint is that method, pointed at the meditative wheel instead of the ego-development ladder, with a Bayesian aggregator instead of an ogive.

---

## 5. Which items/instruments survive contact with wheel bands 4–6 (Awareness → Unbinding)

Verdict table. "Survives" = still carries valid, discriminating signal at that band rather than flooring/ceiling or collapsing to one bucket.

| Instrument | Band 4 (Awareness 31–40) | Band 5 (Awakening 41–50) | Band 6 (Unbinding 51–60) | Net |
|---|---|---|---|---|
| FFMQ | partial (DIF already biting) | ceiling + DIF | fails | Convergent anchor only, low bands |
| MAAS | floors above basic | fails | fails | Skip |
| TMS (Decentering) | good | partial | fails | Optional decentering anchor |
| MODTAS | good (absorption capacity) | good | partial | Include (confusion-pair discriminator) |
| Hood M-scale | mostly floor | partial (introvertive items wake up) | **survives** (ceiling instrument) | Include as 51+ ceiling probe |
| NADA-T / NADA-S | good | **good** | partial (self-report of the ineffable) | **Include — best central-axis fit** |
| TCI-ST / ASPIRES | trait-flat | trait-flat | trait-flat | Optional, one at most |
| MEDEQ/MEDI | good (state depth) | good | collapses (one transpersonal bucket) | Mine for probe wording, not battery |
| **Sentence-completion (STAGES-style, structure-scored)** | **good** | **good** | **good if stems reach far enough** | **Core differentiator — include** |

The pattern is stark: **fixed-item self-report content scales degrade exactly where Waypoint most needs resolution (bands 5–6)**, while **structure-scored free response degrades gracefully** because you can extend the stems and the rubric upward without re-norming a whole instrument. This is the empirical case for the estimator architecture the spec already chose.

---

## 6. The one caveat that dominates everything: response-shift / DIF

Every content self-report instrument here shares one failure mode, and it is the spec's central confusion pair (§3, "insight vocabulary vs lived insight"):
- **The measurement invariance is not there across the range.** Van Dam 2009 (FFMQ DIF) and Ireland 2019 (TMS invariance test) are the two cleanest empirical demonstrations. "Mind the Hype" (Van Dam 2017) generalizes it.
- **Mechanism:** advancing practitioners *recalibrate the meaning of the items* (a 45%-wheel practitioner rating "I notice my thoughts" is answering a different question than a 25% one), so a raw-score increase can reflect either genuine development **or** a shifted response frame — unidentifiable from the score alone.
- **Consequence for Waypoint, stated plainly:** you cannot build the estimator spine out of self-report deltas. You *can* use these scales, once, as convergent anchors with the DIF explicitly modeled (e.g. treat battery scores as noisy, band-conditional evidence, not as a metric with a stable zero). This is precisely why the estimator's spine is **LLM structure-extraction over conversation** (immune-ish to DIF) with self-report as one bounded evidence channel — and why the **interview-vs-dossier delta** (§7 headline number) is the right way to measure the information ceiling.

---

## 7. Recommended one-time calibration battery (spec §7 deliverable)

**Design principles:** convergent validity across the three axes the wheel actually spans (mindfulness/attention, absorption, nonduality/self-transcendence) + one structure-scored free-response protocol that is the real methodological bridge to the estimator. Keep it under ~75 min so it fits one calibration sitting. Calibration-cohort-only; never a product surface (spec §5).

| Slot | Instrument | Version / length | Time | Why it's in |
|---|---|---|---|---|
| 1 | **NADA-T + NADA-S** | Hanley/Nakamura/Garland 2018; 13 + 3 items | ~8 min | Best fit to the wheel's central non-dual axis; trait (slow latent) + state (phase) in one family; α .81–.94 |
| 2 | **MODTAS** | Modified Tellegen Absorption Scale, 34 items | ~10 min | Absorption-capacity anchor; discriminates absorption↔dullness confusion pair |
| 3 | **FFMQ** | 39 items | ~12 min | Dispositional-mindfulness convergent anchor for low-mid range — *scored with DIF caveat, deltas not trusted* |
| 4 | **Hood M-scale** | Research Form D, 32 items (or short form) | ~10 min | Ceiling instrument: the only validated vocabulary reaching 51+ introvertive-mysticism phenomenology |
| 5 | **Self-transcendence trait** | TCI-ST subscale *or* ASPIRES Spiritual Transcendence | ~8 min | Convergent self-transcendence anchor; pick one; NADA-T can substitute if cutting |
| 6 ★ | **Short WUSCT/STAGES-style sentence-completion protocol** | ~12–18 stems, structure-scored (three-dimensions rubric) | ~15–20 min | **The differentiator.** Directly exercises the estimator's scoring method on the same subject; produces the structure-scored gold that the LLM extractor is validated against; ≥5–10 completions gives a stable per-subject read (Bronlet 2025) |

**Total ≈ 55–75 min.** Optional add if time allows: TMS state pre/post a short sit (decentering anchor). Explicitly **excluded:** MAAS (floors), MEDEQ as a *stage* measure (it's a state-depth scale — mine its items for probes instead), PNSE/Finders anything (not an instrument).

**Two design notes the battery must honor:**
1. **The sentence-completion protocol (slot 6) is not just convergent validity — it is a second, independent instantiation of Waypoint's own method** applied by a human-scorable procedure. It is the cleanest place to measure "does the LLM extractor agree with a trained structure-scorer" (target: κ in the 0.7–0.8 band that Bronlet 2025 and STAGES certification set as the bar). Surya calibrates the *meditative* rubric; the STAGES three-dimension frame calibrates the *scoring discipline*.
2. **Score every self-report instrument as band-conditional, DIF-aware evidence,** not as a metric with a stable zero (see §6). The battery's job is to triangulate Surya's gold label, not to be a competing ground truth.

---

## 8. Open items / verification debts for the finder→verifier pass

- `[UNVERIFIED numeric]` Exact inter-rater κ / concordance-with-Cook-Greuter figures in **O'Fallon 2020 Heliyon** (publisher 403 this pass). Pull from the PDF before citing any number.
- `[UNVERIFIED]` **WUSCT/Cook-Greuter** classic inter-rater reliability figures — do not cite a specific coefficient without the manual (Hy & Loevinger 1996) in hand.
- `[UNVERIFIED]` Exact **Van Dam 2009** per-facet DIF result (which facets); the DIF finding itself is verified, the facet breakdown is not.
- `[UNVERIFIED]` **MEDEQ** exact five-level names + item count (30) — reported from secondary sources; confirm against Piron 2001 primary before the marker bank borrows its wording.
- `[UNVERIFIED citation metadata]` MAAS (Brown & Ryan 2003), TMS (Lau et al. 2006), MODTAS-Modified (Jamieson 2005), Cook-Greuter dissertation — all real and standard, but not re-confirmed via Crossref this pass; verify DOIs before the literature dossier locks.
- **Cross-lane handoff:** VCE (Lindahl & Britton 2017) 7-domain/59-experience grid → lanes 2, 4, and the marker bank; ACAM-J phenomenology → milestone markers + seed-corpus lane 6; STAGES three-dimension rubric → estimator design memo (this is the template for the extractor prompt).

---

## 9. Sources (verified this pass)

- Van Dam, Earleywine & Danoff-Burg (2009), *Differential item function across meditators and non-meditators on the FFMQ*, Personality and Individual Differences 47(5):516–521. https://doi.org/10.1016/j.paid.2009.05.005
- Van Dam et al. (2017), *Mind the Hype*, Perspectives on Psychological Science 12(6). https://doi.org/10.1177/1745691617709589
- Ireland, Day & Clough (2019), TMS measurement invariance across meditation experience, J. Clinical Psychology. https://doi.org/10.1002/jclp.22709
- Tellegen & Atkinson (1974), Tellegen Absorption Scale, PsycTESTS. https://doi.org/10.1037/t14465-000
- Hood (1975), Mysticism Scale Research Form D, PsycTESTS. https://doi.org/10.1037/t04627-000 ; short form Spilka/Hood/Gorsuch (1985) https://doi.org/10.1037/t04629-000
- Hanley, Nakamura & Garland (2018), NADA, Psychological Assessment 30(12):1625–1639. https://doi.org/10.1037/pas0000615 ; full text https://pmc.ncbi.nlm.nih.gov/articles/PMC6265073/
- Cloninger, Przybeck, Svrakic & Wetzel (1994), TCI, PsycTESTS. https://doi.org/10.1037/t03902-000
- Piedmont & Toscano, ASPIRES, Encyclopedia of Personality and Individual Differences. https://doi.org/10.1007/978-3-319-24612-3_87
- Piron (2022/2025), MEDEQ/MEDI, Handbook of Assessment in Mindfulness Research. https://doi.org/10.1007/978-3-030-77644-2_41-1
- Lindahl, Fisher, Cooper, Rosen & Britton (2017), Varieties of Contemplative Experience, PLoS ONE 12(5):e0176239. https://doi.org/10.1371/journal.pone.0176239
- O'Fallon, Polissar, Neradilek & Murray (2020), STAGES validation, Heliyon 6(3):e03472. https://doi.org/10.1016/j.heliyon.2020.e03472
- Bronlet (2025), LLM STAGES scoring case study, Frontiers in Psychology 16:1488102. https://doi.org/10.3389/fpsyg.2025.1488102
- Sacchet lab ACAM-J: Cerebral Cortex 35(4):bhaf079 (PubMed 40215476); Potash et al., PMC11642738 / bioRxiv 2024.11.29.626048.
