# Waypoint — Practitioner Stage Assessment (SOME)

**Status:** draft v0.1 — interview-derived, pending research phase
**Date:** 2026-07-16
**Owner:** Fionn (AI Lead). **Domain authority / teacher anchor:** Surya.
**Umbrella:** SOME (Stages of Meditation Eval). The existing SOME app maps *traditions* onto the Wheel; **Waypoint** is the component that estimates where a *human practitioner* sits on the same Wheel.
**Related:** `specs/predictive-profile-harness-v3-design.md` · `specs/waypoint-research/inputs/fable-max-bayesian-profiling-eval.md` (Fable-max review of the Bayesian profiling design — its principles bind here) · `benchmarking/stages-of-meditation-eval/data/wheel-canvas/seed.json` (Wheel Table import) · `archive/old-transcripts/Meditation Path and Poject Overview.txt` (Surya call) · `content/source/meditation-path/0. Stages of Meditation/` (Wheel of Life.png, Stages of Meditation.pdf, PATH TIMELINE.xlsx) · `specs/constitutional-classifier-spec.md` · BodhisattvaBench (judging methodology precedent)
**Seed corpus (Ruben Laukkonen & collaborators — LIFE-affiliated science anchor):** `specs/waypoint-research/inputs/` — *From many to (n)one* (2021), *A beautiful loop* (2025), *Cessations of consciousness / nirodha samāpatti* (2023), *Nothingness in meditation* (Agrawal & Laukkonen), *Whole-brain models of MPE via jhāna* (Vohryzek et al. 2025), *Clear Mind: f-SNR* (2026). These ground the construct in predictive processing — the same frame Surya's model already leans on.

---

## 1. What this is

Waypoint estimates a practitioner's position on the Wheel of Life / stages-of-meditation path from evidence the LIFE system already sees — conversations with Wisdom, calibrated conversational probes, and practice telemetry. It outputs a probability distribution over wheel bands plus a per-pillar profile (sensations / emotions / thoughts / awareness), with stated uncertainty and cited evidence. The estimate is **internal-only**: it feeds Wisdom's teaching adaptation, teacher triage, and validation of the path model itself. It is never shown to practitioners in v1.

Surya's framing (2017→): the path is quantifiable — "only numerically can you have enough detail benchmarking how the meditation path looks like." Waypoint is the instrument that makes that numeric ambition operational for real humans, with error bars.

## 2. Decisions locked (interview, 2026-07-16)

| # | Decision | Choice | Implication |
|---|----------|--------|-------------|
| 1 | Consumers | Wisdom adaptation + teacher triage + research validation | Never practitioner-facing in v1; estimates are hidden AI-notes-class data |
| 2 | Output object | Distribution over wheel bands + per-pillar profile | No single-scalar claims; finest resolution in the 22–45% range where app users live |
| 3 | Ground truth | Triangulated calibration | Surya gold labels + one-time instrument battery + longitudinal behavior; divergences recorded, not overridden |
| 4 | Evidence channels | Passive conversation + active conversational probes + practice telemetry | No standalone questionnaire in-product; probes are woven invisibly into Wisdom conversation |
| 5 | Estimator | LLM marker extraction → thin Bayesian aggregation | Fat prompt / thin deterministic layer; rides model upgrades |
| 6 | Calibration cohort | Team + Fionn now; alpha users ~2027 | Sangha students not in v1 scope |
| 7 | Gold-label protocol | Surya structured interview AND blinded dossier-only rating per subject | The interview-vs-dossier delta measures the information ceiling of any data-only estimator |
| 8 | Cadence | Rolling per-session nowcast + monthly deep consolidation | Mirrors slow-profile / fast-state |
| 9 | Substrate | One Bayesian system with the predictive-profile harness | Waypoint extends the harness's hierarchy: stage is the slowest latent above slow-profile and fast-state; one evidence log, one integrator, two consumers |
| 10 | Safety duty | Stage-aware classifier context + teacher flag from consolidation | Human outreach only via LIFE Team identity, never Wisdom |
| 11 | Research deliverables | Literature dossier + marker bank v0 + probe item bank v0 + estimator design memo | All four required before build |
| 12 | Name | SOME umbrella; component name **Waypoint** | Spec/pipeline/surface named accordingly |

## 3. The construct being estimated

**The Wheel.** 1–100% scale, 10 bands (Emergence 1–10 … Infinity 91–100), as imported from the Wheel Table. Humans occupy roughly **22–62** in practice (Survival → Enlightenment-integration). Expected app population: mostly **22–45** — pre-path (Survival→Wisdom, 22–30), Awareness (31–40), early Awakening (41–50). Unbinding (51–60) onward is live-transmission territory and will be rare in-app; the estimator must still recognize its phenomenology (it carries the highest support needs).

**Pillars.** Development is uneven across the four practitioner toolkits — sensations, emotions, thoughts, awareness (the Stage-2 paths: breathwork / love / clarity / intrinsic awareness). A practitioner can be 2.4-deep in awareness work with untrained emotional integration. A single scalar hides exactly what Wisdom needs for adaptation, so the pillar profile is first-class, with an overall band distribution derived from it plus cross-pillar evidence.

**Position vs phase.** Stage is a slow variable; phenomenological *phase* is fast (glimpse, plateau, dark-night-adjacent dip, integration, intensification). Waypoint tracks both: phase annotations contextualize evidence (a dark-night week is not regression) and drive the safety duty.

**Known confusion pairs** (the marker bank must discriminate, not just detect):
- equanimity vs spiritual bypassing vs dissociation
- dark night / low-meaning phase vs clinical depression
- absorption vs dullness
- insight *vocabulary* vs lived insight — practitioners who have read the map describe experiences in its terms (reading-ahead contamination)
- devotional/mystical intensity vs destabilization

## 4. Output object (per practitioner)

Canonical snapshot written by each consolidation pass:

```
waypoint_snapshot {
  practitioner_id, as_of, evidence_window,
  overall: { band_posterior: p[10 bands], point: 1–100, credible_interval },
  pillars: { sensations|emotions|thoughts|awareness: { band_posterior, point, ci } },
  phase: { state: glimpse|plateau|dip|integration|…, confidence },
  milestones: [ { marker_id, e.g. reverse_breathing | first_glimpse | small_death,
                  status: claimed|corroborated, evidence_refs[], confidence } ],
  rationale: cited free text (quotes + telemetry references),
  data_sufficiency: ok | insufficient   // below threshold → abstain, never guess
}
```

Rolling nowcast updates the same shape between consolidations, marked non-canonical. Full snapshot history is retained (longitudinal trajectory is itself evidence and a validation target).

## 5. Evidence channels (v1)

1. **Passive conversation analysis.** Everything the practitioner already says to Wisdom (chat, voice journaling): phenomenological reports, how experience is described (not just what is claimed), reactivity in the exchange itself, vocabulary trajectory.
2. **Active conversational probes.** Calibrated items Wisdom weaves into natural conversation (sentence-completion-test spirit: score the *structure* of the response, not agreement). Delivery constraints: invisible (never framed as assessment), frequency-capped, never during flagged-vulnerable moments, never repeated verbatim, logged with item-id for scoring. Item bank drafted by the research phase, calibrated and approved by Surya before use.
3. **Practice telemetry.** Practice selection, session counts/lengths, consistency, gate progression (e.g. breathwork sequence position), retention, time-of-day patterns. Weak on inner states, strong on discipline and trajectory; anchors consistency checks (claims vs behavior).

**Explicitly out of scope in-product:** a standalone questionnaire. Established psychometric instruments are used **once, calibration-cohort-only**, for convergent validity (§7) — they never become a product surface.

## 6. Estimator architecture

Fat prompt / thin deterministic layer:

- **Extraction (LLM, Fable-class).** Per session: read transcript + telemetry delta → emit typed **marker events** `{marker_id, polarity, confidence, quote/evidence ref}` against the marker bank. All judgment lives here; prompts live in the unified prompt registry.
- **Aggregation (deterministic).** Bayesian update over the latent position: priors from the path progression model (base rates by tenure, plausible movement speeds — no teleporting from 30 to 55 in a week), per-marker likelihoods from the marker bank (expert-set initially by research phase + Surya; refit once labeled data exists), time decay, pillar coupling, consistency penalties when claims and behavior diverge. Abstention below evidence thresholds. Cold start: prior over the app population range, wide.
- **Cadence.** Cheap extraction every session updates the rolling posterior; monthly consolidation re-reads accumulated evidence end-to-end, reconciles contradictions, writes the canonical snapshot + rationale.
- **Substrate: one Bayesian system.** Waypoint is an extension of the predictive-profiling system, not a sibling. The harness's temporal hierarchy (slow response signatures → medium relationship/practice-arc state → fast session state) gains a slowest level: **wheel position**. Prediction errors spend themselves on fast state first and revise slow latents only through persistent runs of error (hierarchical-Gaussian-filter style) — a practitioner does not teleport from 30 to 55 in a week. Principles carried over from the Fable-max profiling review, binding here too: **(a)** LLM proposes (markers, hypotheses, regime features), the mechanical layer integrates and keeps score — never let the judge add up its own scorecard; **(b)** **partial pooling** across practitioners (population prior → practice-arc archetype → individual), so small-N evidence borrows strength and individuation is measurable drift from the archetype; **(c)** score any prediction at its own horizon against a base-rate null — never reward durability itself; **(d)** keep the marker-extractor and any validation labeler on **different model families** to avoid correlated errors. Extraction shares the harness's evidence/event log; harness reaction-probes double as stage evidence where scoring keys overlap.
- **Design stance: constructs on trial, prediction is ground truth.** Named dimensions (stage bands, pillars, "equanimity") are human-legible compressions of behavioral regularity — priors and interface, not ontology. The bottom of the hierarchy is always behavior-predicting-behavior at explicit horizons, scored against base-rate/persistence nulls; a named construct earns its slot at the slow levels only while it adds predictive skill there (and stays because the humans in the loop — Surya, the safety lane — and the calibration protocol need nameable state to act on and label). The extractor may propose model-discovered candidate structure (new markers, regimes, rival dimensions) into the hypothesis ledger; candidates are promoted by predictive skill, and a learned structure that consistently out-predicts a Wheel construct at its own targets is a recorded divergence → path-model revision conversation with Surya, never a silent swap. Psychometric *instruments* remain calibration-only anchors (semantic drift and ceilings make them unreliable for this population as a product channel).
- **LLM-as-prior & model policy** *(lane-7-verified; full verdict table in `waypoint-research/lanes/7-llm-person-modeling.md`)*. The LLM extractor is itself the tap into a pre-trained person-model — pretraining is the "centuries of observation" prior at scale; the marker taxonomy is a readout layer over its latents, and L1 next-message prediction taps it directly. But it was never trained on our target conditional, and its verbalized confidences are systematically **over**confident (RLHF-amplified; ≈88% stated vs ≈79% correct in published measurements) — so it proposes, the mechanical layer integrates, and extractor confidence enters as an *ordinal feature with an empirically fitted per-marker mapping*, never as a raw probability. **Extraction house pattern (Bronlet 2025):** decomposed structural rubrics, temperature 0, median-of-N runs — GPT-4o reaches weighted κ≈0.78 vs expert raters on Cook-Greuter-lineage sentence completions with zero fine-tuning; **κ 0.7–0.8 vs Surya is the realistic v1 agreement target.** Model policy: **v1 uses frontier zero-shot extractors — no fine-tuning** (nothing to tune on until gold labels exist; N≈10 would swamp the prior), with three targeted open-model exceptions: open-family cross-extractor for correlated-error control; self-hosted trajectory for inner-life data governance; and **measurement stability** — no formal LLM measurement-invariance framework exists in the literature, so ours is defined here: pin estimator model versions like prompts are hash-pinned; every model bump is a recalibration event gated by golden-transcript regressions with pre-registered per-marker equivalence bounds; keep one archived open-weight extractor as the long-horizon reproducible reference. **Tuning ladder once labels exist:** first a LoRA'd extractor on gold marker labels (Mental-LLM precedent: +5–11% over frontier zero-shot with only thousands of labels), later a reaction forecaster; Centaur-scale behavioral tunes need orders of magnitude more data than alpha will produce; nothing tuned ever becomes the integrator. Prior art: Centaur (Binz et al. 2025); Park et al. 2024 (rev., "LLM Agents Grounded in Self-Reports": interview-seeded agents reproduce participants' own answers at 83–86% of two-week self-consistency); USER-LLM/LaMP user modeling; language-based assessment (Kosinski 2013; Park 2015); LLM structural scoring (Bronlet 2025). **No published computational estimator of contemplative stage exists — Waypoint is first**: no external likelihoods to crib, no benchmark to import, which is why the interview-vs-dossier ceiling is the headline v1 number.
- **v1 execution mode:** offline pipeline over exported dossiers (Data-Lab-style exports), not a runtime service. Runtime integration into life-app comes after calibration validates (§12).

## 7. Calibration & validation protocol

**Cohort:** LIFE team + Fionn (now); alpha app users when the alpha group exists (~2027, consent-gated).

Per subject:
1. Compile the **evidence dossier** (conversations, probe responses, telemetry) over a defined window.
2. **Surya structured interview** → gold label: overall band + pillar ratings + milestone judgments + rationale. Interview protocol is a research-phase deliverable (derived from wheel markers; recorded and itself ingested as evidence).
3. **Surya blinded dossier-only rating** of the same subject (order counterbalanced, interview and dossier ratings separated in time). The **interview-vs-dossier delta is the headline v1 number**: the information ceiling of any data-only estimator.
4. **One-time instrument battery** (selection is a research-phase output) for convergent validity — calibration-only.
5. **Estimator run** on the dossier → compare.

**Metrics:** estimator-vs-interview agreement (band-level accuracy, rank correlation, pillar-level agreement), uncertainty calibration (do 80% intervals contain gold 80% of the time), dossier-vs-interview ceiling, longitudinal movement plausibility. Human-vs-estimator divergences are recorded as **open questions with both positions** (house label-divergence convention), not silently overridden in either direction.

**Engineering QA before humans (default, strike if unwanted):** synthetic-dossier discrimination test — personas authored at known wheel positions from the marker bank; the estimator must separate and rank them. QA of the instrument only; never calibration data.

**Known validity threats, tracked from day one:** N is tiny; single-rater anchor (Surya) with no inter-rater check in v1; team cohort is stage-range-narrow and map-contaminated (everyone has read the Wheel); probe reactivity; self-presentation distortion in both directions (spiritual humility deflates, enthusiasm inflates).

## 8. Safety & ethics

- **Internal-only.** No practitioner-facing estimates, no leaderboards ever, no gamification of stage. Estimates are hidden AI-notes-class data (included in GDPR export, invisible in-app).
- **Wisdom use is advisory in v1:** pacing, practice selection, language register, probe scheduling. **No automated hard gating** of content on the estimate until calibration validates.
- **Stage-aware safety.** Waypoint context (band + phase) feeds the constitutional classifier's context so thresholds are stage-aware (it already knows practice stage; this sharpens it — dark-night-adjacent phenomenology around Unbinding reads differently from the same words at 25%).
- **Teacher flag.** The consolidation pass may raise a teacher-dashboard flag ("entering unbinding-adjacent territory; elevated support need"). Flags are supportive-triage signals, not diagnoses; the classifier remains the safety authority for acute risk; human outreach happens under the **LIFE Team identity, never Wisdom** (existing rule).
- **Consent & governance.** Team cohort: informed consent — dossiers expose inner life to Fionn + Surya; a team member can decline or exit with dossier deletion. Alpha users: explicit enrollment consent before any Waypoint processing. KB/data-governance rules apply to any derived content.

## 9. Placement

- **Spec (this file):** `~/LIFE/specs/waypoint-practitioner-stage-assessment.md` — canonical.
- **Research outputs:** `~/LIFE/specs/waypoint-research/` (corpus notes, lane dossiers, deliverables, external reviews, HTML report — precedent: `specs/llm-profiling-research/`). Seed corpus staged under `inputs/`; acquired literature PDFs under `library/` (catalog in `library/LIBRARY.md`, unobtainable items listed in `library/MISSING.md` for manual retrieval).
- **v1 pipeline:** offline scripts + prompts under the SOME repo (exact layout at build time; house patterns apply — registry-routed prompts, no keys in code).
- **Workbench surface (later):** a Practitioners section in the SOME app — cohort dossiers, snapshots plotted on the Wheel Canvas alongside the tradition blocks, divergence review. Hub-SSO'd, admin/teacher-gated.
- **Runtime (post-validation):** life-app server, as a consumer of the shared profile substrate per harness v3.

## 10. Research phase (next step — gated on spec approval)

Multi-agent research pass before anything is built. Fable agents at tiered efforts (max where synthesis/verification quality is load-bearing; xhigh/high for domain lanes; medium/low for gathering and formatting), plus **Oracle (GPT Pro)** for a heavyweight one-shot review and **Codex (effort ultra)** for an independent second-perspective sweep. Finder→verifier convention applies: every empirical claim in the deliverables gets an adversarial check before it lands.

**Lanes:**
1. **Traditional cartographies & how teachers actually assess** — Visuddhimagga ñanas / progress-of-insight interviews, bhumi descriptions, Zen ox-herding + koan checking (sanzen), Neidan attainment signs, Sufi maqamat, Hawkins LoC (and its critiques), plus the internal canon: Surya call transcript, Wheel Table, Stages of Meditation.pdf, PROGRESSION OF THE WORK.pdf. Key question: what do traditions treat as *observable* evidence of stage, and what do they insist only a teacher can see?
2. **Modern psychometrics & scoring methodology** — mindfulness/awakening/nondual/mysticism scales (FFMQ, MAAS, MODTAS, Hood M-scale, NADA, self-transcendence), **Cook-Greuter / ego-development sentence-completion scoring** (closest methodological precedent for scoring humans from language), Lindahl & Britton's Varieties of Contemplative Experience, Sacchet's advanced-meditation research program, jhana/cessation neurophenomenology, Jeffery Martin's PNSE clusters (with skepticism). Key question: which instruments/items survive contact with Stages 4–6 phenomenology, and what do their validity studies teach us.
3. **Statistical machinery** — IRT, latent-state/HMM over longitudinal evidence, Bayesian knowledge tracing, experience-sampling methodology, calibration/abstention design for small-N. Feeds the estimator design memo.
4. **Adverse events & safety duty** — meditation-related difficulties literature; how contemplative and clinical framings distinguish dark night from depression; what a defensible teacher-flag design looks like.
5. **Gaming, contamination & Goodhart** — faking-detection in psychometrics, demand characteristics, reading-ahead vocabulary contamination, and design countermeasures (structure-scoring over content-scoring, behavior-claim consistency).
6. **Contemplative neuroscience & predictive processing (seed corpus)** — deep-read the Laukkonen corpus (`waypoint-research/inputs/`): many-to-(n)one's FA/OM/ND continuum, epistemic depth, cessation/nirodha framework, emptiness-vs-cessation differentiation, jhāna criticality milestones, f-SNR. Mine the reference lists: which cited literature recurs, which is treated as load-bearing, which measurement approaches the network itself trusts — this seeds the psychometrics and statistics lanes with the field's own best-of list.

**Deliverables (all required):**
1. **Literature dossier** — filtered for usable findings; every claim source-cited.
2. **Marker bank v0** — per wheel band × pillar: observable linguistic/behavioral indicators, polarity, initial likelihood weights, plus the confusion table (§3).
3. **Probe item bank v0** — items with scoring keys + delivery constraints, ready for Surya calibration.
4. **Estimator design memo** — state space, likelihoods, priors, update rules, cold start, abstention, consistency checks; Oracle + Codex independent reviews attached.

**Open questions the research must answer:**
- What granularity does evidence actually support — band-level, or finer within 22–45?
- Which established instruments (if any) are worth the calibration battery slots?
- How do ñana-interview and koan-checking traditions structure *verification questions* — and what transfers to probe design?
- What is the defensible mapping between "difficult territory" markers and flag thresholds?
- Where does the Wheel's structure disagree with the strongest empirical work, and does anything in the model need Surya's revisiting?

## 11. Risks

| Risk | Stance |
|------|--------|
| Construct validity (is "wheel position" real & measurable) | Research lane 1–2; triangulation; treat v1 as instrument-building, not truth-claiming |
| Tiny N, narrow stage range in cohort | Report intervals, not points; synthetic QA; expand cohort in 2027 |
| Single-rater anchor | Log rationale; revisit inter-rater when a second qualified rater exists |
| Vocabulary contamination / faking | Structure-scoring, behavior-consistency checks, lane 5 countermeasures |
| Probe reactivity (measurement changes practice) | Frequency caps; probes double as legitimate teaching moves |
| Goodhart (estimate leaks into incentives) | Internal-only, no gamification, advisory-only adaptation |
| Privacy of inner life | Consent, deletion path, AI-notes-class hiding, governance rules |

## 12. Milestones

- **M0** — this spec approved.
- **M1** — research phase (multi-agent) → four deliverables + spec revision.
- **M2** — Surya calibrates marker bank + probe items; interview protocol agreed.
- **M3** — offline estimator v0; synthetic-dossier QA passes.
- **M4** — gold-label study on team cohort (interview + blinded dossier rating + battery + estimator run).
- **M5** — calibration report: agreement, uncertainty calibration, information ceiling. Go/no-go on runtime integration + teacher-flag activation.
- **M6** — alpha-cohort extension (~2027), consent-gated; refit likelihoods on labeled data.
