# Lane 5 — Gaming, Contamination & Goodhart Countermeasures

**Status:** research deliverable v0.1 — 2026-07-16
**Feeds:** marker bank v0 (confusion table + negative-polarity markers), probe item bank v0 (item formats + delivery constraints), estimator design memo (aggregation penalties, credibility weights), consolidation review checklist, governance.
**Spec anchors:** §3 (confusion pairs, esp. "insight vocabulary vs lived insight"), §5 (probe delivery constraints), §6 (consistency penalties, model-family separation), §7 (validity threats: map-contaminated cohort, self-presentation distortion in both directions), §11 (risk table rows: vocabulary contamination/faking, probe reactivity, Goodhart).

**Verification note:** every citation below was checked against a live source during this research pass (PubMed / publisher / repository page) unless explicitly marked UNVERIFIED or SPECULATION. No claim rests on an unverified citation.

---

## 1. Threat model — what actually gets gamed, by whom, in which direction

Waypoint's gaming surface is unusual. There is no external incentive (internal-only estimate, no leaderboard, no gate on content in v1), so the classic high-stakes faking literature (job selection, forensic malingering) over-predicts adversarialness. What remains is subtler and mostly **non-deliberate**:

| # | Threat | Direction | Mechanism | Deliberate? |
|---|--------|-----------|-----------|-------------|
| T1 | **Reading-ahead vocabulary contamination** | inflates (mostly) | Practitioner has studied the Wheel / Laukkonen corpus / pragmatic-dharma maps and describes experience in map terms. The entire v1 calibration cohort is contaminated by construction (spec §7). | No — honest people do this |
| T2 | **Spiritual self-enhancement** | inflates | "I am more aware than most people" is a measurable, *training-correlated* trait (Vonk & Visser 2021). Enthusiasm, identity investment, communal narcissism. | Mixed |
| T3 | **Spiritual-humility deflation** | deflates | Trained modesty norms ("I'm a beginner", non-attainment framing); also genuine under-noticing — cessations are characteristically *under*-claimed (Laukkonen 2023/2025 corpus notes) | No |
| T4 | **Dark-night dramatization / difficulty signaling** | distorts phase + band | Difficult-territory vocabulary is also learnable; support-seeking can amplify it. Faking-bad analogue. | Mixed |
| T5 | **Demand characteristics of Wisdom itself** | inflates + homogenizes | Wisdom *teaches the map*. Practitioners learn what "progress" sounds like from the very agent that assesses them, then produce it (Orne 1962). The assessor is the contamination source — this is Waypoint's single largest structural threat. | No |
| T6 | **Probe learning / item leakage** | degrades items | Repeated or shared probes become known; answers get scripted. The Zen precedent: once koan answers were published (Hoffmann 1975), the static item bank was permanently burned. | Sometimes |
| T7 | **Extractor-side Goodhart** | inflates marker counts | The extraction LLM is asked to find markers, so it finds them (finder bias); reward-hacking its own rubric. | N/A (system-internal) |
| T8 | **Estimate-leak second-order gaming** | inflates | Wisdom adapts on the estimate → practitioner perceives the adaptation → learns which talk unlocks "deeper" treatment. Goodhart through the advisory loop, even with no visible score. | Emergent |

Two asymmetries that must shape every countermeasure:

- **Asymmetric burden:** penalize *unsupported claims*, never *absence of claims*. T3 means silence is weak evidence of absence (cessation under-claiming, humility norms). Countermeasures below are claim-gated: they fire on positive assertions, not on quiet practitioners.
- **Non-adversarial default:** most contamination is sincere. Countermeasures should be framed (and prompted) as *evidence-quality weighting*, not lie detection. A "faking" frame in the extraction prompt will produce false-positive cynicism.

---

## 2. Evidence review by countermeasure family

### 2.1 Structure-over-content scoring (production paradigm) — STRONG

The Loevinger/Cook-Greuter sentence-completion tradition is the closest methodological precedent for scoring developmental stage from free language, and it carries the single most decision-relevant faking result in this lane:

- **Redmore, C. (1976). Susceptibility to faking of a sentence completion test of ego development. *Journal of Personality Assessment*, 40(6), 607–616.** (PMID 16367343). Faking *down* is easy; faking *up* is hard. Instructed fake-good subjects raised WUSCT ego-level scores only modestly, and mainly when given specific information about the scoring theory. You cannot reliably *produce* structural complexity you do not possess; you can only suppress it.
- **Loevinger, J. (1979). Construct validity of the Sentence Completion Test of Ego Development. *Applied Psychological Measurement*, 3, 281–311** cites the faking study as core validity evidence.

Why it transfers: the WUSCT scores *how* a sentence is organized (perspective-taking, complexity of frame), not *what* is claimed. The Waypoint analogue is scoring the **organization of a phenomenological report** — time-granularity, process-vs-thing construal, subject-object structure, spontaneous dereification in syntax ("anger arose" vs "I am angry" — Laukkonen 2025 corpus note) — rather than the presence of attainment claims. The seed corpus independently warns that the *content* vocabulary is learnable but treats report *structure* as the honest signal; the SCT literature supplies the empirical demonstration that structure resists inflation.

Limit honestly stated: Redmore's specific-information condition *did* move scores somewhat. Structure-scoring raises the cost of faking from "know the words" to "internalize the generative grammar" — it does not make faking impossible for a sophisticated, theory-trained subject. In the v1 cohort (everyone theory-trained) this is the live edge case; the mitigations are triangulation (behavior, telemetry, longitudinal texture) and Surya's interview delta (§7 of the spec measures exactly this ceiling).

### 2.2 Overclaiming / foil items — STRONG (and uniquely warning-robust)

- **Paulhus, D. L., Harms, P. D., Bruce, M. N., & Lysy, D. C. (2003). The over-claiming technique: Measuring self-enhancement independent of ability. *Journal of Personality and Social Psychology*, 84(4), 890–904.** (PMID 12703655). Respondents rate familiarity with items of which ~20% are **foils** (nonexistent). Signal-detection analysis yields an *accuracy* index (real vs foil discrimination) and a *bias* index (overall claiming tendency). Critically: **validity held even when respondents were warned about foils or instructed to fake good** (Studies 2–3), and bias generalizes across domains — the same people overclaim everywhere.

This is the only family in the psychometric faking literature whose detection power survives full disclosure of the method — precisely the property Waypoint needs given a map-contaminated cohort and a rule (below, §2.6) that secrecy-dependent countermeasures are brittle.

Waypoint transfer: build **contemplative foils** — plausible-but-nonexistent practice names, stages, and phenomenology ("the descending jhāna gate", "the silver-cord pulse in reverse breathing", a fabricated ñana). Delivered sparsely inside natural Wisdom conversation ("some students report X — is that familiar?"). A practitioner claiming familiarity with foils earns a per-practitioner **self-report credibility weight** that scales down the likelihood of *all* their uncorroborated claims. Two cautions: (a) foils must be checked against every tradition overlay in the internal canon — a "foil" that accidentally names a real Neidan phenomenon is a false positive generator; Surya must clear the foil list (M2); (b) rapport cost if a practitioner recognizes the trap — dose rarely, phrase as genuine curiosity, never confirm the foil as real afterward (Wisdom must not teach falsehoods: have Wisdom naturally move on, or gently self-correct later).

### 2.3 Behavior–claim consistency / collateral information — STRONG (as method), already spec'd

Forensic assessment's most robust principle is triangulation against collateral data rather than any single scale (the organizing principle of the malingering-assessment literature, e.g. Rogers's *Clinical Assessment of Malingering and Deception*, Guilford, multiple editions — cited here as a framework reference). Waypoint already owns the strongest version of this: **practice telemetry** (spec §5.3) is collateral information the practitioner does not narrate.

Concrete checks for the estimator design memo:
- **Claim–gate consistency:** claims of reverse-breathing competence vs actual position in the breathwork gate sequence; "I sit two hours daily" vs session logs.
- **Claim–trajectory consistency:** a claimed milestone (first glimpse, small death) should be *preceded* by the practice pattern that plausibly produces it and *followed* by the corroborating after-effect signature (post-cessation trait residue — Laukkonen 2023 corpus note). Milestones stay `claimed` until the residue shows.
- **Cross-channel reactivity check:** claimed equanimity vs observed reactivity *in the conversation itself* (how they respond when Wisdom challenges, reschedules, or misunderstands them). The exchange is behavioral data, not just report data.

### 2.4 Implausible-presentation detection (MMPI F/F(p) logic) — STRONG in origin domain, transfer is our derivation

The MMPI validity-scale family is the canonical faking-detection architecture:
- **L / K scales** — faking-good: endorsing implausible virtue / denying universal minor faults.
- **F and F(p)** — faking-bad/over-reporting: endorsing symptoms that even genuinely afflicted populations rarely endorse. **Arbisi, P. A., & Ben-Porath, Y. S. (1995). The Infrequency-Psychopathology Scale, F(p). *Psychological Assessment*, 7(4), 424–431** (doi:10.1037/1040-3590.7.4.424) — items endorsed by <20% of *psychiatric inpatients*, so elevation means over-reporting even against a high-pathology base rate.
- **VRIN/TRIN** — inconsistency: paired similar/opposite items answered incoherently.

Waypoint transfers (derivations, not validated instruments — mark as such in the marker bank):
- **F(p)-analogue = rare-attainment plausibility priors.** Phenomenology that genuine practitioners *at the claimed level* rarely report, or attainment claims lacking their gating scaffold: the corpus gives exact rubrics — nirodha-samāpatti claims without 8-jhāna mastery (Laukkonen 2023: default to claimed-not-corroborated), cessation claims with rich experience "during" the gap (Agrawal & Laukkonen: near-neighbor, not cessation), non-dual claims asserting awareness as ontological ground (emptiness mis-assertion is itself a placement signal). These become **negative-polarity markers**: the claim fires *evidence against* the claimed band and *for* the look-alike band.
- **L/K-analogue = too-good-to-be-true texture.** Genuine reports at every band are uneven, pillar-lumpy, and specific about residual difficulty (spec §3: development is uneven by design). Uniformly serene, difficulty-free self-presentation across all four pillars is an inflation flag, not high attainment. Lindahl & Britton's difficulty base rates (corpus digest 1) make "no difficulties ever" statistically improbable for a genuine deep practitioner.
- **VRIN-analogue = paraphrase re-probing.** Semantically equivalent probes, differently phrased, weeks apart (probe bank requirement). Stable phenomenology should replicate under paraphrase; scripted vocabulary often fails to, because the script is anchored to surface forms.

### 2.5 Koan-checking (sassho) — the tradition's own anti-scripting technology — MODERATE (rich precedent, no quantitative literature)

The Rinzai system is a centuries-old adversarial-robustness protocol, and its one documented failure is exactly Waypoint's T6:

- **Hoffmann, Y. (trans.) (1975). *The Sound of the One Hand: 281 Zen Koans with Answers*. Basic Books** (orig. *Gendai sōjizen no hyōron*, pseud. "Hau Hōō", 1916; reissued NYRB 2016). The 1916 publication of the standard answers — previously sold "under the counter" as cribs for dokusan — scandalized Japanese Zen because it burned the static item bank.
- The system's countermeasure was never answer secrecy alone but **sassho (checking questions)**: per major koan, roughly 20–100 lineage-standardized follow-ups probing whether the student can respond *freshly to perturbations* of the case, plus **jakugo** (capping phrases) demanding spontaneous apt expression. A crammed answer passes the koan and fails the third checking question, because the student memorized a point, not a generator.

Design lessons (this mapping is our derivation):
1. **Probe families, not probe items.** Every item in the probe bank is a parameterized template with variation dimensions, and the scoring key scores the *response pattern across variations*, not any single answer. Spec §5's "never repeated verbatim" is necessary but not sufficient — the point is that follow-ups actively perturb ("you said the observer dropped away — who noticed that?").
2. **Depth is tested by adjacency.** Sassho probe the *edges* of an insight (apply it to a new object, invert it, ask what it is not). A scripted report has no edges.
3. **Assume the answer book exists.** The marker bank and probe bank must be designed on the assumption they will eventually leak or be reconstructed by a motivated practitioner (the Wheel is already teachable material). Anything whose validity dies on disclosure is a secondary control (see 2.6).

### 2.6 The coaching caveat — secrecy-dependent detection is brittle — STRONG evidence, binding constraint

- **Rogers, R., Bagby, R. M., & Chakraborty, D. (1993). Feigning schizophrenic disorders on the MMPI-2: detection of coached simulators. *Journal of Personality Assessment*, 60(2), 215–226** (PMID 8473961) and **Storm, J., & Graham, J. R. (2000). Detection of coached general malingering on the MMPI-2. *Psychological Assessment*, 12(2), 158–165** (PMID 10887761): coaching subjects **about the validity scales themselves** substantially degrades detection, while coaching about the disorder's symptoms degrades it much less.

Read across: knowing the *map* (symptom knowledge — our T1) is far less damaging than knowing the *detection strategy*. Consequences for Waypoint: (a) T1 contamination, though universal in the cohort, is survivable; (b) the marker bank's *detection logic* (which patterns are scored as implausible, what the foils are, what the consistency checks compare) is the crown jewel and stays internal — narrower distribution than the Wheel itself; (c) rank countermeasures by disclosure-robustness: production/structure scoring and overclaiming (validated as warning-robust) and telemetry consistency (can't be talked out of) above any pattern-matching validity logic.

### 2.7 Socially desirable responding & spiritual self-enhancement — MODERATE

- **Paulhus, D. L. (1988/1991). Balanced Inventory of Desirable Responding (BIDR)** — two components: *self-deceptive enhancement* (honest but inflated self-view) vs *impression management* (audience-directed). The distinction matters for Waypoint because the fixes differ: IM shrinks when assessment is invisible (Waypoint's probes are invisible by design — spec §5); SDE does not shrink with invisibility and must be handled by evidence-weighting instead.
- **Vonk, R., & Visser, A. (2021). An exploration of spiritual superiority: The paradox of self-enhancement. *European Journal of Social Psychology*, 51(1), 152–165** (doi:10.1002/ejsp.2721; N=533/2,223/965): "spiritual superiority" is measurable, correlates with communal narcissism and supernatural overconfidence, and is *higher* in some training populations — spiritual training does not immunize against self-enhancement; it can feed it. Directly licenses treating explicit self-placement ("I'm quite far along") as near-zero-likelihood evidence, and superiority-comparative language ("unlike most people, I…") as a mild *inflation-risk* flag.
- Calibration battery candidate: BIDR (or the SDE subscale) once, calibration-cohort-only, to estimate each subject's self-report inflation for the gold-label study — cheap convergent data on the credibility weight.

### 2.8 Demand characteristics & the Wisdom-as-contaminator problem — STRONG concept, product-specific mitigation

- **Orne, M. T. (1962). On the social psychology of the psychological experiment… *American Psychologist*, 17(11), 776–783.** Participants infer the study's purpose and drift toward confirming it.

Waypoint's version is structural, not incidental: Wisdom teaches the path, names the milestones, and supplies the vocabulary — then Waypoint scores the practitioner's use of that vocabulary. Without correction this is a self-licking feedback loop that manufactures its own markers (T5) and, post-adaptation, T8.

Mitigation (our derivation; highest-leverage novel recommendation of this lane): an **exposure ledger**. Log, per practitioner, which map terms/milestone descriptions Wisdom (or app content) has surfaced to them and when. The extraction prompt receives the ledger and down-weights vocabulary-dependent markers for terms the practitioner has already been taught, while leaving structure/behavior markers untouched. Spontaneous first use of an un-taught construct ("it's like the anger wasn't mine") is strong evidence; fluent use of last week's lesson vocabulary is nearly none. This is cheap (content exposure is already loggable) and converts the worst structural threat into a measurable covariate. Corollary: probe items must not *teach* the phenomenon they test — score-relevant descriptions live in follow-up space, never in the probe's framing.

### 2.9 Self-report semantic drift (why scale-style self-ratings are not markers) — STRONG

- **Grossman, P. (2008). On measuring mindfulness in psychosomatic and psychological research. *Journal of Psychosomatic Research*, 64(4), 405–408** (PMID 18374739) and **Grossman, P. (2011). Defining mindfulness by how poorly I think I pay attention… *Psychological Assessment*, 23(4), 1034–1040** (PMID 22122674): mindfulness self-report items are interpreted differently by trained and untrained respondents; experienced practitioners can score *lower* because training recalibrates the internal standard (response-shift).
- **Van Dam, N. T., et al. (2018). Mind the Hype: A critical evaluation and prescriptive agenda for research on mindfulness and meditation. *Perspectives on Psychological Science*, 13(1), 36–61** — field-level statement of the same measurement problem.

Consequence: any first-person *rating* ("how aware are you, 1–10?") is a report about self-concept relative to a moving internal anchor, not about stage — and the anchor moves *with* stage, in the deflationary direction (reinforces T3). The extraction prompt must treat self-ratings and self-placements as near-uninformative for position (they retain some value for phase and for tracking the anchor itself). This is also the reason the spec's "no standalone questionnaire" decision is psychometrically right, not just a UX choice.

### 2.10 Linguistic deception markers (LIWC-style) — WEAK; honest assessment: corroboration only

- **Newman, M. L., Pennebaker, J. W., Berry, D. S., & Richards, J. M. (2003). Lying words: Predicting deception from linguistic styles. *Personality and Social Psychology Bulletin*, 29(5), 665–675** (PMID 15272998): 67% classification with constant topic, 61% overall — barely above chance for individual decisions.
- **Hauch, V., Blandón-Gitlin, I., Masip, J., & Sporer, S. L. (2015). Are computers effective lie detectors? A meta-analysis of linguistic cues to deception. *Personality and Social Psychology Review*, 19(4), 307–342** (doi:10.1177/1088868314556539): overall effect sizes **small**, heavily moderated by context; some theory-driven cues supported (liars: fewer sensory-perceptual words, more distancing, fewer cognitive-process references).

Verdict for Waypoint: never gate or penalize on stylistic deception cues alone. One cue family is worth carrying at low weight because it doubles as a *genuineness* cue rather than a lie cue: **sensory-perceptual specificity**. Fabricated/secondhand accounts are thinner in sensory-episodic detail — which converges with the referential-activity literature (next) and with how ñana interviews actually probe ("*where* in the body? *what* happened to the breath?").

### 2.11 Referential activity (Bucci) — PROMISING, UNVALIDATED here; flagged speculation

- **Bucci's multiple code theory & the referential process**: computerized measures — Weighted Referential Activity Dictionary (WRAD) within the DAAP system — quantify how vividly nonverbal/emotional experience is connected to language (validated in psychotherapy-process research; e.g., Mariani, Maskit, Bucci & De Coro, 2013, *Psychotherapy Research*, 23(4), 430–447: linguistic measures of the referential process, English and Italian versions).
- Relevance: high-RA speech (concrete, imagistic, temporally anchored, fluent-then-halting in characteristic ways) marks *contact with lived experience*; low-RA speech (abstract, general, theory-voiced) marks distance from it. A map-contaminated report is, in RA terms, a low-RA recitation of high-RA vocabulary.
- **SPECULATION, clearly flagged:** no study applies RA to meditation reports or to faked phenomenology; fakability of RA itself is untested. Do not import WRAD as a scoring dependency. Do encode the *heuristic* in the extraction prompt: "score whether the description exhibits episodic, sensory, temporally-located specificity versus generic map-language; quote the evidence." An LLM applying this heuristic with quoted evidence is auditable; a dictionary score is not.

### 2.12 Spiritual-bypassing detection — EMERGING scale, strong convergent rubric

- **Fox, J., Cashwell, C. S., & Picciotto, G. (2017). The opiate of the masses: Measuring spiritual bypass… *Spirituality in Clinical Practice*, 4(4), 274–287** (doi:10.1037/scp0000141): SBS-13, two factors — **Psychological Avoidance** and **Spiritualizing** (N=661; Brazilian replication 2018, α≈.86). The construct traces to John Welwood's coinage (early 1980s, transpersonal-psychology literature; his book-length treatment is *Toward a Psychology of Awakening*, 2000 — attribution widely reproduced; primary 1984 article not independently verified in this pass).
- Marker-bank transfer for the equanimity/bypassing/dissociation confusion row — bypassing signature = **spiritualizing frame applied specifically to avoid concrete difficulty**: deflection to doctrine when asked for specifics ("it's all empty anyway" in response to "what did you feel when…"), positive-affect vocabulary co-occurring with topic-avoidance behavior, unresolved reactivity leaking in the exchange while the narrative claims equanimity. Convergent discriminator from the seed corpus (f-SNR note): genuine equanimity *increases* contact with difficult material; bypassing decreases it. So the probe is simple: invite contact with the difficult specific and observe approach vs swerve.

### 2.13 Waypoint-internal Goodhart (the estimator gaming itself)

Distinct from practitioner-side threats; belongs in the estimator design memo:
- **Finder bias (T7):** an LLM instructed to extract markers will hallucinate them at low base rates. Controls: (a) every marker event must carry a verbatim quote/evidence ref (spec §6 already requires); (b) the extraction prompt must present "no markers found" as a first-class, *expected* outcome; (c) QA extraction against **seeded-negative synthetic dossiers** (personas authored to contain *no* genuine markers but plenty of vocabulary bait) — the extractor's false-positive rate on these is a release gate alongside the discrimination test in spec §7; (d) model-family separation between extractor and validation labeler (spec §6d) — keep it.
- **Aggregator Goodhart:** any monotone "more markers ⇒ higher band" rule makes verbosity a confound (chatty practitioners generate more marker-opportunities). Normalize marker likelihoods by conversational volume/opportunity, and let the credibility weight (§2.2) scale claim-type evidence.
- **Advisory-loop leak (T8):** even invisible estimates change Wisdom's behavior, which practitioners can sense and steer toward. Mitigations: slow adaptation constants; adaptation expressed in teaching *register* rather than unlockable *content* where possible; monitor for practitioners whose claimed trajectory outruns their telemetry right after adaptation shifts (a T8 fingerprint).
- **Calibration Goodhart:** once Surya's gold labels exist, prompt-tuning the extractor to agree with Surya on the tiny N is overfitting-by-hand. The house label-divergence convention (record both positions, don't force agreement) is the correct guard — extend it to prompt iteration: no extractor prompt change may be justified solely by moving one cohort member's estimate toward gold.

---

## 3. Ranked countermeasure list → pipeline placement

Ranked by (evidence strength × transfer-directness × disclosure-robustness). **Pipeline stages:** EXT = extraction prompt; PROBE = probe item bank/design; AGG = aggregation (priors, likelihoods, penalties); CONS = monthly consolidation review; GOV = governance/process; QA = synthetic-dossier QA.

| Rank | Countermeasure | Evidence | Disclosure-robust? | Plugs into |
|---|---|---|---|---|
| 1 | **Structure-over-content scoring**: score report organization (dereification syntax, subject-object structure, time-granularity, edges under perturbation), never attainment claims | STRONG (Redmore 1976; Loevinger 1979; convergent with seed corpus) | Yes (mostly) | EXT (core rubric), PROBE (open production formats only — no yes/no, no recognition items), QA |
| 2 | **Behavior–claim consistency**: telemetry cross-checks; milestone claims held at `claimed` until after-effect residue corroborates; in-conversation reactivity as behavioral data | STRONG (collateral-information principle; telemetry is unfakeable-by-talk) | Yes | AGG (consistency penalty — already spec'd; add claim-gate and claim-trajectory checks), CONS |
| 3 | **Contemplative foil probes + per-practitioner credibility weight** | STRONG (Paulhus et al. 2003 — survives warning and fake-good instructions) | Yes (validated) | PROBE (foil family, Surya-cleared, dose-capped), AGG (bias index scales all uncorroborated-claim likelihoods), GOV (foil list secrecy still preferable) |
| 4 | **Exposure ledger**: log taught vocabulary per practitioner; extraction discounts vocabulary-matching markers post-exposure; spontaneous pre-exposure usage upweighted | Structural fix to T1/T5 (Orne 1962 for mechanism; ledger design is ours) | Yes | EXT (ledger as prompt input), PROBE (probes must not teach what they test), GOV (log content exposure) |
| 5 | **Implausibility/negative-polarity markers** (F(p)/L-K analogues): gating-scaffold violations, look-alike signatures (rich experience "during" cessation, ontological-ground assertions), too-good uniform serenity | STRONG in origin domain (Arbisi & Ben-Porath 1995); transfer is derivation, corpus-backed | Partly (degrades if logic leaks — Rogers 1993/Storm 2000) | Marker bank (confusion table rows), AGG (claim fires evidence *for* look-alike band), CONS |
| 6 | **Probe families with perturbation follow-ups** (sassho model): parameterized items, score response-pattern across variations, test insight edges | Rich tradition precedent (Hoffmann 1975; sassho practice); no quantitative literature | Yes by design (assumes leakage) | PROBE (bank format requirement), EXT (score across the family), GOV (rotate surface forms) |
| 7 | **Paraphrase re-probing / response consistency** (VRIN analogue) | MODERATE (validity-scale tradition) | Partly | PROBE (paired items across sessions), AGG (inconsistency discount, phase-aware — genuine phase shifts also change answers; consult phase annotation before penalizing) |
| 8 | **Self-rating deflation rule**: explicit self-placements and 1–10 self-ratings ≈ uninformative for position; superiority-comparative language = mild inflation flag; humility language ≠ evidence of low stage | STRONG (Grossman 2008/2011; Van Dam 2018; Vonk & Visser 2021) | Yes | EXT (likelihood guidance), AGG (near-zero weights), calibration battery (BIDR-SDE once, cohort-only) |
| 9 | **Bypassing/avoidance discrimination probes**: invite contact with the concrete difficulty; score approach vs spiritualized swerve | EMERGING (Fox et al. 2017 SBS-13; convergent f-SNR rubric) | Yes | Marker bank (equanimity confusion row), PROBE, CONS (flag pattern for teacher triage, not band-penalty alone) |
| 10 | **Extractor anti-Goodhart bundle**: quote-anchored events, "none found" as expected output, seeded-negative dossier QA, volume normalization, model-family separation, no per-subject prompt tuning | Methodological (LLM-judge practice + spec §6/§7 extensions) | N/A | EXT, QA (false-positive gate), AGG, GOV |
| 11 | **Sensory-specificity as genuineness cue** (LIWC/RA-derived): episodic, sensory, temporally-anchored detail up-weights; fluent generic map-talk carries ~no weight | WEAK-to-PROMISING (Newman 2003; Hauch 2015 — small effects; Bucci RA — unvalidated here; SPECULATION as a faking detector) | Unknown | EXT (soft heuristic with quoted evidence, low likelihood weights only), never AGG-penalty alone |

**Deliberately excluded:** response-latency/keystroke timing (no contemplative validity data, telemetry noise dominates in chat); bogus-pipeline-style techniques (deceiving practitioners about verification capability — ethically incompatible with a teaching relationship and with §8); standalone validity-scale questionnaires (out of scope per spec §5).

## 4. What this lane asks of other deliverables

1. **Marker bank v0:** every confusion-table row gets *discriminating* probes and negative-polarity entries (rank 5, 9); every claim-type marker carries a `corroboration_required` field (rank 2); vocabulary-dependent markers carry an `exposure_sensitive` flag (rank 4).
2. **Probe item bank v0:** items are families with perturbation dimensions and cross-session paraphrase pairs (ranks 6, 7); include a Surya-cleared foil family with dose caps (rank 3); production formats only; probes never teach the tested construct (rank 4).
3. **Estimator design memo:** per-practitioner credibility weight; asymmetric claim-gated penalties (never penalize silence); volume normalization; phase-aware consistency discounts; seeded-negative QA gate with a false-positive threshold (ranks 2, 3, 10).
4. **Governance:** detection-logic documents (foils, implausibility rules, consistency comparisons) are internal-only, distribution narrower than the Wheel; content-exposure logging lands in the app data model early (ranks 4, and §2.6's brittleness rule).
5. **Calibration protocol addition:** one BIDR (SDE) administration in the instrument battery; record spiritual-superiority-flavored items if lane 2 selects a nondual/awakening scale anyway (rank 8).

## 5. Source list (verified this pass unless noted)

- Arbisi, P. A., & Ben-Porath, Y. S. (1995). *Psychological Assessment*, 7(4), 424–431. doi:10.1037/1040-3590.7.4.424
- Fox, J., Cashwell, C. S., & Picciotto, G. (2017). *Spirituality in Clinical Practice*, 4(4), 274–287. doi:10.1037/scp0000141
- Grossman, P. (2008). *J Psychosom Res*, 64(4), 405–408. PMID 18374739 · Grossman, P. (2011). *Psychol Assess*, 23(4), 1034–1040. PMID 22122674
- Hauch, V., Blandón-Gitlin, I., Masip, J., & Sporer, S. L. (2015). *Pers Soc Psychol Rev*, 19(4), 307–342. doi:10.1177/1088868314556539
- Hoffmann, Y. (1975). *The Sound of the One Hand*. Basic Books; NYRB reissue 2016 (orig. Hau Hōō, 1916).
- Loevinger, J. (1979). *Applied Psychological Measurement*, 3, 281–311.
- Mariani, R., Maskit, B., Bucci, W., & De Coro, A. (2013). *Psychotherapy Research*, 23(4), 430–447. doi:10.1080/10503307.2013.794399
- Newman, M. L., Pennebaker, J. W., Berry, D. S., & Richards, J. M. (2003). *Pers Soc Psychol Bull*, 29(5), 665–675. PMID 15272998
- Orne, M. T. (1962). *American Psychologist*, 17(11), 776–783.
- Paulhus, D. L. (1988/1991). BIDR — SDE + IM subscales (measure documentation; widely reproduced).
- Paulhus, D. L., Harms, P. D., Bruce, M. N., & Lysy, D. C. (2003). *JPSP*, 84(4), 890–904. PMID 12703655
- Redmore, C. (1976). *Journal of Personality Assessment*, 40(6), 607–616. PMID 16367343
- Rogers, R., Bagby, R. M., & Chakraborty, D. (1993). *Journal of Personality Assessment*, 60(2), 215–226. PMID 8473961
- Rogers, R. (Ed.). *Clinical Assessment of Malingering and Deception*. Guilford (framework reference; edition not pinned this pass).
- Storm, J., & Graham, J. R. (2000). *Psychological Assessment*, 12(2). PMID 10887761
- Van Dam, N. T., et al. (2018). *Perspectives on Psychological Science*, 13(1), 36–61.
- Vonk, R., & Visser, A. (2021). *European Journal of Social Psychology*, 51(1), 152–165. doi:10.1002/ejsp.2721
- Welwood, J. — spiritual bypassing coinage (early 1980s); *Toward a Psychology of Awakening* (2000). Attribution widely reproduced; primary article NOT independently verified this pass.
- Seed-corpus cross-references: Laukkonen & Slagter 2021; Laukkonen 2025 (*A beautiful loop*); Laukkonen et al. 2023 (cessations); Agrawal & Laukkonen (nothingness); f-SNR 2026 — per corpus notes in `specs/waypoint-research/corpus/`.
