# Codex Marker Bank Red-Team (gpt-5.6-sol, reasoning effort: ultra, 2026-07-16)

Model: gpt-5.6-sol via OpenAI Codex CLI v0.144.0 | Effort tier: ultra | Date: 2026-07-16
Target docs: waypoint-research/deliverables/marker-bank-v0.md, waypoint-practitioner-stage-assessment.md (esp. §3)
Note: Codex was instructed to review only the two supplied file texts (no filesystem exploration).

---

## Bottom line

This bank is not yet construct-valid as a stage estimator. It currently mixes at least six different things:

- Meditation-path development
- Verbal/narrative sophistication
- Personality and cultural style
- Psychotherapy or coaching exposure
- Compliance with LIFE’s curriculum
- Transient meditation states

The strongest systemic problem is criterion contamination: several “band” markers are simply evidence that someone completed, understood, or can describe the practices assigned to that band. That establishes curriculum mastery, not an independently defined developmental stage.

The supplied bank also stops at band 4. Its promised bands 5–7, phase markers, cross-band guards, and confusion table are absent, although many entries depend on them. Consequently, the bank as supplied cannot perform several of its declared discriminations.

## Ten worst markers

| Rank | Marker | Construct-validity failure | Recommended disposition |
|---:|---|---|---|
| 1 | **M-S4-003 — Neidan verification signs** | Suggestion-prone bodily phenomena, tradition-specific expectations, map-readable canonical sequence, and acknowledged unverified sourcing. “Specificity in canonical order” becomes easier, not harder, after exposure. Heat/pressure may also be adverse rather than developmental. | Remove from stage scoring. Retain as an exposure-sensitive phenomenology/safety note pending teacher verification. |
| 2 | **M-A4-008 — non-dual glimpse** | Highly imitable vocabulary; also produced by depersonalization, psychedelics, panic, awe, sleep disruption, and other altered states. Surprise, instability, and a physiological tail are scriptable and not validity checks. Reproducibility would not distinguish reproducible dissociation from insight. | Phase/milestone claim only; no band update without longitudinal functional integration. |
| 3 | **M-E3-002 — narrow affect vocabulary** | Dominated by language, culture, gender norms, education, alexithymia, neurotype, and willingness to self-disclose. It measures expressive repertoire more directly than emotional development. | Convert to a measurement-capacity modifier; never score toward band 3. |
| 4 | **M-S3-002 — restlessness floor** | Confounded by pain, disability, anatomy, ADHD, anxiety, medication, posture choice, and session design. Inability to sit still is not a pre-path construct. | Accessibility/health covariate only. |
| 5 | **M-S4-006 — posture stability** | Primarily fitness, anatomy, mobility, pain status, chair/floor choice, and training history. Creates a strong disability and age bias. | Remove from stage; use only to adapt practice format. |
| 6 | **M-X3-002 — instrumental practice framing** | Motivation is not attainment. Advanced practitioners can meditate for sleep; novices can seek liberation. Also confounded by pragmatism, privacy, product-entry route, and cultural aversion to spiritual language. | Store as goal/preference data, not stage evidence. |
| 7 | **M-A4-010 — practice consistency** | Measures conscientiousness, time, money, care obligations, app engagement, and streak incentives. Off-app practice is invisible. Its statement that absence is a “loud consistency penalty” conflicts with claim-gated asymmetry and invites socioeconomic bias. | Exposure/adherence covariate only; absence must never lower stage. |
| 8 | **M-T4-001 — thought-as-event syntax** | Canonical ACT/CBT/mindfulness language and explicitly highly imitable. English grammatical style and therapeutic education could drive repeated false positives. Window-level repetition does not make a learned idiom lived insight. | At most weak evidence of learned dereification framing; strip wording before scoring. |
| 9 | **M-T4-002 — beliefs as constructs** | The bank itself calls this “cheap and reading-contaminable,” yet allows weak-to-moderate stage evidence. Common in therapy, philosophy, social science, and educated secular culture. | Curriculum-engagement marker only. |
| 10 | **M-E3-003 — externalized emotional causality** | Risks moralizing legitimate adversity and power asymmetry. Confounded by attribution style, personality, culture, current conflict, and therapeutic socialization. It may systematically underrate people describing genuinely harmful environments. | Remove from band scoring; retain only as a contextual formulation hypothesis. |

Near-bottom markers are **M-X3-003** (openness/education rather than stage), **M-S3-001/M-S4-001** (interoceptive and verbal traits), **M-E4-006** (agreeableness/cultural expressivity), **M-T4-004** (coaching literacy), **M-T4-005** (imagery phenotype), and **M-A4-006** (famous, metaphor-dependent language).

## Confounding and imitability audit

The largest shared method factor is articulate introspective language. It can raise scores simultaneously on **M-X3-001, M-S4-001, M-E4-002, M-T4-001, M-A4-001, M-A4-003, M-A4-005, M-A4-006, and M-A4-008** without any corresponding difference in stage.

Especially easy to imitate from map knowledge are:

- **M-S4-003, M-S4-005**
- **M-E4-002**
- **M-T4-001, M-T4-002, M-T4-003**
- **M-A4-001, M-A4-003, M-A4-005, M-A4-006, M-A4-008**

“Honest difficulty,” surprise, instability, failure-mode detail, and physiological aftermath should not be treated as intrinsic authenticity signals. Once a scoring key is inferred, a coached respondent can deliberately add all four.

Major personality/culture/education confounds include:

- Openness and philosophical education: **M-X3-003**
- Conscientiousness and available time: **M-A4-010**
- Agreeableness/affective expressivity: **M-E4-006**
- Interoceptive trait or somatic training: **M-S3-001, M-S4-001**
- Imagery ability: **M-T4-005**
- ADHD, pain, or arousal traits: **M-S3-002, M-A3-001, M-A3-002**
- Therapy/coaching exposure: **M-E4-001–003, M-T4-001–004, M-A4-001–003**
- Language and narrative education: nearly every “structure” marker, including **M-A4-003**

“Structure over content” is promising, but structural fluency also requires norming. Narrative sequencing, autobiographical memory, interview familiarity, and therapeutic training are not stage-neutral.

## Confusion-pair performance

| Confusion pair | Current markers | Why discrimination remains poor | What is needed |
|---|---|---|---|
| **Equanimity / bypassing / dissociation** | M-E4-005, M-E4-006, M-T4-006, M-A4-005 | A bypasser can learn contact language; dissociation can look coherent and composed; genuine equanimity may be terse or culturally low-affect. Short conversational half-life can reward avoidance. Warmth is a personality marker, not a reliable dissociation exclusion. Referenced guards M-NX-012/013 are absent. | Repeated provocation–contact–recovery observations; memory continuity; felt agency and choice; affective flexibility; delayed behavioral follow-through; functioning across relationships and work; ability to re-engage voluntarily rather than merely terminate a topic. |
| **Dark night / depression** | Only indirect handling in M-T4-003; phase markers are referenced but absent | Practice linkage does not discriminate: practice can precipitate or worsen depression, and depression can be interpreted through dark-night doctrine. The states can coexist. | Duration, pervasiveness, anhedonia, sleep/appetite/energy changes, suicidality, prior episodes, diurnal pattern, functional impairment, context responsiveness, practice-dose relationship, and independent clinical judgment. Treat as joint possibilities, not mutually exclusive foils. |
| **Absorption / dullness** | M-A4-004, M-A4-007 | Both can feel effortless, self-sustaining, and temporally compressed. “Would you have heard a doorbell?” is hypothetical and map-coachable. Self-reported clarity is almost definitional. Stronger absorption may reduce external responsiveness, making doorbell detection an unstable criterion. | Volitional entry/exit, repeatability at chosen duration, immediate post-sit recall, actual rather than hypothetical deviance probes, observable post-sit brightness versus grogginess, and precise onset/offset microsequence. Classify state first; do not infer stage directly. |
| **Vocabulary / lived insight** | Exposure flags plus follow-up on M-T4-001, M-A4-003, M-A4-006, M-A4-008 | Fresh detail and canonical sequence can be reconstructed. The exposure ledger cannot observe books, teachers, forums, therapy, psychedelics, or prior apps. The app itself teaches several scoring templates. | Novel transfer problems, jargon-stripped paraphrase scoring, repeated re-probes from unrelated angles, behavioral predictions made in advance, delayed residue, contradictions across contexts, and scorers blinded to contemplative terminology. |
| **Devotion / destabilization** | None | No devotional markers or operational differential appear in the supplied bank. | Benign devotion: stable sleep/function, humility, reversibility, tolerance of doubt, relational warmth, and lack of urgency. Destabilization: reduced need for sleep, escalating arousal, grandiosity/special mission, command experiences, pressured speech, impulsivity, disorganization, boundary violations, fear, and impaired function. |

The confusion pairs should not be modeled as exclusive alternatives. Depression can accompany dark-night phenomenology; devotion can coexist with mania; absorption can contain intervals of dullness; dissociation can coexist with genuine insight.

## Ten most trustworthy markers

“Trustworthy” here means least bad and best grounded in observable structure or cross-channel evidence—not validated likelihood ratios for overall Wheel position.

| Rank | Marker | Why it is comparatively strong | Validity ceiling |
|---:|---|---|---|
| 1 | **M-A4-003 — spontaneous three-phase structure** | Spontaneous sequence on novel difficult material is more informative than terminology and supports perturbation testing. | Still affected by Mahāsi exposure, verbal fluency, memory, and therapy training. Require recurrence across novel contexts. |
| 2 | **M-E4-005 — staying with aversive contact** | Can be observed in the exchange rather than accepted solely as autobiography; explicitly requires contact and coherence. | Measures a moment of regulatory capacity, not a stable band. Mild conversational discomfort may not generalize. |
| 3 | **M-T4-006 — conversational rumination half-life** | Uses an actual interactional time series and is partly invisible to self-presentation. Requiring initial contact helps distinguish release from immediate swerving. | Should be cross-pillar integration, not “thoughts.” Avoidance and impression management remain substantial foils. |
| 4 | **M-A4-002 — noticing-latency shortening** | A repeated within-person change is more credible than a cross-sectional stage claim. The implied sequence matters more than labels. | Retrospective reconstruction remains possible; use repeated events and baseline-adjusted change. |
| 5 | **M-E4-001 — live-catching with behavioral fork** | Stronger when the report contains an observable consequence such as not sending an email, followed by recurrence. | Also produced by CBT, emotional-intelligence training, and temperament; primarily an emotion-regulation marker. |
| 6 | **M-E4-004 — compassion-target progression** | Telemetry, target order, and persistent difficulty provide cross-channel consistency and resistance to one-off polished claims. | Establishes practice engagement and a specific skill trajectory, not a universal stage. |
| 7 | **M-S4-004 — reverse-breathing competence** | Completed scaffold, terminal practice behavior, and internally coherent failure modes are a relatively strong procedural package. | Valid only for this curriculum’s breathwork mastery. It cannot equate practitioners from other paths or establish overall band 40+. |
| 8 | **M-S4-002 — gate-matched breath mechanics** | Current-gate detail plus plausible logs and failure modes is better than a bare attainment claim. | Again, technique mastery and curriculum compliance—not general contemplative development. |
| 9 | **M-E4-003 — concrete self-acceptance episode** | A specific episode with behavioral aftermath is more useful than self-acceptance vocabulary alone. | Mood, psychotherapy, and ordinary maturation are strong alternatives; needs longitudinal recurrence. |
| 10 | **M-A4-009 — cascade traversal fluency** | Combines logged completion with a report of lived transitions between domains. | The cascade is taught by the app, making this circular evidence of curriculum comprehension. It belongs under cross-pillar skill, not awareness stage. |

Only the first five are promising candidates for broader developmental inference. Ranks 6–10 are more trustworthy as proximal skill or curriculum measures than as Wheel-stage measures.

## Misassigned markers

Several markers should move out of the band × pillar stage model:

- **Goal/orientation covariates:** M-X3-002, M-X3-003
- **Reporting-capacity modifiers:** M-X3-001, M-E3-002, much of M-S3-001 and M-S4-001
- **Accessibility/health covariates:** M-S3-002, M-S4-006
- **Exposure/adherence:** M-A4-010
- **Technique mastery:** M-S4-002, M-S4-003, M-S4-004, M-T4-005
- **Transient state/phase:** M-S4-005, M-A4-004, M-A4-007, M-A4-008
- **Cross-pillar integration:** M-E4-005, M-T4-006, M-A4-009
- **Explicit transition rather than band:** M-T3-002 and M-S4-004

More generally, the stated scale-point anchors—31–34, 36–40, and so forth—have no demonstrated ordering evidence here. People can acquire visualization, somatic granularity, open monitoring, compassion practice, or belief flexibility in very different sequences.

## Missing entirely

In the supplied marker-bank text:

1. **All band 5, 6, and 7 markers** are absent, despite the declared band 3–7 scope.
2. **All actual phase-marker definitions** are absent: glimpse, plateau, dip, integration, and intensification.
3. **The cross-band guard bank is absent.** Referenced markers such as M-NX-005/006/007/009/012/013 do not appear.
4. **The promised confusion table and guard/detection logic are absent.**
5. **Devotion is absent altogether.**
6. **Clinical depression and dissociation are not operationalized.**
7. **Off-cushion functioning is a rule but not a marker system.** There are no defined measures of work, relationships, self-care, ethical conduct, repair after conflict, flexibility, or sustained reduction in suffering.
8. **Cultural, language, neurodivergence, disability, therapy-exposure, and socioeconomic measurement modifiers** are not modeled.
9. **Alternative acquisition routes**—psychotherapy, yoga/bodywork, martial arts, psychedelics, religious formation, artistic training—are mentioned piecemeal but not systematically represented.
10. **Local-dependence controls are missing.** One polished episode could trigger M-S4-001, M-E4-001, M-E4-002, M-A4-002, and M-A4-003 and be multiplied as five Bayesian observations.
11. **Novel-transfer and falsification probes are missing.** Follow-up specificity is proposed, but there are no markers for failed transfer, contradiction, prediction error, or behavioral residue.
12. **A construct-validity control design is missing.** Synthetic dossiers authored from the same bank are circular QA, not validation. Needed controls include map-informed non-practitioners, therapists, bodyworkers, diligent app users without reported insight, multilingual users, chronic-pain/ADHD/aphantasia groups, psychedelic users, and clinically depressed or dissociative participants.
13. **Independent criterion validation is missing.** Surya calibration can align the estimator to Surya’s Wheel judgments, but cannot by itself demonstrate that Wheel position is distinct from curriculum progress or that it predicts anything independently.

Before assigning likelihood ratios, I would split the bank into separate latent families—reporting capacity, exposure/adherence, technique mastery, transient state, trait integration, and safety—and allow only validated trait/integration evidence plus repeated within-person change to update overall stage. Correlated markers from one episode should enter as one evidence cluster, not multiple independent Bayesian hits.
