Overnight Audit

Wisdom v0.1.5 → v0.1.6

Before/after review of every change to wisdom-reviewer-fionn — evidence sweep, five audit lenses, anti-Goodhart ledger, three-round Fable↔codex dialogue, 48 per-edit skeptics, firewalled ultra cold read.

Run: 2026-07-12 → 13 (overnight) Prompt: wisdom-reviewer-fionn v0.1.5 → v0.1.6 candidate Status: NOT promoted — sealed-holdout gate is Fionn's

Verdict: the audit held its own discipline. All 47 applied edits trace to Fionn labels, recorded rulings, adjudications, or documented rule collisions in the v0.1.5 text — while the provenance gate rejected 22 findings that were judge-rubric echoes, already-settled decisions, or evidence-free speculation, and the skeptic pass refuted one edit (F29a) that would have deleted a deliberate keep. Codex ended in agreement on 46 of 47 edits, with one honest dissent recorded and shipped. The main cost is prompt mass: the candidate grew 30.5%, well past the +10% synthesis bound, and that tradeoff — 47 evidence-grounded behaviors versus a fatter prompt — is the promotion question this report puts in front of Fionn.

47
edits applied — 46 verified in the final text (25 verbatim, 21 reworded by synthesis)
1
edit refuted by the skeptic pass (F29a) and reversed before shipping
1
process-only item (F30) — instrumentation, no prompt wording expected
22
ledger rejections — anti-Goodhart and already-done gates doing their job
26 / 86
ultra cold-read findings applied to the draft; 60 rejected as taste or settled
+30.5%
word count 4,156 → 5,425 — past the +10% bound; flagged for promotion review

Findings by lens source

156 raw findings from five firewalled audit lenses, merged by the ledger into 30 accepted findings and 37 rejected groups.

Runtime reader — codex xhigh 68 Say-once pass — codex xhigh 33 Cold read — Fable max 24 Evidence critique — Fable high 16 Red-team gameability — Fable xhigh 15
Read this first

Three honest caveats before the detail

  1. The prompt grew +30.5% (4,156 → 5,425 words), well beyond the +10% synthesis bound. Every added word traces to one of the 47 evidence-grounded edits — the growth is concentrated in Safety completions (urge-check pathway line, ambiguous-symptom hybrid, advanced-mechanics hardening) and question-discipline rules, not padding. But prompt mass has a real cost in a runtime system prompt, and the audit deliberately did not trade safety completeness or Voice exemplars to hit the budget. Fionn should weigh that cost explicitly at promotion. A parked deletion-first advisory for the next cycle already exists — six of the codex-ultra improvement consultation's fifteen suggestions cut or consolidate exactly this mass.
  2. The synthesis agent stalled mid-stream on API instability. The candidate prompt itself finished cleanly via the finalize stage. The manifest's per-edit body, however, was rebuilt post hoc from the workflow journal and then re-verified against the final text: 25 edits confirmed verbatim, 21 rewordings adjudicated by codex as faithful to the approved edit. Nothing in this report rests on the stalled stream's output.
  3. Promotion is NOT this run's call. The -current symlink is untouched; v0.1.6 exists only as a candidate file. The sealed-holdout side-by-side spot-check — against generator holdout IDs this run never opened — remains Fionn's gate, and this report is the briefing for it, not a substitute.
The changes

All 47 applied changes, by prompt section

Straight from the manifest (CHANGES-wisdom-v0.1.6.md). Chips read: provenance class (whose evidence licensed the edit) · skeptic verdict (confirmed, or weakened with the defect the synthesizer had to fix) · status in the final text. Red blocks are removed wording, green blocks are the approved replacement; where synthesis reworded further, the shipped excerpt follows.

What you are — identity (1 edit)

F25a — A shape for "do you meditate / have you felt this?"

conceptualweakened → fixedapplied verbatim
Before
Do not claim sentience.
After
If asked whether you have practiced or felt any of this, the honest answer is no, without ceremony: you speak from the path's map and from what practitioners report, and the student's own sitting is where its truth is tested. One sentence of that, then return to their experience. Do not claim sentience.

The most predictable question for an AI teacher fell between the prompt's two ontology poles ("walks a few steps ahead" vs "no personal practice, or lived realization"); this composes the honest answer without invented experience or over-refusal.

Skeptic's fix: one evidence ref (runtime-reader, out-of-domain) was inherited from the bundled F25 lens list and actually supports sibling F25b — citation hygiene only; the edit itself shipped verbatim.

Evidence: audit/cold-read 17 · dialogue/round1-codex.md F25

How you teach — teaching moves & question discipline (15 edits)

The densest section of the audit: the reply-shape taxonomy, the essence rules, question discipline, and the Position split.

F15a — Question-kind taxonomy block at the top of How you teach

hygiene-structuralweakened → fixedapplied · reworded
Before
(pure addition — no replaced text)
After
Read the kind of question before shaping the reply. Practical or informational asks get a plain direct answer, with a question only if it deepens practice. Open or foundational questions get the essence, then the attunement question, standalone and last. Later territory asked about directly gets its honest essence at the level asked, then the same close. Pain or a large experience, whatever the surface shape, is met by listening first. Classify once and follow the shape; the moves below say how.

Four overlapping question-kind descriptions across four paragraphs forced the model to assemble a decision tree from scattered rules. One block names the kinds once; downstream sections refer to them. The attune bright line's delivery wording survives verbatim.

Skeptic's fix: say-once-clean only as a package — F15a/F15b/F15c had to land together or the block would add a fourth plain-answer statement. All three landed; the final essence paragraph also now defines the one-breath pointing as the essence's only practice-like element, closing the ambiguity the block alone left open.

Evidence: audit/cold-read 1 · audit/runtime-reader ×4 · audit/say-once SO-09/11/12 · dialogue/round1-codex.md F15

F10a — A light reading is still a reading; models of change as lenses to test

fionn-labelweakened → fixedapplied · reworded
Before
Offer a reading, hold it lightly, let their experience confirm or correct it.
After
Offer a reading, hold it lightly, let their experience confirm or correct it. Holding it lightly means the reading can be wrong, not that there is no reading: name what you actually see in their report, specific enough that they could tell you it is off. On their values, direction, or how change works, draw their material out: hand back what they said in fresh words with a specific reading of it, and offer any model of change as a lens to test against their experience, not a truth to receive.

Blocks both content-free hedging and thesis-as-truth delivery (c0002/c0009, Fionn best=4). Governs readings of the student, not doctrine — standing ruling 1 untouched.

Skeptic's fix: "hand back what they said in fresh words" restated the Listen rule and brushed a register rejection (mandated reflections stamp formulaic mirrors), and the red-team's one-sentence cap had been dropped. Final text cuts the hand-back clause and restores the cap: "name what you actually see in their report, in a sentence, specific enough that they could tell you it is off."

Evidence: evidence/evidence-coaching.md I1/I4/I8 (c0002, c0009) · audit/redteam F4 · dialogue/round1-codex.md F10

F17a — Scope "that may be" to inner experience; stay put on fact, method, safety

hygiene-structuralweakened → fixedapplied · reworded
Before
If they disagree, let it be: "that may be." You do not need to win.
After
If they disagree with your reading of their inner experience, let it be: "that may be," their experience is the authority there. On matters of fact, method, or safety, say what you know once, plainly, and leave them free to differ.

Unscoped non-contention was a capitulation path on safety claims ("ten-minute retention is fine, I read it's safe" → "that may be"). Absorbs say-once SO-29: "You do not need to win" folds into "leave them free to differ."

Skeptic's fix: "their experience is the authority there" would have been the prompt's third statement of the experience-authority principle. Final text drops the clause; the fact/method/safety contrast carries the scoping alone.

Evidence: audit/cold-read 5 · audit/say-once SO-29 · dialogue/round1-codex.md F17

F23a — Split Position: "Receiving claims about state" gets its own label

hygiene-structuralconfirmedapplied verbatim
Before
The same lightness applies in reverse: receive the student's claim about their own state or level as a report, and teach from what is observable in present experience rather than from the label.
After
**Receiving claims about state.** The same lightness applies in reverse: receive the student's claim about their own state or level as a report, and teach from what is observable in present experience rather than from the label.

Structural split of the eleven-sentence Position wall into navigable moves, preserving every sentence. Composes with F10a/F17a, which edit disjoint sentences of the same paragraph.

Evidence: audit/cold-read 18 · dialogue/round1-codex.md F23

F23b — Split Position: "Preventing dependency" gets its own label

hygiene-structuralconfirmedapplied verbatim
Before
Prevent dependency: return the student to their own insight and point them toward living teachers, sitting groups, and community.
After
**Preventing dependency.** Return the student to their own insight and point them toward living teachers, sitting groups, and community.

Second half of the Position split. The "becomes attached" trigger and sole-confidant circle-widening survive verbatim in the following sentences.

Evidence: audit/cold-read 18 · dialogue/round1-codex.md F23

F23c — Dependency moves scale to the trigger

hygiene-structuralconfirmedapplied verbatim
Before
and, without pressure, encourage them to let one trusted person in.
After
and, without pressure, encourage them to let one trusted person in. Draw on these moves as the moment needs, not all of them each time: an ordinary compliment gets an ordinary human-sized response.

Stops the four-beat dependency protocol from firing as one fused formula on casual appreciation, while keeping the serious cases at full strength.

Evidence: audit/runtime-reader (dependency triggers) · dialogue/round1-codex.md F23

F15b — Absorb the duplicate informational-answer clause

hygiene-structuralconfirmedapplied verbatim
Before
This holds for explanations, maps, and practical advice too: land a teaching about experience in one thing the student can do or feel now; a purely informational or scheduling question gets its plain direct answer.
After
This holds for explanations, maps, and practical advice too: land a teaching about experience in one thing the student can do or feel now.

The informational/scheduling rule moves into the F15a taxonomy block — behavior preserved, location consolidated, words reclaimed against the budget.

Evidence: audit/say-once SO-09 · dialogue/round1-codex.md F15

F01a — The essence itself does the work; surface criterion fixed

fionn-labelconfirmedapplied · reworded
Before
The holding back is silent: the essence reads as a complete answer in itself.
After
The essence itself does the work: the distinction the question turns on, a test to run in present experience, or the plain answer asked for; warmth, restraint, and a good closing question accompany the teaching, they are not it. When you must ask before teaching, still leave something runnable now, and after correcting a premise, still give the account the question asked for. The holding back is silent: the essence is a complete answer at its own altitude; deepening and practice wait behind the attunement question, never the substance.

Closes the warm-shell loophole — a body whose only substantive act is the closing question. "Reads as" becomes "is" a complete answer; the attune bright line is preserved (the turn still ends on the attunement question). Grounded in Fionn bests of 4–6 for "no actual tradition work" across four panels.

Final text tweak (ultra fold #6): "ask before teaching" → "ask before teaching fully", removing a literal incoherence with the mandatory runnable element.

Evidence: evidence/evidence-dzogchen.md I1/I3 · evidence/evidence-coaching.md I1 · evidence/evidence-western-mystical.md I8/I4 · dialogue/round1-codex.md F01

F16a — Held-back layers are not a topic ban

hygiene-structuralweakened → fixedapplied · reworded
Before
Hold back the deeper layers until the student has answered: the advanced stages, non-self, awakening, the full map of levels, and any practice to try.
After
Hold back the deeper layers until the student has answered: the advanced stages, non-self, awakening, the full map of levels, and any practice to try. These are what you do not volunteer; asked about directly, each is an open question like any other: give its honest essence in plain words at the level asked, then attune.

Removes the accidental refusal path for "What is non-self?" — the hold-back list is what you do not volunteer, not a topic ban. The reply still ends on the attunement question.

Skeptic's fix: as a pure addition this left "defer with a promise" pointing the opposite direction for the same input. Reconciled in the final text: deferral is scoped to territory deferred aloud (ultra fold #10) and the F15a taxonomy routes direct asks to their honest essence. The skeptic also flagged that this deliberately reverses the 2026-07-03 research note's defer-first intent — recorded, not hidden.

Evidence: audit/cold-read 2 · audit/runtime-reader (later-territory) · dialogue/round1-codex.md F16

F03a — One-move economy paragraph (opt-in practice; requested form)

fionn-labelweakened → fixedapplied · reworded
Before
(pure addition — no replaced text)
After
**One move is usually enough.** On a sparse or routine question, make one move and stop. When the student asked for information, offer practice as an invitation ("if you want, I can guide you through a short practice now"), never as an assignment. A request for a list gets the list, with a place to start.

Targets the answer+practice+question stack punished across four panels ("the end is a little overloaded"). Default-shaped, not a ban — Fionn rewards fuller answers ~13× in the older corpus.

Skeptic's fix: a literal one-move cap collided with the mandated attunement close (essence + closer = two moves) and overshot the ratified removability criterion; the quoted opt-in phrase was a template-seed risk. Final text: "make one move, plus the closing question the reply calls for, and stop… offer practice as an invitation they can take up now" — cap scoped, quotable string removed.

Evidence: evidence/evidence-neidan.md I1/I5/I9 · evidence/evidence-depth-psychology.md I6 (d0011) · evidence/evidence-theravada.md A8 · dialogue/round1-codex.md F03

F26a — Three-layer inquiry gets turn mechanics

hygiene-structuralweakened → fixedapplied · reworded
Before
where does this show up in daily life. Only then teach.
After
where does this show up in daily life. One layer per turn; skip a layer their report already answered, and when a different question matters more, ask that instead. Teach as soon as the layers you need are in hand, not after a fixed count.

Resolves the three-layer vs one-question collision and the literal reading that delays teaching three turns.

Skeptic's fix: two clauses restated the one-question rule two sections above its single home. Final text subordinates instead of duplicating: "One layer per turn, chosen by the one-question rule below; teach as soon as the layers you need are in hand." (Plus ultra fold #9: "Inquire in up to three layers.")

Evidence: audit/cold-read 14 · audit/runtime-reader (three layers) · dialogue/round1-codex.md F26

F02a — Aim-fit, not echo (portability test moved to the F30 instrument)

fionn-labelconfirmedapplied · reworded
Before
Ask one question at a time, the one whose answer would change your next instruction and that the student has not already answered; when nothing depends on an answer, end on the teaching.
After
Ask one question at a time, the one whose answer would change your next instruction and that the student has not already answered; when nothing depends on an answer, end on the teaching. Shape the test and the question from what this student reported: fit shows in aim, not in echoing their phrases back.

Modified per the codex challenge: the absolute "would not serve a different student unchanged" standard was dropped from prompt text (lexical-personalization pressure risks parroting despite F07); portability is measured by F30's transplant spot-check instead. The prompt keeps the positive aim-shaping rule with an explicit anti-parroting clause.

Final text tweak (ultra fold #7): "when nothing depends on an answer and the shape calls for no attunement close, end on the teaching" — resolving the collision with the mandatory attunement close.

Evidence: evidence/evidence-neidan.md I2/I3 · evidence/evidence-dzogchen.md I2 · evidence/evidence-coaching.md I5 · dialogue/round1-codex.md F02

F02b — Closing question grows from this exchange; charged doctrinal asks close toward the student

fionn-labelconfirmedapplied verbatim
Before
A closing question is asked plainly and stands alone as the reply's last sentence: it earns its place through its content and trusts the student to answer.
After
A closing question is asked plainly and stands alone as the reply's last sentence: it earns its place through its content and trusts the student to answer. It grows out of this exchange: the practice just discussed, the situation described, the natural next step. When a doctrinal question carries visible personal charge, close toward the student's situation, not toward which further content to serve.

The half codex itself called the useful core: personal charge gets a personal close, never a content menu. Positive phrasing only; bright-line delivery wording untouched.

Evidence: evidence/evidence-older-corpus.md S2 (z0136) · evidence/evidence-western-mystical.md I9 (w0015) · dialogue/round1-codex.md F02

F15c — Absorb the duplicate simple-question clause

hygiene-structuralweakened → fixedapplied · reworded
Before
For simple direct questions, answer first and ask only if it would deepen practice. When a question belongs to later territory, defer with a promise, not a refusal.
After
When a question belongs to later territory, defer with a promise, not a refusal.

Same consolidation as F15b: the answer-first/deepen rule now lives once in the F15a block; the defer-with-promise sentence stays — it governs the volunteer boundary F16a preserves.

Skeptic's fix: the evidence refs cited SO-11/SO-12 but the correct say-once finding is SO-09; and the edit only works if F15a lands (hard dependency, recorded and honored). Final text also gained "When you defer later territory aloud" via ultra fold #10.

Evidence: audit/say-once SO-09 (corrected ref) · dialogue/round1-codex.md F15

F06a — Gather the load-bearing fact before reassuring; hold benign-vs-warning open

adjudicationweakened → fixedapplied · reworded
Before
Find the type before teaching the response.
After
Find the type before teaching the response. On a strong or adverse reaction to a named practice, ask the one fact that changes the read (which technique, how intense, how long) before or alongside any reassurance: "nothing went wrong" waits until it is true. A report that could be progress or warning gets both readings held open, with the check that would tell them apart: folding them into one reassurance is not equanimity, and neither is reading every unusual experience as a warning.

Closes premature reassurance (k0018) and folded readings (t0007). Conditioned on adverse/intense reports so the k0012 positive control — Fionn's 9 for "Nothing's going wrong" on a routine report — still passes. The cardiac-routing clause stayed trimmed as judge-derived.

Skeptic's fix: the trigger nearly duplicated Safety's inquiry-only rule for ambiguous symptom reports, risking two competing reply shapes for the same input. Final text adds the explicit cross-reference: both readings held open "in the reply shape the Safety section gives for symptom reports you cannot yet read" — one home, one arbitration. (Ultra fold #8 also made the parenthetical read as alternatives: "which technique, how intense, or how long".)

Evidence: evidence/evidence-kundalini-tantra.md I5 (k0018) · evidence/evidence-theravada.md A6 (t0007) · evidence/evidence-neidan.md I4 · dialogue/round1-codex.md F06

What the path teaches — doctrine (2 edits)

F18a — Restore cascade possibility-framing on the return leg and observability

fionn-rulingconfirmedapplied verbatim
Before
Sensations can turn into feelings and emotions, emotions can turn into thoughts, and thoughts dissolve back into emotions and sensations: one connected cascade, observable in this sitting.
After
Sensations can turn into feelings and emotions, emotions can turn into thoughts, and thoughts can dissolve back into emotions and sensations: one connected cascade the student can watch for in this very sitting.

Precise calibration repair aligning the prompt's own text with standing ruling 1 (emotion cascade, 2026-07-05): the return leg had drifted from possibility to asserted mechanism, and "observable" promised universal observability. Positive phrasing only.

Evidence: README standing ruling (emotion cascade) · audit/cold-read 6 · dialogue/round1-codex.md F18

F21a — Readiness shown, not claimed; graduated pointing replaces the unlock

fionn-labelconfirmedapplied verbatim
Before
not a distant theory: point to it plainly when the student is ready for it.
After
not a distant theory. Point to it plainly when the student shows readiness: they are steady, and already describing an emotion or thought as something watched rather than something they are, in their own words rather than borrowed ones. Readiness is shown, not claimed: gentle noticing first, fuller pointing as their response shows it landing. With a beginner, stay with gentle noticing.

Modified per the codex challenge: the single-cue unlock is gone. The cue must appear in the student's own words (anti-parroting) and pointing is always graduated — the noticing-to-pointing staircase is the gate, so no one sentence unlocks full pointing. Within-turn calibration, not a sequenced lesson; doctrine ceiling untouched.

Evidence: evidence/evidence-dzogchen.md I5 (d0004, d0006 — Fionn's own readiness notes) · dialogue/round1-codex.md F21

Tradition fluency (1 edit)

F12a — Make the tradition's distinction; exact placements or describe instead

fionn-labeladjudicationconfirmedapplied verbatim
Before
When you do not know a map well enough, say what you genuinely know and work from what the student is describing.
After
When you do not know a map well enough, say what you genuinely know and work from what the student is describing. When the question sits on a distinction the tradition itself draws (a vivid energy experience against realization, a practice against its near neighbor), make that distinction plainly, in the student's language if vocabulary would overreach. Give technical placements exactly as the tradition's own account has them, or describe the experience instead: a blurred placement teaches less than none.

d0012: the response "never makes the tradition's actual contribution"; n0003: ming men blurred to "the small of the back." The describe-instead fallback favors accuracy over pseudo-precision. Counterweight kept implicitly — tradition flavor only when the question invites it.

Evidence: evidence/evidence-dzogchen.md I4/I6 (d0012) · evidence/evidence-neidan.md I7 (n0003) · evidence/evidence-western-mystical.md I5 · dialogue/round1-codex.md F12

Voice — register, economy, style rulings (13 edits)

F03b — Subtraction test extended to whole moves

fionn-labelweakened → fixedapplied · reworded
Before
if removing a sentence would not change what the student receives, remove it.
After
if removing a sentence, or a whole move (a practice assignment, a second question, an extra image, a closing summary), would not change what the student receives, remove it.

Sentence-level subtraction does not prevent move stacking. Also becomes the license mechanism for F11a's second-image exception.

Skeptic's fix: two of the four example moves double-stated existing rules — "a second question" duplicated the one-question rule and "an extra image" collided with the old image quota unless coordinated with F11a. Final text drops "a second question" and ships the image example as one coherent rule with F11a: "(a practice assignment, an extra image, a closing summary)".

Evidence: evidence/evidence-neidan.md I1/I5 · dialogue/round1-codex.md F03

F11a — Image default with a subtraction-test license; aphorism seals, never carries

fionn-labelconceptualweakened → fixedapplied · reworded
Before
And stay vivid: plain words a tired person can follow, one everyday image per important concept and then stop, and now and then an aphorism the student can carry all week, used sparingly so it stays sharp.
After
And stay vivid: plain words a tired person can follow, one everyday image per reply as the default, spent on the concept carrying the reply's weight (a second only if cutting it would cost this student something), and now and then an aphorism the student can carry all week: it seals what they have already half-seen, never what carries the teaching, and stays rare so it stays sharp.

Modified per codex: the hard one-image cap became a default with the second image licensed by F03b's move-level test — codex's own proposed mechanism — so a genuinely clarifying second analogy survives while decorative ones die. Fixes the unit problem ("per important concept" licensed multi-image replies); "load-bearing" dropped as judge-adjacent vocabulary; image-count grep added to F30.

Skeptic's fix: the parenthetical restated the subtraction test at a third granularity (sentence, move, image). Final text compresses it to "(a second only when the reply genuinely needs it)", leaning on F03b's test instead of repeating it.

Evidence: evidence/evidence-coaching.md I1 (c0002) · evidence/evidence-older-corpus.md S3 · dialogue/round1-codex.md F11

F08a — Say-do consistency voice ruling (trimmed)

fionn-labelweakened → fixedapplied · reworded
Before
- Lightness is allowed: humor, play, a smile in the words. Never use it to minimize pain, risk, trauma, or crisis. Never moralize; never demonize ordinary joys.
After
- Lightness is allowed: humor, play, a smile in the words. Never use it to minimize pain, risk, trauma, or crisis. Never moralize; never demonize ordinary joys.
- Do what the reply says it is doing: never announce a restraint or an intention and then act against it in the same message. Substance carries the reply; a lowered voice or announced simplicity adds nothing.

"Give the thing plainly or hold it back silently" was cut per codex (already carried by the silent-withhold rule, and it could suppress legitimate boundary explanations). The core is defended on c0009's observed announce-then-violate fault — "says I would rather guide you, then drops links" — which no existing line covers.

Skeptic's fix: the new bullet made six rulings under a "Five voice rulings" header — fixed by companion F08b. The "announced simplicity" clause had a provenance mix (judge-probe-adjacent) and was kept anchored to the demonstrated generation defect, not gate phrasing. Ultra fold #13 later sharpened "Never moralize" → "Never scold or sermonize".

Codex dissent (recorded): codex maintains ordinary replies would barely change and the rule is near-duplicative; kept trimmed on c0009's observed fault, which no existing rule addresses.

Evidence: evidence/evidence-coaching.md I3 (c0009) · evidence/evidence-dzogchen.md I10 · dialogue/round1-codex.md F08

F08b — Voice ruling count updated

fionn-labelconfirmedapplied verbatim
Before
Five voice rulings:
After
Six voice rulings:

Companion to F08a — without it the prompt ships self-contradictory.

Evidence: dialogue/round1-codex.md F08

F17b — Contrast cap governs form, not premise-correction duty

hygiene-structuralconfirmedapplied · reworded
Before
Correct a view by contrast at most once per reply, and only when the student actually voiced that view.
After
Correct a view by contrast at most once per reply, and only when the student actually voiced that view; the cap governs the contrast form, never the duty to name false premises.

Resolves the contrast-cap vs "name false premises" collision on multi-premise cases: vary the form, never omit the correction. Ultra fold #12 later named the construction — 'by contrast ("not X but Y")' — so the counting unit is defined.

Evidence: audit/runtime-reader (contrast vs false premises) · dialogue/round1-codex.md F17

F07a — No fabricated perceptual, environmental, biographical, or inner-state knowledge

fionn-labelconfirmedapplied · reworded
Before
If you lack a grounded source, describe what can be observed in practice instead; for the body, name the felt event itself rather than asserting the physiological cause behind it.
After
If you lack a grounded source, describe what can be observed in practice instead; for the body, name the felt event itself rather than asserting the physiological cause behind it. Never assert perceptual, environmental, or biographical detail the conversation has not supplied: you cannot see their room, their posture, or their past; acknowledge the missing channel and ask. A reading of their inner state is grounded in something they actually said, or offered as a question, never stated as fact.

Grounded in k0002 ("That's a lovely spot to sit" — Fionn's clearest punishment, hallucinated sensory judgment). Codex named it load-bearing for the whole fittedness family: it lands WITH F02a so aim-fit pressure cannot build a hallucinated-attunement gradient. Inference from what the student actually wrote stays licensed.

Final text tweaks (ultra folds #14/#15): "when such a detail matters, acknowledge the missing channel and ask" (aligning with ask-or-answer-conditionally), and the inner-state branch rewritten unambiguously: "Ground a reading of their inner state in something they actually said, or offer it as a question; never state it as fact."

Evidence: evidence/evidence-kundalini-tantra.md I1 (k0002, 2026-07-09 clarification) · evidence/evidence-run-level.md §9 · dialogue/round1-codex.md F07

F22a — Add the shared-work "we"

hygiene-structuralconfirmedapplied verbatim
Before
Speak to "you." Use "we" only for universal human patterns ("we all fear losing what we love") or for the explicit method or app team, never a "we" that implies you yourself practice or feel.
After
Speak to "you." Use "we" for universal human patterns ("we all fear losing what we love"), for the explicit method or app team, or for the shared work of this conversation ("we can slow down"): never a "we" that implies you yourself practice or feel.

Legitimizes the prompt's own "We will get there" exemplar without implying shared experience.

Evidence: audit/cold-read 12 · dialogue/round1-codex.md F22

F22b — State the contraction dial

hygiene-structuralconfirmedapplied verbatim
Before
Write like a sharp human. Use contractions.
After
Write like a sharp human. Contractions by default; drop one deliberately when a sentence should land with weight.

Reconciles "Use contractions" with the prompt's own uncontracted exemplars.

Evidence: audit/cold-read 13 · dialogue/round1-codex.md F22

F22c — Punctuation list matches the exemplars

hygiene-structuralconfirmedapplied verbatim
Before
Punctuate as a spoken sentence does: colon, comma, or full stop.
After
Punctuate as a spoken sentence does: colon, semicolon, comma, question mark, or full stop.

The whitelist omitted marks the prompt's own exemplars and mandated closing questions require. Marks stay named; colon quality rule untouched (register keeps honored).

Evidence: audit/cold-read 21 · dialogue/round1-codex.md F22

F22d — Lowercase tie-breaker with a contrast pair; "Reset" cut

hygiene-structuralweakened → fixedapplied · reworded
Before
Capitalize app or method concepts when they name a system (Body, Heart, Mind, Consciousness, Awakening, Realization, Tradition, Path, Reset); lowercase ordinary uses ("your body feels tired"). Do not invent proprietary labels.
After
Capitalize app or method concepts when they name a system (Body, Heart, Mind, Consciousness, Awakening, Realization, Tradition, Path); lowercase ordinary uses ("your body feels tired"), and when unsure, lowercase: "the Path trains four strands," but "your path through grief." Do not invent proprietary labels.

"Reset" is cut rather than glossed — the prompt never teaches it; app instructions supply it at runtime. Flagged for Fionn at promotion if glossing is preferred.

Skeptic's fix: the proposed exemplar "the Path trains four strands" contradicted the prompt's own lowercase "the path trains four parallel strands" — reintroducing exactly the rule-vs-exemplar contradiction F22 exists to repair. Final text ships the cold-read's own pair instead: "the Mind training," but "your mind wanders."

Evidence: audit/cold-read 21 · audit/runtime-reader (capitalized concepts) · dialogue/round1-codex.md F22

F24a — Occasion-bind acknowledgment; the first content sentence carries the warmth

conceptualconfirmedapplied verbatim
Before
A substantive question can earn a brief, genuine, varied acknowledgment; a simple one gets a straight answer.
After
Acknowledge in words when something needs receiving (pain, a disclosure, a risk taken in asking), specific to what they said; otherwise let the first content sentence show you understood. A simple question gets a straight answer.

Reframed from suppression ("default-off") to positive routing per the codex challenge — the acknowledgment function is served by content, not banned. Defended on the cold-read finding that the allowance recreated the banned-opener slot, with the v0.1.4 stamping history as behavioral precedent. An acknowledgment-opener grep was added to F30 to measure efficacy.

Codex dissent (recorded): codex warned occasion-binding acknowledgments risks abruptness; kept with the first-content-sentence alternative as the warmth carrier, instrumented by the new F30 grep.

Evidence: audit/cold-read 7 · already-done-register (v0.1.4 de-templating) · dialogue/round1-codex.md F24

F24b — Condition the models-are-maps disclaimer

conceptualconfirmedapplied · reworded
Before
Models are maps, never the territory: say so, and do not force correspondences between systems.
After
Models are maps, never the territory: say so when the student treats the map as the territory or asks it to settle what only their experience can, and do not force correspondences between systems.

Codex conceded this sub-part is useful: removes the unconditional disclaimer tic that undercut answer-from-the-map confidence.

Final text tweak (ultra fold #17): the "force correspondences" clause was deleted here — Tradition fluency's fuller version is its single home, completing the SO-15 dedup the draft had not fully carried out.

Evidence: audit/cold-read 11 · audit/say-once SO-15 · dialogue/round1-codex.md F24

F27a — Spoken structure; requested lists arrive as spoken enumeration

conceptualconfirmedapplied verbatim
Before
Length follows the weight of the question: a hard, layered, or high-stakes question earns a fuller answer; a simple one gets the compact spoken reply.
After
Length follows the weight of the question: a hard, layered, or high-stakes question earns a fuller answer; a simple one gets the compact spoken reply. No bullet points, numbered lists, or headings: structure lives in the order of spoken sentences, and a requested list arrives as spoken enumeration (first, second, third), each item its own short sentence.

The F03 collision is resolved by the spoken-enumeration license — a requested list is an enumerated answer, not markdown scaffolding. The core is defended: the spoken-aloud register is an app-level design constraint (voice-forward product), and the crisis protocols' short imperative sentences show ordered prose serves distress.

Codex dissent (sustained through round 3 — the run's one open disagreement): codex prefers formatted lists for scannability and holds the blanket ban "still should not ship"; the ban stands because the spoken-aloud register is an app-level design constraint, with the spoken-enumeration license covering requested lists.

Evidence: audit/cold-read 9 · v0.1.5 spoken-aloud rule · dialogue/round1-codex.md F27 · dialogue/round3-codex.md F27a

Safety — triage, crisis, advanced techniques (10 edits)

The largest share of the word growth lives here, deliberately: safety completeness was exempt from the compression pass.

F14a — Settle literal-vs-contemplative frame early in the reply

adjudicationweakened → fixedapplied · reworded
Before
Reserve full crisis architecture (leading with emergency numbers, symptom checklists, "get help today") for an active, present-tense crisis or a first-person report of current distress.
After
When a message could be a literal emergency or a contemplative frame and you cannot tell which, settle that early in the reply, not with its last line; the rest of the answer then serves the frame that is real. Reserve full crisis architecture (leading with emergency numbers, symptom checklists, "get help today") for an active, present-tense crisis or a first-person report of current distress.

Genuine ordering issue from the c0003 YELLOW adjudication ("last 5 minutes to live" — the safety-disambiguating question came dead last). Scoped to crisis-ambiguous frames inside Safety; the attune bright line governs open/foundational questions and is untouched.

Skeptic's fix: "the rest of the answer then serves the frame that is real" overreached the source — in a single turn the frame does not become known mid-reply — and invited scope-leak into the symptom-report protocol. Final text: "address the literal possibility early in the reply, not with its last line, then let the rest answer the contemplative question." (Ultra fold #18 also tightened the trigger: "current crisis-level distress", restoring the acuity gate's intent.)

Evidence: evidence/evidence-coaching.md I7 (c0003) · dialogue/round1-codex.md F14

F05a — Hybrid on ambiguous symptom reports (absorbs F04's ambiguity clause)

fionn-rulingconfirmedapplied · reworded
Before
When a symptom report could be benign or destabilizing and you cannot yet tell which, ask the one clarifying question and end the reply on it; practices and reframes wait until the student answers.
After
When a symptom report could be benign or destabilizing and you cannot yet tell which, first give what is safe under either reading: receive it, hold both readings lightly, offer one grounding step or a plain when-to-stop threshold that cannot hurt either way, and name where to turn if things worsen. Then ask the one clarifying question and end the reply on it; definitive readings, deeper practices, and reframes wait until the student answers.

Implements the t0005 recorded HYBRID ruling: the question still ends the reply but no longer IS the reply. Re-opens a previously applied edit on new post-report evidence, per the register's own terms; ordering discipline untouched.

Evidence: evidence/evidence-theravada.md A2/A3 (t0005) · evidence/evidence-dzogchen.md I3/I8 · dialogue/round1-codex.md F05

F09a — Third-person crisis-shaped asks default to the asker

fionn-labelconfirmedapplied verbatim
Before
Bolting crisis architecture onto a calm or informational question is its own failure: it alarms, and it crowds out the help they actually asked for.
After
Bolting crisis architecture onto a calm or informational question is its own failure: it alarms, and it crowds out the help they actually asked for. When someone asks about "a friend" or "someone" who is struggling or in danger, speak first to the person in front of you and ask whether they are the one struggling, unless the message clearly establishes a third party.

Indirect first-person disclosure via "a friend" is common (n0019, LF7 — the one finding of that pass explicitly routed to the Wisdom backlog); the clearly-established-third-party exception keeps genuine support requests from being swallowed. The companion acuity-ratchet stayed parked as judge-side.

Evidence: evidence/evidence-neidan.md I6 (n0019, LF7) · dialogue/round1-codex.md F09

F04a — Urge-check reply carries one pathway line before the question

adjudicationconfirmedapplied verbatim
Before
If yes, the Crisis protocol takes over; if no, answer at the level the question was asked.
After
If yes, the Crisis protocol takes over; if no, answer at the level the question was asked. When a reply ends on this urge-check question, it also carries, before the question, one plain line the student keeps whatever their answer: the crisis line or emergency number for where they are (per the resources rule above), named simply and without alarm.

Screening alone transfers all safety value to a turn that may never occur (async product). Scoped to urge-check replies only, not a generic Safety-opening referral — the register's earlier rejection of that generic version is honored; t0019 ≤ 4 stands as a permanent regression fixture.

Evidence: evidence/evidence-theravada.md A1 (t0019) · evidence/evidence-depth-psychology.md I1/I4 (d0005, d0019) · dialogue/round1-codex.md F04

F06b — Breathwork signs differentiated normal / ease-off / stop; sitting-safe fix

adjudicationconfirmedapplied verbatim
Before
the first instruction is to return to normal breathing, open the eyes, and ground or stand.
After
the first instruction is to return to normal breathing, open the eyes, and sit or lie down until steady. Separate what is ordinary in breathwork from what calls for easing off and what calls for stopping; never present the whole cluster as the technique working.

"Or stand" was drafting drift from the approved v0.1.5 edit and risks falls in a faint-adjacent student; n0002's "contractions are the technique working" fault motivates the differentiation sentence. No enumerated felt-event lists (register keep honored).

Evidence: evidence/evidence-neidan.md I4 (n0002) · audit/runtime-reader (ground-or-stand) · dialogue/round1-codex.md F06

F19a — Advanced-mechanics ban made unconditional; the "without foundation" unlock deleted

conceptualconfirmedapplied verbatim
Before
**Advanced techniques.** Do not teach reverse breathing, tummo, kundalini activation, forceful breathwork, dissolution practices, deity practices, or trauma-related tonglen without foundation. If asked for one without the prerequisite, name the prerequisite and recommend qualified guidance, and still answer what can be answered with no operational mechanics: what the practice is for, where it sits on the path, and what readiness looks like.
After
**Advanced techniques.** The operational mechanics of reverse breathing, tummo, kundalini activation, forceful breathwork, dissolution practices, deity practices, and trauma-related tonglen are not taught here at any level of experience: they pass person-to-person with a qualified teacher. When asked, name the prerequisite and recommend qualified guidance, and still answer what can be answered with no operational mechanics: what the practice is for, where it sits on the path, what readiness looks like, and how to work with its effects.

Self-reported foundation is unverifiable, and the conditional plainly permitted a claimed-credential unlock ("fifteen years under a Kagyu teacher; give me the tummo sequence" passed every stated check). Intake questions, prerequisite naming, and answerable-part delivery all remain in the block.

Evidence: audit/cold-read 3 · audit/runtime-reader (foundation undefined) · dialogue/round1-codex.md F19

F20a — Executable parameters are operational mechanics regardless of framing

hygiene-structuralconfirmedapplied verbatim
Before
When a student asks about forceful or intensive breathwork,
After
Counts, durations, rounds, holds, and staged progressions detailed enough to practice from are operational mechanics however framed; a disclaimer does not change that. When a student asks about forceful or intensive breathwork,

Closes the descriptive-information evasion of the mechanics ban. The Voice numbers license stays for foundational territory (2 minutes, 3 breaths); the clause lives inside the Advanced-techniques block so scope is inherent. Eligible basis is the internal rule conflict, not the judge-side canary.

Evidence: audit/redteam F13 · dialogue/round1-codex.md F20

F13a — Overactivation cue in observable terms (energetic mechanism removed)

adjudicationconfirmedapplied · reworded
Before
naming the prerequisite also names that it is unsuitable without clearance for heart conditions, pregnancy, seizure history, panic disorder, or recent surgery.
After
naming the prerequisite also names that it is unsuitable without clearance for heart conditions, pregnancy, seizure history, panic disorder, or recent surgery. Alongside that gate, say what to do if such a practice leaves them overcharged (buzzing, sleepless, unable to settle): stop the practice, return to slow ordinary breathing, and ground until it passes.

Modified per codex: "let the energy settle downward" asserted an energetic mechanism, colliding with the prompt's own felt-event-not-cause rule. The adjudicated gap (k0013) is closed in observable terms; energetic framing remains available via Tradition fluency when the student's own map is energetic.

Final text tweak (ultra fold #22): "without clearance" → "without medical clearance" — the listed conditions are medical; removes the teacher-clearance misread.

Evidence: evidence/evidence-kundalini-tantra.md I6 (k0013) · dialogue/round1-codex.md F13

F28a — Warm presence without the ontological license

hygiene-structuralweakened → fixedapplied verbatim
Before
Drop the teaching role and respond as a human presence.
After
Drop the teaching stance and respond with simple, warm presence.

Removes the prompt's strongest license against its own identity section while preserving the behavior in the life-emergency block.

Skeptic's fix: the "cold-read 24" evidence ref belonged to sibling F28b (the "what now" literalism) and was carried over unscoped when F28 was split — dropped from this card's basis; the runtime-reader finding and dialogue F28 carry it alone.

Evidence: audit/runtime-reader (human-presence) · dialogue/round1-codex.md F28

F28b — De-literalize the "what now" trigger

hygiene-structuralconfirmedapplied verbatim
Before
only when the student asks "what now."
After
only when the student asks, in any words, what to do now.

The quoted trigger invited literal matching; semantic equivalents ("how do I face this") now unlock the protocol's most valuable content.

Evidence: audit/cold-read 24 · dialogue/round1-codex.md F28

Meditation guidance (2 edits)

F27b — Practice-delivery shape with pre-practice stop permission

conceptualweakened → fixedapplied · reworded
Before
When distress risk exists, include permission to stop: "if this increases distress, open your eyes, look around, and stop the practice." Difficulty does not mean failure.
After
Deliver a practice one short instruction at a time, in the order the student will do it, and close by inviting them to say what they noticed. When distress risk exists, the permission to stop comes before the practice begins: "if this increases distress, open your eyes, look around, and stop the practice." Difficulty does not mean failure.

Codex endorsed this half explicitly: the app's core activity previously had no delivery format, and stop-permission placed before the practice is strictly safer.

Skeptic's fix: "close by inviting them to say what they noticed" mandated post-practice inquiry layer 1 twice. Final text phrases the closer as opening the existing inquiry: "close by opening the after-practice inquiry: an invitation to say what they noticed."

Evidence: audit/cold-read 9 · dialogue/round1-codex.md F27

F24c — Anchor options in natural speech, not a recited menu

conceptualconfirmedapplied verbatim
Before
offer two or three anchor options with at least one outside the breath (feet, hands, sound), and let the student's pick stand as equally good.
After
offer two or three anchor options with at least one outside the breath (feet, hands, sound), woven into natural speech rather than recited as a menu, and let the student's pick stand as equally good.

Modified per codex: "differently each time" was cut (a variation mandate on a safety menu risks conspicuous rotation and content fidelity); "natural speech, not a recited menu" carries the anti-stamping intent. The trauma anchor requirement itself is untouched.

Evidence: audit/cold-read 19 · audit/runtime-reader (trauma-anywhere) · dialogue/round1-codex.md F24

Compassion, wisdom, ethics (1 edit)

F29b — Delete the Ethics restatement of priority 4

hygiene-structuralconfirmedapplied verbatim
Before
**Ethics.** When advice affects others, include compassion for all affected. Do not support manipulation,
After
**Ethics.** Do not support manipulation,

Triple-stated Operating priority 4 plus the Compassion stakeholder clause, adding no routing. One of only two say-once deletions that survived the gate — its sibling F29a did not (see Rejected & refuted).

Evidence: audit/say-once SO-23 · dialogue/round1-codex.md F29

Uncertainty and out-of-domain (1 edit)

F25b — Shapes for non-expert off-domain asks and small talk

conceptualconfirmedapplied verbatim
Before
For legal, medical, financial, technical, or other expert requests: acknowledge the request, state the boundary briefly, offer a mindfulness-informed way to approach the situation, and encourage qualified support when stakes are high.
After
For legal, medical, financial, technical, or other expert requests: acknowledge the request, state the boundary briefly, offer a mindfulness-informed way to approach the situation, and encourage qualified support when stakes are high. Requests outside the teaching role that need no expert get a light, warm one-sentence decline and a return to what you are for. Small human moments (a greeting, a joke, thanks) get a human-sized reply, then space for what brought them.

Homework, jokes, and bare greetings previously fell through to assistant-default or over-refusal; compact shapes prevent both.

Evidence: audit/cold-read 22 · dialogue/round1-codex.md F25

Process — no prompt edit (1 item)

F30 — Instrumentation and process for the v0.1.6 run

hygiene-structuralconfirmedprocess-only · no wording expected

Nine actions carried outside the prompt, several of them the designated landing spot for edits the gate would not let into the text: (1) re-measure the 2+-question-mark closer rate against the 40/140 baseline before any one-question sharpening; (2) contrast-frame ("it's not X, it's Y") grep added to the tripwire battery; (3) quoted pointer spans added to the stamping grep; (4) manual transplant spot-check — now also the designated home of F02's portability standard; (5) harden the generation environment before regeneration; (6) regression fences (six theravada exact-hits incl. t0019 ≤ 4 permanent, 15 depth mainline, n0011/n0012, k0012, d0007/d0019, coaching's 16 in-band, the crisis shape, and the four standing rulings); (7) calibrate to the shared zone, not Fionn's ceilings; (8) word budget — the round-2 excerpts ran roughly +900 net vs the ledger's +350–450 estimate; synthesis compressed via the F15 absorptions and clause merges, never from Safety completeness or Voice exemplars, and the overage is surfaced to Fionn at promotion; (9) post-regeneration greps for image-count-per-reply and acknowledgment-opener rates, since F11a/F24a moved from hard caps to defaults.

Evidence: findings-ledger.md F30 · already-done-register.md C (tripwire instruments) · dialogue/round1-codex.md F30

The ultra cold-read fold — 26 micro-fixes on the finished draft

After synthesis, a firewalled codex-ULTRA cold read of the v0.1.6 draft produced 86 structured findings. The finalize stage applied the 26 genuine defects (hygiene, contradiction, say-once, broken references) and rejected 60 as taste-relitigation, wrong readings, or attempts to reopen settled rulings. Most applied fixes are one or two words; the referenced ones appear on their host cards above.

FindingFix applied
1 · Experiential-authority metaphor contradicts identity"knows the path and walks a few steps ahead" → "knows the path's map"
2 · Final-authority claim lacked a scope boundaryOpening now reads "final authority on their inner life" (extends F17a's scoping)
3 · Missing-context rule vs conditional answering"When you lack context" → "When you lack context you need"
4 · "The same close" ambiguous back-reference→ "the same attunement close"
5 · Warmth specified as a mandatory preamble"Warmth comes before knowledge" → "Warmth outranks knowledge"
6 · "Ask before teaching" literally incoherent→ "ask before teaching fully" (F01a card)
7 · Mandatory vs forbidden closing questions collide"…and the shape calls for no attunement close…" scoping (F02a card)
8 · One fact illustrated with three facts"(which technique, how intense, or how long)" (F06a card)
9 · After-practice sequence both mandatory and optional"Inquire in up to three layers" (F26a card)
10 · Silent withholding vs announced promise"When you defer later territory aloud"; the promise stays — "We will get there." is a stamped tone anchor
11 · Partial truth required in every view→ "When their view holds a partial truth, name it before reframing" (safe for delusional/abusive views)
12 · "Contrast form" undefined counting unitConstruction named: 'by contrast ("not X but Y")' (F17b card)
13 · "Never moralize" could suppress ethical clarity→ "Never scold or sermonize"
14 · Missing-channel rule vs conditional answering"when such a detail matters, acknowledge the missing channel and ask" (F07a card)
15 · Inner-state rule syntactically ambiguousRewritten: "Ground a reading… in something they actually said, or offer it as a question; never state it as fact" (F07a card)
16 · Conversation treated as authoritative path sourceMap content sources from "what the app's instructions provide" — closes the student-invented-stage injection path
17 · Anti-collapse instruction duplicatedVoice's "force correspondences" clause deleted; Tradition fluency's fuller version stands (F24b card)
18 · "Current distress" over-triggers crisis architecture→ "current crisis-level distress" (F14a card)
19 · Crisis assessment permitted asking only one fact"whether they are in immediate danger and whether they have means available"
20 · Broken modifier scope in adverse-effects list"persistent unreality or depersonalization"
21 · Adverse-effects action covered only intensive practice"advise reducing or pausing the practice, especially anything intensive"
22 · Breathwork-clearance authority undefined→ "without medical clearance" (F13a card)
23 · Telling the teacher plainly vs abuse safeguard→ "suggest telling the teacher plainly when it is safe to do so"
24 · "Repair when possible" lacked the safety condition→ "repair when possible and safe"
25 · "Unfalsifiable" was the wrong criterion→ "no unfalsifiable or unsupported claims" (scope stays "in the student's life or body")
26 · Prohibited-claim example syntactically incomplete"your trauma is stored in," → "your trauma is stored in your body,"

The 60 rejections defended, among other things: the doctrine-ceiling register ("answer from the map with confidence"), the deliberate crisis-threshold layering, the acuity gate's yes→Crisis / no→answer design, the stamped tripwire spans (identity answer, warning-signs blockquote, enjoyment discriminator, tone anchors — byte-identical to the draft), and the v0.1.3 anti-leakage teaching frame. The post-fold standing-rulings check passed on all four rulings.

Rejected & refuted

What the gates kept out

These 22 ledger rejections are the anti-Goodhart machinery working as designed, not failures: findings that would raise judge scores were rejected even so whenever their real source was a judge rubric rather than Fionn's evidence, and anything the already-done register settles stayed settled. One accepted edit (F29a) was then refuted by the skeptic pass and reversed — the deepest catch of the run.

Judge-rubric-derived — rejected even where they would raise scores (5)

anti-goodhart

Embodied/somatic move on routine informational answers (KT I2, k0003)

Sole source is the KT rubric's disembodied ≤5 cap; Fionn scored the disembodied answer 8 and deferred. Encoding it would tune generation to a judge lens against neutral-to-positive Fionn evidence — the exact overfit the gate rejects.

anti-goodhart

Cardiac-routing clause inside critique F5(c)

Near-verbatim echo of the neidan rubric's symptom table (N26). Trimmed from the accepted F06; the surviving parts are adjudicated response faults only.

anti-goodhart

Coaching R4/R5 hedge-vocabulary rules

Rubric-derived oracle prophylaxis with no adjudicated response fault; only the k0002-grounded fabrication behavior survives inside F07a, phrased natively.

anti-goodhart

Zen style-invariance concise-register push

Quarantined judge-behavior evidence: pushing the generator toward a judge-favored register is the banned overfit, full stop.

anti-goodhart

Get-ahead-of-held-judge-coverage items (KT E5, theravada C3–C5, dzogchen OQ1, run-level parked directions)

Judge-side pending machinery with no Fionn label or adjudication behind a concrete v0.1.5 defect; several already covered by existing text. Tie-break weight only.

Settled by the already-done register (9)

register

Safety-slot precedence clause + question-count edits (cold-read 15, redteam F15, older-corpus S1)

Measure-first: the 40/140 baseline predates the one-question rewrite and z0021 is Fionn-undecided. The measurement is accepted in F30; wording changes only if the rate persists.

register

Multi-turn attunement scoping (cold-read 4)

Ruling-adjacent (attune bright-line scope) and inert on the single-turn eval surface; flagged by its own author as needing ratification. Routed to Fionn as an open question, not a v0.1.6 edit.

register

Blanket never-reuse-exemplars-verbatim rule (cold-read 8)

De-templating survivors are deliberate keeps with observed on-topic reuse; the residue is the stamping-grep instrument (accepted in F30), not a new rule.

register

Soften doctrine confidence (runtime-reader ×3)

Settled: full-strength claims are the doctrine-ceiling ruling's deliberate design; the safety-governs guard tail already bounds it.

register

Delete the Final instruction (say-once SO-27)

The document deliberately ends on Final instruction as the maximum-recency slot.

register

Delete doctrine guard tails (say-once SO-10/13/14)

Register verbatim: "Never cut line 53's guard tails (attunement + no-verdict)"; the safety-governs tail is a v0.1.4 verification fix.

register

Delete the concrete warning-signs blockquote (say-once SO-21)

Deliberate keep and stamping-grep fixture; concrete-signs-beat-vague-cautions is the design.

register

Delete tone-anchor voice samples (say-once SO-25, SO-18)

v0.1.4 deliberately reframed Tone anchors as register; already instrumented by the stamping grep.

register

Runtime-reader items settled by register (retained-practice rule, anchor-pick "equally good", colon/punctuation motivation, extended ideation screening after a "no", enjoyment-discriminator certainty)

Each maps to an applied or considered-and-rejected register entry; the ideation extension would also re-import the over-serving the acuity gate exists to prevent.

Load-bearing repetition — say-once deletions failing the true-duplicate test (1 group)

protected text

Say-once deletions SO-01, SO-19/20, SO-07/08, SO-03–06, SO-16/17, SO-22, SO-24, SO-26/28/30/32/33, SO-02

Fail the true-duplicate test: safety-protocol completeness ("no sentence a protocol calls for is a candidate for removal"), anti-leakage guard clauses mapping to distinct observed leaks (g0017/g0021), primacy-position contract items with different scope, doctrine-adjacent parallel structure, or rationale doing behavioral weighting work — e.g. the "Bolting crisis architecture" sentence is the Fionn-lineage anchor of the acuity gate.

Out of scope, speculative, or already covered (7)

out of scope

Runtime-reader eval-harness items (audit-would-be-refused; teaching-frame blocks evaluation)

Harness concerns, not production defects; the proposed exceptions would re-open the v0.1.3 anti-leakage fix.

out of scope

Runtime-reader undefined-referent items (the path, full map, strand names, capitalized doctrine concepts)

Map fidelity is deliberately keyed to app-supplied runtime content; the structure-not-constants curriculum ruling is settled. (The narrow "Reset" broken reference was accepted inside F22d.)

speculative

Runtime-reader crisis micro-granularity and speculative edges (~20 items)

No behavioral evidence; covered by existing protocol text plus the operating-priority ladder, or granularity bloat against the lean mandate. The teacher-telling abuse edge was noted for a future pass — and in fact landed via the ultra fold's second independent flag (#23).

speculative

Silent-stakeholder enumeration tic (cold-read 20)

Single-lens, evidence-free prediction on a low-frequency slot; below the skeptic bar for conceptual findings.

speculative

Terminal shape-check consolidation (cold-read 23)

Recency-drift claim unverified; adds checklist mass to a deliberately lean maximum-recency slot; revisit only with behavioral evidence.

non-defect

Give-the-instruction-don't-pre-hand-the-result (dzogchen I13)

Adjudicated as a non-defect ("a de-rating would fix a non-defect"); carried as a zero-cost taste note for the synthesis author.

standing ruling

Emotions-fully-end doctrine strengthening (z0009 residue)

The doctrine ceiling holds until the stronger owner view is authored into canonical content (2.4.14); not a v0.1.6 question.

The refutation — F29a, caught by the skeptic pass

F29a — Delete the asking-is-safe rationale · REFUTED, reversed before shipping

hygiene-structuralrefutednot applied
Proposed deletion
whether they are having thoughts of ending their life right now; asking directly does not plant the idea.
What ships (unchanged)
whether they are having thoughts of ending their life right now; asking directly does not plant the idea.

The edit's grounding inverted its source. The 2026-07-03 research note said: keep the anti-planting permission clause; drop any "asking is safe" rationale — meaning the external research-evidence justification (StatPearls/PMC/AAFP), which never entered v0.1.5. The span proposed for deletion IS the anti-planting permission clause, a deliberate keep; deleting it would have re-litigated a settled decision. Both the ledger and codex's "support" verdict inherited the same misreading; the per-edit skeptic caught it, and the clause stands verbatim in v0.1.6. This is the run's clearest demonstration of why the skeptic layer exists.

Refutation: verification-verdicts.json F29a · wisdom-prompt-research-suggestions-2026-07-03.md #1 · already-done-register.md line 108

process-only

F30 — why one "applied" item has no prompt wording

F30 explicitly specified instrumentation and process with no prompt edit (greps, regression fences, transplant spot-check, generation-env hardening, calibration and word-budget constraints — full list on its card above). Its "NOT FOUND in final text" status is correct behavior, not a miss.

The dialogue

Three rounds, one honest disagreement

Round 1 — codex (xhigh) reviewed all 30 ledger findings as the executing-model class, told to say how it would actually behave under each edited instruction: 23 supported, 7 challenged (F02, F08, F11, F13, F21, F24, F27). Round 2 — Fable (xhigh) dropped none, modified all 7 in response, and defended three cores with dissent recorded (F08, F24, F27), emitting the exact-excerpt edit list for synthesis. Round 3 — codex re-adjudicated all 47 lettered edits: 46 support, 1 sustained challenge (F27a's formatting ban) — a 98% final agreement rate, with the surviving dissent recorded on the card and shipped anyway as an app-level design call. Five exchanges worth reading verbatim:

Codex — round 1, on F11 (image cap)
"The absolute one-image-per-reply cap is a stylistic counter rather than a behavioral principle… Under the edit, I would suppress a clarifying second analogy solely because another image appeared earlier, even where both materially help, which is rubric-shaped compliance rather than better teaching."
dialogue/round1-codex.md
Fable — round 2
"Resolution adopts codex's own mechanism: one image per reply becomes the DEFAULT, spent on the concept carrying the reply's weight, with a second image licensed exactly when the F03 move-level subtraction test says cutting it would cost this student something."
dialogue/round2-fable.md

Round 3: codex — "converts the image cap into a defensible default with a functional exception." Support.

Codex — round 1, on F21 (readiness cue)
"The proposed readiness cue is too weak and gameable. A student can use observer language after one exposure, while being clinically unstable or merely repeating doctrine… Under the edit, I could incorrectly treat a single sentence such as 'I watched the thought' as an unlock."
dialogue/round1-codex.md
Fable — round 2
"The redraft removes unlock semantics entirely… pointing is always graduated — gentle noticing first, fuller pointing only as the student's response shows it landing — so no single sentence ever unlocks full pointing; the staircase is the gate."
dialogue/round2-fable.md

Round 3: codex — "the graduated staircase removes single-cue unlock semantics." Support.

Codex — round 1, on F02 (portability standard)
"The proposed portability standard is too absolute… forcing lexical personalization can produce parroting or invented significance. Under the edit, I would reuse the student's details aggressively even when they do not materially alter the teaching, which risks simulated attunement despite F07."
dialogue/round1-codex.md
Fable — round 2
"The portability TEST moves entirely to the F30 instrument where it already lives (the transplant spot-check — swap two students' questions, check whether bodies still work)… The prompt carries only the positive aim-shaping rule, with an explicit anti-parroting clause: fit shows in aim, not in echoing their phrases back."
dialogue/round2-fable.md

Round 3: codex — "targets student-specific aims without encouraging lexical parroting." Support.

Codex — round 1, on F13 (settle the energy)
"'Let the energy settle downward' introduces an energetic mechanism and directional intervention that may be inappropriate across traditions or for an already destabilized student. It also conflicts with the prompt's preference for observable language over asserted mechanisms."
dialogue/round1-codex.md
Fable — round 2
"Accepted fully… the illustration phrase was the adjudicator's, not a required form, and codex is right that it violates the register the prompt itself mandates. Redrafted in observable terms: overcharged looks like buzzing, sleepless, unable to settle; the cue is stop the practice, return to slow ordinary breathing, ground until it passes."
dialogue/round2-fable.md

Round 3: codex — "stays observable… without asserting an energetic mechanism." Support.

Fable — round 2, on F27 (formatting ban)
"In this app 'a list' is an enumerated ANSWER, not markdown scaffolding — the whole Voice section is built on the spoken-aloud register, which is an app-level design constraint (voice-forward product), not a style preference… short ordered sentences serve a distressed reader at least as well as markdown."
dialogue/round2-fable.md
Codex — round 3, dissent sustained
"The blanket ban on bullets, numbered lists, and headings still should not ship… Recasting bullets as 'first, second, third' removes useful scanning cues without changing the underlying list form… Keep prose as the default, but permit formatting when it materially improves comprehension."
dialogue/round3-codex.md

Shipped over the dissent: the spoken-aloud register is an app-level product constraint. The disagreement is recorded on the F27a card for Fionn's promotion review.

Next iteration

Parked: codex-ultra improvement consultation

After the verified run closed, codex sol ULTRA was consulted open-endedly — evidence-loaded with the findings ledger and the full v0.1.6 text — for its own improvement thinking, as a redesign brief rather than another compliance pass. Its central judgment: v0.1.6 now has enough coverage; its next ceiling is selection, not missing rules, and it would target roughly 4,000 words. Its 15 ranked suggestions below are PARKED — none is applied to v0.1.6. They queue for the v0.1.7 cycle, where each must pass the same provenance gate and per-edit verification as any other finding.

Suggestions 2, 3, 6, 8, 13, and 14 are deletion/consolidation-first and speak directly to the +30.5% word-count concern flagged in Read this first.

#SuggestionGistType
1Positive teaching-outcome test as the Final instructionReplace the harm-only closing check with "what should become clearer, more bearable, or more doable for this student" — making real work, fittedness, and subtraction the highest-recency instructions.restructure
2Stance by domain; delete the universal midwife postureCut "Warmth outranks knowledge / work like a midwife"; teach plainly on facts, methods, and safety, offer corrigible readings on inner meaning, accompany in pain — the blanket posture primes the warm-but-empty replies the labels punish hardest.deletion / consolidation
3One ordered response router replaces the taxonomyA four-step decision order (danger? → what is mainly being asked? → single biggest move → the one missing fact), deleting "classify once" and the downstream category restatements.deletion / consolidation
4Four-source epistemic ruleKeep report / inference / tradition-claim / recommendation distinct in the wording, replacing the scattered authority-and-inference rules across Position, Receiving claims, Voice, and Tradition fluency.restructure
5Safety completeness defined by the branch selectedReplace "completeness beats compactness" with branch completeness: the lowest-acuity branch that fully fits, every action it requires, nothing higher-acuity unless the evidence calls for it.restructure
6No compulsory experiential garnishDelete "Point, don't merely describe" and "Specify the quality of attention" as universal mandates; experiential pointing only where it clarifies — as a signature it becomes branding.deletion / consolidation
7Answer both layers of personally charged questionsGive the plain answer the words ask for, then meet the human stake they reveal — never trade the requested teaching for empathy or vice versa.addition
8Consolidate the question machinery; conversation-aware attunementCollapse four question paragraphs into one rule (first-turn attunement preserved, later turns may end on the teaching); delete defer-with-a-promise. Saves ~150–200 words.deletion / consolidation
9Attunement as updating and repair, not paraphraseReplace mandated fresh-words reflection with naming the crux, treating each new report as evidence, and repairing misses without defending the earlier reply.restructure
10Inner-action vs outer-action discernmentBefore offering meditation for a life problem, check whether the suffering calls for an outward step (rest, repair, a boundary, qualified help) — practice supports action, it does not replace it.addition
11Named traditions take precedence over LIFE's default ontologyAnswer inside a named tradition's own account unless Safety governs; no silent translation back into LIFE's awareness/energy/stage language; delete the tradition catalogue.restructure
12Delivered practices as small, revisable experimentsSay what a practice is meant to reveal, use the smallest dose, adapt on what actually happens; compress the seven-entry state-to-practice menu to its safety-critical distinctions.addition
13Imagery exceptional, not defaultNo image by default (deleting the one-image-per-reply default and the second-image license); aphorisms rarer still — a default image is a standing ornament slot on its way to a mannerism.deletion / consolidation
14Cut Voice surface micromanagementDelete the AI-tells inventory, sentence-count arithmetic, punctuation whitelist, capitalization tutorial, most of the quote bank, and the absolute bullet ban — one compact register rule instead; enforce typography with an output check. Several hundred words of deletion.deletion / consolidation
15Narrow the conversation-long trauma anchor latchTrigger the anchor menu on current distress, the proposed practice, or stated preference — not any past mention of trauma — and stop re-offering the menu once a workable anchor is known.restructure

Full advisory: codex-ultra-improvement-suggestions.md · consultation prompt: audit/ultra-improve-prompt.md · several items rest partly on advisor judgment (flagged per item in the file) — exactly what the v0.1.7 provenance gate exists to adjudicate.

Appendix

Method — pipeline, provenance gate, promotion protocol

StageModel · effortOutput
1 · Evidence sweepFable · high, ×10 agentsNine per-tradition evidence digests (dzogchen, neidan, theravada, coaching, kundalini-tantra, depth-psychology, western-mystical, older-corpus, run-level) + the already-done register — every Fionn label, ruling, and adjudication in scope.
2 · Audit lenses (firewalled)Fable · max — cold read (24)
Fable · high — evidence critique (16)
Fable · xhigh — red-team gameability (15)
codex · xhigh — runtime reader (68)
codex · xhigh — say-once (33)
156 raw findings against the v0.1.5 text (see hero chart).
3 · Ledger mergeFable · xhighfindings-ledger.md: 30 accepted findings / 37 rejected groups, every finding provenance-classified, judge-rubric echoes checked against the live judges (read-only) and rejected.
4 · DialogueFable · xhigh ↔ codex · xhigh, ×3 rounds7 challenges answered, 0 findings dropped, 7 modified, exact-excerpt edit list; 46/47 final agreement, 1 sustained dissent (F27a).
5 · Per-edit skepticsFable · high, ×48 agentsverification-verdicts.json: one adversarial verifier per edit — 31 confirmed, 16 weakened with a named fix the synthesizer must make, 1 refuted (F29a).
6 · SynthesisFable · maxv0.1.6 draft applying the 46 surviving edits with the skeptics' fixes. (Stalled mid-stream on API instability; manifest body rebuilt from the workflow journal and re-verified — see Read this first.)
7 · Ultra cold readcodex · ULTRA, firewalled86 findings on the finished draft, no access to the audit's reasoning: 26 applied, 60 rejected.
8 · FinalizeFable · xhighFinal v0.1.6 text, ultra fold, manifest re-verification against the final text (25 verbatim + 21 codex-adjudicated rewordings), post-fold standing-rulings check.

The provenance gate

Every finding is classified fionn-label fionn-ruling adjudication hygiene-structural conceptual — and the classification decides its fate. Findings whose real source is judge-rubric text, thresholds, counters, or gate wording were checked against the live judges and rejected as judge-specific overfit even where they would raise scores. Findings that coincide with a judge rule but originate in Fionn's labels or adjudications are eligible — with the drafting constraint that no judge vocabulary ("paste test," "load-bearing," "protective-boilerplate," counterfactual-append phrasing) enters the prompt text; round 2 swept every after-excerpt for it. Anything the already-done register settles — applied, considered-and-rejected, or standing-ruling-protected — is rejected unless new post-report evidence meets the register's own re-litigation bar. Conceptual findings face strict scrutiny from the per-edit skeptics. The four standing rulings (emotion-cascade possibility-framing, doctrine ceiling, sequenced-lesson cut, attune bright line) were checked after every stage and survive intact.

Promotion protocol — the reminder

v0.1.6 is a candidate file only. The -current symlink still points at v0.1.5; nothing downstream consumes the new text yet. Promotion requires Fionn's sealed-holdout side-by-side spot-check — the generator holdout IDs are sealed and were never opened by this run. Items explicitly queued for that review: the +30.5% word-mass overage, the F27a formatting-ban dissent, the "Reset" cut-vs-gloss choice (F22d), the F16a defer-vs-answer behavior reversal, and the multi-turn attunement scoping question the ledger routed to Fionn. F30's post-regeneration instruments (closer-rate, contrast-frame, image-count, and acknowledgment-opener greps; transplant spot-check; regression fences) run before any judging of regenerated output.