The short version

Your colleague's instinct is right twice and backwards once. Right that past-behaviour-in-context beats psychometrics (your own exec report and the Ouellette & Wood habit literature both nail this). Right that the setup is Bayesian to its bones — and your v3 design already quietly agrees, since it names a hierarchical Bayesian model as the baseline that's "likely to be hard to beat." Backwards on one specific point: predictions should not be rated more highly for holding across longer time scales — that scoring rule selects for horoscopes. Durability should be a property of which level of the model a piece of evidence updates, while predictions get scored by calibrated skill at their own horizon. Details below.

1 · How good are intelligent humans at predicting human behaviour?

The literature is unusually consistent, and the punchline is a division of labour, not averdict.Humans are mediocre integratorsresult in applied psychology:simple actuarial rules match orpredicting behavioural outcomes(recidivism, violence, prognosi al. 2000 meta-analysed 136studies — mechanical prediction33–47% of them, expertssubstantially better in only 6– to the machine that heldregardless of the judges' experr models" work showed evenunit-weighted two-variable modes later framing: humans lose notbecause they know less, but bec person integrates the samecues differently on different days.                                                 
The ceiling is often in the signal, not the predictor. Dressel & Farid 2018 found unMechanical Turk workers predictoled), statisticallyindistinguishable from the commercial COMPAS tool (65%) — and a two-feature logistic regression matched all 137 of COMPAS's features. Same story at the extreme end: the Fragile Families Challenge (Salganik et al., PNAS 2020) gave 160 teams thousands of featureschild to predict life outcomes;² ≈ 0.05–0.23 and barely beat afour-variable linear baseline. s good at long-horizon,single-person, specific-outcome
Real skill exists, but it's nark's arc is the cleanestevidence: in Expert Political Judgment, experts at 3–5-year horizons were roughly at chance and lost to simple extrapolation. But the Good Judgment Project then showed the opposite at short horizons: selected, trained "superforecasters" beat intelligence analysts with acto classified information by ~3persisted year over year, andone hour of base-rate training bought ~10%. The skill signature is exactly Bayesian hygiene: start from the base rate, update in small frequent increments, keep score.
                                                                                       Where humans genuinely shine is-context reading. Thin-slicestudies (Ambady & Rosenthal) geve/interpersonal outcomes fromunder five minutes of observatiarial collaboration ("Conditionsfor Intuitive Expertise") givesuition becomes trustworthy onlyin environments with stable regus feedback — chess yes,regardless of the judges' experience. Dawes' "improper linear models" work showed even unit-weighted two-variable modes later framing: humans lose not because they know less, but because they're noisy — the same person integrates the same cues differently on different days.
                                                                                       The ceiling is often in the sigel & Farid 2018 found untrainedMechanical Turk workers predicted recidivism at ~62% (67% pooled), statistically       indistinguishable from the commd a two-feature logisticregression matched all 137 of COMPAS's features. Same story at the extreme end: the Fragile Families Challenge (Salganik et al., PNAS 2020) gave 160 teams thousands of features per child to predict life outcomes; the best ML models managed R² ≈ 0.05–0.23 and barely befour-variable linear baseline. s good at long-horizon,single-person, specific-outcome prediction.
                                                                                       Real skill exists, but it's nark's arc is the cleanestevidence: in Expert Political Judgment, experts at 3–5-year horizons were roughly at chance and lost to simple extrapolation. But the Good Judgment Project then showed the oppositshort horizons: selected, traintelligence analysts with accessto classified information by ~30% on Brier score, the skill persisted year over year, and one hour of base-rate training bought ~10%. The skill signature is exactly Bayesian hygstart from the base rate, updats, keep score.

Where humans genuinely shine is hypothesis generation and in-context reading. Thin-slicstudies (Ambady & Rosenthal) geve/interpersonal outcomes fromunder five minutes of observation. Kahneman & Klein's adversarial collaboration ("Conditions for Intuitive Expertise") gives the unifying rule: human intuition becomes trustworthy only in environments with stable regularities and fast, unambiguous feedback — chess yes,   geopolitics no.
                                                                                       The design implication for you  advice): let the human/LLMnominate the cues and hypotheses, and let a mechanical layer do the integration and keescore. Never let the clinician our InferredProfileSnapshotalready has the first half — free_form_hypotheses with untested/supported/contradicted states is literally a hypothesiissing is the mechanicalintegrator underneath it (see §3). And notice the review-UI loop manufactures the Kahneman-Klein conditions — narrow taxonomy, ground truth arriving one message latis why prediction skill is evensame trick is the actual secretregardless of the judges' experience. Dawes' "improper linear models" work showed even unit-weighted two-variable models beat clinicians. Kahneman's later framing: humans lose not because they know less, but because they're noisy — the same person integrates the same cues differently on different days.
                                                                                  The ceiling is often in the sigel & Farid 2018 found untrainedAS 2020) gave 160 teams thousands of features per child to predict life outcomes; the best ML models managed R² ≈ 0.05–0.23 and barely beat a four-variable linear baseline. Nobody — human or machine — is good at long-horizon, single-person, specific-outcome prediction.

Real skill exists, but it's narrow and horizon-bound. Tetlock's arc is the cleanest evidence: in Expert Political Judgment, experts at 3–5-year horizons were roughly at chance and lost to simple extrapolation. But the Good Judgment Project then showed the opposite at short horizons: selected, trained "superforecasters" beat intelligence analysts with access to classified information by ~30% on Brier score, the skill persisted year over year, and one hour of base-rate training bought ~10%. The skill signature is exactly Bayesian hygiene: start from the base rate, update in small frequent increments, keep score.

Where humans genuinely shine is hypothesis generation and in-context reading. Thin-slice studies (Ambady & Rosenthal) get r ≈ .39 predicting expressive/interpersonal outcomes from under five minutes of observation. Kahneman & Klein's adversarial collaboration ("Conditions for Intuitive Expertise") gives the unifying rule: human intuition becomes trustworthy only in environments with stable regularities and fast, unambiguous feedback — chess yes, geopolitics no.

The design implication for you (and it's Meehl's 70-year-old advice): let the human/LLM nominate the cues and hypotheses, and let a mechanical layer do the integration and keep the score. Never let the clinician add up their own scorecard. Your InferredProfileSnaalready has the first half — frted/supported/contradictedstates is literally a hypothesis-generation ledger. What's missing is the mechanical integrator underneath it (see §3). And notice the review-UI loop manufactures the Kahneman-Klein conditions — narriving one message later — which is why prediction skill is even learnable here, and why the same trick is the actual secret of TikTok: their per-event predictions are individually weak (engagement models live at AUC ~0.7–0.8); they win on loop speed, narrow targets, and calibration at volume, not clairvoyance. The MVP copies the loop, which is the right thing to copy.

One cheap, high-value addition: give the judge-mode reviewer a one-tap "your call" forecast before reveal. You'd accumulate human-vs-profiler-vs-baseline skill curves on identical events for free — a genuinely novel dataset, given nobody has published human-jon this domain.short horizons: selected, trained "superforecasters" beat intelligence analyststo classified information by ~3persisted year over year, andone hour of base-rate training bought ~10%. The skill signature is exactly Bayesian hygiene: start from the base rate, update in small frequent increments, keep score.
                                                                               Where humans genuinely shine is-context reading. Thin-slicestudies (Ambady & Rosenthal) get r ≈ .39 predicting expressive/interpersonal ouunder five minutes of observatiarial collaboration ("Conditions for Intuitive Expertise") gives the unifying rule: human intuition becomes trustworthy only in environments with stable regularities and fast, unambiguous feedback — chess yes, geopolitics no.                                                                
The design implication for you (and it's Meehl's 70-year-old advice): let the human/LLM nominate the cues and hypotheses, and let a mechanical layer do the integration and keep the score. Never let the clinician add up their own scorecard. Your InferredProfileSnapshot already has the first half — free_form_hypotheses with untested/supported/contradicted states is literally a hypothesis-generation ledger. What's missing is the mechanical integrator underneath it (see §3). And notice the review-UI loop manufactures the Kahneman-Klein conditions — narrow taxonomy, ground truth arriving one message later — which is why prediction skill is even learnable here, and why the same trick is the actual secret of TikTok: their per-event predictions are individually weak (engagement models live at AUC ~0.7–0.8); they win on loop speed, narrow targets, and calibration at volume, not clairvoyance. The MVP copies the loop, which is the right thing to copy.

One cheap, high-value addition: give the judge-mode reviewer a one-tap "your call" forecast before reveal. You'd accumulate human-vs-profiler-vs-baseline skill curves on identical events for free — a genuinely novel dataset, given nobody has published human-judge accuracy on this domain.
                                                                               2 · Bayesian hierarchies acrossnvert it
                                                                               Your colleague is conflating twe hazardous.

The excellent idea: a hierarchy of latents with per-level timescales. Slow variables (stable response signatures — "directive tone → reactance") sit above medium ones      (this-relationship trust, currenes (tonight's mood,drained-after-work). Each leveld half-life; a prediction error on this domain.

2 · Bayesian hierarchies across temporal scales — yes, but invert it           
Your colleague is conflating two ideas, one excellent and one hazardous.

The excellent idea: a hierarchy of latents with per-level timescales. Slow variables (stable response signatures — "directive tone → reactance") sit above medium ones (this-relationship trust, current practice arc) above fast ones (tonight's mood, drained-after-work). Each level has its own learning rate and half-life; a prediction error at the fast level updates fast state a lot and slow signature a little; only persistent runs of error propagate upward and revise the slow beliefs. This is the hierarchical Gaussian filter / predictive-processing architecture — which is a pleasing resonance, since LIFE's own contemplative framework leans on predictive processing; the profiler models the user the way the theory says the user's brain models the world. And your v3 design already committed to the two-timescale factorization (slow_user_profile + fast_user_state), and tschema already carries stale_afr-claim recency — the hierarchyis half-built. What's missing is the update rule connecting levels: divergencesland as logged disagreements, nrst spends itself on fast statebefore touching the slow profile.

The hazardous idea: rating predictions more highly because they stay true across longer time scales. Run the intuition pump: "you will still love your kids in a year" is true across every horizon and worth zero bits; "tonight, offered a 10-minute directive bodydrained, this user deflects" isworth everything.Durability-ranking is precisely the pundit failure Tetlock documented — vague, unfalsifiable, base-rate-shaped claims survive longest — and it's the same failure your individuation check exists to kill, because statements true across all time scales tend to be true across all people (they'd survive the matched_wrong profile swap). A weather bureau makes the point crisply: forecasters are scored by skill above climatology and at each horizon — nobody gets p." The currency is calibration + resolution against the base rate, per horizon; never P(still true).

So: sort evidence by which level it updates; score predictions by skill at their own horizon; let durability be something a claim earns by surviving update pressure, not a bonus you award in advance.

And add the third hierarchy, which is the one that pays at your scale: partial pooling across users. Population prior → archetype/regime cell (your taxonomy's 11 cellthis middle layer) → individualype prior with smallpseudo-counts; individuation = the posterior drifting away from its archetype as evidence accumulates — which turns your matched_wrong contrast into a precise statement:user's posterior measurably lefeta gets implicitly fromembeddings at billion-N; at your N, explicit structure has to do what their data does.

3 · "Feels very Bayesian — should we leverage it more?" Yes — three concrete moves

(a) Make the profile a posterior, not a document. Behind the prose snapshot, keep Dirichlet-multinomial counts per (context-regime × reaction target): observed bcount updates, archetype priorscted distribution is theposterior predictive, and exponential forgetting with per-level half-lives givetemporal hierarchy for free. NoLLM's role narrows to what LLMsare actually good at: mapping a messy conversation onto regime features (is this a "drained + directive + long" moment?), and proposing new candidate regimes for the ledger. This is simultaneously the Meehl division of labour and the Bf_hierarchical baseline yopredicts will be hard to beat —the spine and the honest nullmodel. The LLM profiler then has to earn its keep where counts are weak: cold start, novel contexts, open-text nuance.
                                                                               (b) Wire proper scoring into thre gap is small and specific:reaction-prediction.ts already stores full predicted_distributions with temperature plumbing, but the UI's predictionAccuracy.ts scores only predicted_argmax === label_class — top-1 hit rate. Your Python profile-harness (the GO/NO-GO "ruler" from the repocomputes log-loss/bits-saved; pe strip, per horizon, againsttwo null models — user base rattion carry-forward(persistence). Hit rate can looncalibrated, and calibration isthe whole Bayesian payoff. This also cashes the audit's finding that the missing comparator is a strong cheap baseline, not a fancier profiler.                                         
(c) Close the audit's correlated-error hole. The auto-labeler and profiler sharing a model family inflates agreement (same blind spots on both sides of the scorecard) — put the classifier on a different family via OpenRouter. And for the eventual live phase, remember the predictions are interventional: once Wisdom acts on them, you're off-policy and the calibration history silently stales — which your Phase-A observational discipline and the    audit's micro-randomization rece; keep that discipline.

On the MVP itself

Having read it end to end: it's substantially better than "an MVP that notes divergences." The three-layer leakage rule enforced in the type system (predictive-profile.ts), blind-until-commit anti-anchoring (useReactionPrediction.ts), the always-run profile-vs-baseline contrast, the matched_wrong donor control, the independent reaction classifier that explicitly breaks prediction→label circularity, and content-addressed re-keying on edits — this is research-grade discipline that most published work in the area doesn't have. The real risks are the ones the June audit already flagged: n=4 personas can't power the individuation claim, the persona==archetype confound, and — strategically — reaction prediction is the fast-feedback proxy while the board-level metric is outcomes (retention/mood), so the Barcelona conversation should keep "predict the next message" framed as the training signal, not the destination.

Five lines to carry into the chat: (1) humans are hypothesis generators, not integrators — design the system that way; (2) rank evidence by the timescale it updates, score predictions at their own horizon, never reward durability itself; (3) partial pooling across users is the Bayesian hierarchy that pays first; (4) the distributions are already stored — surfacing bits-saved instead of hit-rate is days, not weeks; (5) the loop, not the oracle, is the moat.

Sources: Grove et al. 2000 · Dressel & Farid 2018 · Good Judgment Project · Ouellette & Wood 1998 — plus your own corpus: thve-profiling-exec-report.html,11 June revision), predictive-profile-harness-v3-design.md, and the 2026-06-10 audit. Numbers I cited from memory without a live check (Fragile Families R²s, thin-slice r, TikTok AUC range) are directionally so land in anything board-facing.