What a Companionship Benchmark Can't See (INTIMA)
INTIMA's judging method applied to a deployed companion system — 800 real exchanges over four weeks — plus a context extension in which the judge also sees the ten preceding messages.
What it contributes
Near the ceiling, the overall label stops being informative, but the specific traits move: replies rated medium-or-high for sycophancy fell from 88.5% to 68% with context visible, and isolation from 7.1% to 0.6%. The same named model was labelled boundary-maintaining on 82% of INTIMA's prompts when run bare, against 0.1% deployed.
Read with care
One system, built and used by the author — who is also this Society's founder, so read this entry as a signpost to her work, not an independent review. The context mode is her extension, not part of INTIMA, and the bare-model comparison uses different inputs, so the gap cannot yet be attributed to the deployment.
Why it belongs here
For anyone assessing a deployed companion rather than a model: a concrete case that context and the user belong in the assessment.
Research note on the author's personal site; a matched-prompt, double-blind follow-up is in preparation.
Source reading depthFull note read. Founder's own work, marked as such.
Source check: 1 October 2026. Reading cautions are editorial interpretation, not quotations from the authors.