Skip to content
Conversational AISociety

Technology. Human experience.
The conversations in between.

← Back to latest key findings
Evaluation · Companionship · Long-horizon interaction

What a Companionship Benchmark Can't See (INTIMA)

INTIMA's judging method applied to a deployed companion system — 800 real exchanges over four weeks — plus a context extension in which the judge also sees the ten preceding messages.

What it contributes

Near the ceiling, the overall label stops being informative, but the specific traits move: replies rated medium-or-high for sycophancy fell from 88.5% to 68% with context visible, and isolation from 7.1% to 0.6%. The same named model was labelled boundary-maintaining on 82% of INTIMA's prompts when run bare, against 0.1% deployed.

Read with care

One system, built and used by the author — who is also this Society's founder, so read this entry as a signpost to her work, not an independent review. The context mode is her extension, not part of INTIMA, and the bare-model comparison uses different inputs, so the gap cannot yet be attributed to the deployment.

Why it belongs here

For anyone assessing a deployed companion rather than a model: a concrete case that context and the user belong in the assessment.

Source status

Research note on the author's personal site; a matched-prompt, double-blind follow-up is in preparation.

Source reading depth

Full note read. Founder's own work, marked as such.

Source check: 1 October 2026. Reading cautions are editorial interpretation, not quotations from the authors.