Skip to content
Conversational AISociety

Technology. Human experience.
The conversations in between.

← Back to latest key findings
Safety · Design · Evaluation

INTIMA: A Benchmark for Human-AI Companionship Behavior

A framework covering 31 companionship behaviours through 368 targeted prompts, with responses classified by their relational stance.

What it contributes

Makes companionship-reinforcing, boundary-maintaining and neutral responses explicit evaluation categories. It offers a starting vocabulary for examining relational behaviour across models.

Read with care

These categories describe model behaviour, not measured harm or benefit to a person. Prompt-based evaluations cannot by themselves establish the effects of a continuing relationship.

Why it belongs here

For teams making relational boundaries and emotional support part of their evaluation plan.

Source status

Linked to the August 2025 arXiv version. Later publication status not established in this sweep.

Source reading depth

arXiv abstract and dataset record checked. This source note does not establish peer-reviewed publication status.

Source check: 1 October 2026. Reading cautions are editorial interpretation, not quotations from the authors.