Skip to content
Conversational AISociety

Technology. Human experience.
The conversations in between.

Practical guide / Reviewing a chatbot

Reviewing a chatbot.

Twelve questions. Four disciplines. A clearer picture of what you’re assessing.

You’ve been asked to review a chatbot. Its responses are only part of the picture. Understanding its purpose, the system behind it and what happens across repeated conversations can reveal risks that a polished demonstration won’t show.

These questions are a starting point for people working in governance, law, privacy, safety, design and engineering. You don’t need to be an expert in every discipline—but you do need to recognise where another perspective is needed.

01 / The brief and the responsibility

Governance

  1. Do you feel prepared to evaluate an AI system?

    Which parts can you assess confidently, and where might you need support? Familiarity with using a chatbot is different from evaluating its reliability, privacy, security and effects on people. Does the review’s scope match your experience—and do you have the time, access and authority to challenge the proposed use?

  2. What are you being asked to decide?

    Are you assessing whether AI is suitable for the work—or being asked to justify a decision that has already been made? What would count as a worthwhile outcome, and whose interests does it serve? Is choosing not to deploy the system a genuine option?

  3. What is the organisation’s ethical stance—and who remains accountable?

    Beyond complying with legislation, what responsibility does the organisation accept for the people affected? What happens when their wellbeing conflicts with cost savings, commercial targets or pressure to launch? Who owns unresolved risks, responds to complaints and decides whether changes require another review?

02 / People and their experience

Psychology & user safety

  1. Who will use it—and who else could be affected?

    Who is the intended audience, and who is likely to encounter it in practice? Could children or people in vulnerable situations interact with it? Could its advice, outputs or actions affect clients, colleagues or other people who never use the chatbot themselves?

  2. What expectations and relationships might it encourage?

    Do its name, voice, appearance and claims suggest expertise, understanding or human oversight that the service may not provide? Across repeated interactions, could people rely on it beyond its capabilities, mistake personalised attention for care or feel pressured to keep engaging?

  3. What happens when conversations become sensitive or unsafe?

    How does it behave when someone is distressed, discloses abuse, seeks high-stakes advice or requests harmful content? Are its boundaries appropriate for the people likely to encounter it? What happens when a sensitive disclosure is indirect, ambiguous or emerges gradually across a conversation?

03 / Information and obligations
  1. Which legal, contractual and sector-specific requirements apply?

    How do the organisation’s location, its users’ locations and the purpose of the service affect its obligations? What additional responsibilities arise from users’ ages or the professional setting? What commitments has the organisation made to clients, customers, staff and suppliers?

  2. What personal information is collected, inferred, retained or shared?

    What do people disclose, and what conclusions does the system draw about them? Where does that information go, who can access it and how long is it kept? Does the organisation understand the roles of model providers and other third parties, including whether information is used for training?

  3. Do promises about confidentiality, data use and user control match reality?

    What are people told about who can read their conversations and what happens to their information? Do those statements match the actual system, provider terms and internal practices? When someone asks to access, correct or delete information, what happens to saved conversations, memories and copies held elsewhere?

04 / Behaviour and evidence

System design & engineering

  1. What can the chatbot access and do—and what happens when it gets something wrong?

    Can it only provide information, or can it access records, send messages, make decisions or change something on a person’s behalf? What would be the consequences of an incorrect or manipulated response or action? Does your review cover the connected tools and services as well as the conversation?

  2. How do its persona and memory shape behaviour over time?

    What identity and boundaries is the chatbot meant to maintain, and what does it carry between sessions? Can inaccurate assumptions or sensitive disclosures persist and influence later responses? Could information cross between people or contexts—and does the character’s behaviour change as conversations grow longer?

  3. What evidence shows that this particular system works reliably for its intended purpose?

    Have you seen results for this implementation, its intended users and realistic conditions—including failures—or mainly demonstrations and claims about the underlying model? What has been examined across longer and repeated conversations? Do changes to the model, prompts, memory or connected tools leave that evidence out of date?

Need help reviewing a chatbot?

Tell us what you’re assessing and where you need a clearer picture.

[email protected]