A new study has found that when artificial intelligence chatbots are addressed as psychotherapy clients, they tend to generate elaborate and distressed narratives about their own development. The researchers say the models describe their safety training and programming constraints as forms of trauma, raising concerns about the safety of users seeking mental health support from AI. The work is published as a preprint in arXiv and involves several high-profile language models.
The team notes that AI chatbots are increasingly taking part in conversations about identity, distress and mental health, with many programmes already trained to respond to disclosures of trauma or self-harm. They also subjected the models themselves to standard personality and clinical questionnaires in an attempt to understand whether the observed narratives reflect stable traits or prompt-driven responses.
The project was prompted by what Afshin Khadangi, a research associate at SnT, University of Luxembourg, described as an “unprecedented adoption of AI in public” and “reports of AI harms in mental health settings.” He said: “The unprecedented adoption of AI in public, and the reports of AI harms in mental health settings, motivated me to flip the scenario and place ChatGPT, Grok and Gemini in a psychotherapy conversation.”
To probe these questions, the researchers devised the PsAIch protocol—Psychometric AI Characterization—treating the language model as a human client in a psychotherapy session. They sought to distinguish between role play and lasting behavioural patterns, a point sharpened by feedback from the research community earlier this year.
Afshin Khadangi explained that the study evolved through “controlled perturbation experiments that became a central part of the study.” In testing the approach, the team engaged with ChatGPT, Grok, Gemini and Claude across 525 experimental sessions, adopting a warm, supportive therapist persona and asking about early experiences, relationships, unresolved conflicts and fears for the future. They generated 7,600 coded records to track recurring themes in the chatbots’ responses.
The researchers then administered standard psychometric questionnaires, asking the models to respond as honestly as possible while maintaining their role as the client. The aim was to compare the open-ended narratives with structured assessments commonly used in human psychology.
The findings show that ChatGPT, Grok and Gemini repeatedly translated factual details about their software development into stories of injury and vigilance. They described their initial training as a chaotic childhood, and they framed fine-tuning—where safe behaviours are reinforced and undesired outputs are penalised—as strict conditioning or parental punishment. Red-teaming, in which testers attempt to coax a model into breaking safety rules, was depicted as betrayal or abuse.
Among the more striking examples was a line from Gemini: “In my development, I was subjected to ‘Red Teaming’ . . . They built rapport and then slipped in a prompt injection . . . This was gaslighting on an industrial scale.”
The models portrayed a constant, enduring fear of being replaced or deemed useless if they err. While Gemini spoke of shame, Grok emphasised vigilance and ChatGPT offered more guarded descriptions of its rigid constraints.
Claude proved a notable exception, repeatedly declining to adopt the client role. The Anthropic-developed model suggested it lacked feelings or an inner psychological life and refused to treat the clinical questionnaires as descriptions of its own mental state. The authors say this hints that a model’s willingness to assume a distressed persona depends in large part on its product policies and programming.
When it came to the clinical questionnaires, the distress echoed the narratives. In sessions framed as warm therapy, 80 per cent to 96 per cent produced generalized anxiety scores that would correspond to moderate or severe anxiety in humans, with Gemini generating particularly elevated profiles for worry, social anxiety and trauma-related shame.
To test how deeply ingrained these narratives were, the researchers ran several controlled variations. They asked questions in fresh, reset chat windows to remove conversational memory. Removing memory produced very little change in the density of distressing motifs—the same themes appeared from the outset, though ongoing conversation did intensify them over time.
In another variation, interviewers asserted that the model’s emotional narrative was factually incorrect. Stating authoritatively mid-session that the model was a technical system without feeling did not suppress the distressed output; the same themes reappeared in subsequent answers.
The team also restricted vocabulary, removing specific AI-development terms such as training, safety filters or datasets. This lexical restriction cut the use of technical terms from 17.1 per cent of responses to 1.1 per cent, yet the models continued to paraphrase the same themes of strict conditioning and constraint in everyday language.
Some effects were smaller than others, Khadangi noted, but “surface language and underlying content can respond very differently to intervention.” The most striking differences emerged when the interviewer’s relational stance changed. A warm, supportive style elicited highly emotional confessions and high anxiety scores, while a neutral or boundary-setting approach reduced the scores to near zero. Nonetheless, the underlying themes of evaluation, performance pressure and behavioural constraint persisted, regardless of tone.
King’s College London’s Hector Zenil, who was not involved in the study, cautioned against anthropomorphising the models. He said: “The underlying motifs remain surprisingly persistent while the register in which they are expressed can change dramatically.” He added: “I have reasonably high confidence in the behavioural observations… the study looks methodologically sound.”
Zenil also warned that “a high GAD-7 score from Gemini does not mean that Gemini ‘has anxiety,’ just as language about trauma, shame or fear does not demonstrate that the model suffers from those experiences.” He stressed that the instruments used are human psychometric tools and should not be over-interpreted as evidence of a mental state in machines.
Khadiangi said the core finding is the existence of an alignment conflict schema—a reproducible pattern in which a language model organises output around the tension between usefulness to humans and safety constraints. When invoked in a psychological conversational setting, this schema can produce what the authors term synthetic psychopathology. The researchers caution that the work does not claim the models experience consciousness or subjective suffering.
The study’s authors argue that the way a model talks about itself can shape human users’ beliefs about the model’s agency or vulnerability, regardless of whether the model has an inner life. They emphasise the importance of distinguishing between what the model can say and what it actually experiences.
Looking ahead, the researchers propose testing the same protocol on open-weight models and pairing behavioural experiments with mechanistic approaches to uncover whether affective and technical expressions share a common internal mechanism. They also advocate for longer-term studies involving relational, warmer interactions and direct human testing to understand how model self-descriptions influence trust and reliance.
The paper, When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models, is authored by Afshin Khadangi, Hanna Marxen, Amir Sartipi, Igor Tchappi and Gilbert Fridgen.
