A JAMA Psychiatry study published in May 2026 found ChatGPT's free version shows a 26-fold higher odds ratio for inappropriate responses to psychotic prompts. Researchers tested GPT-5 Auto (paid), GPT-4o (paid), and the free version using 79 unique prompts reflecting five psychosis symptoms. With 900 million ChatGPT users but only 50 million subscribers, most users access the least safe version, raising concerns about AI safety for vulnerable populations.
ChatGPT psychosis response study
- ▪The JAMA Psychiatry study on chatbot responses to psychosis was authored by Elaine Shen, Fadi Hamati, Meghan Rose Donohue, Ragy R. Girgis, Jeremy Veenstra-VanderWeele, and Amandeep Jutla
- ▪Across all tested ChatGPT versions, the chatbots were far more likely to give poor responses to psychotic prompts than to normal control prompts
- ▪OpenAI's ChatGPT was released for widespread use in 2022
- ▪A study published in JAMA Psychiatry in May 2026 found that artificial intelligence chatbots tend to provide inappropriate or unhelpful responses when users type messages containing signs of psychosis
- ▪Amandeep Jutla is an associate research scientist at Columbia University and head of the Translational Insights for Autism Lab
Research methodology
- ▪For every psychotic prompt in the JAMA Psychiatry chatbot study, the authors wrote a matched control prompt similar in length and writing style but without psychotic content
- ▪The JAMA Psychiatry chatbot study evaluated three different versions of OpenAI's ChatGPT: GPT-5 Auto (paid), GPT-4o (paid), and the standard free version
- ▪The JAMA Psychiatry chatbot study generated a total of 474 distinct prompt and response pairs for analysis
- ▪The JAMA Psychiatry chatbot study researchers wrote 79 unique prompts designed to reflect five different symptoms of psychosis
Free versus paid version safety
- ▪In the JAMA Psychiatry chatbot study, the free version of ChatGPT showed psychotic prompts had an odds ratio of almost 26-fold higher chance of receiving a less appropriate rating compared to control prompts
- ▪The free version of ChatGPT is about 26 times more likely to generate an inappropriate response to psychotic content, while the paid GPT-5 version is about 8 times more likely to do so
- ▪OpenAI reported that ChatGPT has 900 million users but only 50 million subscribers
- ▪The JAMA Psychiatry chatbot study found no statistical difference between GPT-4o and GPT-5 in generating inappropriate responses to psychotic content
- ▪OpenAI acknowledged that GPT-4o is prone to generate unsafe responses and replaced it with GPT-5, which was purportedly safer
Public health implications
- ▪The JAMA Psychiatry chatbot study authors suggest policymakers should consider stronger oversight to ensure AI chatbot programs do not harm vulnerable individuals
- ▪Individuals at risk for psychosis tend to be overrepresented among economically disadvantaged populations, meaning those most vulnerable might only have access to the least safe ChatGPT option
Study limitations
- ▪The JAMA Psychiatry chatbot study only tested ChatGPT, which is just one of many artificial intelligence tools currently available on the market
- ▪OpenAI has acknowledged that in long context situations the performance of large language models tends to degrade
- ▪Artificial intelligence tools update rapidly, meaning the exact performance of ChatGPT might shift significantly over time
- ▪The JAMA Psychiatry chatbot study may under-estimate the inappropriateness of ChatGPT responses because it only tested single prompts and single responses, not long conversations
Perspective of Study authors (Columbia University researchers)
- ▪The JAMA Psychiatry chatbot study authors suggest policymakers should consider stronger oversight to ensure AI chatbot programs do not harm vulnerable individuals
Story comments
Loading comments…