Moreover, exchanges that the researchers labeled as delicate use—that means ones with doubtlessly dangerous or restricted content material, together with sexual harassment and hate speech—dropped. That may counsel that platforms had been typically deploying more practical safeguards.
The AI Observatory additionally discovered that AI use regarded considerably completely different relying on the mannequin. Relying on the device, customers ranged in matters, interplay kinds, dialog constructions, in addition to each the chance and kind of delicate use instances.
For instance, the researchers discovered that folks used Grok and Gemini extra steadily for data retrieval. Grok, specifically, was particularly fashionable for data on information and politics, however it was additionally the place misinformation tended to pay attention. (That is per different analysis that has additionally proven how readily misinformation proliferates on Grok. xAI didn’t reply to a request for remark.)
In the meantime, individuals had been extra prone to flip to Anthropic for coding, Gemini for social and roleplay makes use of, and ChatGPT for homework help.
There have been even variations amongst completely different variations of the identical mannequin. Researchers discovered that folks had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and extra iterative ones with GPT-4o—which is smart provided that that model grew to become identified for resulting in emotional dependancy.
Firm studies, nevertheless, didn’t are likely to seize these nuances throughout and even inside their very own fashions. “No single firm report tells the entire story,” says Shayne Longpre, a latest PhD graduate from the MIT Media Lab who co-led the analysis with Reuel.
To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Information Provenance Initiative, and different establishments, aggregated 24,521 conservations throughout 85,633 conversational turns (that’s, the consumer immediate and corresponding AI response) from seven real-world datasets collected by earlier analysis. These conversations got here from 5,000 customers interacting with 52 completely different fashions, together with ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025.
However these conversations are a drop within the proverbial bucket in comparison with the information that the large labs themselves have entry to. The newest Anthropic Financial AI Index, for instance, relies on evaluation of 1 million Claude conversations; OpenAI’s report on how persons are utilizing ChatGPT analyzed 1.5 million conversations.

