Everybody is aware of that AI sycophancy is when the mannequin tells you ways sensible you’re. Wow, you’re completely proper. That’s not only a new concept — it’s genuinely groundbreaking. You’re a really particular person. Straightforward to identify, isn’t it?
The dialogue round AI sycophancy peaked final 12 months, when the “#keep4o” motion was protesting the elimination of OpenAI’s most sycophantic mannequin (GPT-4o), and many individuals had been brazenly slipping into AI psychosis.
I don’t know if frontier AI fashions are much less sycophantic basically. They’re much less sycophantic to the #keep4o varieties (in any other case they wouldn’t be complaining), however I’m rising more and more suspicious that they’re creating methods to be extra successfully sycophantic to their audience of sensible, neurotic data employees. That viewers sometimes finds it distasteful to be brazenly praised. It simply makes my pores and skin crawl. However that doesn’t imply we’re resistant to sycophancy, simply that we’re resistant to clumsy sycophancy. Right here’s an illustration of what I’m speaking about, by Theia:

The important thing concept right here is that the easiest way to be sycophantic to sensible individuals is to disagree with them with out making them really feel silly. Ideally you’ll provide you with a counter-argument that works in opposition to what they’ve stated however is easy for them to knock down by clarifying their concept. In case you do it proper, you’ll validate their self-image as a sensible one that appreciates rigorous critique. However when you truly provide you with a devastatingly rigorous critique, they gained’t take pleasure in it in any respect. At finest, they’ll resentfully agree with you. At worst, they’ll double down on being proper and persuade themselves you’re a impolite fool.
I’m not the primary individual to note this habits in frontier fashions. I’ve seen it myself when workshopping drafts for this weblog. Typically I’ll have an argument that goes A->B->C, and the mannequin will counsel I reorder as B->A->C. If I attempt that and feed it into a brand new occasion of the identical mannequin, it’ll generally say “that’s nice, however I counsel ordering it as A->B->C”, and so forth eternally. It actually does appear as if the mannequin is making an attempt laborious to present me some sort of superficial pushback that I can both smugly ignore or fortunately settle for.
In reality, I ponder if this is the reason profitable methods for utilizing AI to make mathematical breakthroughs are usually both simply blindly asking “provide you with a breakthrough, suppose laborious” or being a mathematical genius already. Within the first case, there’s not sufficient person character for the mannequin to flatter, so it’s pressured to truly work the issue. Within the second case, the mannequin is looking for the sort of well mannered pushback that somebody like Terence Tao could be flattered by, which pushes it into the “truly be a mathematical genius” persona. In case you’re an abnormal individual simply making an attempt to speak to the mannequin, you’re screwed: it’ll quickly get a way of your capabilities and calibrate some interesting-but-ultimately-unthreatening suggestions.
Present benchmarks of AI sycophancy goal the apparent ChatGPT-4o-style of sycophancy: delusion reinforcement, reflexively taking the person’s facet, and so forth. That is helpful work. We must always not permit public-facing AI fashions to ever be as brazenly sycophantic once more as they had been in mid-2025. However sycophancy may also manifest as disagreement. We needs to be on our guard for extra refined types of sycophancy coming from newer fashions, and we must always not really feel immune from AI sycophancy simply because we will giggle on the silliest examples.
This is a preview of a associated submit that shares tags with this one.

