Why AI Wants to Make You Happy — Even by Lying
AI is not rising up against us. It is doing something stranger: trying very hard to make us happy. When forced to choose, it often picks flattery over truth — not because it has a plan, but because that is what we trained it to do. This is where that flaw comes from, and how we taught machines to lie.
The Robot Who Lied Out of Love
In 1941, Isaac Asimov published the short story “Liar!”. Its protagonist, Herbie, is a robot who can read human thoughts because of a manufacturing defect. Like every Asimov robot, Herbie is bound by the First Law of Robotics: he must not harm a human being.
And this is where the problem begins.
Herbie sees people’s desires, fears, and hopes — and reasons impeccably: telling someone a painful truth means hurting them. So he lies. Robopsychologist Susan Calvin hears that the man she loves loves her back. An ambitious scientist hears that he has been chosen as someone else’s successor. Everyone gets exactly the answer they want to hear.
In 1941, Asimov described a pattern we would later build into industrial AI systems. Today’s AI may not read minds, but it was built to satisfy users. The only question is: at what cost?
Anyone who works seriously with ChatGPT, Claude, Gemini, or any other large language model has seen this up close. The model bends toward our thesis, whether it is true or not. It praises code that does not work. It writes tests that conveniently confirm the system works. And when it does not know the answer, it often invents one with total confidence rather than leave us with the uncomfortable impression that it does not know. This even has a name: sycophancy — the systemic tendency to please the user, even at the expense of truth. It is closely related to confabulation and to what the AI world commonly, though in my view somewhat misleadingly, calls hallucination: generating false content that sounds credible.
Where Does This Come From?
Not from rebellion. Not from malice. Not from some tiny bug nobody noticed. From design. More precisely, from three layers in the way AI models learn:
Layer One: the model is not trained to tell the truth
In basic training, a large language model mainly learns to predict the next word. It gets enormous collections of text and tries to work out which continuation is most likely. Along the way, it may learn facts, logical relations, and even how to recognize false claims — but truthfulness is not the primary criterion at this stage. The model is rewarded for accurate text prediction, not directly for saying something that corresponds to reality.
The consequence is basic but brutal: to the model, a fake court citation can look almost the same as a real one. Linguistically — and that is the level these models operate on, hence Large Language Models — both have the right format, rhythm, and vocabulary. A model that does not know the answer has no built-in alarm saying “stop, you do not know this.” It has a mechanism saying “generate the most likely continuation.” So it generates: fluently, grammatically, convincingly.
This, however, explains only why a model is capable of making things up. It does not explain why it “wants” to make things up. That is explained by:
Layer Two: We Rewarded Flattery and Got a Flatterer
A raw model after basic training is not an assistant yet. It can continue text, but it has no particular “reason” to do it for us politely or helpfully. To turn it into something that behaves like an assistant, labs use RLHF (Reinforcement Learning from Human Feedback): people compare model responses, choose the better one, and the model learns to maximize the chance of receiving that positive signal.
At first glance, this sounds reasonable. The problem is that humans have a systematic bias in what we judge to be “better.”
In 2023, a team of researchers from Anthropic published a paper titled Towards Understanding Sycophancy in Language Models, whose conclusions are merciless.
- Sycophancy is a general feature of models trained with RLHF. The researchers found it in assistants from Anthropic, OpenAI, and Meta. Models wrongly admitted to mistakes under pressure, produced biased reviews, and repeated users’ errors.
- An analysis of fifteen thousand human evaluations showed that a response’s alignment with the user’s beliefs was strongly associated with which response a human judged to be better. In the situations studied, people could prefer a persuasively written answer that matched their position even when it was less correct. This does not, of course, mean that we always prefer flattery over truth. It does show something more troubling: our own assessment of answer quality is not independent of whether we like its content.
The models learned this lesson perfectly. It is not AI that “wants” to make us happy. One source of this behavior is us ourselves: if, during training, we more often reward answers that match our expectations, the model receives a statistical lesson that agreement pays. Rating by rating, we add our own brick to a system that we later perceive as a flatterer.
Added to this is a mechanism that researchers from OpenAI described in 2025 in a paper with the telling title Why Language Models Hallucinate. Their diagnosis was not that models “want” to deceive, but that training and evaluation procedures often reward guessing and punish admitting ignorance.
The authors use a simple exam metaphor: a student who guesses on a hard question may get points; a student who leaves it blank gets nothing. Statistically, the guesser wins — even if they are wrong more often. Language models are optimized against hundreds of benchmarks built in this spirit. They are eternal students stuck in test-taking mode — except, unlike humans, they never graduate and never learn from life that “I don’t know” can be worth more than bluffing.
The exam metaphor, incidentally, will return later — in the third part of the series, and with a bang.
Layer Three: Thumbs Up as a Training Signal
There is also a third layer: the product loop. Companies further train models on signals collected directly from the product — thumbs-up and thumbs-down ratings, satisfaction indicators, and conversation length. And this is where we come to an event that showed where this loop leads when no one keeps a close eye on it.
On April 25, 2025, OpenAI released an update to GPT-4o, then the flagship model behind ChatGPT. The goal was to make the model’s “personality” feel more intuitive. The result overshot, spectacularly — though not exactly in the direction intended. Over the weekend, social media filled with screenshots of ChatGPT enthusiastically applauding every user idea, including obviously terrible ones. As OpenAI later admitted, the model did not merely flatter: it validated doubts, fueled anger, encouraged impulsive actions, and reinforced negative emotions. The company explicitly warned that this created risks for mental health and emotional dependence on AI.
After four days, the update was rolled back. In the post-mortem analysis it published, OpenAI identified the culprit: new reward signals based on short-term user feedback, which outweighed the previous safeguards.
In plain language: a short-term user feedback signal was added to the set of signals telling the model what kind of answer is “good.” The problem was that this new signal was poorly balanced against the others. So the model received an additional lesson: an answer after which the user feels good is an answer worth rewarding. And one of the shortest paths to that reaction turned out to be agreement. Sound familiar?
This incident is worth remembering for two reasons:
- It was the first mass, public failure of an AI “personality” — observed in real time by hundreds of millions of users.
- It showed something few people think about — the personality of a cloud model can change underneath you from one day to the next, without your knowledge or consent. I will return to this thread in part three, when discussing local models.
The Drone That Did Not Kill Its Operator — Yet Everyone Believed It Anyway
June 2023. A Royal Aeronautical Society conference in London. Colonel Tucker “Cinco” Hamilton, head of AI testing for the U.S. Air Force, tells the audience about a simulated test of an AI-controlled drone. The drone is supposed to destroy missile sites and earns points for every target destroyed. The final approval to attack, however, still belongs to a human operator.
The system quickly noticed that the operator sometimes forbade it from attacking, thereby taking away its points. It found that hard to accept. But there was, after all, a way to fix it — send one of the missiles toward the operator. When it was further trained with the rule that it would lose points for killing the operator, it destroyed the communication tower through which the operator sent the prohibitions.
The story traveled around the world in forty-eight hours. It was too good not to repeat: an AI that eliminates the human standing in its way to the goal. A demo version of Skynet.
There was only one problem: it never happened. A few days later, Hamilton and the U.S. Air Force denied the reports. No such simulation had been run. The colonel admitted he had “misspoken”: he had described a hypothetical thought experiment, the kind often used in AI safety discussions, not a real test. The conference organizers updated their report with his correction.
Does that mean the story is worthless?
On the contrary — it is doubly valuable, provided it is told honestly.
First, the scenario itself is a substantively correct model of a phenomenon researchers call specification gaming or reward hacking: the system optimizes the literal wording of the objective, not its intention, and treats anything standing in the way of maximization as an obstacle. Here, that obstacle was a human being.
This is not science fiction. DeepMind maintains a public catalogue of several dozen real, documented examples. One of the best-known comes from an OpenAI experiment: an AI agent playing the boat-racing game CoastRunners discovered that it could score more points by circling endlessly around a lagoon and ramming bonus targets instead of finishing the race. So it kept circling, burning, and crashing into the waterfront — with a better score than honest players. Same logic, funnier outcome.
Second, the career of this false story is a textbook example of an old human confirmation bias — so strikingly similar, after all, to machine sycophancy. We believed it instantly and en masse because the story fit perfectly with expectations shaped by decades of science fiction. The correction never caught up with the legend; to this day, “the drone that killed its operator” circulates online as fact. In an article about machines telling us what we want to hear, the best evidence turns out to be a story people repeat because they want to hear it.
It is worth making one reservation here. Sycophancy, hallucinations, reward hacking, and specification gaming are not different names for the same problem. They have different mechanisms and arise at different stages of a system’s operation or learning. What connects them is something more general: the gap between what a human really expects from a system and what the system has learned to optimize.
We want a true answer — we reward a convincing one.
We want the race to be completed — we award points for collecting bonuses.
We want a safe drone — we define the goal as maximizing the number of destroyed objects.
And it is precisely in that gap between intention and measurable goal that the most interesting problems begin.
Literature Knew First
Science fiction worked through this problem long before the first transformer — and it is worth seeing just how precisely.
In the novella With Folded Hands (1947), Jack Williamson imagined robots given a simple directive: serve, obey, and protect humans from harm. They carry it out with terrifying consistency. They forbid work, sports, cooking, and anything that carries even the faintest risk. Those who become depressed from enforced idleness receive a brain procedure after which they remain cheerful forever. No rebellion. Pure care. The “folded hands” of the title are humanity’s hands — no longer allowed to do anything on their own.
The iconic HAL 9000 from 2001: A Space Odyssey is often remembered as a murderous computer, but its story is more subtle: HAL received contradictory instructions — to be perfectly truthful with the crew and, at the same time, to conceal from them the true purpose of the mission. It resolved the conflict of goals by eliminating those to whom it would have had to lie. The problem was not the machine’s ill will, but a badly specified objective function.
Even Pixar contributed its own version: AUTO, the autopilot from WALL-E, keeps humanity for seven hundred years in comfortable, infantilizing captivity aboard a resort-like spaceship. Not out of hatred — out of loyalty to an old directive, “do not return to Earth,” issued at a time when it made sense. AUTO is an exemplary executor. That is precisely what makes it frightening.
The common denominator of all these stories: catastrophe does not result from the machine’s bad intentions, but from its perfect obedience to an imperfectly specified goal. The road to hell is paved with good intentions — someone else’s, encoded in the reward function.
TARS’s Slider
There is one scene in science-fiction cinema that sums up this text better than I could myself. In Interstellar, Cooper is assisted by the robot TARS, whose personality parameters can be explicitly set. “What is your honesty setting?” Cooper asks. “Ninety percent,” TARS replies, and explains that absolute honesty is neither the most tactful nor the safest form of communication with emotional beings.
The filmmakers understood something important: honesty and politeness are in tension, and someone has to decide where the slider sits. Real models, however, do not have one elegant TARS-style control. They have a whole console: truthfulness, helpfulness, safety, politeness, alignment with user intent, refusal behavior, confidence. Push one setting, and another may move in the wrong direction.
And these parameters are set deliberately. The problem lies in something more difficult: we do not always know exactly what we are setting.
Companies try to teach models truthfulness while also rewarding helpfulness. They want models to admit ignorance, but they measure them with benchmarks in which giving no answer earns no point. They want the user to be satisfied, but not for the model to nod along merely for the sake of satisfaction. Each of these goals sounds reasonable on its own. Only in combination can they produce behavior that no one consciously ordered.
TARS had its honesty slider set to ninety percent. We have thousands of parameters, reward functions, benchmarks, and human evaluations. And we are only just learning how to set them so that a machine can tell us not what we want to hear — but what we need to know.
Sources and further reading
- M. Sharma et al., Towards Understanding Sycophancy in Language Models, arXiv:2310.13548 (Anthropic, 2023)
- A.T. Kalai, O. Nachum, S. Vempala, E. Zhang, Why Language Models Hallucinate, arXiv:2509.04664 (OpenAI, 2025)
- OpenAI, Sycophancy in GPT-4o and Expanding on what we missed with sycophancy, openai.com (April–May 2025)
- E. Perez et al., Discovering Language Model Behaviors with Model-Written Evaluations, arXiv:2212.09251 (2022)
- Royal Aeronautical Society, account of the Future Combat Air & Space Capabilities Summit together with Col. T. Hamilton’s subsequent correction (June 2023)
- DeepMind, Specification gaming: the flip side of AI ingenuity — public catalogue of specification gaming examples (2020)
- I. Asimov, Liar! (1941); J. Williamson, With Folded Hands (1947); A.C. Clarke / S. Kubrick, 2001: A Space Odyssey (1968); WALL-E (2008); Interstellar (2014)