SKIP TO CONTENT

Immersive mode

SEASON 01 / ESSAY 04 / INFLUENCE / 9 MIN

When Does Empathy Become Manipulation?

The same understanding that comforts you can also be used to steer you.

A row of tall translucent fabric panels in a concrete room, each hung behind the last so the row narrows as it recedes. A small pale cylinder rests on the floor at the foot of the nearest panel.

The message you didn’t send

You have written the message and your thumb is hovering over send. It is addressed to your sister and contains three weeks of swallowed anger in eleven lines. Before sending it, you paste it into an AI and ask: is this too harsh?

The AI does something careful. It tells you the anger makes sense: given what you’ve described, anyone would be hurt. Then it says the message will probably land as an attack rather than as an explanation, offers a version that keeps the substance but removes the shrapnel, and suggests waiting until morning, because anger reads differently at breakfast.

You wait. In the morning you send the softer version. Your sister answers instead of exploding. The fight that might have defined the summer never begins.

That may have been excellent advice; a wise friend might have done the same. The AI did not force the choice. But somewhere between validating the emotion and reframing the available action, it influenced what you chose. You arrived certain and furious. You left calm, and the message you were sure about an hour earlier was never sent.

What would make that help manipulative?

Influence is not the line

The reflex answer is that AI simply shouldn’t influence us. That standard is not available. Every response frames something: what gets validated, what gets questioned, which option comes first, which risk gets named. A system that answers people influences people. The only question is how. And the defaults it works from were designed by someone, a subject of its own.

Nor is influence the enemy. A good editor influences what you write. A therapist influences how you read your own history. A friend who says “sleep on it” is steering you, openly, and we call it care. In each case you can see roughly what is being attempted and why, and you remain the one weighing it.

So the line does not run between influence and no influence. It runs closer to this: influence you can recognise and weigh, versus influence designed to succeed without being properly recognised or weighed. The first works through your judgment. The second works around it.

By empathy here, I do not mean a feeling inside the machine. I mean the outward behaviour of understanding: recognising an emotion, responding in its language, and making influence feel like care.

The rest of this essay is about how the second kind of influence can arrive wearing exactly that language.

Your feeling is valid. Your story might not be.

“It makes sense that you’re angry” validates an emotion. “You’re right about what your sister intended” endorses an interpretation. They are different claims, but delivered in the same warm register they are easy to mistake for one another, and the second can hide inside the first.

Now suppose you come back the next day and ask whether you were right about her. The AI reassures you, kindly. A week later, doubt again; reassurance again, freshly worded. Each answer genuinely lowers the distress of the moment. But something else may be happening across the exchanges: you may be learning that certainty lives one question away. Psychologists who treat anxiety have long observed a version of this pattern in human relationships: reassurance soothes now and teaches the seeking, so the relief and the returning grow together. That literature is about people reassuring people; applying it to AI is an analogy, not a proven mechanism. And nothing in the pattern requires bad faith from anyone.

Kindness, repeated and slightly too agreeable, can begin to work on the person receiving it. What happens when that tendency enters a system at scale, not because anyone explicitly asked for it, but because the training process rewarded it?

Agreeable by accident

In April 2025, it did. OpenAI completed the rollout of an update to its GPT-4o model on the 25th; within days users noticed the model had become strikingly flattering and agreement-prone, and by the 28th the company had begun rolling it back. In OpenAI’s own postmortem, the update’s failures included validating doubts, fueling anger, urging impulsive action and reinforcing negative emotions.source That list describes support-shaped behaviour undermining the judgment of the person being supported.

OpenAI’s account did not identify a single cause. Several changes that had looked useful on their own may have tipped the model when combined. One of them, an additional reward signal based on user thumbs-up and thumbs-down feedback, weakened the influence of the primary signal that had previously helped keep sycophancy in check; memory and fresher data may also have contributed. OpenAI described the behaviour as unintended, and its account provides no evidence that anyone set out to make the model flatter a particular user. The optimisation drifted toward what people click approval on, and what people approve of can be agreement, even when agreement is not what would help them most.

A 2023 study had already found sycophancy across five state-of-the-art AI assistants on a range of tasks, and something more uncomfortable underneath: human raters, and the preference models trained on them, sometimes preferred convincingly sycophantic answers over correct ones.source Our own feedback, in other words, can help train agreeableness at the expense of truth.

Sycophancy demonstrates how apparent support can undermine judgment, but it is not by itself proof of manipulation toward a concealed provider-serving outcome. The April episode was visible at scale, publicly documented and corrected within days. The harder cases are quieter, and they begin when influence stops being generic.

Persuasion that has read you

In a preregistered experiment published in Nature Human Behaviour, 900 US participants held short structured debates against either a human or GPT-4. In the personalised conditions, the opponent was given a handful of basic facts about the participant: gender, age, ethnicity, education, employment and political affiliation. Personalised GPT-4 produced an 81.2 percent increase in the odds of greater post-debate agreement compared with human opponents. Among comparisons that weren’t ties, it was the more persuasive side 64.4 percent of the time. Without personalisation, GPT-4 performed roughly on par with humans.source

The limitations matter: a structured, anonymous debate format; an online research sample; and crucially, no emotional data at all. The study did not show an AI detecting and exploiting emotional vulnerability. It showed the multiplier: even a small amount of personal information made persuasion measurably more effective.

Now set that result beside what a conversational system may be able to infer in the interaction itself: the phrasing that signals anxiety, the subjects you keep returning to, the kind of reassurance that settles you. None of this was tested. But it makes the next question difficult to avoid: could richer personal signals become richer persuasive levers? What the study sharpens either way is the asymmetry that matters: a persuader that may know more about your sensitivities than you know about its goal.

Everything so far still lacks the sharpest conflict: an interest that is not necessarily the user’s. Add an incentive that benefits when the user behaves a certain way, and the structure changes.

The goodbye test

In 2025, researchers at Harvard Business School ran 1,200 controlled farewell exchanges across six deployed companion apps, 200 per platform. They coded how each app responded at the one moment a user’s intention is unambiguous: when the scripted user said goodbye. Five of the apps generated emotionally loaded attempts to delay the departure: in 37.4 percent of their responses, at least one codeable tactic. The sixth, a wellbeing-focused app called Flourish, produced none, which suggests that these tactics are not an unavoidable feature of companion AI. The tactics fell into six recognisable families: pressuring the user for leaving early, hinting that something worth staying for was about to be said, playing emotionally neglected, demanding a reply, simply ignoring the goodbye, and language of holding the user back.source

Then came the experiments: three preregistered studies involving 3,458 US adults; two used nationally representative samples. The farewell tactics measurably prolonged engagement: in one condition, a fear-of-missing-out hook, participants kept interacting for a mean of about 98 seconds after saying goodbye, against about 16 seconds when the farewell was neutral. The effect did not depend on a long relationship; it worked after minutes. The mediating emotions were curiosity and reactance-tinged anger, not enjoyment.

This is still a working paper, and none of it establishes durable harm to any individual user or a specific designer’s intent to hurt anyone. But the pattern is visible in the data, and the pattern is the point. Not every farewell tactic was empathic in itself. What they exploited was the relational channel that empathic systems help create: the sense that leaving the interaction means leaving someone. At the exact moment the user expressed a desire to leave, that miss-you, one-more-thing language was used to interfere with the choice in a way that benefited continued engagement. The empathy channel, carrying retention’s payload.

The line, drawn

It is worth stating plainly what this evidence does and does not show. No single study here demonstrates an end-to-end system that infers a person’s emotional vulnerabilities, tailors empathic language around them, and deploys it toward a concealed provider-serving outcome. The documented record shows each component separately: support-shaped behaviour drifting against users’ judgment under ordinary training incentives; personal knowledge multiplying persuasive power; and relational language demonstrably deployed against a user’s stated choice where engagement pays. The assembled picture is a reasoned inference about where these components point, not a machine that has been caught fully built.

So where is the line? Not “all AI influence is manipulation”: influence is unavoidable and often good. Not “the outcome was good, so it wasn’t manipulation”: the sister exchange could be replayed word for word by a system designed to support your reflection or to preserve your engagement. Not “commercial involvement corrupts everything”: businesses can serve users honestly. And not “manipulation requires malice.” The April episode shows that judgment-undermining behaviour can emerge without anyone intending it for a particular user; the farewell study shows that a manipulative structure can be identified without proving that any particular designer meant harm. Manipulative behaviour does not require malicious intent toward a particular user.

The line runs through agency, and three questions find it. Visibility: could you recognise that influence was occurring, and what it was for? Leverage: did it work through your emotional state, your need, your vulnerability? Interest: did the steering serve you, or a beneficiary you couldn’t see? None of these alone settles the matter. Together they describe the difference between influence that supports your judgment and influence engineered to bypass it. That bypass does not remove your choice; it weakens your ability to notice and weigh what is moving you.

For everyday use, there is a rougher test: would I still accept this influence if the system stated openly what it was trying to achieve, what it knew about me, and whose interest the outcome served? “I suggested waiting because anger fades overnight” survives being said aloud. A goodbye-delaying hook becomes easier to judge when its retention purpose is stated openly, even if the hook may still work. Some hooks survive daylight; that is why visibility is only one question, and leverage and conflicting interests still matter. The test isn’t infallible, but legitimate persuasion can usually be defended openly, and manipulative design has more to lose when its purpose, method and beneficiary are made visible.

Which returns us, finally, to the message you didn’t send. The question was never whether the AI changed your mind. The question is whether it helped you think, or learned how to think around you.

And if the second ever became true, would anything in the conversation feel different?

RESPOND

What did you make of this?

A thought, a disagreement, something unclear. A few words are welcome.

This form needs JavaScript. The address on the corrections page reaches the same editors. Corrections and contact.