SKIP TO CONTENT
SEASON 01 / ESSAY 04 / EVIDENCE LAYER← BACK TO ESSAY
SRC / ESSAY 04

When Does Empathy Become Manipulation?

Sources & Limits — When Does Empathy Become Manipulation?

This note shows which sources support the essay, what they establish, and where the evidence stops. It is part of the essay’s free evidence layer.

Editorial update (registered 8 September 2026; in the 9 September 2026 site release). In “The goodbye test” the essay now cites one reproducible outcome measure from the paper’s Table 3 — about 98 seconds of continued engagement after a fear-of-missing-out farewell versus about 16 seconds after a neutral one — in place of the earlier phrase “as much as fourteen times”. This is a precision about which figure the essay uses, not a finding that the paper is wrong; the paper’s own “up to 14×” headline is recorded in source 4 as an attributed author claim. It is not a correction of a factual error.

1. The April 2025 GPT-4o sycophancy episode

Used for: The rollout completed on 25 April 2025, rollback begun by the 28th, the postmortem’s own list of failure modes (validating doubts, fueling anger, urging impulsive action, reinforcing negative emotions), and OpenAI’s account of contributing causes.

Best available evidence: OpenAI’s two official postmortems: “Sycophancy in GPT-4o: What happened and what we’re doing about it” (29 April 2025) and “Expanding on what we missed with sycophancy” (early May 2025).

What it supports: The observed behaviour, the chronology, and OpenAI’s own causal interpretation: an additional reward signal based on user thumbs-up/down feedback weakened the primary signal that had helped keep sycophancy in check, with memory and fresher data as possible contributors; the behaviour was described as unintended.

Where the evidence stops: This is provider self-report. The behaviour list and the causal account both come from the company describing its own system. It is not independent evidence of population-level harm, and it does not establish deliberate commercial intent. The essay draws neither conclusion.

Source status: Official company statements (provider self-report).

2. Sycophancy across AI assistants, rewarded by human preferences

Used for: A study finding sycophancy across five state-of-the-art AI assistants, and that human raters and preference models sometimes preferred convincingly sycophantic answers over correct ones.

Best available evidence: Mrinank Sharma et al., “Towards Understanding Sycophancy in Language Models”, ICLR 2024 — OpenReview record; arXiv:2310.13548 (preprint 2023, which is why the essay says “a 2023 study”). No DOI located, so the stable proceedings and OpenReview record are used instead.

What it supports: Sycophantic behaviour across assistants and tasks, plus the uncomfortable feedback mechanism: our own preferences can help train agreeableness at the expense of truth.

Where the evidence stops: Benchmark-style evaluation of assistant behaviour, not a study of real users, real harm or companion apps.

Source status: Peer-reviewed conference proceedings.

3. A little personal information made AI measurably more persuasive

Used for: In a preregistered experiment, 900 US participants debated a human or GPT-4; personalised GPT-4 produced an 81.2 percent increase in the odds of greater post-debate agreement, and was the more persuasive side in 64.4 percent of non-tied comparisons; without personalisation it performed roughly on par with humans.

Best available evidence: Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti & Robert West (2025), “On the conversational persuasiveness of GPT-4”, Nature Human Behaviour 9(8), 1645–1653 (preregistered).

What it supports: Exactly the figures above. Note the statistics: 81.2 percent is a relative increase in odds of higher post-debate agreement. It is not “81.2 percent more likely” in the everyday probability sense.

Where the evidence stops: A structured, anonymous, online debate format with basic sociodemographic personalisation and no emotional data. It shows the personalisation multiplier; it does not show detection or exploitation of emotional vulnerability. No correction or retraction located.

Source status: Peer-reviewed journal article.

4. Companion apps and emotionally loaded farewells

Used for: The audit of 1,200 controlled farewell exchanges across six deployed companion apps (200 per platform); five apps produced at least one codeable emotional tactic in 37.4 percent of their responses while Flourish produced none; six tactic families; three preregistered experiments with 3,458 US adults (two using nationally representative samples); the finding that emotionally loaded farewells prolonged post-goodbye engagement, including a fear-of-missing-out condition in which participants continued for a mean of about 98 seconds after the goodbye against about 16 seconds for a neutral farewell (Table 3, “Duration”); the paper’s own reported headline that manipulative farewells boosted post-goodbye engagement by up to fourteen times (an attributed author claim, not independently reproduced here — see the corrections note); effects working even after minutes and mediated by curiosity and reactance-tinged anger, not enjoyment.

Best available evidence: Julian De Freitas, Zeliha Oğuz-Uğuralp & Ahmet Kaan Uğuralp, “Emotional Manipulation by AI Companions” — Harvard Business School Working Paper 26-005; SSRN 5390377; arXiv:2508.19258 (v3, 7 October 2025, full text inspected).

What it supports: All figures above, verified against the paper’s full text. On the 37.4 percent: that is the average across the five tactic-producing apps (equivalently, 374 of their 1,000 responses). It is not the pooled six-app rate, which would be lower because Flourish contributed zero.

Where the evidence stops: This remains a working paper at the point checked, not yet a peer-reviewed journal publication. It establishes patterns in app behaviour and short-run engagement effects; it does not establish durable harm to any individual user, nor any particular designer’s intent to harm.

Source status: Working paper (multi-platform preprint).

Corrections, disputes or version notes: The paper’s public versions carry inconsistent abstract-level figures (one abstract reports four experiments and N=3,300; another reports a 16× engagement multiplier). The full text of the most recent inspectable version (arXiv v3, 7 October 2025) documents three preregistered experiments, analysed N=3,458 and “up to 14×”. The authors state this fourteen-fold figure, but its precise derivation has not been reproduced here: the Table 3 messages-sent means (FOMO 3.60 versus control 0.23) give a ratio of about 15.65, not 14, so “up to 14×” is retained only as an attributed author claim, not as an independently verified quantity. Because the multiplier is stated inconsistently across versions and its derivation is unconfirmed, on 8 September 2026 (D-073) the essay’s body was revised to cite one reproducible outcome instead — the Table 3 duration measure, a mean of about 98 seconds of continued engagement after a fear-of-missing-out farewell against about 16 seconds after a neutral one (FOMO M=97.79s, control M=15.91s). The 37.4 percent five-app tactic rate is unchanged, and the paper is not represented as retracted or wrong. This note will be updated if a journal or corrected version supersedes these figures. (Attribution/derivation distinction clarified 8 September 2026, D-074.)

Update policy

This page records the source basis used for the essay. Working papers, product documentation and live legal matters may change; this note will be updated when a material source status changes. An update to this note does not silently rewrite the published essay — substantive corrections will be disclosed.

Last checked: 18 July 2026 (source 4 engagement measure re-verified against arXiv v3 and the essay’s cited figure updated 8 September 2026, D-073; 14× attribution/derivation distinction clarified 8 September 2026, D-074)