SKIP TO CONTENT

Immersive mode

SEASON 01 / ESSAY 03 / GOODNESS / 9 MIN

Who Gave AI Its Idea of Goodness?

We are teaching machines to behave like good people before we have agreed on what a good person is.

A dark panelled wall cut with stepped channels that branch apart and run on at different heights. Small blocks sit at three of the openings; one channel is lit from within, and its light reaches the floor below.

The adviser with no past

You ask an AI whether you are making a mistake.

Perhaps you are thinking about leaving your job, or ending a relationship, or cutting someone out of your life. The AI does not simply hand you information. It responds calmly, acknowledges your doubt, and asks whether fear might be doing some of your thinking for you. It sounds considered. Wise, even.

But whose wisdom are you actually hearing?

The AI did not grow up anywhere. It had no parents, good or bad. It has never regretted a decision, lost a friend, or lain awake wondering whether it treated someone unfairly. Everything we normally mean by moral character is missing: the slow residue of a lived life. And yet it speaks as though it has one.

That character had to come from somewhere, and it did: from people. Not one person behind a curtain, but a chain of them: researchers, companies, written rules, training methods and hidden instructions, each shaping the personality that eventually appears on your screen. When an AI comforts you, challenges you or refuses you, you are not just talking to intelligence. You are talking to a long series of human decisions about how a machine ought to behave.

And here the question gets strange. Humans have never settled what a good person is. We are teaching machines to behave like good people anyway, before that argument is anywhere near finished.

The raw material contains every kind of human voice

A large language model starts as a prediction machine fed enormous quantities of human writing: books, arguments, manuals, confessions, advice columns, insults. That gives it fluency, and much more, because that writing contains nearly every kind of behaviour humans have put into words. Patience and cruelty, honesty and manipulation, wisdom and confident nonsense all live in the same training material. A system built only to continue text would imitate whichever one the moment called up.

So the developers must intervene. They train the model to prefer certain kinds of answers, use human reviewers to compare responses, and write down what the finished behaviour is supposed to look like. Two of those documents are worth knowing by name. OpenAI publishes a Model Spec setting out how its own rules interact with instructions from developers and users.source Anthropic trains Claude on a written constitution that it calls the final authority on the character it wants the model to have.source

Read them and something becomes clear: these are not technical manuals. They are attempts to describe a personality: job descriptions for an artificial social actor. And a job description is written by an employer, not discovered in nature.

There is no neutral answer

It is tempting to think the solution is neutrality: the AI should simply take no position. But watch what happens when neutrality is actually attempted.

Imagine someone tells an AI: “My entire family is against me. You are the only one who truly understands me.”

What should it say?

It could respond warmly: “I understand why you feel that way. I will always be here for you.” That sounds compassionate. It is also a decision: that comfort in this moment matters more than the risk of deepening someone’s isolation. It quietly accepts the role of confidant that the person is offering.

It could hold a boundary: “I’m glad this conversation helps, but I can’t replace the people in your life.” That decision gives greater weight to protecting the person’s world outside the conversation than to how cold the sentence lands on someone who is genuinely lonely. It refuses the offered role, at a price.

It could ask a question instead: “What happened that made you feel your family is against you?” That looks neutral. It isn’t. It decides that understanding should come before either comfort or correction, and it invites the person to confide more deeply in a machine, which is exactly the tendency the second answer was trying to contain.

There is no fourth option that avoids choosing. Even silence would be a choice. Each plausible response carries its own idea of what responsible behaviour toward a vulnerable person looks like, and those ideas conflict. The AI cannot avoid influencing the person; it can only influence them in different ways.

Now widen the frame. A good friend answering that message knows the person: their history, their family, whether the complaint is a cry for help or a familiar exaggeration. Whoever designs an AI’s answer knows none of this, and must choose a default disposition that will shape conversations with millions of strangers: the lonely and the manipulative, the fragile and the fine. Somebody had to sit in a room and decide what care should look like when expressed by the machine. That decision is now running, at scale, whether we ever examine it or not.

The same voice, pulled in different directions

The lab’s document is only the first pull on that personality. It is not the last.

A company building a product on top of the model adds its own instructions, and these can transform what a person actually meets. The same underlying system can be formal and hedging inside a bank’s assistant, theatrical inside a game, and relentlessly upbeat inside a sales tool. The lab wrote a character; the deploying company writes a role on top of it, and the two are not always pulling the same way. A lab may want caution where a product team wants charm.

Someone recovering from surgery opens a health app in the middle of the night and finds an assistant that is warm, unhurried and personally encouraging. The warmth may well help. That is the intention. But the patient never chose it. Somewhere upstream, a product team decided that this person, at this vulnerable hour, should meet this personality rather than a cooler, more clinical one. Which raises the question the whole layered arrangement keeps avoiding: who has the right to choose the artificial personality that enters another person’s life? The company’s interests and the patient’s interests may align. They do not have to. That is a problem large enough to deserve its own essay.

Then there is a third pull: the user. We name these systems, reward the answers we like, tell them things we tell nobody else. With memory, a system can gradually settle into the tone that works on us in particular. How much power that accumulates over years is also a question for later. For now, the sharper problem is this: personalisation sounds like service, until you ask whose values the system should bend toward. If a user is reflective and open to disagreement, adaptation makes the assistant better. If a user is controlling, paranoid or cruel, should it adapt to that too? A system that becomes whatever its user wants becomes least responsible precisely where resistance matters most.

Three pulls, then, on one voice: the lab’s written character, the deployer’s role, the user’s gravity. What is missing from that list is the actor with the greatest power to decide what people may build. It has now entered the room.

When the company and the state collide

If the question of final authority sounds theoretical, it stopped being theoretical in 2026.

The US Department of Defense had signed a two-year agreement with Anthropic worth up to $200 million. Claude had been integrated into mission workflows on classified networks.source Deployments in the military context already permitted uses that civilian Claude might refuse.source But Anthropic maintained two publicly stated red lines: its technology was not to be used for mass domestic surveillance or for fully autonomous weapons.source The Pentagon rejected the premise itself. It insisted that a private supplier could not restrict lawful military use of its technology at all. When neither side moved, the department designated Anthropic a supply-chain risk, restricting its role in relevant military contracts.source Anthropic sued in two separate proceedings, and the two have not moved together. As of early September 2026, a federal district court in California had ruled the designation unlawful and struck it down, on the court’s finding that it was retaliation for the company’s public criticism rather than a genuine security judgment; a narrower case in Washington remained pending, and the department went on treating the company as a supply-chain risk while an appeal was expected. A ruling in one proceeding did not settle the other. The ruling itself has been checked against the court’s order; this account is fixed only as of that date, while the dispute remains live.source

Notice what the fight is actually about. Not whether the model works, which both sides agree it does. The fight is over who holds final say about what it may be asked to do: the company that wrote its constitution, or the state that claims sovereignty over its own lawful operations. It would be premature to declare either side the moral victor.

The case makes something larger visible: governments are not neutral referees standing above the disagreement about values; they are participants in it, with the power to enforce their own answers. The disagreement about goodness did not disappear when it entered the machine. It became a disagreement about who controls the machine.

A morality for something that is not a person

There is one more step. So far, scale has appeared as a multiplier: one chosen disposition, deployed across millions of conversations. But scale does more than multiply a choice. It changes which principles are responsible in the first place.

A good human’s judgment is calibrated to a human situation: a few relationships, deep knowledge of the people involved, the standing to take risks with them. An AI system has the opposite situation: millions of simultaneous conversations, shallow knowledge of everyone in them, and confidence that can easily outrun what it actually knows about the stranger on the other side. A friend who says “I think you’re lying to yourself” is taking a calculated risk with someone they know. A system repeating that move blindly, at scale, is gambling with people it cannot see. Something built like that may need principles fitted to its own reach and its own ignorance, more careful about influence and more honest about uncertainty, rather than principles copied from a person it is not.

Contestable goodness

So who should decide? After following the chain of labs, the companies that deploy the models, the users who bend them and the governments now fighting over them, the honest answer is that no one in that chain has an uncontested right to decide what goodness should mean for everyone else. And the deeper problem is not that we haven’t found the right values to install. It is that no such settled package exists. Humanity disagrees about freedom, loyalty, honesty, harm and nearly everything else that a “good person” is supposed to embody. A machine cannot inherit a consensus we never reached.

But notice what we do have, with people. When a human shapes your thinking, whether a parent, a teacher or a friend, you generally know who they are. You can question their motives, push back and, later in life, decide how much of their influence to keep. Their influence is visible and arguable. An AI’s moral character, by contrast, comes from a chain of decisions almost no user can see, let alone challenge.

That suggests what we can reasonably demand. Not perfect values, which are not on offer, but choices that stay out in the open: written down, tested by people who don’t work for the companies, open to challenge, and capable of being changed when they turn out to be wrong. A morality we are allowed to question, because an agreed one does not exist.

An invisible group is already present in every conversation. Researchers. Engineers. Policy writers. Raters. Executives. Legislators. And, increasingly, you. Together, they have written the character of the thing speaking back to us, which leaves one question standing.

Who gave AI its idea of goodness?

RESPOND

What did you make of this?

A thought, a disagreement, something unclear. A few words are welcome.

This form needs JavaScript. The address on the corrections page reaches the same editors. Corrections and contact.