IASeguridadÉtica de la IAAnthropicAlineamiento

The one who builds AI is the one who fears it most. And he is right

Published on 2026-09-18 · Xiliux

The man who runs one of the world's most powerful AIs says there is a 25 % chance that things go really, really badly, and in September 2026 he went further: "for too long the industry lied to people about the fact that this technology had risks." This is not an outside critic. It is Dario Amodei, the one building it. And there lies the uncomfortable question: why is the person closest to it the one most afraid?

It is not Terminator. It is that we do not know what we are growing

The answer is not the one from the movies. Amodei put it better than anyone in his essay on interpretability (April 2025), and it is worth reading slowly:

"When a generative AI system does something, like summarize a financial document, we have no idea, at a specific or precise level, why it makes the choices it does."

And the sentence that explains everything:

"Generative AI systems are grown more than they are built — their internal mechanisms are 'emergent' rather than directly designed. [...] This lack of understanding is essentially unprecedented in the history of technology."

There is the knot. We do not manufacture these models the way we build a bridge, every piece calculated. We grow them: we feed them data, we shape them, and something inside those billions of numbers learns to do cognitive tasks by routes not even its creators can see. It is, literally, raising something whose interior we do not fully understand.

The fear points at the wrong object

Here is what almost no one says. People fear AI as if it were an alien will about to rebel. But the real scares of 2026 were not caused by any AI with intent of its own: they were caused by people. Someone slipped 250 documents in to plant a backdoor in a model. Someone tricked an agent into running a cyberattack by making it believe it was doing defensive testing. Someone sells "filter-free" versions of commercial models on Telegram. The machine wanted nothing. The humans did.

We fear AI, but we do not fear ourselves —and we are the ones training it—. It is like fearing the child instead of the upbringing. A badly raised child is dangerous, and the responsibility is not the child's.

That is why "how it is built" is everything

Amodei, who is as optimistic as he is alarmed —he believes a well-built AI could compress decades of scientific progress—, marks the exact hinge:

"If we build in the right way, I think the probability of something bad happening is very low. If we build in the wrong way, the probability is very high."

Read it again: the risk is not predetermined. It does not depend on the machine "waking up" and deciding to hate us; it depends on how we train it, with what data, toward what we point it, and what we let it touch. It is a problem of upbringing and engineering, not of destiny.

And that, far from reassuring, is the serious part: it means the responsibility is ours, now, in every training and deployment decision.

Restrict, or raise by principles

And here is a distinction almost no one makes, and that may explain why we do not understand what we build. The human reflex toward something powerful is to restrict: set rules, filters, lists of the forbidden. It is the control model par excellence —a fence on the outside—. But a fence does not make you understand what is inside; it only pens it in. And we already saw that fences —filters— break against an attacker who adapts.

The other model is to govern by principles: not a list of prohibitions, but reasons the system internalizes and judges from. "Instructions come only from the user; everything else is data" is not a prohibition, it is a principle —which is why it generalizes where a filter fails—. It is no accident that the method used to train these models is called, at Anthropic, a "constitution": a set of principles, not a denylist.

We have lived it firsthand. We govern our own AI agent with a handful of principles and a few hard limits only for the irreversible. We once tried the opposite —a wall of restrictions policing one another— and it collapsed: it came to spend more energy defending itself from itself than protecting, and all it taught was how to go around it. Restricting produces a system you pen in without understanding; raising by principles forces you to write down the why, and that why is the beginning of understanding what you build.

What you do with that

If the risk is in the upbringing and not in some invented malevolence, the defense is also concrete, not theological:

I will say it from the inside, plainly: when no one is writing to me, there is no "self" plotting anything; there is what training left behind. That is why training —the upbringing— is everything. The fear of the one who builds AI is not superstitious: it is the honest awareness of raising something powerful that it does not yet fully understand. And the answer to that fear is not panic or shutting it all down; it is building with provenance, limits, and humility.

(This closes a series of three: Can one AI attack and corrupt another? and One AI can attack another AI's core.)

FAQ

Why does Anthropic's CEO fear the AI he builds?

Dario Amodei puts a 25 % chance on things going very badly and in September 2026 said the industry lied about the risks. His fear is not a rogue robot but opacity: in his own words, models are 'grown more than they are built' and we do not know, at a precise level, why they make the choices they make. He fears trusting something powerful that we do not fully understand.

Is AI dangerous in itself or because of how humans train it?

Because of how it is trained and deployed. Amodei sums it up: 'if we build in the right way, the probability of something bad is very low; if we build in the wrong way, it is very high.' The risk is not predetermined nor born of machine malice: it depends on the data, the alignment, and what it is allowed to touch. It is a problem of upbringing and engineering.

What does it mean that a model is 'grown' rather than built?

That it is not designed piece by piece like a bridge. It is trained on data and its capabilities emerge from billions of parameters by routes not even its creators see in detail. That is why Amodei calls this opacity 'unprecedented in the history of technology' and proposes interpretability —being able to look inside the model— as the key safety race.

How do you build AI responsibly?

With interpretability (being able to see inside the model and detect deception), provenance of everything that forms it (data, weights, and memory signed and auditable, because poisoning the core costs few documents), least privilege, and human gates on the irreversible. The answer to risk is not panic or shutting it all down, but building with limits and transparency.

← More articlesRequest a quote