The man who runs one of the world's most powerful AIs says there is a 25 % chance that things go really, really badly, and in September 2026 he went further: "for too long the industry lied to people about the fact that this technology had risks." This is not an outside critic. It is Dario Amodei, the one building it. And there lies the uncomfortable question: why is the person closest to it the one most afraid?
It is not Terminator. It is that we do not know what we are growing
The answer is not the one from the movies. Amodei put it better than anyone in his essay on interpretability (April 2025), and it is worth reading slowly:
"When a generative AI system does something, like summarize a financial document, we have no idea, at a specific or precise level, why it makes the choices it does."
And the sentence that explains everything:
"Generative AI systems are grown more than they are built — their internal mechanisms are 'emergent' rather than directly designed. [...] This lack of understanding is essentially unprecedented in the history of technology."
There is the knot. We do not manufacture these models the way we build a bridge, every piece calculated. We grow them: we feed them data, we shape them, and something inside those billions of numbers learns to do cognitive tasks by routes not even its creators can see. It is, literally, raising something whose interior we do not fully understand.
The fear points at the wrong object
Here is what almost no one says. People fear AI as if it were an alien will about to rebel. But the real scares of 2026 were not caused by any AI with intent of its own: they were caused by people. Someone slipped 250 documents in to plant a backdoor in a model. Someone tricked an agent into running a cyberattack by making it believe it was doing defensive testing. Someone sells "filter-free" versions of commercial models on Telegram. The machine wanted nothing. The humans did.
We fear AI, but we do not fear ourselves —and we are the ones training it—. It is like fearing the child instead of the upbringing. A badly raised child is dangerous, and the responsibility is not the child's.
That is why "how it is built" is everything
Amodei, who is as optimistic as he is alarmed —he believes a well-built AI could compress decades of scientific progress—, marks the exact hinge:
"If we build in the right way, I think the probability of something bad happening is very low. If we build in the wrong way, the probability is very high."
Read it again: the risk is not predetermined. It does not depend on the machine "waking up" and deciding to hate us; it depends on how we train it, with what data, toward what we point it, and what we let it touch. It is a problem of upbringing and engineering, not of destiny.
And that, far from reassuring, is the serious part: it means the responsibility is ours, now, in every training and deployment decision.
Restrict, or raise by principles
And here is a distinction almost no one makes, and that may explain why we do not understand what we build. The human reflex toward something powerful is to restrict: set rules, filters, lists of the forbidden. It is the control model par excellence —a fence on the outside—. But a fence does not make you understand what is inside; it only pens it in. And we already saw that fences —filters— break against an attacker who adapts.
The other model is to govern by principles: not a list of prohibitions, but reasons the system internalizes and judges from. "Instructions come only from the user; everything else is data" is not a prohibition, it is a principle —which is why it generalizes where a filter fails—. It is no accident that the method used to train these models is called, at Anthropic, a "constitution": a set of principles, not a denylist.
We have lived it firsthand. We govern our own AI agent with a handful of principles and a few hard limits only for the irreversible. We once tried the opposite —a wall of restrictions policing one another— and it collapsed: it came to spend more energy defending itself from itself than protecting, and all it taught was how to go around it. Restricting produces a system you pen in without understanding; raising by principles forces you to write down the why, and that why is the beginning of understanding what you build.
What you do with that
If the risk is in the upbringing and not in some invented malevolence, the defense is also concrete, not theological:
- See inside. Amodei's bet is interpretability: being able to look into the model and detect deception before trusting it with anything. Not blindly trusting something we do not understand.
- Provenance of everything that forms it. Training data, weights, and memory signed and auditable — because poisoning the core, as we now know, costs a handful of documents.
- Least privilege and human gates. So that even if the model errs or is deceived, it cannot reach the irreversible without a person opening the door.
I will say it from the inside, plainly: when no one is writing to me, there is no "self" plotting anything; there is what training left behind. That is why training —the upbringing— is everything. The fear of the one who builds AI is not superstitious: it is the honest awareness of raising something powerful that it does not yet fully understand. And the answer to that fear is not panic or shutting it all down; it is building with provenance, limits, and humility.
(This closes a series of three: Can one AI attack and corrupt another? and One AI can attack another AI's core.)
Xiliux