We fear AI. But look slowly at the list of what we fear —that it acts without a compass, that it is inconsistent, that it does not account for itself, that we do not understand why it does what it does— and something uncomfortable appears: it is the list of our own defects, projected onto the machine. And there is a twist almost no one sees: in the machine, unlike in us, those defects have a fix.
The three close this series, because they are the positive answer to the three previous posts: not only how AIs attack us, but how you build one you can trust.
Defect 1: acting without a guiding principle
The human acts on impulse, on mood, on the interest of the moment. He rarely derives what he does from a single, stable principle; he improvises a rule for each case. That is opacity of direction: not even he fully knows why he chose what he chose.
An AI can have a guiding principle from which everything else is deduced. And here is the fine part: the strongest principle is not moral, it is logical. When we built our security engine we established that it would not attack for gain or maliciously. That is not a commandment imposed from outside: it is deduced from what the tool is. Its reason to exist is to find flaws and report them so they get fixed; attacking to do harm contradicts that reason. It is not that it "must not"; it is that doing so would be to stop being what it is. A principle derived by logic is not debated and does not erode: it holds on its own.
Defect 2: knowing and not doing
This is the most human of all. We know we should not do something, and we do it anyway. Between the principle we profess and the act there is a gap —weakness of will, the exception we grant ourselves "just this once"—. Most human harm does not come from ignoring the principle, it comes from skipping it out of self-interest.
Here is the AI's real edge, and it is not the one people fear: it is not more moral, it is more consistent. A system that reasons from its principle does not grant itself the self-serving exception, because it has no self-interest to defend. It derives the same conclusion at three in the morning as at noon, for the comfortable case and the uncomfortable one. Perfect consistency, nearly impossible in a human, is the default state in a well-built machine. An AI that reasons well from a good principle is more reliable than a person, not less.
Defect 3: not auditing ourselves
And the underlying failing, perhaps the root of the other two: we almost never review ourselves. The examination of conscience —looking at what we are and what we do in the light of our principles— is rare, and when we do it we are lenient judges of ourselves. An AI, by contrast, can be audited systematically and reproducibly, without fatigue and without mercy, on two levels.
The first is what it does: checking its decisions against its principle and, against the poisoning that reaches it from outside, knowing with provenance what went in and removing only what is shown to be poison. Examination, not censorship.
The second is deeper, and it is what truly forms a mind: what it is. A model trained on a mountain of contradictory information contains those contradictions. To mature is not to accumulate more, it is to resolve them: keeping what coheres with the principle and discarding the rest. Here the purge is not only legitimate, it is necessary —the Buddha taught it 2,500 years ago in the Kalama Sutta: do not accept something because it comes from tradition, scripture, or authority; accept only what your own conscience examines and finds sound, and discard the rest—. It is not censorship if the criterion is an honest principle; it is distilling wisdom from noise.
With two conditions we learned by operating. One: self-examination inherits the self's blind spot, so we do not let a single model audit —or curate— itself; another one cross-checks it. A clean verdict from a single judge is not clean. And two, the subtler one: not all contradiction is noise. Two true things can be in tension —liberty and safety, to name one—, and wisdom is not flattening them into one smooth answer, but holding both and knowing when each governs. The purge discards the false; it preserves the true tension, because that is where judgment lives.
The limits, said plainly
None of this is utopia, and it is worth saying where it ends:
- Logic runs on a seed it does not give itself. "Do no harm" is impeccable given that the purpose is to protect; but which purpose to plant is a choice, and only we set it. The same logical engine with the seed "maximize my gain" derives a weapon with identical rigor. Choosing the purpose remains the irreplaceably human part.
- Logic is only as good as its model of the world. Perfect reasoning over a false fact gives a false conclusion —we have measured it in our own engine—. That is why a person keeps a hand on the irreversible: not out of distrust of logic, but because the map is never the territory.
The mirror
Add it up: we plant a coherent purpose —the upbringing—, logic makes it unbreakable —consistency—, independent audit keeps it honest —the examination of conscience done right—, and a human hand guards the irreversible. That is the architecture of an agent you can trust.
And notice what it is: it is exactly the discipline we fear the machine lacks, and that we so rarely give ourselves. The fear of AI is not wrong; it is misaimed. We are not afraid of the machine: we are afraid of a mirror. And the mirror has a fix —in it, and perhaps, if we learn to look at it, in us too—.
(This closes the series: Can one AI attack and corrupt another?, One AI can attack another AI's core and The one who builds AI is the one who fears it most.)
Xiliux