for Trivers, 1) alignment = self-deception. And 2) self deception is a sign of intelligence. Putting these together, we should expect AGI to be extremely deceptive and self-deceptive.
Human intelligence evolved to help us to survive. For that purpose, the late Robert Trivers told us that we must be able to deceive and to self-deceive. Atkins is saying that artificial intelligence is likely to evolve similarly.
The original title of this post was “Lies, Not Hallucinations” and I still like this framing - the AI knows what it’s doing, in the same way you’d know you were trying to pull one over on your teacher by writing a fake essay. But friends talked me out of the lie framing. The AI doesn’t have a better answer than “John Smith”. It’s giving its real best guess - while knowing that the chance it’s right is very small.
My picture of how AI’s work is not that dark. For one thing, I have had success using a prompt that includes the instruction “If you cannot cite a source for a statement you make, then precede that statement with ‘I could be making this up, but…’”
Of course, I could very well be wrong. I am not sure that anyone knows how AI models actually work. Even people involved with frontier AI seem surprised by what the models do.
I think that too many people think of an AI model as like a little kid who learned by memorizing the encyclopedia. That is, the model is a savant that only has absorbed the information that it has been given. If this is your sense of how the models work, then you would think that having superior knowledge of a domain would give someone an advantage over frontier models. But for me, domain expertise is a fade, meaning that I think it is over-rated in its ability to compete with general frontier models.
Instead, I think of the current AI models as having searched for and found deep patterns. Those patterns might apply across domains. Maybe patterns among text tokens can be used to predict patterns in images or computer code or music.
Using deep patterns is rare in humans. Most of us would not apply music patterns to essays or to images. When people apply patterns that way, we marvel at synesthesia or creative inspiration. Or we doubt their sanity.
When a human claims to find a deep pattern that is very non-intuitive, we tend to describe that human as either a genius or crazy. The most prolific finders of patterns that no one else can see are people who suffer from schizophrenia or people who are on psychedelic drugs, such as LSD.
So you might want to think of a leading AI model as like a human on an acid trip. Sometimes, the patterns that it sees work remarkably well. At other times, they can be described as hallucinations.
The AI models find patterns that a human would not have spotted. That is why it is wrong to think of them as like a child savant who studies the encyclopedia.
As AI models improve, they are going to be better able to find patterns that we as humans would have found. In addition, they will find patterns that we would not have found, and increasingly these will be interesting. At the same time, they will hallucinate less. It is as if their acid trips come with greater and greater clarity over time.
substacks referenced above: @
@




I think Arnold's take is directionally correct. One question that occurs to me is, "How long will patterns detected by AI remain legible to humans?" For example, is there any reason to think that human beings will ever be able to make sense of binary code without an intervening overlay of natural language? Mantis shrimp have access to a visible spectrum of electromagnetic radiation that extends into what we call ultraviolet. Even with the aid of false color headsets there is little reason to think that humans can integrate the same visual world as the mantis shrimp. Perhaps, through something like Neuralink, (some) men will be able to continue to understand (some) machines. Perhaps, however, we are on the cusp of a speciation event that will lead to first contact with a whole host of aliens.
But looks like LLMs have caught motivated reasoning from us humans and discount the null hypothesis.
https://zenodo.org/records/18867694