Flying Blind: Unexplained Mysteries of AI and Why They Matter
At a recent family gathering, AI risk was a major topic of conversation. I repeated the central argument from many AI scientists: we are trusting a technology we created, but whose resulting capabilities we cannot accurately predict or always explain. That answer seemed baffling to many, so I attempted to elaborate, but my efforts were less than satisfying.
Mankind has been harnessing technology that we haven’t fully understood for centuries. Things like general anesthesia (ether), steam power (before thermodynamics), selective breeding, pharmacology (aspirin), and metallurgy (steel). Anesthesia is the most striking of these: we have been putting people under since 1846, and there is still no settled account of how it produces unconsciousness. We made it safe through monitoring and protocol, not theory. While this may be somewhat comforting, my instincts are that we may not be so lucky with AI, which is a moving target that may not be easy to safely nail down. What’s different about AI is the breadth of its capabilities and its ability to control real-world resources.
Across all my research and recent thinking, I’ve realized how important is for us all, as consumers of AI and world citizens, to understand more about the mysteries behind it. So, I’ve identified four fundamental mysteries around AI’s capabilities, which can help to structure the way we think about it and guide us on how to move forward.
Generalizing Knowledge
In machine learning, classical theory says that a model with far more internal parameters (adjustable settings) than training examples (data) would simply memorize what it was given but would perform poorly on any task it has not seen before. The easiest way to understand this theory is that a model with far more capacity than the data picks up random quirks and errors (noise) in the training data and begins to associate meaning with that noise, which impacts its ability to correctly interpret previously unseen data.
Modern AI models violate this theory badly. During training, a model’s performance on data it has never seen can sit flat (not improve) for a long stretch after it has already memorized the training set perfectly, and it appears to be doing nothing but memorizing. Then, with continued training on that same material, the model can abruptly gain the ability to generalize, enabling it to apply what it knows to different contexts and answer questions it was never explicitly trained on.
This delayed burst of generalization capability is referred to as grokking, a term from the famous Heinlein science fiction book Stranger in a Strange Land. In the book, it refers to consuming an experience so thoroughly that an observer merges with the observed. Grokking has been cleanly demonstrated on small, narrow problems but does not explain how large models generalize. It is an example of the mystery, not a solution to it.
Why a model that classical theory says should be memorizing rather than generalizing instead lands on solutions that generalize is the foundational gap in our understanding of how AI works. Our understanding of everything else sits on top of this.
Patterns That Are Hard to Understand
A trained AI model is organized as a series of neurons, which are simple computational units. It is tempting to assume each one holds a concept, but there are nowhere near enough of them: an AI needs to represent far more distinct concepts than it has neurons. So it packs those concepts across many neurons, storing each concept as a pattern rather than in a single location, and every neuron participates in many unrelated concepts. Researchers call this superposition.
This makes AI hard to understand, since the value of a single neuron tells you nothing about the underlying concept directly. And trying to change the model’s weights to correct one behavior may have unintended consequences, since unrelated concepts are mapped through the same neuron values. (Weights are the numbers that measure the strength of connections between neurons.) The truth is that you can read the value of every neuron and still not know what it computes. Researchers have made real progress pulling interpretable features back out of superposition, but the raw neuron values themselves remain opaque.
We Can’t Predict What AI Can Do
We can train an AI model on a given corpus of information, but we don’t know what exact capabilities it will have until after it is created and we can measure them. The list of abilities gets written after the fact.
There is one thing we can forecast well: how wrong the model will be on average at predicting text. That number falls along a smooth, fittable curve as we add data, parameters, and compute.
But abilities tend to show up abruptly, at apparent thresholds. Some researchers argue those thresholds are partly an artifact of pass/fail scoring, since switching to a finer-grained metric often smooths the jump out. Either way, nobody can say in advance which capability is enabled at a given training scale.
The worrisome thing to me is that capabilities may exist that we didn’t think to measure after the fact, hence the need for thoughtful and thorough testing of both AI models and the applications that leverage them.
In-Context Learning Is Unexplained
In AI, there are two distinct processes: training, where the AI learns, and inference, where it answers questions. During inference, we only partially understand how the AI learns from examples in the prompt and then answers questions based on them. The surprising part is that nothing in the model changes when this happens. Not one weight is updated. The model is identical before and after, and when the conversation ends, whatever it picked up is gone. And nobody knows how much of the effect is genuine rule inference versus merely selecting a behavior the model already had.
It’s why prompting is folklore instead of engineering. There is no theory to help guide the type and quantity of examples you need to include with your prompt before the AI will catch on and be capable of fulfilling your particular request.
Where Do We Go From Here?
You’ll remember the examples I gave in the intro: anesthesia, steam power, selective breeding. The historical parallels tell us something useful about what to do next with our imperfect understanding of AI. In every one of those cases, what made the technology safe to use was not eventual theory but disciplined measurement: boiler inspections, clinical trials, anesthesia monitoring standards, materials testing. Every one of those is a form of systematic testing, and each was invented precisely because the underlying mechanism was unknown.
What concerns me is that the measurement problem with AI is harder than any of those. A boiler has a handful of inputs, while an AI model has an unbounded input space and can behave differently depending on how a question is phrased. A patient under anesthesia has a small, known set of failure modes, and the silent one (intraoperative awareness) is exactly what anesthesia monitoring was built to catch. AI failures are frequently silent and can look like perfectly reasonable output. And a drug trial has a clear finish line: you know before you start exactly what result would count as success. With AI, you often have to discover what the system can do before you can even define what you are testing for.
None of this is an argument for waiting until the science catches up. It argues for treating testing as a permanent discipline rather than a temporary bridge to a theory that may never arrive. We still don’t know how anesthesia works. We just got very good at monitoring it.
I track how these mysteries play out in practice (what researchers are learning and what it means for testing AI systems) in the State of AI QA Newsletter. [Subscribe to get the next issue.]