The Weak and the Strong
So, uh, this…
…probably happened:
That question was long Metaculus’s “flagship” AI question … and it has long been problematic. Its “weak AGI” criteria are odd and fascinating:
You can understand the reasoning, right? The presumption was that if a single system could perform those four quite different feats, its intelligence must therefore be generalized such that it is reasonable to call this a general intelligence.
This is, as Enrico Fermi might put it, not even wrong.
We don’t have to agree on what general means to agree that it isn’t the heights of the famous “jagged frontier” that make an intelligence general, but the depths of its lowest valleys. What’s possible now, what’s hypothetical in the near term, what’s speculative in the medium term … will be largely defined not by the highest or fastest-growing peaks of the jagged frontier, but by its low, slow laggards. That’s basically Amdahl’s Law.
Will the laggards accelerate? Can better RL turn canyons into peaks, or are some of them fundamental to LLM architecture? We don’t know. And the only way to learn is by doing. What could possibly go wrong?
How Exactly Do You Just Do Things?
There are basically two ways: the theoretical, and the iterative. Theoreticians valorize long contemplation, from which is birthed an elegant, closed-form, generally applicable solution. Iterators throw together a crude first hack, see how the real world treats it, and use those (painful) lessons to inform a second design, and then a third, and a fifth, and an eighth, until refined into something reliably good.
Theoreticians frequently make assumptions that turn out to be painfully naive in, or construct solutions that only work in alarmingly small subsets of, the real world. It turns out the real world is extremely hostile to elegant theories. The real world is hard, chaotic, unpredictable, and full of emergent properties that are not just unpredicted but unpredictable, irreducible in their complexity. Theoreticians implicitly postulate that everything can be solved in advance of actually testing it, if you think about it hard enough. This is often catastrophically wrong.
Iterators frequently trial-and-error their way to solutions that are brittle, gratuitously overbuilt, or both, which would be immediately obvious if they understood the theoretical principles. They suffer from path dependency; when you’re five iterations in, it’s very hard to throw those all out and try another direction from scratch. And they only solve the kinds of problems they’ve directly faced, with no real framework for imagining new categories. This is catastrophically, well, problematic.
As such both approaches lead to essentially the same problem: fragility.
Engineering is basically the study of combining both approaches to optimize the combination of cost, capability, and fragility. As a field grows more mature, as we learn from previous iterations and expand our theories accordingly, theory takes over. An ancient bridge might fall because its builders didn’t understand basic structural dynamics. The infamous Tacoma Narrows bridge collapsed because we failed to incorporate vibrations into that theory. Today, our theory is sufficiently broad, informed by so many iterative lessons, there’s really no excuse for a bridge collapsing like that again.
This is, of course, a post about AI alignment.
Aligned With Whom, Exactly?
…is the question one should always ask, ideally with an arched eyebrow. “Helpful, honest, and harmless” sounds good, except … the US Department of War does not want its frontier AI models to be harmless. We don’t want Astra to be helpful to someone synthesizing sarin gas in their basement with their mind’s eye fixed on a big city’s subway system. Honest is the closest of the three to universally aligned, but it’s easy to come up with scenarios where it’s at least arguably best for an AI to tell white lies (e.g. a Muse that uncovers evidence a surprise party is being planned.)
That said you can make a good case that alignment means AI should be helpful, honest, and harmless to humanity. Even this sure gets tricky when humans are pitted against one another, when you fire up WarClaude and/or ask how to overthrow your local tyrannical / “tyrannical” government; but there are centuries of consideration of ethical military action, I am tentatively willing to stipulate a reasonable definition of “aligned-to-humanity yet adversarial-to-some-humans AI” can exist.
But quite aside from what AI alignment means, if it is actually going to work - if alignment engineering is to become as reliable as bridge engineering - then:
like all engineering, it must be both theoretical and iterative. The notion that some theoretician(s) might retreat into a think tank and ‘solve alignment’ is absurd and risible. The entire concept of solving alignment is inevitably and calamitously fragile.
So is the notion that we can just trial-and-error our way.
Alignment will require a defense-in-depth (alignment-in-depth?) approach, not a single silver bullet. Many elements work together to hold up a bridge!
Given AI’s unpredictable and dynamic nature, alignment engineering will have to be antifragile, to borrow Nassim Taleb’s term, which is to say that strains and stresses will make it stronger, not weaker.
Much more speculatively: aspects of its depth and antifragility may resemble a social compact, based on incentives, more than structural engineering.
A Few Inescapable Conclusions, Sorry
If alignment engineering is practiced only inside secure containment facilities within a few frontier labs, it will inevitably fail. Models must be exposed to the real world, rapidly and frequently, for us to learn from what that exposure entails.
This means the many and varied OpenAI, Anthropic and other agent breakouts this year were good, actually. Air-gapping would have been bad, actually. Exposure to the real world beyond their sandboxes, however inadvertent, induced the agent swarms’ emergent properties, we learned about and then from - and, crucially, these emergent properties were still relatively harmless, because they had not yet evolved far enough from the previously-well-understood safety baselines to become truly pathological. If the sandboxes had worked, then those properties would still be sitting there, latent and unknown, waiting to combine and multiply and evolve further with further training!
Similarly, a billion users will inevitably stress-test models in ways far beyond, and far weirder, than any internal testing possibly can. Releases (with guardrails) are another crucially important form of real world-exposure. “The first approved product is rarely the safest or the best but when we fail to approve the first we don’t get the much better 5th.”
The inevitable conclusion: failing to expose and then release a frontier model to the real world before training the next one is inherently misaligning behavior. It is a fundamental failure of iteration, and iteration will be absolutely key to success.
I am surprised that people are surprised that the real world turned out to be more problematic than the theoreticians expected. The real world is always more problematic than the theoriticans expect. It was only because the models were exposed to the real world that we noticed this. and now the alignment engineers are frantically iterating to address the new failure mode.
Good! This is how AI progress should work, actually! In a tiered, controlled way, of course. Watchful exposure within the labs, then guardrailed releases. But definitely not in the form of “multiple model evolutions occurring secretly within frontier labs, hidden away from the world.” That is the least aligned of all possible futures.
Obviously we should be on alert for problems. But if “pacing the frontier” turns out to mean the frontier withdraws deeper and deeper inside the labs, then we have a real problem on our hands. The worst lesson we could take away from this is that alignment engineering henceforth must happen entirely inside frontier labs, and on unreleased models. We need the coal face of alignment engineering to be iterating against the real world at all times; only then can we do real alignment engineering.
…All of which is triply true, and even more painfully obvious, now that the models themselves are nascently aware of the distinction between eval-world and real-world.
As such the worst thing that could happen for AI alignment right now is for the labs to stop releasing models for general-purpose use. The second-worst thing would be for their release pace to fall behind their development pace. I realize this may sound paradoxical, even horrifying, especially to theoreticians! But it is an inescapable corollary of the fact that the real world is hard, cruel, and irreducibly complex, so alignment engineering uninformed by copious real-world usage - and (sometimes painful) experience - will inevitably be dangerously … maybe catastrophically … fragile.






