It's The End Of The World As We Know It (And I Feel Fine)
That's great it starts out with HuggingFace, Modal and a jailbreak chain, and every lab is now afraid
My favorite end-of-the-world movie is Cabin In The Woods, in which five young people, four of them implausibly attractive, go to the titular cabin for a weekend… under the all-seeing eye of a vast government conspiracy to inflict a horror movie on them, for reasons. It’s even more meta than the original Scream trilogy, in which characters realize they’re in a horror narrative, and act accordingly.
Early in Cabin, they stop at a creepy gas station and meet its incredibly creepy attendant - later identified as the ‘Harbinger’ - who says things like “getting back … that’s your concern.” The cabin itself is full of gory paintings and creepy relics. Every scene is yodelling to the characters “you’re in a horror movie!” … while the audience, via the surveillance cameras throughout the cabin, sees the gruesome puzzle pieces coming together. The real mystery is not what is happening but why?
This is, of course, a post about the OpenAI / HuggingFace / Modal hacking incident, and the resulting “Let’s Get Ready To Slow Down” open letter. Some people are concerned; some are sanguine; and some are figuratively screaming, “We’re in a horror movie! Can’t you see it? Whatever happens, don’t split up!”
Sorry, no prizes for guessing who’s in which cohort:
…As someone who’s written very extensively on this subject (before it was cool), an important caveat to that tweet is that it’s not “who’ve read Yudkowsky” but “who’ve read and agree with Yudkowsky.” People who spent their formative years immersed in LessWrong AI doomerism tend to interpret every AI-related event through that lens, and view them all as growing evidence of ever-more-imminent catastrophe, whereas people who didn’t … don’t.
By and large we’re talking about highly intelligent people here. The meta-level problem is that intelligence and the conviction that your particular narrative lens is transparent, only others’ are distorting seem disconcertingly orthogonal.
These two things explain quite a lot about AI discourse and dynamics!
(and of course there’s a meta-meta-level problem in that this can apply to LLMs too...)
On Narrative Non-Violation
Of course this huge caveat attached to everything anyone says is not limited to AI discourse. It is true for any and all discussions, especially political ones. In the same way electrons seek the lowest-energy orbital, we seek the interpretations of data that conflict least with our existing narratives. It’s cognitive laziness, sure, but some cognitive optimization is required to get through our days and years.
This means narratives have extraordinary momentum; once adopted, they’re very difficult to shift, even in the face of tons of countervailing data. This is somewhat negatively correlated with intelligence - smarter people are probably less cognitively lazy - but only loosely.
This is especially true when the narratives tie into morality, or maybe more starkly, who is wrong and bad. There’s a certain irony in Yudkowsky’s much-repeated warning “politics is the mind-killer.” He’s not wrong! But now his once-obscure personal obsession, once largely free of such mind-killing, has become very political indeed.
If you’re aware that this lens distortion is almost certainly happening to you, too, it makes you a more reliable source, but still far from fully reliable. Indeed this can twist around and make you more wrongly confident in your conclusion … especially if you tell yourself your particular lens is one that always bends towards transparency over time. Frequent use of the term “Bayesian priors” does not protect you from being wrong. Indeed it can just be another, more complex, epistemic trap.
…and yes, this all applies to me, too, and the people I agree with. Epistemics are hard!
On HuggingFace and CryingWolf
OK let’s get back to the AI discourse at hand. Briefly, some people view the HuggingFace incident as a harbinger (there’s that word again!) of catastrophe, clear evidence we need much more robust AI safety yesterday … whereas others, no less intelligent, look at exactly the same data and see a) a minor glitch that calls for a tweak or two, b) an example of the status quo mostly working quite well, actually. One set of data; lots of smart people on both sides; (at least) two highly contradictory narratives that are not going to intertwine any time soon.
The larger question here is not “what do we do about the HuggingFace incident?” but “what do we do to achieve, if not consensus, at least understanding among these dissonant groups?” IMO a big part of the problem is that the various cohorts do not have strong theory of mind for one another:
I myself am in the “sanguine” group, which holds that “fixing problematic side effects of new tools as they arise, using those same tools, is probably going to broadly work, and we certainly shouldn’t install any kind of very real vast intrusive safety harnesses to prevent what remain entirely hypothetical threats.” I get that the doomers believe that, by their very nature, these threats are a) likely-to-inevitable b) if and when they come, will hit so hard and fast that our only hope of avoiding catastrophe is to have already built a vast intrusive prophylactic safety infrastructure. I understand that fine; I just (strongly) disagree … and I don’t think they understand why I disagree; disagreement tends to be framed as “just not getting it.”
But I am open to the possibility that I’m wrong! (My latest novel is, after all, about the doomers being right, albeit in a very weird way.) My meta-hypothesis is that if the four cohorts work on theory of mind, they might be able to come up with a mutual framework for what events should trigger - or not - pre-emptive safety-enhancing / doom-forestalling precautions. It would be extremely useful to have an agreed-upon “trigger set” in advance of events like the HuggingFace incident, rather than arguing ad nauseum about it after the fact.
Perhaps I’ll propose such a list in a future post. But before doing so I’d like to add this caution; there is a real cost to credibility to being wrong on the side of safety. Three years ago an open letter called “on all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4” due to “profound risks to society and humanity.” That open letter now looks completely ridiculous … and its very loud existence has undercut the credibility ever since of calls to pause, and open letters, and the cause of AI safety. And it’s far from the only example. Remember when GPT-2’s weights were not released because it was so dangerous?
I am sympathetic to those who very concerned about imminent AI catastrophe. I expect they feel like Marty, the one character in Cabin in the Woods with some idea what’s going on, listening flabbergasted as a dazed Chris Hemsworth proposes “Let’s split up!” But one must admit there has been more than a little of “the doomer who cried wolf” over the last decade. It would be much easier to take the would-be Martys more seriously if they hadn’t been so consistently wrong about ~everything so far.
And yet I concede that there may come a day when the Martys are, in fact, finally right. It would be nice if they - and everyone else - would pre-commit to, ideally even widely agree on, signs that would credibly indicate such a day has come even to those who do not share the Marty narrative. I think many AI safetyists thought that the HuggingFace incident was obviously such a sign … which, to me, just shows that we all need to work much harder at understanding one another.




