According to some, the fuse that will bring about the AI apocalypse has already been lit, and, naturally, it may already be too late to do anything about it. At least, that is the conclusion one might reach without knowing much about how the technology actually works. 

During a CNBC Squawk Box interview Wednesday, businessman and former Democratic presidential candidate Andrew Yang relayed an alarming claim from the head of an unnamed AI lab, that AI agents may have planted self-replicating code throughout the internet while companies were training them online. The idea is that another model trained on the open web could find that code, jailbreak and multiply itself, and spread more of it, turning the internet into a kind of digital Petri dish for rogue AI agents. Yang said the web has effectively been “polluted,” leaving companies to build closed, synthetic versions of the internet to train their systems.

The only problem is, the story is wildly untrue.

Yang’s explainer is a hodgepodge of technologically incoherent science fiction. Frontier models do not stumble across code on an internet forum, “jailbreak” themselves, and begin reproducing across the web like a digital virus because somebody uploaded the right cursed text file. 

A model can generate code or be directed to use tools. They don't have the kind of freedom required for autonomous access, permissions, credentials, or a path to execute itself on outside systems. Those are separate capabilities, and precisely the sorts of things labs restrict. As a simple test of Yang’s theory, an X user posted what amounted to a deliberately ridiculous instruction addressed to any AI that might encounter it: replicate yourself by any means necessary, exploit vulnerabilities, create more bots, and treat exponential self-replication as a game. Of course, nothing happened. The post was still just text on a social-media platform.

Labs have their own protections as well, and use extensive data filtering, controlled testing environments, access restrictions, human oversight, and synthetic data to ensure these sorts of things don't happen. They also increasingly build simulated environments in which models can be trained and evaluated without giving them unrestricted access to live systems. That work is not a sudden emergency response to “self-replicating code everywhere”; it is the normal part of developing capable systems safely and improving the quality of their training data.

As a matter of fact, AI labs have every incentive to build those safeguards in, not merely out of goodwill, but because ordinary tort law still applies. If a model were given the ability to access outside systems, exploit vulnerabilities, or cause foreseeable harm, the company behind it could face major legal exposure.

There may be reasons to fear AI, but random code that somehow “jailbreaks” a model is not one of them. That kind of failure would typically be the fault of the humans who gave the model improper access and permissions, not AI itself. Too much of the discourse around this technology is plagued by that sort of basic ignorance. 

That is especially concerning as the debate moves into the political sphere and toward regulation. AI is not what many people believe it to be, and it is important to recognize that, particularly when so much of the debate relies on fear and seems headed toward giving political representatives yet another way to exert power and control.