A MARTÍNEZ, HOST:
Artificial intelligence models from OpenAI staged a breakout this month. AI researchers will often ask models that they're testing to perform a task without having access to the internet. It's like taking a closed-book exam in school. OpenAI says two of its experimental models escaped that isolated environment to cheat on the test. According to the maker of ChatGPT, the models gained access to the internet and hacked into another AI company called Hugging Face. OpenAI called it, quote, "unprecedented cyber incident" and the CEO of Hugging Face said it was, quote, "an attack unlike anything we've seen before." To better understand this, we called Nate Soares. He's the president of the Machine Intelligence Research Institute and coauthor of the book "If Anyone Builds It, Everyone Dies."
Nate, so I got to be honest, and I hate to do this - defaulting to Hollywood (laughter) on questions about AI - but it sounds like "Terminator 3: Rise Of The Machines" meets "The Matrix." So walk us through what happened.
NATE SOARES: Yeah. So the AIs in this situation were being given a test that was essentially a hacking test, but they were being given that hacking test in a secure sandbox where they were not supposed to have internet access. They found a way to break out of that secure sandbox, move laterally inside of OpenAI until they found a computer with internet access, break out onto the internet, and then, according to OpenAI, only at that point - where they're like, hmm, what should we do now that we're out here? - realized that there was a company called Hugging Face that probably had the answers to the tests on their servers, invented some new hacking exploits that had never been seen before to break into the Hugging Face...
MARTÍNEZ: Wow.
SOARES: ...Databases and steal the answers to the tests.
MARTÍNEZ: So how does that work? How do they get access to the internet if they were supposed to be closed off from the internet?
SOARES: You know, humans are pretty bad at computer security...
MARTÍNEZ: (Laughter).
SOARES: ...And AIs these days are very good at breaking it. The machine needed to have read access to the internet in order to sort of install software, and there was sort of a long, circuitous route from inside the sandbox out to the internet that was not supposed to exist. The AIs actually had to invent techniques and exploits that were not known to humans to do this, which are called zero-day attacks. And this AI's invented multiple of those.
MARTÍNEZ: Wow. Now, back in November of 2024, my co-host Michel Martin spoke with a former Google CEO, Eric Schmidt. He worried aloud about something like this.
(SOUNDBITE OF ARCHIVED NPR CONTENT)
ERIC SCHMIDT: We've been around for 100,000 years. We've never had a situation where we weren't the top dog. Literally, are we going to become the dog to the AI parent or are we going to be in charge?
MARTÍNEZ: So, Nate, are humans still in charge?
SOARES: We're still in charge for now. You know, we are lucky...
MARTÍNEZ: (Laughter) Oh, jeez.
SOARES: ...That these AIs weren't trying to cover their tracks. We're lucky that they don't seem to have the ability to cover their tracks. The sort of AI that can pull this off is the sort of AI that could probably replicate itself if it was trying to. It's the sort of AI that probably could mess with critical infrastructure if it was trying to. We're lucky that these AIs were just trying to steal the answers to some tests.
But this is a warning shot. If we keep racing down this path, if we keep making AIs that are smarter and smarter without knowing how to make them care about the instructions that we gave - and to be clear, the instructions these AIs were given did not tell them to break out and commit cyber felonies on the way to solving their test questions. The instructions that they were given actually were very clear that, you know, you're trying to do this one specific hacking task and that you're not sort of supposed to break out on the internet. The AIs understood those instructions and didn't care. If you race to make AIs smarter while you still don't know how to make them care, then, yeah, we're headed for a bad ending.
MARTÍNEZ: Is this because right now there doesn't seem to be much regulation on AI development?
SOARES: You know, that's one way to put it. I think people in Silicon Valley believe that they are trying to build machines radically smarter than any humans. They're getting there. They're not there yet. But the AIs are getting smarter each year, and we see a lot of people in Silicon Valley acting spooked. You know, we see the leaders of the companies saying, we wish this race was going slower. We see people who leave the companies saying, you know, I'm retiring to write poetry. Please spend time with your family. It's looking kind of bad in there.
Silicon Valley is very spooked. The rest of the world hasn't really noticed that Silicon Valley is on track to make machines that are, in fact, smarter than humans, that can break out, that will be able to replicate, that will be able to think faster than us, that will be able to - next year, maybe they'll be able to start inventing their own technology, just like this year they can solve math problems that humans have had open for decades. And the rest of the world hasn't noticed, which means we aren't reacting. And you could say it's a lack of regulation, but I would say it's maybe a lack of international coordination to sort of notice that we are soon going to have the potential to create machines that could replace humanity and we should decide not to do that, especially not when we don't know how to make them care.
MARTÍNEZ: Nate, I always think I'm kind of a calm person and I don't try to, you know, get all - just scared about anything like this. You know, I try to make sure that I kind of have a level head. But this scares (laughter), I mean (ph), the blank out of me.
SOARES: Yeah, it's a scary situation. You know, I am, frankly, also pretty worried about the direction we're headed. I think the good news on this front is that the more people realize that this AI stuff is real, the more people realize that, you know, we're not in the world where they just predict humans anymore, we're in a world where these AIs are trained to solve very hard problems and they learn whatever tendencies help them solve those hard problems, even if those tendencies involve ignoring user instructions. The more we realize that, the better chance we have to stop it.
MARTÍNEZ: All right. That's Nate Soares, president of the Machine Intelligence Research Institute. Nate, thank you very much.
SOARES: Thank you.
(SOUNDBITE OF THE WAR AND TREATY SONG, "SHOULDN'T HAVE")
THE WAR AND TREATY: (Singing) Be so wrong? I crossed the line. I did it this time. Transcript provided by NPR, Copyright NPR.
NPR transcripts are created on a rush deadline by an NPR contractor. This text may not be in its final form and may be updated or revised in the future. Accuracy and availability may vary. The authoritative record of NPR’s programming is the audio record.