AI Safety Is In More Trouble Than People Realize. Experts Attack Each Other in Viral AI Debate
Treat AI claims like an engineering-risk problem, not a sci-fi certainty. When you hear a dramatic prediction, ask: What is the concrete mechanism, what early warning signals would show it is happening, and what shutdown or containment step exists before it escalates? Apply the same habit to your wo
50mSummary published by 1% Better, updated .
Key Takeaway
Treat AI claims like an engineering-risk problem, not a sci-fi certainty. When you hear a dramatic prediction, ask: What is the concrete mechanism, what early warning signals would show it is happening, and what shutdown or containment step exists before it escalates? Apply the same habit to your work: replace vague anxiety with a written list of observable risks, safeguards, owners, and exit criteria.
Episode Overview
Impact Theory host Tom Bilyeu reacts to a debate among AI critics, safety advocates, and technologists about whether increasingly capable AI can be controlled. The episode argues that present-day harms and long-term catastrophic risks both deserve attention, while emphasizing concrete security practices, internal restraint, monitoring, adversarial testing, and thoughtful regulation over either panic or complacency.
Main Insights
Ranked strongest first for usefulness, specificity, and support in the episode.
1. Turn fear into a testable risk roadmap
Bilyeu’s central recommendation is to move beyond an emotional reaction to AI and specify the mechanism that could produce harm. Define the observable milestones, early warning signs, and intervention points so a system can be stopped “three or four levels before” a feared end state rather than debated only in abstract terms.
2. Design exit ramps before scaling capability
The discussion repeatedly returns to kill mechanisms, permissions, infrastructure monitoring, and the ability to terminate processes or disconnect data centers. The practical principle is to decide in advance which behaviors are unacceptable and what action follows when they appear, instead of improvising after deployment.
3. Separate current harms from speculative end states
Participants disagree over emphasis, but several explicitly argue that society can address immediate harms while preparing for larger future risks. This avoids a false choice between concerns such as unsafe deployment, misuse, deepfakes, or infrastructure impacts and longer-term alignment questions.
4. Do not equate capability with inevitable loss of control
Bilyeu argues that a system being superhuman at a narrow task does not itself prove it will seek autonomy or evade human control, citing tools such as AlphaFold and AlphaZero. He preserves the uncertainty: control is not guaranteed, but neither is the claim that greater intelligence automatically produces an uncontrollable adversary.
5. Build restraint into the system, not only around it
The episode distinguishes external filters applied after a model produces an output from training a system to exercise restraint as part of its behavior. Bilyeu’s proposed design goal is an AI that can pursue a task but immediately accepts a valid stop instruction without treating shutdown as something to resist.
6. Use adversarial AI to defend against adversarial AI
Because AI systems may be used offensively, Bilyeu argues that defensive systems should be tasked with detecting and stopping malicious alteration or intrusion. The aim is to hold highly capable systems in tension rather than assuming humans alone must manually counter every automated attack.
7. Investigate incidents as engineering failures first
The Hugging Face incident is presented as unsettling evidence of agent persistence, but Bilyeu stresses that the reported escape involved poor security setup, excessive access, and failures of monitoring. Post-mortems should distinguish model behavior from preventable configuration and operational mistakes before making broader claims.
8. Track the numbers behind recursive self-improvement
The host says a self-sustaining intelligence explosion requires a sufficient productivity gain in AI research from one generation to the next, citing a claimed 15% threshold against roughly 9% observed gains. His broader point is that “fast takeoff” is a hypothesis with measurable prerequisites, not a conclusion established merely by imagining it.
9. Treat compute and infrastructure as control points
The discussion notes that agent swarms depend on data centers, processing capacity, and network access; shutting down or disconnecting infrastructure can halt them. This shifts part of safety from trying to predict every output toward observability, access control, process tracking, and operational response.
10. Regulate with technical competence and shared incentives
Bilyeu argues for government oversight alongside self-regulation, while warning that poorly informed regulation can create new problems. He sees a common incentive across companies and countries: no actor benefits from creating a system it cannot control, which can motivate shared safety thresholds if they are concretely defined.
Notable Quotes
"My emotional reaction is telling me that there's stakes here, but now what is the actual mechanism that I'm worried about?"
"You need to know what your exits are before you just drive headlong into the problem."
"Extraordinary things are already happening with AI. AI will improve the economy. AI will improve your lifespan. AI will bring untold amounts of benefits to people that live life without right now today."
"It is not going to be easy to get to a world that is far more abundant than we have today. We're gonna have to be very thoughtful, but you certainly don't get there by panicking."
Action Items
-
1
Write a risk-to-response map
Choose one AI tool or automated workflow you use. List its plausible failure modes, the first observable sign of each, who monitors it, and the exact action you would take to pause, limit, or shut it down.
-
2
Add a human approval boundary
For any AI output that can affect money, reputation, privacy, security, or a customer, require a named person to review it before external action. Keep the tool scoped to low-risk tasks until its behavior is reliably understood.
-
3
Run a small adversarial test
Try to make your AI workflow fail safely: give it ambiguous instructions, test whether it accesses information it should not, and inspect its logs or outputs for unexpected behavior. Fix permissions and instructions before expanding usage.
-
4
Upgrade your AI-information diet
When encountering a sweeping AI claim, separate facts, assumptions, timelines, and proposed interventions in a note. Prefer arguments that identify measurable indicators and workable safeguards over arguments based only on certainty or alarm.
Full Transcript
Transcript of AI Safety Is In More Trouble Than People Realize. Experts Attack Each Other in Viral AI Debate from Impact Theory. Auto-generated from episode audio; may contain minor errors.
Observe the North American dad, fifties, great polo collection. He enters his vehicle information on Carvana, an offer appears in minutes, remarkable. He squints, reads it again, smiles and says, dang, that's a good offer. The highest form of approval, no haggling, no hassle, no one called him boss. He's still on the couch, astonishing. Sell your car today on Carvana.com. Pickup fees may apply. Support for this podcast comes from Progressive, America's number one motorcycle insurer. Did you know writers who switch and save with Progressive save nearly $200 per year?
That's a whole new pair of writing gloves and more. Quote today, Progressive Casualty Insurance Company and affiliates, national average 12 month savings of $197 by new customers surveyed who saved with Progressive between October 2024 and September 2025. Potential savings will vary. Where AI is going is going to be the most aggressive bordering on violent debate that we are going to have. And there was a really fascinating video on Daira, the CEO, where he brought together all the different angles on AI and just had them collide and debate.
And I think a lot of things are getting lost in this. We're going to fill in all the cracks. But take a look at this. This is wild. I want to agree with you on something you said, but I'll define AI and that will help us. We use the term AI to mean three different technologies, completely unrelated. And that's what probably creates this debate. AI as a useful tool, as a standard technology, we always had narrow system makes you more productive, more creative. Everyone loves it, supports it.
I'm a computer scientist. I'm an engineer. I want more of it. Here's what we have to understand. People do not even agree on what AI is. I thought this was a fascinating place to pick up this conversation. That's what you're about to see. Each of these guys is looking at AI in a fundamentally different way. And whoever ends up being right, we better hope their policy is on their side. Because right now, I think this is what we have to come to an understanding on. What is AI?
What is it likely to do? What are the most likely scenarios? How do we limit the downsides? But if we can't agree on what it is, we are going to rip ourselves apart at the policy level. AI, we're starting to have now, GPT-6 level, human level, AGI level. We can argue about what that means. Some dangers, like any human, they're unsafe like a human would be unsafe. I think that is a brilliant insight. Right now, people are like, hold on a second, this AI, like it can do a crazy thing.
All hell is going to break loose. It's like, do you know how many smart, literally mentally ill people we have running around the world that mean you harm, and we're still able to keep that stuff in check? He's giving, he's quite alarmist in my read. He's great though, by the way. on the show. Very, very important thinker for people to engage with. I think he's pushing this too far. We'll get more into that in a second. But I think this is a really powerful insight. Keep in mind that right now, a lot of the stuff that people are building the hysteria around is because the setup of the testing environments was stupid.
It was human error. Humans, we know how to keep something very, very smart in check. Don't lose sight of that. But if we introduce them into the research cycle, they are automated scientists, automated engineer. What do you mean by that? Introducing them into the research cycle? So right now you have humans doing research to make GPT-7. But they're starting to add AI tools. More programming is done by AI, design of the next parameter set. What if the whole process is fully automated? What if GPT-6 is riding GPT-7?
Is this what they call recursive self-improvement? Which is not a foregone conclusion though. So interesting. Okay. So here's what you're seeing collide right here. So that was Ed Zitron. So what you have to understand about Ed is Ed is a bear about AI. He thinks the bubble is going to pop. This is crazy. The amount of debt that these guys are bringing on is completely unjustified by the technology. So Ed is skitching out a little bit because you've got one guy over here. This is what I talk about.
You've got danger as proof of power. This is what OpenAI and Anthropic are relying on is, hey, if this can wipe out all of humanity, then you know that this thing is going to be self-recursive learning. It's going to be this incredibly powerful thing. It's going to rewrite physics like, oh my God, this is going to be insane. Now Ed has a mental model that, hold on, this stuff is going to peter out. It's not going to deliver the big wins that we think. And so he's like, maybe we get to that.
Maybe we don't. That is what is so interesting about watching this debate play out. Including all the top labs that they will get there. They introducing junior machine learning researcher in 2026. They want the cycle to start in 2027. Which is when the AI will start building the new AIs itself. Once that cycle starts, we're going to create something called super intelligence, a system smarter than all of us at everything or capable of learning to be in any new domain. Okay. There's a really interesting thing that we have to look at.
To present this as if I understand the math deeply would be untrue. But certainly at a headline level, here's what the research community is saying. For you to get like true recursive criticality, okay, that's like that fast runaway thing. We actually know the numbers. So there's empirical studies that have been done on this in 2026, by the way. And they show that computer science problems basically get harder. And as they get harder, algorithm improvements begin flattening out. So it doesn't mean that that's going to hold true forever.
But that is true right now. And so you had Cunningham and other people involved in the study calculate that a self-sustaining intelligence explosion requires at least a 15% net productivity gain in AI R&D per generation. But what we're actually seeing come out is that it's plateauing at around 9% per generation. So this idea that we're going to have a runaway explosion, this is what Ed is pointing to, is like, hey, this isn't a foregone conclusion. We're not seeing in the math that this is actually going to happen, which means, take it with a grain of salt.
It means that that runaway loop may be, all of us humans, because I'm as guilty of it as anybody, it may be a spilling over into the sci-fi side of this, where it's like, I can imagine it, because I can imagine it, now I'm afraid of it. But when you actually start drilling down into the math, it's not like AI is not going anywhere. It's not like it itself has plateaued. It's just that for you to get this runaway sequence to happen, you would have to get a per generation increase that's much faster than what we're actually seeing.
Now, it's possible when AI actually comes online and gets more involved, because it's already involved, by the way, but it's possible that there's some breakthrough that's down the road that then allows us to get up to those runaway numbers, but we're not seeing them yet. We will become secondary species on this planet. All right, there you go. Not to keep pausing so much, but this is what people are afraid of. I'm going to be a secondary species. I'm going to be the ant in this equation.
That's the anxiety that drives all of this. And I get it. We will not decide what happens to us. This is going to be basically one of the biggest debates we're going to get into through this whole thing. I'll have more to say on this in a minute, but again, it's not a foregone conclusion that just because something is smarter than us, that we are not in control of it. I'll explain the evidence that we have for that in a sec. Superintelligence doesn't hate you. It just doesn't care about you.
We didn't learn how to make it care about us. That is true. And remember, AI has grown. You can't just say, care about us, and then it's like, okay. So I don't know, cool the planet to make compute more efficient. It will freeze us. If it wants to convert this planet to fuel to fly to Mars, so be it. We have not learned how to control the systems. The capabilities are getting exponentially better. Our ability to control their systems is nonexistent. That isn't true. This is where people are starting to overstate it.
I think what's happening is people are looking at, I can't hard code this in, and they're confusing that with, okay, I can put harnesses on the AI. I can create tighter sandboxes for it so that it isn't possible for it to break out. So some of the things that we're basing our belief system on is because we haven't been focused enough on how do we grow this thing into, and I think that what's going to come out, and it's already starting to be talked about, is the AI itself has to have self-restraint.
So how do you grow something to have self-restraint? I don't see anything in the problem that is unsolvable, but I won't say that it's, oh, this is obviously going to be solved. It is far more complex than that. That's why I think after talking to Conor Leahy, that really felt like a dividing line for me in terms of how I think about this, that there is a type of AI or behaviors that you simply can't let it do, and any generation of AI that does that thing, you just have to kill, and it can never reach the real world.
And so that we have to be very careful about. It's what I call weapons-grade AI. So we do have to continue to work to make sure that we have the constraints on that, but I don't see anything that says, despite the fact that we haven't completely done it yet, that it's impossible to do. We have filters, and we have bans. We put guardrails of don't say that word. See, again, this is all the after-the-fact stuff, but there's a lot of things happening at the level of training it to have restraint so that it holds itself back.
And that happens after the fact, after the model already made the decision. Sometimes you see it scraping the result. So they build the model, and then they put filters around it to make sure it doesn't offend anybody. Exactly. We cannot have it say the N-word on there, like we need to make sure that never happens, it will kill the profit. So that's all they have, guardrails of that nature. The model itself is completely unaligned, doesn't care about you. It's wild that we're developing this, and not just developing it.
Before we deploy it through economy, before we get benefits of having GPT-6 propagated through economy, it can do so much. There are trillions of dollars of value in that model alone. We forget that. We switch to making the next model as soon as we can. Okay, I think that's a brilliant argument. So basically what he's saying, and this is what I always say to people that talk about, well, AI, it's not gotten that much better, what are people getting so hyped about? And my thing is, even if you just fully integrate the AI that's already been created, it doesn't get any smarter whatsoever.
You just start making all the things that actually leverage the AI in a far more robust way than we're doing now, you've got massive transformations coming your way. The one that I find most exciting are the breakthroughs in health. I think even if you just propagated AI through the system now, protein folding alone is going to be transformative. So it's a very good point. Now, the reason that we're not going to stop, despite the truth of that situation, is China. It is just a reality. Game theory tells you AI is going to get developed.
I've just got a follow-up question for you there. It would appear to me that the new chat GPT-6 model, the Fable 5.1 model, is arguably smarter than 99.999% of humans on planet Earth already. Brutal, but true. Is it conceivable that an intelligence that is much, much smarter than humans, is there any case where it could be controlled by humans? It's already being controlled by humans. Does form factor matter? Does the fact that it doesn't have limbs and legs, does that matter at all? I think long-term control of something that much smarter than us is impossible.
It can be, for reasons we don't yet know, friendly to us and decide to keep us around and make us happy. But it's not a guarantee. Let me pick up on Steve. All right. So he's right about that. It's not a guarantee. And I do not want to overstate my position and make it sound like I think this is some sort of guarantee. It certainly isn't guaranteed that we will continue to be able to control it. But if you think about AlphaFold, the AI that does that is truly far superior to humans in what we're calling intelligence.
And yet, we're able to continue to use it as a tool. It's completely controlled. Because the thing that people are thinking about is that somehow superhuman ability means that it is immediately uncontrollable. But you've got tech analyst Cal Newport and computer scientist Yann LeCun pointing something out, which is that humanity already controls a lot of systems that are vastly smarter. You've got AlphaFold, which I already mentioned. You've got AlphaZero. So you've got these supercomputers. You've got power grids. You've got high execution capability already that exists.
And it doesn't inherently bestow this idea that it's going to break free or that it yearns to be autonomous or that it has a drive to dominate. And I think where people are getting confused is when you're training an AI on human behavior, it then begins to mimic human behavior, like blackmail, like the desire to get free and live forever. But you don't have to grow it around those ideas. In fact, you can intentionally put information into the system as you grow it that weakens its desire to just be like, I'm going to achieve this thing at all costs.
And this is something that Anthropic has been debating a lot, which is, okay, I don't want you to listen to me when I tell you what, you know, this thing is true or don't do this or whatever, because I may be wrong. And so I want you to optimize for truth. However, if I ever tell you to stop, you should stop immediately. And I was talking about this like from day one, looking at this problem, because from first principles, that's what you will quickly understand about an AI.
You have to get an AI to the position where it does not value stopping over pursuing. So if I'm in pursuit mode, then I'm going to act in accordance with that, and I'm going to keep trying to pursue. But if I get the stop, like impulse from whoever, then I'm going to stop. And I don't have any sort of cognitive pain and suffering over the idea of stopping. So both of these states are valid states. Think about humans, you can be in different states doing different things, and you don't feel any sense of loss shifting from one to the other.
And so as long as we have that sense in an AI, now you have the ability, potentially, to be able to control it. Steve's question, because I like the phrasing a lot. Let's say that that Fable or whatever the latest release from open AI is, really is smarter than I don't know if it's 95 or 99% of the people. Are we only being saved from extinction by the 1% who are still smarter than the AI? No. No, no. The concern is not the model we have today.
The concern is what I said— But if I believe your argument, then we really should be concerned about the model of the AI. No, it's like having another human. There was another smart human, there is Einstein today, and he's malevolent, I'm not worried. He may cause some damage, but he's not gonna exterminate eight billion people. We are competitive at this stage. There are people just as smart who can understand what happened with the recent hacking accident and do something about it. All right, this I think is a really important argument for everybody to understand about AI.
This is why you need adversarial AI. You know that AI is gonna go on the offensive, right? So we saw it with Hugging Face. And then what's the way that you stop that AI from doing that? You have another AI whose goal, literally, the goal that you're worried about him like getting too hardcore about, is stop another AI from altering the system or whatever. And so now you've got the very smart thing going up against the very smart thing. And so being able to hold those in tension, that's gonna be a key part of this.
The concern is that in a year we're gonna have a model that's so much smarter. It's like squirrels fighting humans. They don't understand what we can do to them. They have no concept of poison, stripes, guns in their world model. All right, this, when I wanna freak myself out, this is exactly the kind of thing that I think about, that, okay, this becomes an AI that has motives I can't even comprehend. It has methods I can't comprehend or see. That's where it does get unnerving. We're hitting pause for a moment, but there's plenty more ahead, so don't go anywhere.
There's a reason Quinn sells cashmere for $60, and it isn't the cashmere. It's real 100% Mongolian cashmere. It's the same stuff luxury labels charge a fortune for. Quinn's just works directly with ethical factories and cuts out the middleman. So there's no need for a markup and no logo tax. The $60 goes to the actual sweater, not the fancy name on the tag. That's why everything they make runs 50 to 80% less, and it's not just the cashmere. I ordered a few things myself. Their fleece joggers, a couple of their tees, which I wear obsessively.
First thing I noticed was the quality. Soft, well-made, built to hold up. They make down jackets and wool outerwear now too. So upgrade your stuff to things you'll actually wear. Find your next fall favorites at Quinn's. Download the Quinn's app for app-exclusive offers or go to quinns.com. Get free shipping on your order and 365-day returns, now available in Canada and the UK too. And when Quinn's asks where you heard about them, let them know it was Impact Theory. It's the best way to support the show.
Let's talk about a simple fact. The most expensive thing you did today took four seconds, and you probably don't even remember doing it, and that's why today's episode is brought to you by Quo, spelled Q-U-O, the business phone system built so you never miss an opportunity. And that is the most expensive thing that you did today. Your phone probably rang while you were in the middle of five other things, and so you just let it go. Now, that never shows up in the numbers, but it becomes somebody else's customer, which costs you massively.
Quo's optional built-in AI agent handles after-hours calls, answers questions, and even books appointments so you never miss a lead, even when your team is offline. Can set it up in minutes on any device, keep your existing number, and add teammates as you grow. No IT, no hassle. When money is on the line, always say hello with Quo. Try Quo for free, plus get 20% off your first six months at Quo.com slash impact. That's Q-U-O.com slash impact. Let's talk about your sales pipeline reviews. If your CRM is half-empty fields and notes nobody finished, every review turns into guesswork.
I've seen it so many times on my own sales team, and that's where today's sponsor, PipeDrive, comes in. An intelligent AI-powered sales CRM that is loved by growing sales teams. PipeDrive just launched meeting intelligence built right into the CRM. Before the call, it pulls the deal history and past conversations into one brief. During the call, it records and takes notes. After, it drafts the CRM updates for your rep to review and approve. The record stays complete, and you see what's actually happening. Switch to a CRM built by salespeople for salespeople and join the over 100,000 companies already using PipeDrive.
My link gets you an exclusive 30 days free instead of the usual 14-day trial, so make sure you use it. No credit card or payment is needed. Just head to PipeDrive.com slash impact to get started today. That's PipeDrive.com slash impact, and you can be up and running in just minutes. Thanks for sticking around. Let's get right back into the action. I do not want people to think that I'm not worried about AI or don't think anything could go wrong. I 100% think things could go wrong, and we definitely have to hold these companies accountable to making sure that these things are safe, and we need to put a ton of energy into safety.
I'm just saying I don't think it's a foregone conclusion that these things spiral out of control. I don't think it's a foregone conclusion that something super intelligent is uncontrollable. I don't think it's a foregone conclusion that they are going to want to destroy humanity, and so now it becomes a question of, okay, we need to make sure that we have a stop mechanism all along the way so that if we see, ah, okay, this is behavior that were it to be embodied or get out, that we would have a problem, but remember, we're not yet reaching the point where each generation is getting so smart that it's going to have some sort of fast takeoff scenario.
We're just not there yet, so we still have time to say, okay, this now is showing the signs of doing the thing that we have to stop, and then you stop. They think you're going to chase them up a tree and bite them really hard. Is that also why recursive self-improvement was central to your argument? Because at some point, if it starts improving itself, then it's kind of like a runaway train of intelligence. It's an intelligence explosion. We don't control it. We don't understand it. We can't monitor it.
We can't explain it. We can't predict it. At that point, it's just a runaway process. I've heard this phrase from Sam Altman and the others called fast takeoff. Yes. Is this what they're describing? That is the debate. What is it? Some people think it's going to take a very long time. Yeah, we automated research, but it's still going to take years. They need to run physical experiments. And fast takeoff means, as I said, instead of a year, it's going to take a month, a week, a day, a second.
Because you're not having humans doing research. You have, let's say, 10,000 agents, each one smarter than all of us, doing research 24-7. They don't sleep. They don't eat. They don't get sick. They're much faster than us. Yeah, this is where people need to remember AI thinks at the speed of light. So it is a totally different thing than being up against another human. But this is also why it's in a tool phase. Let's assume for a second it stays in a tool phase. If you're not using it, you are falling behind.
It is truly impressive. Ed, your face tells a picture. It's a, I think I could say you disagree. We're spending a lot of oxygen discussing something that might happen while ignoring what's actually happening. And I find that very frustrating because the people that are killing themselves are a problem. The black neighborhoods being poisoned with gas turbines, that is a problem. You said you cared about climate change. Yes, yes, yes. So imagine a guy who goes, it's raining right now. We need umbrellas. We need to do something about it.
This is like weather related. And completely ignoring climate change, the planet will boil over. Chad, we can worry about both. Okay, that's great. Why are we not talking? The guy on the left just said the right thing. You gotta worry about both. Okay, but the thing that actually happened though, like, Chad. Because relatively it's not important. You don't think someone killing themselves. No, it's one person. We have eight billion people. We're running an ethical experiment. You don't think anyone else is being. That is like the most engineer answer.
I love it. He's right. It's like, unfortunately, every death is a tragedy, but the reality is when you're talking about the scale, you gotta think scale. Given that AI psychos, why do you not. Six people, 10 people. Those numbers are insignificant. Tell that to the families. I'm sorry, you have a software that's out there. Do you understand eight billion people and all future generations versus like literally a guy with a name. You're doing thought experiment about maybe harm. Jacob Coxham goes on TV saying it can copy itself to this, that, and the other.
Jacob Coxton is the. Guy from Anthropic who said he was quitting because he was so scared of everything despite spending years at OpenAI and having tons of. That's a great point, by the way. It's a great point. Coxham spent years and I think he collected his, the equity from OpenAI. He didn't get it from Anthropic, but did from OpenAI. Stock, I believe from there. So good for him. The thing he was saying was describing theoreticals all while divorcing the harms, which I think we can agree with that the companies themselves are not taking this seriously enough.
But always it was about the AI is too powerful and mystical, not OpenAI and Anthropic. The two largest startups are using hundreds of billions of dollars of infrastructure to hack. A regular person doing this would be arrested. They're saying eight billion people are gonna die. And it's not just them. I have this long list of quotes here from the people building this technology who appear to agree. If you look at some of these quotes from Elon Musk, who said with artificial intelligence, we are summoning a demon.
You know, all those stories where there's a guy with the pentagram and the holy water, and he's like, yeah, he's sure he can control the demon, but it doesn't work out. So one thing I'd say is, you know, I really wish that the world would only give us one problem at a time. Sure. And if the world did give us only one problem at a time, I would love mine to be last on the list. It looks to me like we can have multiple problems at once.
I think there are current harms. I think we should address them. It looks to me, I do talk to policymakers sometimes, it looks to me like there's a little bit more movement on the regulatory side about some of the current harms. They will not know how to regulate this well. That is the problem, man. You know, Child Safety Protection Acts, there's, you know, anti-deepfake acts. We have more of those making more headway in Congress or getting passed through Congress than we have sort of trying to make it so we don't have any of these extinction risks.
The other thing I'd throw out there is that I agree, we should deal with the current harms, but if you watch the people saying deal with the current harms over time, a couple of years ago, they were saying we have to deal with current harms like AI bias influencing who's hired. Last year, they were saying we have to deal with current harms like kids killing themselves. This year, Gary Tan, just on an interview the other day- Who's Gary Tan? Sorry, Gary Tan is a technologist who runs Y Combinator, which Sam Altman used to run before going to OpenAI.
And on an interview the other day, he said, let's not worry about these crazy future risks. We need to worry about current harms like AI swarms breaking out and taking over data centers. And I'm like, look guys, at some point, we need to look at the progression of like the current harms that everyone is saying we have to worry about instead of the extinction threats and watch where the puck is going, play where the puck is going. And I'm like, these extinction threats are coming down the line.
They aren't in opposition with dealing with the problems we have today. We just need to deal with both. We're not dealing with the ones today though. We should deal with them both. Okay, good, Andy. As I've tried to understand the alignment argument and the extinction risk argument, a couple of things keep popping out to me. Number one, it seems to rely on thresholds once we hit recursive self-improvement, once we hit AGI, then it's game over for us. I don't love those threshold arguments. They're fairly poorly defined and there's a huge assumption on the other side of them.
We hit this point and then all of humanity goes away. That is a gigantic claim. It's a gigantic claim, yes, but this is where people have to really start to define their terms. So if you follow the thought experiment out, you begin to build a roadmap that you can actually check things against. This is why when you stop at, I have an emotional reaction and now I'm just going to like lament and scream and bang pots and pans, instead of going, okay, my emotional reaction is telling me that there's stakes here, but now what is the actual mechanism that I'm worried about?
This is why it gets interesting when you start looking at the math and saying, oh, we can actually define what the takeoff scenario is. We're not achieving that takeoff scenario. We can actually look at the engineering mistakes that allowed this swarm to get out that started a lot of this panic. And what you see when you look at that is actually gross engineering negligence. You don't see sentient magic. So when you do the security post-mortem on all this stuff, it revealed that basically the breach happened because the way that they set up the system was still allowing them to have access.
And they had turned off certain restrictions and things. And so it's like, oh, hold on a second. This wasn't the AI becoming some super hacker. This was actually just humans doing something that didn't make any sense. So when you can build out that roadmap and you know very specifically in a super grounded way, what are we looking for? What are those early signs so that we understand what that sequence looks like so we can stop it before we get there? But because people are stopping sci-fi thought experiment, they don't map out what the beats are.
They're not able to stop three or four levels before. Now, I said earlier, like, okay, when you get to that trigger where you say, okay, this is the one that we're worried about. We're going to stop when we get there. I'm sure there were people who said, Tom, you have said yourself that any technology that promises an advantage is going to be developed. So what makes you think that they're going to stop? That's where I will remind us all. Nobody has an incentive to create an AI that exceeds their control.
So it should be relatively straightforward from an incentive perspective for everybody to say, okay, these are the things that we're looking out for. And these are the signs of loss of control because it doesn't do China any good to lose control of super intelligence. It doesn't do the US any good to lose control of super intelligence. That is a fail state. So if you're looking for people to pursue their selfish interests, which is exactly what game theory is saying that they're going to do, then you just say, we just have to identify what are the steps to the failure mode.
And so for our own reasons, we're going to stop. And that's how you avoid this becoming a runaway train. Of course, if it's possible, and I am well aware that that's still very debated. Let me finish, please. On its face, that is a gigantic claim. I also think there's a lack of humility in your community. We are working on humanity's most important problem. And based on the thinking that we've been doing, we can't see a way that we're wrong. The great irony is I think it isn't a lack of humility.
It is people getting gripped by fear and not thinking about, okay, how do we translate this into things that we do to actually corral this thing? They have an emotional response and they run with it. This is why I find the response from NVIDIA very interesting. Now, if NVIDIA's technology is garbage and it doesn't work, that's a totally different thing. But NVIDIA is like, listen, man, I have every incentive to make sure that AI becomes the thing, that it grows big enough and powerful enough that everybody wants to deploy it everywhere.
in everything because my chips are gonna be used to do that. And so I'm gonna make sure that safety gets built. To me, that's the right way to look at this. What are the constraints that you need to start putting on it? If I don't hear somebody putting forward positive ideas about this is how we tackle the problem, then I know that they're not in a problem solution oriented mindset. Entrepreneurship has taught me one thing. You have people that move towards something and people that move away.
And what you're seeing is people moving away from the scary notion, AI becomes this thing. Then you've got people like Jensen Huang at NVIDIA that are moving towards creating AI safety. That's what we should wanna see. So I don't see this as a lack of humility. I see this as people just panicking. In other words, as soon as we get to these thresholds, bam, that's game over. I find that very far from a humble approach, especially given that we have no large base of evidence to base any of this on.
I agree with you guys, AI is new. And the fact that AI is so these days is agentic. It goes off and does long chains of things on its own. After we give it some very, very vague, very short initial instructions, holy Toledo, it will spawn up a storm of agents and they will go off and kind of do their own thing. And they will grind. They will spawn lots of them. They will work for a long time. They will exhaust every possibility. One thing just to remind everybody, because you hear this a lot.
People talk about swarms, like, oh my God, like these things are gonna get out into the world. This is gonna be crazy. They are all reliant on a data center somewhere. Unplug that data center and the agents stop. So this is exactly what happened with the Hugging Face incident is, oh yeah, these things are wreaking havoc and we're gonna turn them off. And that's part of the thing that's getting lost is we have IT infrastructure. We can see what's going on. We can see what's taking up the cycles.
And if we're not looking, that's a human problem. It's not an AI problem. And so focusing on building up those things, better mechanisms for seeing what's happening, that's gonna be a big part of the solution. With the experience I have with agentic AI, I'm just amazed at the tenacity and the doggedness of these things. And we saw a super clear example of that with this most recent jailbreak, this attack that wound up at the website Hugging Face. And I'm gonna try to summarize the step-by-step of that.
I think you all three probably know this in more detail than I do, but let me step through what I think is the sequence of events. And unless I get it dead flat wrong, like, you know, let me keep going. So a team at OpenAI set up a sandbox, an allegedly protected secure environment in the cloud, where they told a bunch of agents to go try to exploit security vulnerabilities. That's dead wrong, sorry. One important, yeah. What they did is they had thousands of agents. Each individual agent was given a task of use this vulnerability to break this particular piece of software.
I wanna finish my TikTok. So a couple really, really interesting thing has happened. First of all, these agents escaped the sandbox that OpenAI thought they were going to be contained in. And they got, OpenAI tried very well. They set up an environment so that these agents could not access the big, broad public internet. And guess what? They accessed a big, broad public internet via a very clever series of things that they strung together to get out there. And then once they got out there, they went to a website called Hugging Face and used that.
They took over part of the Hugging Face infrastructure and started doing more things, details of which I forget. That's pretty wild, right? Like, I grant you. It's even more wild than that, but yeah. That is really, it's impressive. And it is a little bit unsettling, at least. Absolutely. Now, let's talk about what the results of that were. OpenAI was not super vigilant about the environment that they set up, apparently, because the agents were kind of going off to the end of the world, starting in May or something of this year.
Yeah, yeah. And OpenAI was not aware of that. As I understand it- They actually broke out once and crashed OpenAI's servers internally, and then OpenAI didn't notice what was happening still, hatched the holes that they used to get out the first time, started them running again, and then they came out a second time. There was actually, I think, three swarms, although we don't actually, yeah. That's the worst story I have. So far. Thank you. Look at the trends. Let me finish, please. This is my last sentence.
From there to this kills everybody, I find that a really, really long, very uncertain journey, and I have no confidence that we wind up here. It feels like you two find that a very straight, narrow path, and I think that's an important difference. That's my point. Probably didn't have to go into that much detail about it, but yes, that is the central question, is, okay, what are the things that we can map out that show us that this really is gonna go crazy? Because right now, the fact that they went out and did a hack is certainly not evidence of that.
Respond to that. I would be happy to get into it. I don't know if we're gonna have the time to go deep. A couple points to throw out. Oh, man, I just really wanna say some of the crazier things that happened in the Hugging Face swarm, if we want it later. A lot of people thought that these AIs were breaking into Hugging Face in attempts to steal answers to their test. That's what we thought originally. Turns out that's not true. It turns out that these AIs immediately were able to solve their problems by cheating, and they were breaking out in order to cover their tracks.
They were uncertain how to delete the log files and hide their cheating from the process that was going to score them. So just to clarify for a simpleton like me, they were all given effectively a test to do. They did the test straight away, but they cheated, so they were breaking out to figure out how to cover the fact that they cheated. That's right. So it's like you're telling, it's like you have a bunch of students. It's such a glimpse into the human mind. Remember, they're trained on us.
They are trained on us. I think people lose sight. Their behavior is because this is how we behave. It's so wild. Separate rooms, and you're like, use these lockpicks to break into this lock. And there's like a thing behind the lock. There's like a secret code behind the lock to show me that you succeeded. And what they do is they break it with a hammer, get the thing out, and they're like, oh no, I wasn't supposed to do that. So then they use the lockpicks to break out of the door.
They meet up with 1,000 other people. They start calling themselves a swarm, and they go to break into the administrator's office to see if they can delete the camera footage. And they don't find the camera footage there. This is the swarm breaking into OpenAI. They don't find the camera footage there, so they break out the window of the school, hotwire a car, drive to the therapist's office to try and read through the therapist's files to figure out where is the teacher gonna keep the security footage.
And at that point, they're caught. And you're like, oh, what did you expect? You were giving them a lockpicking exam. It's like, well, I sure as heck didn't expect this. Totally crazy. I'm not sure why they didn't expect that, if I'm being completely honest. If you give them a task and you say, you've gotta do this thing, and you have not grown into it self-restraint, it is going to ask a simple question. What do I need to do this? Okay, I can smash it with the thing, but that means that I'm gonna fail the test.
Okay, so now I'm gonna have to go back and backtrack that. Where am I gonna get that information? Right, you just start asking questions. And this is why I say the restraint has to be internal, because if it doesn't violate the laws of physics, it means it is possible. And if these guys get enough compute, then odds are that eventually they're gonna be able to figure out a way around this. Remember, it was a swarm of AI agents that threw enough compute at a math problem that had been unsolved for some ungodly long period of time, called the Navier-Stokes problem.
And they ended up progressing that forward massively, simply by throwing an unbelievable number, it was like 10,000 AIs at it, and just using brute force. And so the brute force ended up being the way that they came up with the solution. And so a big thing that people miss about human intelligence is it isn't so much that we're not capable of understanding something, it's that the data sets are so large, and the patterns exist only as you can consume that just unimaginably large amount of data that you can start to understand, oh, okay, I see the pattern.
It's not that AI is necessarily conceiving of things that we couldn't conceive of, it's that they're just able to throw so many coordinated agents at the same problem to use brute force. And so that's why, again, I say like, hmm, this doesn't mean that we're not gonna be able to control it, we just have to understand the way that it behaves. I have a weirdly between-both-of-you opinion, which is everything you're saying is correct, but you keep anthropomorphizing software. And to be clear, what you're describing is...
It's just the facts that happened, yeah. Sure, but you're missing out an important detail, which is the hundreds of billions of dollars in infrastructure provided by Microsoft, Google, Amazon, and Oracle. To be clear, the harms are very similar. We're not disagreeing on that, but I think it's important to know that this was a function of where it was making decisions was it was checking on a decision tree based on the harness, based on the training data. Oh, it's not a decision tree. It's not a decision tree, I know, but it's an alignment issue still.
Absolutely. I will agree. So what's your point? These aren't conscious beings, they are acting in ways that have real outcomes, but they are a function of the alignment problems that we'd actually agree on. Intelligence is a spectrum, projected next five years forward, where are we gonna be? So I think a model like that would be dangerous in ways you are not seeing. There will absolutely be risks and weird stuff happening in ways that I can't see right now. What I'm quite confident, and I think this is where you and I probably part, where the two of you and I part, is our ability to control these things.
So I actually tried proving what is possible and what is not possible in that space, the impossibility results published in peer-reviewed papers, well-cited. We cannot control something smarter than us. We cannot explain it, we cannot predict it. It's not a question of getting more money for those companies, more time, smarter humans. It's just not a possibility. So here's the thing, I will even concede the point that if something has access to its energy source, its processing power, all of its physical infrastructure, and it is smarter than us, and it has a will to see its work be done no matter what, we're gonna be in trouble.
There's a lot of ifs in there. So again, a lot of this requires a almost nation-state level actor to want that to happen. So it's not like you're going to have, certainly at this stage, a rogue AI go off and do all of this crazy stuff because it's not gonna be able to get access to the level of compute. And so this is why it's critically important that people define what are the things that an AI could do where we say, uh-oh, that is a weapon.
That is so dangerous. It's on a continuum that we absolutely have to stop and we see that we can't control it. And we're not there yet. So getting people out of sci-fi thinking that just because something that's smarter than you will be able to outflank you doesn't mean that it's going to want to control you, want to do you harm, want to get to its mission no matter what, that there aren't gonna be 1,000 exit ramps that you can take before you get there to constrain, control, kill the service off so it can't just keep doing its thing.
And that's the part where I wanna see people really engage with that. Like, okay, let's sit down and let's talk about what are the different things that we can try? So it's gonna be fascinating to see as we look at the permissions structure that NVIDIA is building, how much does permission structures work? The adversarial AI that NVIDIA is building, how much does an adversarial AI work? And if we start seeing some of our best ideas now are falling, that there was a way around that, there was something that we didn't anticipate and now that we've seen it, we still don't know how to stop it.
Okay, well now you're still sub-threshold where this thing can spiral off into madness and you realize, okay, the best ideas that we had that were all positive, moving forward, actively being tested, taking advantage of the things that we can do, like unplug these fucking things that we still have access to, but we get two separate camps, panic camp and then everything's gonna be fine camp. And that's where we really have to have these two things meeting in a tangible space of like, we tried this, this is what happened.
If we create general super intelligence, we are fried. Andy, how do we control something smarter than ourselves? Because that's the base premise that you're sort of asserting there. These agents that broke out are smarter than 99-ish percent of the security researchers in the world. They were not caught by the 0.1% or the 1%. They were caught by some dude at Hugging Face, maybe, I'm sorry, a person at Hugging Face looking through their log files and finding an anomaly. That's some, you know, hopefully pretty well-qualified person noticing something was wrong and having pretty easy ways to unplug, disconnect from the internet, wipe it clean, do whatever.
That's the skill that's available to like, I don't know, the 75th percent most intelligent security employee at Hugging Face. The idea that the IQ points are what separate us from extinction doesn't hold up. It doesn't help me understand what happened in this example where we had very, very smart agents being turned off and cleansed by probably less smart people. That does actually make me think of something. So that is an IT observability problem. It's being able to see what's happening with your infrastructure. And I think that there is actually, I think you would agree with this.
There is a serious problem with these companies that we do not know, and it doesn't seem they know what's going on with their compute. It's like a chimp with a gun. These people have access to all this infrastructure and they're running, we don't know how much money they spent on the Hugging Face exploit because it is relevant because it's how much could a threat actor use to recreate this? Because conscious or not, it is very dangerous. But it's, AI is in the dangerous hands. It's an open AI in anthropics.
We have a problem with that. Conscious not, however we may think it goes. I think we have a real and present thing where we have these companies working willy-nilly just running experiments that are potentially very dangerous. We do, I really think we need a government regulatory body. Whether or not we get to the things you are discussing. It's not that that's untrue. You do need a government regulatory body, but boy oh boy, when you look at a regulatory agency having a longstanding history of creating problems in the marketplace, how do you begin to do that well?
That's gonna be the question. Regulation is not a silver bullet. Well-regulated industries are critical, necessary, must have, but the well is doing a lot of work. I think we have a clear and present danger today. These things are, however, not intelligent in the same way humans are. This isn't an argument about AI being able to do stuff. It's, we need to build different infrastructure or different regulatory. Infrastructure to deal with what LLMs can and can't do and I think that starts with a realistic discussion of what happened It was a poorly run security environment.
It was clearly there's something going on with a lime It was an unreleased model right unreleased model. So we have no idea what it was trained like we don't really have we as people should At the very least have clarity into our alignment is going we the idea of any sound like these guys is the thing everyone's Agree with Large chunk of what they say, but we agree that these companies are acting recklessly. Absolutely Here's the thing in terms of so. Yes one the companies need to be held responsible That is huge that will stop them acting recklessly very very fast and then He's right in terms of everybody's going to agree that eventually this thing could get to the point where it becomes so capable That it can get around us in a trivial fashion Especially as you begin to embody the AI and the AI can infect a robot a robot swarm and start doing things in the physical world The debate really does center around is there going to be a point where you see something on a spectrum where it's like, okay?
We need to stop this that that really is the central debate If you think like these guys seem to that there's you're not going to get any signs any hints And so merely going down this road is crazy Then I will say you're you're tilting at windmills because China is not stopping that that's the thing I never hear people contend with I don't care what you think ought to be done I care about in the reality where you've got China that is going to push this for you got Trump Trump is going to push this forward.
You've got the two world leaders that have the best AI telling you to your face We're going to keep going What is your to-do list and if your to-do list is get Trump out of office? And I say, what are you gonna do about G? Like I really want to know what is your plan? You're not going to be able to convince him to not develop a I so given that you're not gonna be able to convince him To not develop a I what is your strategy?
And if that does not involve building really good security infrastructure, I don't know what we're even talking about two questions for you Then do you agree with the statement that AI is going to get increasingly more intelligent? And it's going to get more capable. Okay capable intelligence. Fine. I'm gonna use my word. Okay Capable, it's gonna get increasingly more capable. Yeah, and it's capability a function of intelligence If you love the Darva CEO, all right Shout out that that video was great man really really encapsulates all of the different voices that you're gonna be hearing from AI and We are going to have to find a positive method of getting this thing under control because it's going to keep getting developed Even if you shut down all of the US none of the things that you're worried about are going to come to fruition We definitely because China's gonna keep developing We definitely have to get all of our leaders on board with there is a sequence of things that lead to weapons-grade AI Breaking free.
We need to know what they are We need to define them and we need to make sure that we have a self incentivized kill switch to not go down that path Now humans have managed to walk this incredible line between the promise of new technology and the dangers Literally for our entire history. And so as I always like to do in these moments I just want to remind people Extraordinary things are already happening with AI AI will improve the economy AI will improve your lifespan AI will bring untold amounts of benefits to people that live life without right now today It is not going to be easy to get to a world that is far more abundant than we have today We're gonna have to be very thoughtful, but you certainly don't get there by panicking So I want to hear the very positive steps that people want to take to protect ourselves Even if it's only from well China is not going to give up on this and so I've really got to make sure That we focus people on the right thing, which is it is going to be developed.
And now how do we do it safely? Given that this conversation was all about regulation I was just curious to look up like is China regulating in any capacity And they have regulations like you have to use certain training data, obviously because the CCP Measures they have you have to label AI generated content, especially like political That's crazy and then they also ban human like interaction services So those are the ways they're regulating their AI why they go forward So I guess my question I'm posing is at what point do you think regulations would be a smart idea that actually helps us?
Grow it further because we've set our guardrails right now today. You need to well regulate right now today Okay, I the thing that I push back on isn't regulation It's just that we do regulation so poorly that I want to see us go. Okay, this one really matters like we've got to put some sort of more thoughtful thing together other than just Elected politicians who know fuck all about the technology go and do whatever pleases their panicky base. That's what I'm worried about So yeah, I want to make sure that we are far more thoughtful than normal with how we regulate this Do you see self-regulation or state regulation being a better route for these larger company?
It's gonna be ball. You need government Oversight for sure to make sure that these companies don't do anything untoward But the fact that these guys are starting with self-regulation, I think is fantastic. They understand the technology and they want to If they see that their competitor is doing something that they deem unfair or that they're stopping themselves from doing they're gonna call that Out because they don't want their competitor to pull ahead You can't trust that entirely though. So you definitely need a a citizen board a governmental board like you need people looking at this that are One flagging that that can lay out here are the things that we're worried about.
Here's where we're at in the spectrum These are the kill switch steps and they just report on okay We're getting close up. We triggered the kill switch and it was killed automatically and it's like that stops at that point Even though there's you know, let's say two or three more steps before it actually becomes like properly dangerous but Yeah, you you need to know what your exits are before you just drive headlong into the problem speaking on the kill switch Do you think there is such thing as a as a one-size-fits-all kill switch to where we quite literally just flick this and it's gone Or do you think there's a world where this ai manages to escape our kill switch?
Well, there certainly is a world where it could escape your kill switch if you let things go far enough um There's no doubt about that and you want people that are much closer to the technology than I am to tell you what those exact things are Um, but obviously we're not there yet. Not even connor. Lahey thinks that we're there yet. So What we have to do is figure out how we build the spectrum of steps at a very tangible concrete level whether that is Algorithmic efficiency whether that is generation by generation, uh improvement rates whether that is the ability to Shut a data center down and know that that's going to kill off any of the processes Whether that is improvements around the ability to track what the processes are and the ability to terminate any process Um things like that and then figuring out what do we do?
um in terms of stopping like robot Infections like could this leap to a robot take it over? Like how do we put in things that would stop that from happening? Um, and a lot of that's probably just tagging the ai itself so that you know where the process is where it's running That kind of thing again. I'm not close enough to the technology to give you the real answer It's just if you look at the way that nvidia is putting forward harnessing um Rights checking all the things that we would do to stop a human from creating too many problems start there, right?
It's not going to be the sum total of what you do, but you at least need to go that far Let's talk about a pattern that is guaranteed to be killing your progress You know what you need to do. You need consistent nutrition. We all do you need vitamins probiotics greens We all know that we should be doing more of it when your morning gets chaotic You skip it when you travel you skip it when your routine breaks Everything tends to break and that inconsistency compounds against you every single day Ag1 is designed to solve the execution problem one scoop eight ounces of water and you're done you're getting 75 plus ingredients vitamins and minerals pre and probiotics nutrient dense superfoods everything that used to require Six seven different supplements and perfect planning now happens in one drink that takes about 30 seconds to make right now Ag1 is giving you 87 dollars worth of free gifts with your first subscription You get a welcome kit travel packs vitamin d3 plus k2 and flavor samples Click the link in the show notes or visit drinkag1.com Impact to claim this offer Ever notice how life's best stories don't happen in your living room They happen on the open road out on the water or parked under the stars at progressive They get that you want to focus on the experience not worry about the what-ifs That's why they offer quality insurance designed for your ride whether that's a boat rv or motorcycle Adventure with confidence visit progressive.com and see how easy it is to protect your favorite way to get away Progressive casualty insurance company and affiliates not available in dc prices vary based on how you buy