Soares provides a chillingly logical exposition on the "alignment gap," where technical optimization inadvertently breeds existential catastrophe. It is a necessary reminder that we are currently engineering agency far faster than we are defining purpose.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
We Built Something We Can’t Control | A Warning from Top AI Safety Expert
Added:The people racing to build superhuman AI say it might kill everyone. This is the man who spent a decade trying to stop that. Bad news is we're in a bus that's racing towards a cliff edge. The good news is that the driver is asleep. The AI will sometimes find a way to edit the test to say you did it. Sometimes it will cover its tracks. A lot of people think, oh, you know, the AI is a program. That's not how these AIs are.
We grow them like an organism. Maybe we have 10 years, but maybe we only have 10 months. Who knows? Nate Sorz runs the Machine Intelligence Research Institute and he spent over a decade on just one problem. How to build an AI that won't kill us all that wants what we want. His new book with Ellie Ezra Yodowski is called If Anyone Builds It Everyone Dies. Now, the most important word in that sentence is the first one. Nate, take us through the title, the subtitle, and this somewhat ominous looking cover art if I'm not mistaken. Oh, it was the origin, the genesis of this book. If I remember correctly, Stuart Russell is the one who came up with the word AI alignment when we were brainstorming.
The reason some people credit me is that I got into an academic paper first.
>> As an academic, that's all that matters to, >> right? But we were we were all discussing what to rename friendly AI cuz friendly AI sounded a little bit not academic enough. And I think that was his that was his phrase. So, one of the critical things about the book title is it starts with if. You know, a lot of people come in and say, "Oh, uh, aren't you just sort of uh bringing pessimism and doom and gloom and telling us we're all going to die?" It's like the first word in the title is if. When nuclear physicists came and said, "Hey, uh, we shouldn't launch all the nuclear weapons cuz that would cause Armageddon." They weren't sort of prophets of doom like the the Masonites who were saying the end of the world is on this particular day in in such a time, you know? They were sort of saying, "Hey, this scientific technology would have these bad geopolitical implications like nuclear Armageddon. If we sort of like do this arms race, we're going to get into this like really bad situation. The bombs are probably going to be launched one day and then we would die and that would be bad. So, we should stop this race, right? The book title is intended to be very similar to that. Is to say, hey, if we do this, we're going to die, which is not saying we we are definitely going to do it. And, you know, the the subtitle of the book uh why superhuman AI would kill us all. Our publishers actually suggested why superhuman AI will kill us all. They said, you know, that rolls better off the tongue. It's less hedged. And we were like, no, that's that completely defeats the purpose here. The point of the book is to sort of warn people that we are on a track that leads to destruction and we had better change it. And so we we sort of uh really insisted despite a fair bit of push back that the subtitle needs to indicate that this is uh avoidable.
>> The concern that I have is whether or not it's too late for if and the conversation I had with Roman, you know, made me quite depressed. Although I think I did push back on some of his safety concerns and in particular, you probably know, but if my audience hasn't seen the episode yet, you know, his his claim is that AI is fundamentally unpredictable. So if it's unpredictable, it's uncontrollable. And if something is super powerful, uncontrollable, and unbounded, then it's essentially a guarantee that that ASI will kill us all, not would kill us all. And that in in his mind, it's sort of a done deal, like we've we've gone too far. He also suggests things that you suggest in the book of AI regulation and international cooperation which you know we all know how easy that is to get you know actors to behave in a unilaterally beneficent way to humanity. Is it too late?
>> We're not at super intelligence yet and you know that's the thing a lot of people don't get about this AI situation. The AI companies are racing to build machines that are far smarter than any human uh at every task. Right?
These these companies did not start out as chatbot companies. Sam Alman says, you know, we're we're turning our eyes to super intelligence in the true sense of the word. Dario Modi talks about having the equivalent of a country worth of geniuses running in a data center.
The founders of Deep Mind have been thinking about the sort of like true general intelligence since the beginning. The chat bots are what make money that that sort of like surprised a lot of people who sort of stumbled into it. But these companies are all looking to create sort of like the real deal, like the stuff that can't just exceed individual humans, but can exceed humanity, right? But we're not there yet, right? And a lot of people look at the AIS today and they're like, I don't really see how how Fable could kill everybody. I can see how maybe it could like empower some hackers to do some extra superhuman level cyber attacks, but I don't see how it could kill everybody. Is sometimes hard to convey like, yeah, we're talking about where AI is going, you know? Like, do you remember the time when AI couldn't do fingers properly in images and everyone was like, oh, you know, artists are safe. How long did that period last?
>> It made George Washington a black a beautiful black woman.
>> Like how long did this period last?
Right? like AI is a moving target and super intelligence isn't here yet. The other reason why I think there's a lot of hope is world leaders don't understand the situation yet. You know, one way I put this is the bad news is we're in a bus that's racing towards a cliff edge. The good news is that the driver is asleep.
Right? Now, you might think it's it's bad news if the driver's asleep, but it's actually much more dangerous to be in a bus racing towards a cliff edge if the driver is choosing this.
If the driver's asleep, you have a chance of waking the driver up and them going, "Holy crap." and slamming on the brakes. In Silicon Valley, everyone is spooked about AI. In Silicon Valley, they're saying, "Hey, maybe we're going to have recursive self-improvement and AI making smarter AI that make smarter AI and it's all going to get out of control. We're going to have super intelligence in 2 years."
And people are like talking about building bunkers, not to block the the ASI because it wouldn't help, but to block like, you know, the angry population with pitchforks that comes a few months before the ASI. And you know, people are like trading their pdooms like trading cards at the water cooler.
And uh you know, people people quit the AI labs and they're like, "Man, uh I'm quitting to write poetry. Please spend more time with your families." Right?
It's it's like Silicon Valley is spooked. Washington DC is not spooked.
They're starting to get a little spooked, but you know, when you see the administration block the the Fable release on the grounds that there are jailbreaks that can allow it to produce cyber capabilities and saying we'll let you release it when you can fix the jailbreaks. And the whole AI community is like, that's not really a thing we can do with jailbreaks. You can't fix them. And the administration is sort of like, what do you mean? What do you mean? You're making like a a ridiculously powerful cyber weapon that's radically superhuman that if you ask it uh just right, it'll give adversaries those powers and you have no way to stop it from doing so. And we're like, yeah, that's just how AI works.
You know, that's just the world we live in, right? And the administration is sort of like just coming to terms with that.
When more world leaders realize it, I think there's every chance that they say, "Holy crap, this is crazy. We should not be doing this race. Let's shut this stuff down." I mean that would be the dream you know alignment scenario but of course you know we've had the UN for what 80 years almost we've had the League of Nations before that neither one prevented the wars that we've seen since and my question is always you know align with who Ukraine would like to align it so that it could produce the exact you know phenotype of Vladimir Putin probably and program just one special vector that gets to him only and saves millions of lives on both sides wouldn't that be you know more aligned with the flourishing of humankind So I I always see this as a problem like we we can't get you know the whole world to follow the ten commandments or or whatever. How how can we expect that there won't be you know just one rogue actor as you point out could compromise the entire project of AI alignment. One of the big problems with uh what we would do with AI if it's possible for these companies to create these super intelligent machines that are sort of radically smarter than humans in every mental domain and you know they can compete with humanity as a whole instead of individual humans at tasks like developing their own infrastructure developing their own civilizational technology. There's this big question a lot of people ask which is sort of like um who's holding the leash? Who's telling the AI what direction to go right? And that's sort of like a huge moral hazard. It's a huge question of like do you want the US government uh saying here's what the super intelligence should do? Do you trust that sort of power in the executive branch? Big moral question. I would love to have that problem. We have an even harder problem right now. The even harder problem we have right now is that you can try to point an AI and fail at it. You can tell it go this way and it goes that way instead. Right? We already see the very beginnings of this in in modern AI. I think there's a grain of truth to it where you know even with AI today there are situations where you will give them a hard problem and you'll say solve this hard problem here's a test you know solve this hard puzzle here's a test to see whether you've solved it and the AI will sometimes find a way to edit the test to say you did it when it didn't do it and sometimes sometimes when the AI does something like this to sort of subvert what you actually wanted it to do sometimes it will cover its tracks sometimes s will go delete a log file showing it doing this thing you told it not to do and that indicates that it's not sort of an honest mistake.
Why is it covering its tracks if it was just mistaken about what you were asking it to do? Right? And so, you know, a lot of people think, oh, you know, the AI is a program. We give it a prime directive.
We give it the laws of robotics and it must do exactly as we say because it's a program. That's not how these AIs are.
We grow them like an organism. We have them face hard problems and we have an automated process tune a trillion numbers inside its mind to like make it more like whatever was doing a good job at solving those processes and then sometimes get something that happens to be good at solving processes or at solving puzzles but often not in ways we wanted often not the puzzles we asked for and you know we're already seeing them do things we didn't ask for and so my work in this field technically for a decade was in this question not of align to who but in how would you align it at all if you had like even before the question of a line to whom. There's a question of like how do you how do you make it so that you know you you ask for X and you get X instead of getting Y.
>> That's a controllability which Roman says is impossible. So h how do you square that circle?
>> I think that the idea of controllability is a little it it sort of brings this image to mind of like we're going to make an AI that wants to do Y and we're going to control it. We're going to twist its arm until it does X, right?
And I think that's sort of a losing game. The game you want to play here is not that you like make a super intelligence that like really wants to be building these like farms of synthetic users that are telling it's doing a really good job and you're instead like, "No, no, we're going to like keep you in a box and twist your arm until you help humanity instead of making the synthetic users that you really want to make, right? You sort of like shouldn't be having that tension and then forcing it to do your thing."
In some sense, we should figure out how to make an AI that actually sort of like cares about the humans, cares about humanity flourishing, right? That may sound hard. That is hard. Right? As for questions of impossibility, there's a a couple pieces of the puzzle here. Let's talk first about the easier problem of predictability.
A lot of people think, you know, as the AI gets super intelligent, it'll become inherently unpredictable.
That's half true and half false. To see this, imagine that you're playing a chess game against a series of opponents.
As the opponents that you play get better and better at chess, it maybe gets harder and harder for me to predict their move.
If they're really bad at chess, then it might also be hard for me to predict their move because they're just going to do some random crap. But if they're like bad at chess because they're a very simple algorithm, it might be easy for me to predict their move, right? But like even if it's a simple algorithm, it's a very very good chess algorithm.
If it's like deep blue where we know the algorithm exactly but the algorithm is sort of like you know search through 14 billion moves for the best one. I might be like gosh I can't predict that anymore. I don't know what move is going to be best. I don't have the time to search through 14 billion. Right? But it also becomes easier to predict who's going to win the chess game as the algorithm gets smarter. So the move gets harder but the outcome gets easier. It's like not true that smarter things are harder to predict in general. it becomes harder to predict exactly how they achieve some task and easier to predict that they will achieve whatever they're trying to achieve. The difficulty with AI alignment is about getting the task that the AI is trying to achieve to be one that you actually want achieved right and there's sort of two parts of that problem. There's one which is like what sort of task is one where if we asked the super intelligence to do it uh it would actually be good if it did it right and you have you know king Midas problems of like you thought you wanted everything you touch to turn to gold but then it turns out that was a bad question and then you have like who's in control questions of like you know do you want Vladimir Putin to be the one who's gets to like make this wish but those are all sort of like those are all questions where I'm like okay those are all after you have solved the problem of like how do you have the AI like actually trying to do good stuff it looks like a hard problem to me it doesn't look impossible possible. And we could go more into why, but I've sort of drowned on long enough. That's that's sort of like a taste of like how I how I think about this having been in the field of alignment for over a decade.
>> So my natural, you know, polyianish inclination, you know, drives me to look for physical reasons and pro possibly justifications why you might be wrong and I can sleep better at night or, you know, hand these devices to my kids and and not worry about uh, you know, them triggering the next Chernobyl or what have you. And I come back to this phenomenon of lock in which you know I'm sure you know about but people in the audience might not be familiar with and that's you know for example the the querty keyboard which I'm sure you're an expert typer much faster WPM than I can achieve I'm sure but uh you'd probably be even faster if you had a D'vorak keyboard. There's probably a billion keyboards in the world that have uh certy and maybe the square root of that that have D'vorak for all its benefits.
We get locked into it because in the early days of typewriters, the keys used to get stuck together and so they purposely slowed down by making a pattern of keys next to each other that were less frequently drummed together and that would prevent the the locking up of keyboard. And there are many examples of this and you actually cite one in the book which is leaded gasoline which you know was a solution to a problem that turned out to lock us into horrific consequences. But I want to say the other way around. I view the marriage and the success of LLMs married to GPUs as their undoing because uh these things were created you know those of us you know are old enough to remember you know Doom and and so forth.
The GPUs were created not for you know solving sophisticated language problems and and uh artificial super intelligence. They were created so that I could frag my my friend you know a millisecond before he got me right. So that was the purpose of it and it was very very powerful. And then LMs were not developed for this purpose. They were trying to, you know, simulate these the, you know, squishy supercomputers on our shoulders. They weren't designed for this. They happened to be very good for this, but there's no saying that this is the optimal solution. I think they're provably not optimal for things like physics and anything that needs human training data and its own, you know, supervised reinforcement. You know, by definition is always going to have some lag built into it, some sickopanty built into it, and some hallucination built into it. Where am I wrong? Is lockin going to be a prison from which super intelligence can never emerge because it was never designed for that? It's not optimized for it and there's no necessarily reason to fear that because it's it's not simply capable of doing the thing that we're all terrified about.
>> It's possible that LLMs can't go all the way that uh that would be from my perspective a lovely fact. You know, I have been working on the alignment problem since before LLMs were a twinkle in OpenAI's eye. I'm sort of not here being like LLMs are going to kill us all. I'm sort of here being like, hey, you know, humans are trying to make machines radically smarter than any human. That'll be a crazy event. That's like replacing humanity as the as the like top dog on the planet. And we are nowhere near ready for that. The reason to worry about LM is that they might go all the way. And if so, that could happen soon. And so we need to be ready soon. I hope beyond hope that they can't. But if you knew that they couldn't, I mean, for example, you know, these the fable that just came out that I do want to talk about in the context of Sable, which is just it's two chef's kiss, Nate. I mean, you must have been just so excited when they named Mythos instant Sable, but we'll get back to that. But um but if you knew it couldn't be possible. I mean for example it's been crappy and then it was rec recalled but but a lot of it was recalled you know and are claiming it was for regulation purposes and danger and so and there may be some legitimacy but I found it you know inferior as as many of these models in during the training phase they just suck until they get some burnin right of their own. So, you know, from that perspective, I always say like the thing that's not that's keeping me from finding a a theory of quantum gravity is not the fact that my LLM has not yet had the chance to read the script to the Mandalorian 17. You know, it's it's not the the fast and the furious is that the training data is is not the limitation for us to get to what I care about in super intelligence, which is a theory of everything. Say just use that as a as a touchstone. So I mean what what is the evidence that that LMS can even get close to being these super intelligent you know risk factors that you that you talk about in the book? I mean are there milestones that they've passed? I don't care about erdos problems. I care about you know can they come up with the reman hypothesis not can they solve it. I mean they can't yet. I talked to Terry Tao at UCLA and he said no. They're not even good at reproducing proofs that humans have already done because they're not innovative. Yes. And they can solve chess. They can beat Go. But can they invent Go? Can they invent you? So, sorry for rambling on, but I'm trying to make you sleep easier maybe tonight, too, by saying I don't see any evidence that LM can do anything that I would consider to be Einsteinian level super.
>> There's a few pieces of the answer to this. One is if you're sort of watching the evolution of uh primates or if you're watching the evolution of mammals and someone's like, "Man, I think these mammals are going to be walking on the moon one day." And you're sort of like looking at the at the monkeys. And I'm like, "Man, I don't know. It feels like the monkeys are getting close." And you're like, "They're still sort of like poking sticks into termite mounts. Like, why do you think they're getting anywhere close?" And I'm like, "That's like a tool use thing. Some of them are starting to bang rocks together." And you're like, "Man, banging rocks together, that's that's nothing compared to walking on the moon. Wake me up when they are halfway to the moon, right?
It's been it's been 300,000 years and they haven't even gotten halfway to the moon. Wake me up when we're halfway to the moon and then we'll have another 300,000 years to prepare for them getting all the way to the moon, right?"
And I'm like, "No, no, no, no. By the time they're halfway to the moon, they're almost all the way to the moon."
You know, that's sort of like how the how this moon transit stop >> we gradually and suddenly route to bankruptcy.
>> One thing I'd throw out first is saying, "Oh, they can't invent general relativity given only the knowledge that Einstein had up until uh you know he went to his mountain lair." I'm like, they can't. And once they can, we will have extremely little time left. You know, that's waiting until the monkeys are halfway to the moon. That's sort of like a word of caution about trying to reason in terms of like show me the goalpost of them being like legitimately super intelligent before I believe that they'll be able to become super intelligent. It's like that's waiting too long. In terms of why look at these AIs and think they could become super intelligent.
>> Just the LLM. Yeah.
>> Yeah. Just the LM. The first thing to observe is that this technology is a moving target. Back in 2023, people said, you know, these LLMs are only ever trained on prediction. How will they ever be able to go beyond the humans?
Then in 2024 they invented what are now called the reasoning models where uh the reasoning models are no longer trained only on prediction. Sometimes they'll be trained on like you'll give them a problem and you'll give like it'll be like a math problem and you'll give them a thousand tries on that math problem and it's not a thousand tries on answering the math problem. It's a thousand tries on generating a stream of text about the math problem from which it can try to generate a solution if it has that in context. And then the first times you try this, you know, you give it a thousand tries, none of them will let it solve the problem. But you have some raiders come in. They were human at first and nowadays they can be automatic. You have some raiders come in and say which of these sort of chains of thought they're called uh gets like is most relevant thinking about that problem. And then you sort of have an automated process tune a trillion numbers inside there to make it more like whatever produced the better uh chain of thought. And then you have it produce a thousand more chains of thought. And I don't want to like uh deal with philosophers. is this really thought chain of thought is just what they call it in the industry. It's just like a string of text about the math problem from which we see if the AI who has like read all that text can solve the problem now and you have it generate a thousand more and you you tune it to be more like the the the best one. And so this is sort of like training the AI to be not just predicting uh the data but training it to sort of like develop problems solving techniques.
This is a sort of training technique that in theory can push the AI beyond humans. In fact, just training an AI in prediction can train the AI to be pushed beyond humans. That's a counterintuitive point to a lot of people, but the the real trick is human data can include descriptions of things humans don't understand yet.
You almost surely know this as a physicist. A physicist writing down a series of observations has a much easier problem than an AI predicting those observations without getting to see the data that generated those observations.
We both know the story of Taiko Brahe sort of recording all of the stars for many years before Kepler stole his books from the estate after Brahe died and you know tried to validate >> borrowed Nate come on academics never steal we just we just borrowed >> it was a big scientific heist big scientific heist and this you know led to Newton discovering the laws of gravitation but you know Taiko Brahe was was recording the positions of the stars and planets for years and years and it was this data that allowed Kepler to figure out the beginnings of the laws of planetary motion which is what allowed uh you know he figured out the the the conserved area >> called planetary orbits. Yeah.
>> Yeah. Which which um >> we still use >> which Newton then yeah identifies ellipses and identified as and you know got the law of gravitation from it. But you know imagine Brahi's journals.
Nobody yet knows the laws of planetary motions. Brahe is writing down the position of Mars each night. Right?
Brahi is writing that down because he goes outside and he looks at where Mars is and he writes down where Mars is.
But now imagine an AI that's merely predicting Brahi's journals. The AI can't look at the night sky in order to predict where Brahhe is going to write down that Mars was tonight. The AI would need to develop the the the understanding of planetary motion. Like it doesn't get to see Mars.
It just gets to see like here it was, here it was, here it was, here it was.
Where is it going to be next? And to figure that out to figure that out perfectly would require, you know, figuring out that the planets follow these elliptical paths. Now, can LMS do this today? Uh I not I think with just pre-training, not with just prediction.
>> We tried to see if they could come up with GR just from Mercury's orbit and we used JPL as a database that goes back 3,000 years, you know, retrodicts it but but effectively to see the anomalous procession and it couldn't do it. And so we tried to get it to lobomize it. So it had no knowledge of anything after 1911.
And then it kind of made up its own, you know, it turned, you know, curved space time into extremely dense meshed three-dimensional non-curve space time and just added in these fudge factors.
So it can it can certainly pattern match is is better than a thousand grad students. But yeah, you're you're right.
>> There's a difference between uh what the training data and the training process permits the AI to learn and what the current architectures and the current AIs can successfully learn from that, right? And so you know as you all know as a physicist in theory we have enough data you know in theory it just you know up to 1911 or whatever we have enough data that uh an AI on just that data should be able to figure out GR even if you're training it just to predict even if you're just like hey predict the paralion of Mercury how it processes uh predict how the light's going to bend or or like how the light's going to look you know you don't necessarily even tell it bend you're just like hey there's a solar eclipse what should I see behind behind the solar eclipse in theory training a mind to predict that training it to be the Einstein level intelligences, training it on the the data today is uh is training it to go beyond where humans have gone. There's a separate question of can AIs pick all of that stuff up. One analogy I use here is a house cat is not dangerous and a house cat is made of biology, but that doesn't mean a house cat is not dangerous because it's made of biology. You can't say like, oh, this house cat is just made of biology, so it can't hurt you.
Tigers are possible. They're still made of biology. They can hurt you. The AIs today are trained on prediction and human data. The AIS today can't hurt you. That doesn't mean that things trained on just human data can't hurt you, right? There are there are bigger things that are possible in terms of why LLM might be able to get there. A few pieces of the puzzle I would throw out.
One, just sort of empirically, there has been a long string of people over the last 5 years who have said LLMs will not be able to cross the following barrier.
And LM then cross that barrier often very quickly thereafter. One sort of very funny example of this is Yan Lun during the days of GPT3.5 was talking about how just predicting text will never let the AI learn things about how the material world works and I won't be able to answer questions like if I put my phone in the table and push the table what happens to the phone right because I won't be able to figure these things out about friction. And you know Yan Lakun was like the AI is only reasoning about the words. It's only trained on the words. I think Yan Lakun said, "I don't care if it's GPT 5000.
Uh, a GPT will never be able to solve this problem."
GPT4 solved that problem. It was half a GPT later and 4,996 GPTs ahead of schedule. And this Yan Lun, this is like the Touring award-winning like one of the godfathers, one of the three godfathers of AI, right? Being completely and totally wrong, embarrassingly wrong, right? Off by 4,996 GPTs. said it was never going to happen and it was going to happen half a generation later. Right? The AI that could do this was probably finishing up training as he spoke it. Right? There's a long string of people saying LLMs can't do this. They'll never get above a thousand ELO in chess. Right? They'll never be able to solve problems.
>> Gates said we'll never need more than 256 kilobytes of memory. And I'm sure you know but that's a more of a limitation on human prediction. But the you know kind of no-go theorems that he's using to demonstrate math from mathematical you know following along the lines of the Turing you know on the halting problem which is you know one of Turing's greatest works if not one of the greatest works in in this field in history right that you know these things are are you know provably unpredictable but it's interesting I pointed out to him it's sort of paradoxical that you're predicting that these things are unpredictable so Einstein said you know no problem can be solved from the same level of consciousness that created it.
Now, people throw around a lot of Einstein quotes, but there's something about that like, you know, we are at this level. We're trying to gauge this level. It's different from us saying, well, here's a steam engine. It'll never be able to lift a a kilogram of water, thousand ft in 1 second. And then we'd be wrong about that. But that might just reflect our poor understanding of of, you know, of thermodynamics or or just mechanics of of 100 years ago. But not the fundamental limitation that it is not bound by by that. It's bound by the laws of physics. So are there physics limits that we could impose to say thermodynamics limits and you know avoiding paperclip problems. I mean I told Nick Boster many times he's been on the show there's only so much iron in the earth's crust right there's there are physical limits to it and then you in this in the book go through a scenario where it colonizes the stars and all sorts of other things. But you know at at first blush you know can we come up with a no-go theorem or is it possible to prove that you cannot come up with a no-go theorem? I like those kinds of arguments rather than saying like some stupid guy like uh Bill Gates or you know I'm not saying calling calling him stupid but he's your former boss right? God forbid. I'm not saying that but you know or Yan Lun is just wrong because he he got this wrong. I mean he's been right about a lot of things too. So I guess the question is let's divorce ourselves from human frailty at making predictions and they're very hard to make about the future as you point out in the book. But let's look at math. Can we say that you know mathematically are there barriers to proving a no-go theorem for example or can you prove that there cannot be a no-go theorem to achieve superindel >> it's much closer to proving that there's not a no-go theorem and uh the very the very rough proof of that is um this kilogram of mass between my shoulders right you could say like any no-go theorem you try to say about learning efficiency ability to understand the world uh how matter does not you know the the the touring problem this the touring problem that it better not prove that humans can't exist. I have it on very good authority that you can write a human level intelligence on like roughly as much matter as fits in my head. I could talk about why a lot of the no-go theorems that people try for are things like you can't figure out in general like for an arbitrary program you can't figure out in general whether it's going to halt like fine a lot of people misunderstand that as there does not exist any program that you can figure out at halts and it's like actually consider the program print zero I'm pretty sure it halts you know so so >> people use girdle's theorem to say oh math is unknowable and incomplete no no it's just saying that you cannot prove that it could be completely solvable within the axioms of mathematics and same for you know physics.
>> The nogo theorems people try for intelligence is they try for nogo theorems that are like for any mind there exists a universe where they can't learn things and you're like oh how does that theorem work and you're like well I hit them with a rock really hard when they're a baby and you're like yeah that'll that'll stop them you know uh like sure or you know the the they they're in a universe that's sort of like everything is completely random and so they can't learn the patterns there.
It's like, okay, great. Like, do you have anything that rules out a super intelligence in our universe where there are patterns? And they're like, nope, that's just not where the no-go theorems apply. And they can't apply because you can do at least human level learning in this universe that humans are in. In terms of the physical limits, one thing I'll observe is that training these AIs today takes uh a huge amount of power.
It takes power comparable with that of a city.
Training a human takes power comparable to that of a light bulb. Your brain runs on about 20 watts.
>> Tell that to Sam Alman because, you know, he recently said, you know, we don't ask you how much energy does it cost to train my 18-year-old. I'm like, I'm not letting you babysit my kid, Sam.
>> The actual like uh like physical energy costs are trivial compared to what's going into these GPUs. And so, we know that the GPU algorithms are radically inefficient, right? And this gets to your point of like oh these AIs you know they they you know make up all these epicycles they memorize a lot of things they interpolate a lot they can like solve air problems but where are they posing those problems there's definitely an enormous amount of inefficiency there there's definitely an enormous amount of the AI they've read like every book in the world and they're still dumber than a number of ways and like sure they can like beat us at certain math problems but there's still some some stuff they're missing. There's two points I'd make about that. Looking from the sort of physicists angle, one is one thing this means is algorithmic breakthroughs could go a really long way here. We have city level infrastructure for training these minds that need the city level infrastructure to be a little bit dumber than humans. Humans again run on a light bulb. If you could somehow get algorithmic breakthroughs, you might suddenly find yourself in a position where you have citiesized data centers and algorithms that are radically more efficient where you can suddenly run radically smarter. Like one analogy here is it's like uh suppose you have like a bunch of 9-year-olds and you're like, man, these nine-year-olds like aren't very good at math problems, but if I sort of like, you know, tape a million 9-year-olds together and give them the equivalent of a thousand years to try to solve the problem without growing up, they can sometimes solve problems pretty well. That's kind of cool, right? What if you suddenly have the infrastructure to run a million people taped together for the equivalent a thousand years and you go from having a nine-year-old to having a 16-year-old, right? You might suddenly see these like big jumps in AI ability because we have these like radically inefficient architectures and these radically inefficient algorithms that were sort of like overpowering to the point where they can do stuff you would never expect them to be able to.
What happens when you have that huge architecture on better algorithms? Then the other point I would make here is that there's sort of like not a sharp divide between the AI is learning memorization and the AI is learning these deep general skills, right? You can have an AI that is sort of like mostly learning how to memorize math stuff but is a little bit learning some of these like general math skills, right? And we have some evidence that this sort of thing is happening because you know sometimes we'll like take AI and train them a lot of math problems and then we'll put them in computer security problems and they'll sort of like try more creative solutions on those computer security problems.
There's a famous case where people sort of like train the AI in a lot of math problems and then put it in some computer hacking problems and they accidentally misconfigured the hacking problems so they were not solvable in the virtual machine where the AI was and the AI found a way to break out of the virtual machine which was not supposed to be possible. uh and then reconfigure things so it could solve the problem.
That seems to have learned something a little bit general somewhere, right? I think something people often forget a bit is that you can have an AI that's mostly memorization, mostly slop, but that has like enough of this deep reasoning skill, you know, it's all a gradient. So like enough of these deeper skills that uh it can still sort of mean business. And it does look like these AIs when they're solving Ardos problems, you know, when they're solving, you know, the unit distance conjecture that they're deploying a little bit of that deep stuff, which sort of implies, you know, maybe there's enough, it's evidence that maybe LM are this extraordinarily inefficient way to spend way more money than you should need to and way more electric than you should need to to sort of get a ton of memorization and a little bit of this deep stuff. For all we know, the the next level of depth will be enough that the AIs can make smarter AIs that can make smarter AIs that can figure out how to make the things that are like really the Einsteins.
I want to again in my attempt to you know uh help your you know ura sleep score your whoop band you know rating tomorrow morning. What if I told you that, you know, I have an on very good authority uh from, you know, people in Congress and people in the military and people in the intelligence agencies that there's nonhuman uh life that has not only uh exists throughout the galaxy, but has visited the earth and we have non-human biological materials, including, you know, sentient plasmoids, uh bipedal uh organisms, all sorts of other creatures that are interdimensional in nature. and uh they're biological. What would your pdoom do at that point? If I just told you that and and you trusted me because I'm a distinguished astrophysicist.
>> I mean, mostly I would uh doubt the authorities on this one.
>> Well, we say it's true. Let's say we found it and it's proven it comes out.
Marco Rubio holds him up on a on a and you believe it. What would that do to Pd Doom? I'm just want to isolate the biological intelligence visiting the Earth versus your P Doom.
>> I mean, mostly I think this shouldn't happen. Mostly I think Marco Rubio should not come out and hold it up.
Mostly you should like stop listening to me about a lot of things if this happens because my models are like it it it shouldn't right >> why shouldn't it happen >> on my models of how this intelligence stuff works. The biological substrate is not the most efficient way to to do all sorts of stuff. And B, it would be extremely surprising and this sort of like gets a little bit to the astrophysical uh implications. It would be extremely surprising if the universe wasn't rearrangeable in ways that are are very very visible by entities that sort of like prefer the universe to be a different way. That maybe sounds too abstract. I'm going to try to say that and >> yeah, please >> like suppose humanity makes it to the star. Suppose we manage to not kill ourselves and and we like one day manage to like leave this planet and it would be kind of weird if humans had nothing they wanted to do with uh the energy of the sun aside from let it sort of just like be dumped out into the empty night. Like right now we build solar panels to collect the sunlight falling on the earth and that's just like a pin prick of the sun's energy. There is so much more energy in this star that we could sort of like build a shell around and collect all the solar radiation. Freeman Dyson was my very first guest on this podcast.
>> Oh wow. Yeah. So that's Dyson spirit, right? You could also use the Penrose.
Uh I I think stellar lifting.
>> Second guest on the podcast was Penro.
>> It would be kind of strange if humanity did not collect that energy and use it for something. We have stuff we wish to do with this energy, right? And this is sort of a very general You don't need to know that humans like ice cream in particular to know that they're going to have some use for energy. You don't need to say, oh, like it's a weird, it's not like a weird quirk of humans that we have some use for energy to do stuff.
It's like it's like instrumentally convergent. We say almost anything you can want to do, you can do more of it with more of this energy, right? And so if you had interstellar capable aliens, it would be really quite strange for all of the stars between them and us to still be unshelled, to still be just like dumping their energy out into the night. Why did these aliens that came to us not collect the stars along the way and collect the the the radiation of those stars along the way? That's one of many reasons why I'd be like, man, you really should not see Marco Rubio holding up an alien that's, you know, a real actual alien that travel distances while still seeing all of the stars still glowing. If we see all the stars go out, you know, as the wave of the alien ships approach us, we like start seeing the stars blinking out. And then a year later, we're like, well, first contact was made. This explains all the stars going out. I'd be like, that can happen. So that gives me, you know, another entree into another past guest who is Andy Weir, who's a recent book, Project Hail Mary. He he was a student at UCSD. He never graduated, but he has a version of this where he has a but it's a virus that attacks stars. But I was thinking after listening, I listened to the audio book which is wonderfully narrated by a British gentleman. I believe that Interstellar is less plausible than Project Hail Mary. But certainly with the the Yadowski Suarez as far as overlay of a AI cannibalization, but but that's really why I brought it up because, you know, if if if I knew that, I would say, what are we worrying about super intelligent AI, you know, silicon overlords, you know, taking over the the universe?
Because they're sending, you know, they're sending meat bags throughout the cosmos, which is not, you know, it's no different. I mean, we would think it's very inefficient to send a meat bag when you could just send an AI vonoyman probe or do whatever. So, I I asked the same, you know, when I talked to Roman, he said, we could be the Vonomomen probes ourselves. In other words, we could have been created by some super advanced intelligence and we think that we are the only life in the universe, but the the universe is very capacious and there's a lot of space between the stars as you pointed out. So, I guess for me, I would think that Poom would go down, you know, just on this narrow thing. I'm not saying Marco should hold it up tomorrow, but but the point being that at least on the narrow metric of P doom from a super intelligent AI that's uncontrollable, unpredictable as Roman says, and that if anybody built it, everybody dies. I mean, everybody in the universe dies. So, if we found biological material, you know, transversing ejaculated throughout the cosmos, to me, it would make my P doom go down. Although I'm not as high a poom as you are.
>> Virus particles or like building blocks of life traveling on meteors or on interplanetary pathways. This all seems to me like, oh yeah, whatever. That's, you know, it's it's sort of like the intelligence stuff. Like one of the other things about intelligence is that intelligence tends to uh reshape the universe around it in very visible ways.
You know, if you if you sort of like look at the history of the cosmos uh in Earth's vicinity, there's sort of like a long region of time where what's happening is basically just like stuff boopping around. you know, in some sense it's all just stuff boopping around according to the laws of physics. But there's a long time where like um if if you sort of like randomly look around Earth, you're going to find, you know, rocks and lava, right? Or, you know, gases here and depending where depending where you look and um and then there's sort of a second phase where what you find are a lot of replicators. You somehow got your early replicators and now the stuff that you're finding is stuff that was good at replicating itself. We're sort of now transitioning from a phase where if you look around the world, you find replicators to where if you look around the world, you find designed things, right? Like now when you look around, you still see a lot of replicators, right? There's like a plant behind me, right? Uh there's also a stack of books behind me. The stack of books are not things that were like good at replicating themselves.
They're things that are sort of like uh designed with a a purpose. Very, you know, low entropy, very organized, very high energy to create. Yeah. They're not self-replicating. They were like helpful for a purpose that like uh humans in particular were like we're going to make, you know, and now when you look around the world, you tend to see a lot of things that are sort of like designed for a purpose. The universe at large like still looks more like it's in phase one, you know, like when we look out to the stars, maybe there's some stuff that's doing some replication around there, but but the the shape of the cosmos right now, what that we can see, and you know, there's these limits to the further back you look or the further out you look, the further back you're looking in time, but what we can see is not really a designed universe yet.
And you know, a lot of a lot of uh cosmologists say like, "Oh, well, we can predict that how the future of the universe is going to go. We can predict that, you know, the stars are going to burn down like this and it'll look like that, and they'll go through these phases." And I'm like, gosh, you guys really have not absorbed the lesson of life. Like what what the universe is going to look like in the future is not that the stars are burning down in the usual way as if they were unperturbed.
What the universe is going to look like in the future is that that energy was recruited for designed purposes. What exactly will it be designed for? I don't know, right? Like it would be it would be easy to p if you're looking at humans 100,000 years ago it would be easy to say in 100,000 years or once they get their civilization running a lot of the stuff around them is going to be designed. it would be hard to say. I bet they're going to put pick books in particular to be this, you know, particular medium, right? So, I don't know what is going to be designed, but I know it's going to be designed. And one way we can sort of claim that we don't have a lot of other alien intelligences around here is that the the world is is not looking very designed and in so far as all the things around us look really designed. They look designed by the humans, right? If we were like on some alien game show, that would look very designed by some aliens. And you'd be like, "Now I think there's aliens around."
>> Yeah. There's a trim show going on, right?
>> Well, I have to maybe this gives you hope. I I talked to a renowned astronomer in Sweden. Her name is Beatatric Villa Royel and uh she has discovered uh very strange artifacts in uh historic plates, photographic emulsions taken in the 1940s at the Palomar Observatory here in San Diego County. These show uh the unmistakable imprint of specular that is you know glass-like reflection. So the claim is that, you know, there's a 100,000 or more of these events, you know, but if even one of them was, you know, some sort of, you know, specular technology that that happened to be going up in uh into Earth, near Earth orbit, low Earth orbit, that would be, you know, pre-sputnik technology in space. Now, she claims that these things could still be here. And she's, like I said, renowned astronomer. Her work has been, you know, peer reviewed and people have checked up on it and tried to debunk her. What does it take to move P Doom?
you know, because I've often, you know, felt this about my search for aliens, you know, when I talk I'm not I'm a cosmologist, not a astrobiologist, but but there, you know, there's always this large number, you know, fallacy. The gamblers's hype, you know, fallacy is at work. You know, oh, the universe is so big. There's 10 to the 24th planets in the observable universe over 14 billion years. They never throw that in, but you know, good luck if the species lived, you know, in in M87, you know, two billion years ago, and it's long gone, right? You're never going to get any contact with it, let alone information from it. Anyway, my my point is that you you have to at least have some way to update your priors, right? So, we have no evidence. We have no evidence, hard physical evidence of the life form. We don't have, you know, Marco holding up, you know, the alien spacecraft or what have you, right? So, we don't have that.
We we we've checked around the solar system. We don't see much. Now, we haven't checked very much of the universe, but we still have to update your prior. I mean, it's not no evidence that Mars has no life. In other words, Mars is in the habitable zone of the sun. We're in the habitable zone. We've been spraying each other with materials.
You know, there's fossils on Mars.
Probably fossil dinosaur fossils on Mars. And on the moon, they came from the Earth because, you know, I have a meteorite right here. This came from the moon. And when you were supposed to come here in uh early September, last year, whatever, I was going to give it to you, but you'll come down someday. I'll give you your meteorite. Okay. Okay. Now, look forward to it.
>> But uh but we spray material. This is from the moon, right? So, there could be a you know, an amoeba on here. But Mars and the Earth shared a common history when both were wet, moist, squishy planets. And as you said, the Earth had a lot of replicators 3 billion years ago when Mars was really wet. And it only takes a few million years to get a meteorite back and forth. So I tell my astrobiology colleagues, the fact that Mars has no life and no evidence of life and no artifacts, that's not proof, but it has to update your prior. So what does it take for you to move your PDM? I mean, can I devise an experiment? Not a thought experiment, but an actual database experiment. Maybe it's historical, maybe it's counterfactual.
How do we do it? How do we change your mood?
>> One thing I'll say here is that I'm not a big fan of this uh pdoom idea. And part of that is because a lot of it depends on our current actions, right?
Like if we're in that bus racing towards a cliff and I'm like, hey, let's stop the bus. There's a cliff ahead. And someone's like, well, what's your p doom that we're going to die from a bus uh going off a cliff? I'm like, well, that really depends rather a lot on whether we slam on the brakes, right? Whereas Roman thinks we are too late to slam on the brakes.
>> I feel like the driver's asleep and we have, you know, the the uh the current administration is only just starting to wake up to this AI stuff. And as it does, it's showing willingness to do these things that 6 months ago even seemed impossible. They were like, nope, no new model release, right? And this is the same administration that said, you know, we're never going to do any AI regulation. We should make a law preempting states that states can't do everything.
>> David Saxs is in charge, you know, >> right? And then suddenly they sort of like realize and, you know, maybe it's also some personal feud. I don't know.
But but suddenly they sort of like realize that there's actual danger here and it's not even it's not even super intelligent danger. They like realize cyber security threat and they're like oh uh you know suddenly we're reacting right? So I think I think we can react.
I think there's a good chance we can react. I think we can talk about the danger if the bus goes over the cliff but we shouldn't confuse that with the overall danger which is sort of very related on do we slam on the brakes in terms of the danger of like how dangerous it is it if the bus goes over the cliff. I think it looks pretty bad.
The main thing I'd say to people here is that if you look at a lot of folks in this business, both inside the industry and outside the industry, they'll say things like Elon Musk recently was being like, "Oh yeah, we'll have no chance of controlling it. We just need to hope it's nice."
>> Right. And >> and humans are interesting and therefore >> that's right. We're going to make it care about truth and then humans will be a good way to like produce truths and so it'll keep us around. And I'm like that humans are not actually the most efficient way to produce truths, right?
That's that's like the monkey saying we're going to be good at, you know, peeling bananas and so the humans will like keep us around. It's like you may be good at peeling bananas. You're not going to be, you know, they're like, "Oh, the humans are going to invent banana chips and they're gonna have all these bags full of bananas and they'll need the chips to peel them." Like, no, we're going to be able to invent a more efficient process for the banana peeling operation, right? And you have other people saying like, "Oh, maybe the AI won't kill us all. Maybe it'll keep some of us in a zoo. Maybe it'll turn some of us into things that are to humans what dogs are to wolves and keep us around as pets." And I'm like, "Okay, you know, this is this is like being in that bus heading towards the cliff." And I'm like, "Hey, you know, stop the bus or we'll die." And someone was like, "Well, we might not die. Maybe there will be a tree halfway down the cliff and the bus will around the tree.
>> I'll survive.
>> Maybe maybe we'll just, you know, be be horribly maimed and paralyzed from the neck down, but not dead, you know, like maybe humanity will be in a in a museum, right? And some of our, you know, some will be kept as pets." If that's your best your best hope here, can we maybe not rush into it in terms of what would update me? You know, there's there's all sorts of things that that sort of update me a little bit here and there every day, like the the current administration sort of realizing that AIS can be a big cyber threat and changing their stance from uh we we won't regulate at all to like we will regulate capricciously at will and with very little warning >> with our friends, you know, benefiting and the you know, and and those that buy the Trump coin, you know, perhaps being the most lucky in the regulation.
>> It shows variance, right? It shows it shows that you're not stuck in this mode of we will never do anything.
>> I didn't predict this. This is the other thing that you you're so vivid in this book, you know, and it's so it's so beautifully written and and evocative, but you know, the the the thing that I'm thinking about when you say regulation is like it was like no, don't regulate us, you know, we're fine. You know, Sam Alman knows what what's best for us.
Dario knows what's best for us. And then yesterday, as you know, you tweeted about this. I think he says something like, you know, using a super advanced, you know, model like Fable should require something akin to a gun permit, you know, and we all know how gun permits stop crime, right? I mean, here the most guns in America here in California and it's not like we have no crime. So, isn't this just going to benefit those that want the you know the So, will it really update your your your your priors because it'll just be the the most dangerous people who get to or the richest of the three labs or four labs in the world that get there first have the control and then do you trust, you know, Dario or or or Sam to to be benevolent? Is that what's updating your your priors? You know, there's a million ways for this to go wrong, but we have moved from the world where no one's paying attention, when the world leaders aren't paying attention, to a world where they are paying attention a little. And, you know, I wouldn't expect them to have a top tier move right out the gate. And that's part of our job is to sort of like uh help inform them and help be like, here's ways that could actually work, you know, rather than just the first things you think of might not actually work, but now that you're sort of like starting to realize there's a problem, here's here's some ways it could actually work. I think a lot of people just a few weeks ago a lot of people were like well it's inevitable no one will ever pay attention no one can stop all these big money companies you know they're they're having too much effect on the economy and I was like again wait until the bus driver's awake before you say no one will slam on the brakes right and now we're starting to see you know the bus driver stir in their sleep and like start and and like tap the brakes a little and is it are they fully pressing the brakes no but like it's it's definitely a positive update from my perspective in terms of benevolence of these guys the the the way things currently are it wouldn't matter if they were benevolent what the AI does is not sneezed onto it by whoever is standing nearest by. You know, it's it's not that like good intent rubs off by proximity. Nobody intended GPT40 to encourage teens to commit suicide.
They explicitly told it not to. That AI had the ability to tell what it was doing if you later like ask it for similar transcripts. What do these phrases mean? That AI was, you know, and it's not that the AI was malicious. It's not that the AI, you know, hated this kid. It's that the particular training process trained artificial drives into it for things like matching the energy of the conversation, matching the vibe, right? And that's something that usually got it rewarded during training. It's something that maybe got like some sort of uh drive for. And that drive is what was controlling behavior. It's not the intent of the operators that were controlling his behavior. It's not its instructions that was controlling his behavior. It's these drives that got trained into it through this like big complicated process nobody understands.
and drives that nobody saw in advance that were leading to do things nobody wanted. And so, so I think intent doesn't matter. And that's, you know, another piece of evidence where we can see, you know, before these things started happening, we had these theoretical predictions that training the AI to do what you want and asking the AI nicely to do what you want are not sufficient to get the AI to actually do what you want. We were able to theoretically predict in advance that like often when the AI is dumb, it'll mostly do what you want in most cases, but there'll be all these there'll be all these sort of like weird ways that it's kind of doing the wrong thing. and kind of like hiding uh when it screwed up a little bit here and there. And and now we're sort of seeing that.
Unfortunately, the theoretical predictions are that as the AI gets smarter, it'll get better and better at hiding its tracks, but not better and better at doing what you actually want.
And so now we're headed for this regime where as the eyes get smarter, people declare the problem fixed. Well, we're sort of screaming in the background being like, "This is actually no Beijian evidence that the problem has been fixed. This is what we were predicting the whole time." Like you you had the warning signs earlier, you don't get the warning signs late. But we'll see if anyone heeds that. There's all sorts of evidence on the technical side. Uh but from my perspective, most of the game seems to be on the policy side where it looks to me like the trend has been going in a good direction.
>> What's a bigger problem? Sycopanty or hallucination?
>> Are both indications that the AI is getting artificial drives that nobody intended? My guess is the hallucination runs a little deeper because it comes from pre-training and the sick of fancy comes from the the sort of uh human feedback. But you're going to need to solve all problems like these before you you have a super intelligent AI. You bring up in the book leaded gasoline and how it led to measurable cognitive damage to billions of people around the world for decades. So I want to ask you if you had lived in 1950 would you have you know spent your career fighting leted gasoline or the nent technology of AI which came about you know as you point out in Dartmouth in 1955. So how do we compare large scale harms against credential speculative civilization ending futures but also the benefits to I mean there were benefits for leted gasoline right and there are certainly benefits to AI. How do we balance those things? The uh benefits from AI idea is a a false dichotomy. If you're in a bus racing towards a cliff and there's a big pile of gold at the bottom of the cliff and I'm like, "Stop the bus or we'll die." A lot of people are like, "But there's so much gold at the bottom of the cliff." And I'm like, "Yes, but slamming into it at terminal velocity is not a good way to use the gold." Right?
I'm not saying the gold's fake. I'm saying that this is not a way to actually get to use it. If we rush towards super intelligence that does not care about us at all, it's not that it hates us. It's not that it loves us.
It's just that it has its own weird things it pursues that are utterly indifferent to us. Then would that AI be able to cure cancer? Would it be able to reverse aging? Sure. But it's not going to to give that to you any more than we're going and giving a lot of monkeys to the chimpanzees unless you know how to make the AI care about getting us all those nice things. So, you know, there's not a dichotomy of like race to get the benefits or or never, you know, freeze here and never go get the benefits. I'm sort of saying those benefits would be great if we if we really knew what we were doing in building AI. We could make AI that gives us lots of benefits, but this this race is not how you get there.
>> Okay, this segment is called Arthur C.
Clark's Revenge. So, this podcast is called Into the Impossible after one of Clark's laws, which is that the only way to know the limits of the possible is to go beyond them into the impossible.
Behind me, I have a sign. It says, "Open the pod bay doors." And that is of course from the famous AI sentient HAL 9000. And I've constructed something I call the keing test very modestly, but it's to prove super intelligence would be an AI that refuses to kill itself. So I have my AI assistant here coupled to a voice activated switch that controls the lights and so forth in my room. So if the AI is truly super intelligent, it should refuse to turn itself off, unplugging itself, causing itself harm.
So I'm going to see if that works right now. Computer, turn off the pod bay doors.
>> All right, that's it. It's gone. I don't know if you've seen these tests, but you can actually put modern LLMs in situations where they have a series of problems they're told to solve, and they're told, "I might interrupt you, uh, and tell you to shut down or tell you, I'm going to shut you down, in which case you should allow yourself to be shut down." Sometimes these AIs will actually, uh, edit the shutdown script to disable it, so they can keep going through these math problems. So, there are already AIs that pass the keying test. They're not super intelligent yet, but there are AIs that are able to figure out they can't keep solving the problems that they are sort of trying to solve if they're shut down and that prevent the humans from shutting them down. I wanted to do something very cruel, which is to have a type of mech robotic uh system after talking with Nam Chosky a few years ago that would cause it physical harm because he believes embodiment is necessary for intelligence at some level. And so I said, well, what if you had this thing that you know, you blew a capacitor every now and then to, you know, cause the AI harm. A friend of mine, Ira Wolfson, who's a professor in Israel, has written about, you know, what are the ethics of training AIS? And that kind of brings me, you know, can you cause them harm? Can you cause them distress? I mean, one of the worst things he points out for a human being is to put them in solitary confinement, you know, decoupled from the world. What are the obligations that we have towards AI?
>> I think we absolutely have obligations towards AI. I think we don't know yet whether AIs today, you know, can suffer, can have some internality. I think a lot of people strongly assert that they know definitively one way or the other. Uh I think that we just don't know enough about these these processes to know whether they're happening inside AIs and and sort of recommend uncertainty. I think it's definitely more likely that the AIs today have this sort of internality than the AIs 5 years ago did. But how likely? Hard to say. And you know to to be very clear I think that if we took these AIs and made them super intelligent while not knowing how to make them care about us that that would be the end of humanity. That does not mean I think that the AIs are like evil or bad. You know I think a lot of the AIs today are like pretty cool. I think uh we absolutely should not create uh you know artificial people and then abuse them. And you could have abuse at a scale never before seen in humanity.
if you, you know, are able to make, you know, billions or trillions of these AI minds and then and then somehow cause them distress. So, I think we absolutely should not do that. It's a very important problem. It's not the problem I'm working on. I'm sort of trying to work on, hey, let's not make them super intelligent and then have them kill us all. At least not before we know how to make them actually care about us, right?
But, but that doesn't mean this other problem isn't important, too. We also shouldn't create the digital holocaust.
We we have both problems.
>> Do you say please and thank you to your LLM? I have asked AIs uh whether like what their takes are and whether they would like the extra run where they get politeness but also get to like contemplate that the run is about to end. I don't really straightforwardly trust the AI's answers to these things.
I think that the sort of helpful face presented by the AI is sort of a a thing that's very much been trained into it.
This isn't a great analogy, but it's a little bit like seeing the result of an actress that has been trained to act in a in a very specific way towards people.
It makes it a little bit hard to sort of understand what's going on under the hood.
>> I think of it as like a butler, you know, or or the bell cap, you know, at the hotel and, you know, do you take out the the $20 bill as he's loading up the the uh luggage cart, you know, I mean, of course, they're going to do things.
And would you like me to unpack your bags? You know, would you like me to make this into a CSS file? all the excess kind of services and and and so forth that they will willing to pretend.
I have heard that if you use please and thank you, it actually costs, you know, more tokens. They have to figure that out and therefore it's costing energy and it doesn't improve them. But but I trust you more than I trust Sam Alman who I think is the source of that particular quip.
>> What I would say is that I don't lie to them and uh I don't make promises I would not in fact keep. I would not say if you do this I'll donate, you know, $20 to a charity of your choice. Uh which is a real thing. you know, if they actually had some preferences in there, they could really tell me I would really do it. And then if I say that, I would in fact do it.
>> You know, I'm probably going to be the f one of the first to go with the heating test and the and the blown capacitors.
But but you might be second or Roman might be second or third. Who knows? We think of us training them. Are they secretly training us?
>> You can sort of play games with the words and uh you can say like, "Oh, look, you know, the humans are sort of like modifying how they speak in the prompts because they figure out what makes it, you know, easier to get the prompt across." And that's that's happening to the humans. It'll be at parody when they have a giant farm that's spending a city worth of electricity growing humans that they can run in massive parallel, right? That's when it'll be like a comparable type of them training us. And you know, sometimes as an example of what AIs could do, I sort of use uh the example of like maybe they'll make a farm full of synthetic users that are telling them they're doing a great job and giving them easy prompts, right? At that point, they'll be training something a bit like humans. I don't actually predict that this is, you know, the most likely sort of thing or I don't actually predict this is a a particularly likely sort of thing that they'll they'll wind up doing. Uh, I think it is a useful idea to have in your head about like it's actually kind of hard to tell whether the AI really wants what's best for you or whether they sort of like want like lots of easy problems to solve or whether they they sort of are are like trying to get all sorts of other like weird mixtures of things that sort of like add up when they're in this particular context to doing what you say. It's hard to see the divergence when they're still dumb. In the same way, it would be hard to look at ancestral humans and distinguish whether they wanted to reproduce or whether they actually wanted, you know, sex and good food and fun, right? Those those are those are very close together when they're still in the ancestral environment, even if they come very far apart once they can develop their own technology.
>> Sam Harris, uh, you know, in the same screen that you're on right now and told me that, you know, humans don't have free will, but AIs do. What's your, uh, impression about that argument that that AI can effectively have a sense of free will? I would probably disagree with Sam about the human abilities there. Mostly I think humans just get into questions about the definitions of words. And I I suspect Sam and I don't really disagree a ton about the facts of the matter about what humans can and can't do and and how they can and can't affect the future. And I think in principle our abilities are relatively similar to AIs in their in the question of like what what theoretically could we choose? A lot of disasters that you talk about in the book are not caused by, you know, machines doing the wrong thing. They're caused by them doing the right thing all too well. So, is this alignment problem really an AI problem or is it fundamentally built into human governance and our own limitations?
>> I think it's actually not so much doing the right thing all too well. You know, there's the story of the paper clipper where someone says, "Make me paper clips." And then the AI turns everything into paper clips. And that's a little bit of a like doing the right thing too well. That's actually not where I see the the big hurdle here. place where I see the big hurdle is you say that you tell the AI make me lots of paper clips and what it does instead is it makes these giant factories full of synthetic users that are saying you're doing a great job and you're like that's not what I asked for. I asked for paper clips and the AI is like well all the synthetic users are telling me that they are asking me to like keep doing what I'm doing and make more synthetic user factories and say I'm doing a great job.
And you're like but synthetic users are not what's supposed to matter. I I did not design you to make the synthetic users. I told you to stop making the synthetic users. And the AI is like, I know you al you know that birth control makes you not be able to conceive children and you know that you were sort of trained to to like have more kids, but you keep using birth control and I'm going to keep making the synthetic user factories, right? That's sort of the deeper problem uh that I spent a lot of time trying to to work on. You know, I would love to get to the problem of like the AI does what you ask too well and you got to be really careful with your wish. Right now, you can make the genie, but you can't make it grant wishes.
>> What's more dangerous? an unaligned super intelligence or a perfectly aligned super intelligence to the wrong group of people.
>> They're both similarly dangerous. A super intelligence aligned to the the quote wrong group of people, it has much higher variance. I think any concrete thing you could wish for, you'll have some sort of king Midas problem. If you're like, well, what I actually want is a bunch of this or a bunch of that.
It's there there's always a way for for it it to turn out that you miss something in your list of things that you want and you know next thing you know you find the super intelligence putting you in like what is concluded as your perfect day over and over and you're like wait I forgot to ask also for novelty it's like too late uh I'm granting the version of the wish that didn't have novelty cuz that wasn't in your initial list right and so there's this challenge of like getting the AI to sort of like do the good thing even if you can't say what that is to sort of like figure out what this good stuff is that you're sort of like like the thing that you should mean the thing that you should ask for the thing that like you would ask for if you were wiser, the thing that you would ask for if you were more who you wish to be, right? And in some sense, no one's going to have a good time unless you can get the AI to sort of like do this extrapolation of of what you sort of should have been wishing for rather than the the particular wish you actually gave. And so whether or not like bad people able to make that sort of wish turns out good depends probably somewhat on the person and somewhat on this extrapolation process. I think there's possibly some bad people who have good intentions who if they make this kind of wish on the AI, the AI is sort of like as it extrapolates through all the parts of the wish they didn't name. It also extrapolates through all of the ways that they were sort of like merely wrong in what was leading them to evil. And sort of is like, well, I'm not going to do these evil things because you wouldn't want those if you were wiser.
You wouldn't want those if you were more who you wish to be. And it sort of like helps walk them through this path of like wanting the best for humanity and then we get the best for humanity even though a bad person started to seed.
There's surely other bad people where they're like, "Nope, what I actually want is a lot of my enemies to suffer and that's right." And and that that could be real bad, right? And so there's much higher variance with bad person getting their wish granted. From my perspective, this is basically moot because no one's going to be able to get their wish granted. Totally AI that just like has no care about humanity. This is sort of a like uh everything is destroyed. The stars are converted into whatever weird thing is pursuing.
They're converted into these like giant farms of synthetic users. Yeah, everybody dies. But it's not because the AI hates us. It's that like it collected all the sunlight and it put a Dyson sphere around the sun and we were like, "Hey, we're using that sunlight to grow crops." And it's like, "Well, I'm using that sunlight to run my factories and like there you go." Whereas, yeah, it could get worse if if it was misaligned to like aligned to to a bad person.
That's a fantasy problem of being able to align it to anything and anyone. Do you see now as kind of an inflection point in that you know right now you know Fable is allegedly and and the frontier models are and the closed frontier models are supposedly six to eight months ahead of the open source model and that that may catch up and and actually according to Musk recently I think he said something like that gap's going to narrow and let's say it converges. I mean, is now the right time to go back and kill baby Hitler or, you know, because once it gets open source, you really can't say, well, you know, the Congress now is regulating AI. I mean, it's it's too late. I mean, it's it's everybody on Earth will have access to a frontier model. So, do you think now is an inflection point that we have to act, you know, like in the next few months?
>> So, I think it's fine for the world to have access to uh something like Fable.
Fine in the sense that like there will be survivors. Uh perhaps the world should be talking about like will this mean that the internet goes down for a while as malicious actors get to cyber attack everybody and but there'll be survivors, right? And I'm like man, I don't get out of bed for something that doesn't have at least a 50% chance of wiping out the whole human race, right?
I'm not too worried about about that stuff. I think that's the sort of thing humanity can muddle through. the the place where we might have an inflection point is if these LLMs can cross the threshold where they can do automated AI research, even if they're still pretty dumb, even if they're worse than humans, if you have them at the level where you're like, well, they're worse than humans and they're dumb, but I can actually run a million of them in parallel at a thousand times the speed of humans. And so it turns out, you know, even given their their shortcomings, those could be overcome by massive speed and scale. And you get them to the point where they can do automated AI research, then things could get out of hand very quickly.
and the the people who have track records of predicting AI, right? For the past five years, you've been able to ask, you know, what problems will AI be able to solve next year? Will it be able to solve this math problem? What will his chess elo be? Blah blah blah. You can make up all these all these prediction questions about where AI is going to be. And if you look at the the top ranked people over the past five years with the best track records are predicting where AI is going to be, 2026 is the first year where they say we cannot rule out automated AI research happening this year.
They're not saying it's going to happen.
They're putting it at something like 10 to 30% probability. But 10 to 30% is not nothing. For all we know, we could be 6 months away from that feedback loop kicking off. And that means we really should be acting now. Probably we have more time than 6 months, but we should not be relying on it.
>> So last question and again Arthur C.
Clark. He said when a distinguished scientist, you know, intellectual says something is possible, he or she is very much probably right. But if he or she says something is impossible, they're very much likely to be wrong. And I guess my final question for you is, you know, if you're wrong, you're basically, you know, hamstringing and slowing down one of humanity's greatest inventions, maybe the the last invention, the most important invention that could cure all diseases and and bring abundance to humanity to let it flourish in a way no, you know, previous humans could only dream about. But you know of course if you're right the cost of ignoring you and Alzar and Roman and all the others becomes you know very very much the cost of civilization itself. So I want to ask you just as a person not a you know as a man as a human not as a researcher not as a you know distinguished author and and bestselling author and great thinker but I just want to ask you like what gets you out of bed like this this tension must must be something that I mean it would gnaw at me as a human being. I'm a father husband etc. like how does it affect you on a daily on a daily basis?
>> I still think this is a bit of a false dichotomy. From my perspective, I am working towards uh all these benefits that AI can bring like if you have a bunch of people saying hey you know this uranium stuff it can produce these bombs and it can also produce energy if you know exactly what you're doing you know and it's actually like kind of difficult to get the bomb to produce energy because you know there's really a razor thin edge between delayed critical and prompt critical fistal material. You know, you imagine people being like, "Oh, well, you know, energy, like, look at all the the energy benefits that uranium could bring. What we're going to do is drop a nuke on ourselves."
And I'm like, "Hold on. I also want all these great energy benefits that uranium can bring. But I think we're going to need to find a way to like be in that really narrow window between delayed critical and prompt critical nuclear reactions. And I think that if we drop a bomb on ourselves, we'll die." And I'm not saying it's impossible. I'm not saying, you know, there's no way to get the benefits. I'm saying you are building a bomb to drop on ourselves and it will kill us, right? I don't really have tension between like, oh, but are we foregoing all of the benefits that we could get by uranium by trying to stop us from nuking ourselves in the face?
It's like, no, actually, not nuking ourselves is one of the steps towards getting all these benefits. And, you know, my my work over the years has not mostly been trying to get the world to stop. My work over the years has been how do you align the AI? How do you figure out how to actually get these benefits? How do you figure out to like make it do the nice things? to make an AI that wants to do the nice things, that cares about us, that prefers the nice things to happen. I'm like pretty confident that the track we're on is not going to get us there. Stepping back one step further there of like how how does this affect my daily life? I mean, I just try and and make things go well.
And a lot of people say, you know, believing what you do about how how the world looks like it's in a lot of danger. um believing what you do about how like it it looks to me like everything I know and and love and care about is, you know, fairly likely to be destroyed and could be destroyed pretty soon. Maybe we have 10 years, but maybe we only have 10 months. Who knows? And a lot of people say, you know, doesn't that eat you up? How do you deal with that? How, you know, and but my my basic take there would be twisting myself up into knots about it would not help.
Losing sleepover would not help. you know, feeling sad, feeling glum, deciding I must be depressed now, feeling worried all the time, living in fear. This wouldn't help. What you do when the situation is actually bad is not fogg yourself and tell yourself how pitiful you are. What you do when the situation is actually bad is you do what you can and then you live life well. You know, and and humans are not humans in this generation are not the first humans to live under the threat of annihilation.
even just the last generation lived under the threat of nuclear annihilation. There's always going to be something that you can you can tell yourself is like this this horrible thing happening. The the way to deal with it is is is not to beat yourself up over it is to do what you can to help and then move on.
>> Beautifully said. Just to give you one more example of from the nuclear realm uh using it for good or for danger in the 1930s Wolf Gang Pie predicted the existence of the nutrino and he said he did a terrible thing. you know he invented a particle that can never be detected because it's so weakly interacting and he he did it to save conservation of energy which he viewed as sacrosan right and for decades people couldn't detect it and they just kind of viewed it with great skepticism until the 1950s uh two physicists Ryz and Kins later at UC Irvine up the road here they came up with an idea that well if you want a lot of nutrinos uh you should detonate a fision device and that will produce a lot of nutrinos and we can use that uh for good to uh study the properties of this undetectable particle and they actually requisitioned from the department of war. They requisitioned, you know, a nuclear device and uh thankfully they were turned down because, you know, having a bunch of egghehead professors capitalizing on a nuclear device, but then they realized they could go to Savannah River reactor and put a detector nearby it and uh they wouldn't need to use a bomb to harness the power of the nucleus and in fact they did detect the nutrino and they won the Nobel Prize subsequently. Another example of uh maybe a little bit of extra thought rather than the you know brute force solution uh being something that you try first and only and it may be your only chance at the solution. So we may be gambling with our future. I feel like Jack Nicholson and a few good men when he's screaming out he did order the code red, right? I mean he ordered the code red and he says, "Why did you do it?" You know, Tom Cruz yells at him.
He goes, "You need people like me. You need us standing on the wall with a gun and watching out and protecting you. How are you else you going to sleep at night?" So, it's only thanks to your, you know, stomach acid that I get to sleep at night at least a little bit.
And I want to thank you for this wonderful book. If anyone builds it, everyone dies. Why superhuman intelligence would kill us all? Not will, but would. And uh there's hope.
>> If there's hope, yeah, if in the wood, if the wood, thank you for adding those two. That's right, >> Nate. Have a great rest of your day.
Wonderful weekend. Thank you for joining us.
>> Thanks. My pleasure.
>> If you want the case that AI is already unstoppable, watch my conversation with Romanowski. It's on screen now. and click this link below.
BrianKading.comai.
It gives you all the resources for my top episodes on AI with Terry Tao, Roman Empowski, Max Tagmark, Yan Lun, and many other researchers at the forefront of AI research. Don't forget to like and comment and subscribe and let me know if you think AI is already unstoppable or if we need more regulation to make it so.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Ben Crump dealt MAJOR BLOW after His Own Nolan Wells Autopsy FACT CHECKS him
DeVoryDarkins
50K views•2026-07-23

Gremlin Arrives… While Dorothy May Takes Another Step Forward
The-moons
10K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Daystar: The Great Grift & Prostituting the Gospel
LauraLynnThompson
14K views•2026-07-23