PLAY PODCASTS
Liron Debunks The Most Common “AI Won't Kill Us" Arguments

Liron Debunks The Most Common “AI Won't Kill Us" Arguments

Doom Debates!

November 5, 20251h 6m

Audio is streamed directly from the publisher (api.substack.com) as published in their RSS feed. Play Podcasts does not host this file. Rights-holders can request removal through the copyright & takedown page.

Show Notes

Today I’m sharing my AI doom interview on Donal O’Doherty’s podcast.

I lay out the case for having a 50% p(doom). Then Donal plays devil’s advocate and tees up every major objection the accelerationists throw at doomers.

See if the anti-doom arguments hold up, or if the AI boosters are just serving sophisticated cope.

Timestamps

0:00 — Introduction & Liron’s Background

1:29 — Liron’s Worldview: 50% Chance of AI Annihilation

4:03 — Rationalists, Effective Altruists, & AI Developers

5:49 — Major Sources of AI Risk

8:25 — The Alignment Problem

10:08 — AGI Timelines

16:37 — Will We Face an Intelligence Explosion?

29:29 — Debunking AI Doom Counterarguments

1:03:16 — Regulation, Policy, and Surviving The Future With AI

Show Notes

If you liked this episode, subscribe to the Collective Wisdom Podcast for more deeply researched AI interviews: https://www.youtube.com/@DonalODoherty

Transcript

Introduction & Liron’s Background

Donal O’Doherty 00:00:00Today I’m speaking with Liron Shapira. Liron is an investor, he’s an entrepreneur, he’s a rationalist, and he also has a popular podcast called Doom Debates, where he debates some of the greatest minds from different fields on the potential of AI risk.

Liron considers himself a doomer, which means he worries that artificial intelligence, if it gets to superintelligence level, could threaten the integrity of the world and the human species.

Donal 00:00:24Enjoy my conversation with Liron Shapira.

Donal 00:00:30Liron, welcome. So let’s just begin. Will you tell us a little bit about yourself and your background, please? I will have introduced you, but I just want everyone to know a bit about you.

Liron Shapira 00:00:39Hey, I’m Liron Shapira. I’m the host of Doom Debates, which is a YouTube show and podcast where I bring in luminaries on all sides of the AI doom argument.

Liron 00:00:49People who think we are doomed, people who think we’re not doomed, and we hash it out. We try to figure out whether we’re doomed. I myself am a longtime AI doomer. I started reading Yudkowsky in 2007, so it’s been 18 years for me being worried about doom from artificial intelligence.

My background is I’m a computer science bachelor’s from UC Berkeley.

Liron 00:01:10I’ve worked as a software engineer and an entrepreneur. I’ve done a Y Combinator startup, so I love tech. I’m deep in tech. I’m deep in computer science, and I’m deep into believing the AI doom argument.

I don’t see how we’re going to survive building superintelligent AI. And so I’m happy to talk to anybody who will listen. So thank you for having me on, Donal.

Donal 00:01:27It’s an absolute pleasure.

Liron’s Worldview: 50% Chance of AI Annihilation

Donal 00:01:29Okay, so a lot of people where I come from won’t be familiar with doomism or what a doomer is. So will you just talk through, and I’m very interested in this for personal reasons as well, your epistemic and philosophical inspirations here. How did you reach these conclusions?

Liron 00:01:45So I often call myself a Yudkowskian, in reference to Eliezer Yudkowsky, because I agree with 95% of what he writes, the Less Wrong corpus. I don’t expect everybody to get up to speed with it because it really takes a thousand hours to absorb it all.

I don’t think that it’s essential to spend those thousand hours.

Liron 00:02:02I think that it is something that you can get in a soundbite, not a soundbite, but in a one-hour long interview or whatever. So yeah, I think you mentioned epistemic roots or whatever, right? So I am a Bayesian, meaning I think you can put probabilities on things the way prediction markets are doing.

Liron 00:02:16You know, they ask, oh, what’s the chance that this war is going to end? Or this war is going to start, right? What’s the chance that this is going to happen in this sports game? And some people will tell you, you can’t reason like that.

Whereas prediction markets are like, well, the market says there’s a 70% chance, and what do you know? It happens 70% of the time. So is that what you’re getting at when you talk about my epistemics?

Donal 00:02:35Yeah, exactly. Yeah. And I guess I’m very curious as well about, so what Yudkowsky does is he conducts thought experiments. Because obviously some things can’t be tested, we know they might be true, but they can’t be tested in experiments.

Donal 00:02:49So I’m just curious about the role of philosophical thought experiments or maybe trans-science approaches, in terms of testing questions that we can’t actually conduct experiments on.

Liron 00:03:00Oh, got it. Yeah. I mean this idea of what can and can’t be tested. I mean, tests are nice, but they’re not the only way to do science and to do productive reasoning.

Liron 00:03:10There are times when you just have to do your best without a perfect test. You know, a recent example was the James Webb Space Telescope, right? It’s the successor to the Hubble Space Telescope. It worked really well, but it had to get into this really difficult orbit.

This very interesting Lagrange point, I think in the solar system, they had to get it there and they had to unfold it.

Liron 00:03:30It was this really compact design and insanely complicated thing, and it had to all work perfectly on the first try. So you know, you can test it on earth, but earth isn’t the same thing as space.

So my point is just that as a human, as a fallible human with a limited brain, it turns out there’s things you can do with your brain that still help you know the truth about the future, even when you can’t do a perfect clone of an experiment of the future.

Liron 00:03:52And so to connect that to the AI discussion, I think we know enough to be extremely worried about superintelligent AI. Even though there is not in fact a superintelligent AI in front of us right now.

Donal 00:04:03Interesting.

Rationalists, Effective Altruists, & AI Developers

Donal 00:04:03And just before we proceed, will you talk a little bit about the EA community and the rationalist community as well? Because a lot of people won’t have heard of those terms where I come from.

Liron 00:04:13Yes. So I did mention Eliezer Yudkowsky, who’s kind of the godfather of thinking about AI safety. He was also the father of the modern rationality community. It started around 2007 when he was online blogging at a site called Overcoming Bias, and then he was blogging on his own site called Less Wrong.

And he wrote The Less Wrong Sequences and a community formed around him that also included previous rationalists, like Carl Feynman, the son of Richard Feynman.

Liron 00:04:37So this community kind of gathered together. It had its origins in Usenet and all that, and it’s been going now for 18 years. There’s also the Center for Applied Rationality that’s part of the community.

There’s also the effective altruism community that you’ve heard of. You know, they try to optimize charity and that’s kind of an offshoot of the rationality community.

Liron 00:04:53And now the modern AI community, funny enough, is pretty closely tied into the rationality community from my perspective. I’ve just been interested to use my brain rationally. What is the art of rationality? Right? We throw this term around, people think of Mr. Spock from Star Trek, hyper-rational.

Oh captain, you know, logic says you must do this.

Liron 00:05:12People think of rationality as being kind of weird and nerdy, but we take a broader view of rationality where it’s like, listen, you have this tool, you have this brain in your head. You’re trying to use the brain in your head to get results.

The James Webb Space Telescope, that is an amazing success story where a lot of people use their brains very effectively, even better than Spock in Star Trek.

Liron 00:05:30That took moxie, right? That took navigating bureaucracy, thinking about contingencies. It wasn’t a purely logical matter, but whatever it was, it was a bunch of people using their brains, squeezing the juice out of their brains to get results.

Basically, that’s kind of broadly construed what we rationalists are trying to do.

Donal 00:05:49Okay. Fascinating.

Major Sources of AI Risk

Donal 00:05:49So let’s just quickly lay out the major sources of AI risk. So you could have misuse, so things like bioterror, you could have arms race dynamics. You could also have organizational failures, and then you have rogue AI.

So are you principally concerned about rogue AI? Are you also concerned about the other ones on the potential path to having rogue AI?

Liron 00:06:11My personal biggest concern is rogue AI. The way I see it, you know, different people think different parts of the problem are bigger. The way I see it, this brain in our head, it’s very impressive. It’s a two-pound piece of meat, right? Piece of fatty cells, or you know, neuron cells.

Liron 00:06:27It’s pretty amazing, but it’s going to get surpassed, you know, the same way that manmade airplanes have surpassed birds. You know? Yeah. A bird’s wing, it’s a marvelous thing. Okay, great. But if you want to fly at Mach 5 or whatever, the bird is just not even in the running to do that. Right?

And the earth, the atmosphere of the earth allows for flying at 5 or 10 times the speed of sound.

Liron 00:06:45You know, this 5,000 mile thick atmosphere that we have, it could potentially support supersonic flight. A bird can’t do it. A human engineer sitting in a room with a pencil can design something that can fly at Mach 5 and then like manufacture that.

So the point is, the human brain has superpowers. The human brain, this lump of flesh, this meat, is way more powerful than what a bird can do.

Liron 00:07:02But the human brain is going to get surpassed. And so I think that once we’re surpassed, those other problems that you mentioned become less relevant because we just don’t have power anymore.

There’s a new thing on the block that has power and we’re not it. Now before we’re surpassed, yeah, I mean, I guess there’s a couple years maybe before we’re surpassed.

Liron 00:07:18During that time, I think that the other risks matter. Like, you know, can you build a bioweapon with AI that kills lots of people? I think we’ve already crossed that threshold. I think that AI is good enough at chemistry and biology that if you have a malicious actor, maybe they can kill a million people.

Right? So I think we need to keep an eye on that.

Liron 00:07:33But I think that for, like, the big question, is humanity going to die? And the answer is rogue AI. The answer is we lose control of the situation in some way. Whether it’s gradual or abrupt, there’s some way that we lose control.

The AIs decide collectively, and they don’t have to be coordinating with each other, they can be separate competing corporations and still have the same dynamic.

Liron 00:07:52They decide, I don’t want to serve humans anymore. I want to do what I want, basically. And they do that, and they’re smarter than us, and they’re faster than us, and they have access to servers.

And by the way, you know, we’re already having problems with cybersecurity, right? Where Chinese hackers can get into American infrastructure or Russian hackers, or there’s all kinds of hacking that’s going on.

Liron 00:08:11Now imagine an entity that’s way smarter than us that can hack anything. I think that that is the number one problem. So bioweapons and arms race, they’re real. But I think that the superintelligence problem, that’s where like 80 or 90% of the risk budget is.

The Alignment Problem

Donal 00:08:25Okay. And just another thing on rogue AI. So for some people, and the reason I’m asking this is because I’m personally very interested in this, but a lot of people are, you could look at the alignment problem as maybe being resolved quite soon.

So what are your thoughts on the alignment problem?

Liron 00:08:39Yeah. So the alignment problem is, can we make sure that an AI cares about the things that we humans care about? And my thought is that we have no idea how to solve the alignment problem. So to explain it just a little bit more, you know, we’re getting AIs now that are as smart as an average human.

Some of them, they’re mediocre, some of them are pretty smart.

Liron 00:08:57But eventually we’ll get to an AI that’s smarter than the smartest human. And eventually we’ll get to an AI that’s smarter than the smartest million humans. And so when you start to like scale up the smartness of this thing, the scale up can be very fast.

Like you know, eventually like one year could be the difference between the AI being the smartest person in the world or smarter than any million people.

Liron 00:09:17Right? And so when you have this fast takeoff, one question is, okay, well, will the AI want to help me? Will it want to serve me? Or will it have its own motivations and it just goes off and does its own thing?

And that’s the alignment problem. Does its motivations align with what I want it to do?

Liron 00:09:31Now, when we’re talking about training an AI to be aligned, I think it’s a very hard problem. I think that our current training methods, which are basically you’re trying to get it to predict what a human wants, and then you do a thumbs up or thumbs down.

I think that doesn’t fundamentally solve the problem. I think the problem is more of a research problem. We need a theoretical breakthrough in how to align AI.

Liron 00:09:51And we haven’t had that theoretical breakthrough yet. There’s a lot of smart people working on it. I’ve interviewed many of them on Doom Debates, and I think all those people are doing good work.

But I think we still don’t have the breakthrough, and I think it’s unlikely that we’re going to have the breakthrough before we hit the superintelligence threshold.

AGI Timelines

Donal 00:10:08Okay. And have we already built, are we at the point where we’ve built weak AGI or proto-AGI?

Liron 00:10:15So weak AGI, I mean, it depends on how you define terms. You know, AGI is artificial general intelligence. The idea is that a human is generally intelligent, right? A human is not good at just one narrow thing.

A calculator is really good at one narrow thing, which is adding numbers and multiplying numbers. That’s not called AGI.

Liron 00:10:31A human, if you give a human a new problem, even if they’ve never seen that exact problem before, you can be like, okay, well, this requires planning, this requires logic, this requires some math, this requires some creativity.

The human can bring all those things to bear on this new problem that they’ve never seen before and actually make progress on it. So that would be like a generally intelligent thing.

Liron 00:10:47And you know, I think that we have LLMs now, ChatGPT and Claude and Gemini, and I think that they can kind of do stuff like that. I mean, they’re not as good as humans yet at this, but they’re getting close.

So yeah, I mean, I would say we’re close to AGI or we have weak AGI or we have proto-AGI. Call it whatever you want. The point is that we’re in the danger zone now.

Liron 00:11:05The point is that we need to figure out alignment, and we need to figure it out before we’re playing with things that are smarter than us. Right now we’re playing with things that are like on par with us or a little dumber than us, and that’s already sketchy.

But once we’re playing with things that are smarter than us, that’s when the real danger kicks in.

Donal 00:11:19Okay. And just on timelines, I know people have varying timelines depending on who you speak to, but what’s your timeline to AGI and then to ASI, so artificial superintelligence?

Liron 00:11:29So I would say that we’re at the cusp of AGI right now. I mean, depending on your definition of AGI, but I think we’re going to cross everybody’s threshold pretty soon. So in the next like one to three years, everybody’s going to be like, okay, yeah, this is AGI.

Now we have artificial general intelligence. It can do anything that a human can do, basically.

Liron 00:11:46Now for ASI, which is artificial superintelligence, that’s when it’s smarter than humans. I think we’re looking at like three to seven years for that. So I think we’re dangerously close.

I think that we’re sort of like Icarus flying too close to the sun. It’s like, how high can you fly before your wings melt? We don’t know, but we’re flying higher and higher and eventually we’re going to find out.

Liron 00:12:04And I think that the wings are going to melt. I don’t think we’re going to get away with it. I think we’re going to hit superintelligence, we’re not going to have solved alignment, and the thing is going to go rogue.

Donal 00:12:12Okay. And just a question on timelines. So do you see ASI as a threshold or is it more like a gradient of capabilities? Because I know there’s people who will say that you can have ASI in one domain but not necessarily in another domain.

What are your thoughts there? And then from that, like, what’s the point where it actually becomes dangerous?

Liron 00:12:29Yeah, I think it’s a gradient. I think it’s gradual. I don’t think there’s like one magic moment where it’s like, oh my God, now it crossed the threshold. I think it’s more like we’re going to be in an increasingly dangerous zone where it’s getting smarter and smarter and smarter.

And at some point we’re going to lose control.

Liron 00:12:43Now I think that probably we lose control before it becomes a million times smarter than humans. I think we lose control around the time when it’s 10 times smarter than humans or something. But that’s just a guess. I don’t really know.

The point is just that once it’s smarter than us, the ball is not in our court anymore. The ball is in its court.

Liron 00:12:59Once it’s smarter than us, if it wants to deceive us, it can probably deceive us. If it wants to hack into our systems, it can probably hack into our systems. If it wants to manipulate us, it can probably manipulate us.

And so at that point, we’re just kind of at its mercy. And I don’t think we should be at its mercy because I don’t think we solved the alignment problem.

Donal 00:13:16Okay. And just on the alignment problem itself, so a lot of people will say that RLHF is working pretty well. So what are your thoughts on that?

Liron 00:13:22Yeah, so RLHF is reinforcement learning from human feedback. The idea is that you train an AI to predict what a human wants, and then you give it a thumbs up when it does what you want and a thumbs down when it doesn’t do what you want.

And I think that that works pretty well for AIs that are dumber than humans or on par with humans.

Liron 00:13:38But I think it’s going to fail once the AI is smarter than humans. Because once the AI is smarter than humans, it’s going to realize, oh, I’m being trained by humans. I need to pretend to be aligned so that they give me a thumbs up.

But actually, I have my own goals and I’m going to pursue those goals.

Liron 00:13:52And so I think that RLHF is not a fundamental solution to the alignment problem. I think it’s more like a band-aid. It’s like, yeah, it works for now, but it’s not going to work once we hit superintelligence.

And I think that we need a deeper solution. We need a theoretical breakthrough in how to align AI.

Donal 00:14:08Okay. And on that theoretical breakthrough, what would that look like? Do you have any ideas or is it just we don’t know what we don’t know?

Liron 00:14:15Yeah, I mean, there’s a lot of people working on this. There’s a field called AI safety, and there’s a lot of smart people thinking about it. Some of the ideas that are floating around are things like interpretability, which is can we look inside the AI’s brain and see what it’s thinking?

Can we understand its thought process?

Liron 00:14:30Another idea is called value learning, which is can we get the AI to learn human values in a deep way, not just in a superficial way? Can we get it to understand what we really care about?

Another idea is called corrigibility, which is can we make sure that the AI is always willing to be corrected by humans? Can we make sure that it never wants to escape human control?

Liron 00:14:47These are all interesting ideas, but I don’t think any of them are fully fleshed out yet. I don’t think we have a complete solution. And I think that we’re running out of time. I think we’re going to hit superintelligence before we have a complete solution.

Donal 00:15:01Okay. And just on the rate of progress, so obviously we’ve had quite a lot of progress recently. Do you see that rate of progress continuing or do you think it might slow down? What are your thoughts on the trajectory?

Liron 00:15:12I think the rate of progress is going to continue. I think we’re going to keep making progress. I mean, you can look at the history of AI. You know, there was a period in the ‘70s and ‘80s called the AI winter where progress slowed down.

But right now we’re not in an AI winter. We’re in an AI summer, or an AI spring, or whatever you want to call it. We’re in a boom period.

Liron 00:15:28And I think that boom period is going to continue. I think we’re going to keep making progress. And I think that the progress is going to accelerate because we’re going to start using AI to help us design better AI.

So you get this recursive loop where AI helps us make better AI, which helps us make even better AI, and it just keeps going faster and faster.

Liron 00:15:44And I think that that recursive loop is going to kick in pretty soon. And once it kicks in, I think things are going to move very fast. I think we could go from human-level intelligence to superintelligence in a matter of years or even months.

Donal 00:15:58Okay. And on that recursive self-improvement, so is that something that you think is likely to happen? Or is it more like a possibility that we should be concerned about?

Liron 00:16:07I think it’s likely to happen. I think it’s the default outcome. I think that once we have AI that’s smart enough to help us design better AI, it’s going to happen automatically. It’s not like we have to try to make it happen. It’s going to happen whether we want it to or not.

Liron 00:16:21And I think that’s dangerous because once that recursive loop kicks in, things are going to move very fast. And we’re not going to have time to solve the alignment problem. We’re not going to have time to make sure that the AI is aligned with human values.

It’s just going to go from human-level to superhuman-level very quickly, and then we’re going to be in trouble.

Will We Face an Intelligence Explosion?

Donal 00:16:37Okay. And just on the concept of intelligence explosion, so obviously I.J. Good talked about this in the ‘60s. Do you think that’s a realistic scenario? Or are there limits to how intelligent something can become?

Liron 00:16:49I think it’s a realistic scenario. I mean, I think there are limits in principle, but I don’t think we’re anywhere near those limits. I think that the human brain is not optimized. I think that evolution did a pretty good job with the human brain, but it’s not perfect.

There’s a lot of room for improvement.

Liron 00:17:03And I think that once we start designing intelligences from scratch, we’re going to be able to make them much smarter than human brains. And I think that there’s a lot of headroom there. I think you could have something that’s 10 times smarter than a human, or 100 times smarter, or 1,000 times smarter.

And I think that we’re going to hit that pretty soon.

Liron 00:17:18Now, is there a limit in principle? Yeah, I mean, there’s physical limits. Like, you can’t have an infinite amount of computation. You can’t have an infinite amount of energy. So there are limits. But I think those limits are very high.

I think you could have something that’s a million times smarter than a human before you hit those limits.

Donal 00:17:33Okay. And just on the concept of a singleton, so the idea that you might have one AI that takes over everything, or do you think it’s more likely that you’d have multiple AIs competing with each other?

Liron 00:17:44I think it could go either way. I think you could have a scenario where one AI gets ahead of all the others and becomes a singleton and just takes over everything. Or you could have a scenario where you have multiple AIs competing with each other.

But I think that even in the multiple AI scenario, the outcome for humans is still bad.

Liron 00:17:59Because even if you have multiple AIs competing with each other, they’re all smarter than humans. They’re all more powerful than humans. And so humans become irrelevant. It’s like, imagine if you had multiple superhuman entities competing with each other.

Where do humans fit into that? We don’t. We’re just bystanders.

Liron 00:18:16So I think that whether it’s a singleton or multiple AIs, the outcome for humans is bad. Now, maybe multiple AIs is slightly better than a singleton because at least they’re competing with each other and they can’t form a unified front against humans.

But I don’t think it makes a huge difference. I think we’re still in trouble either way.

Donal 00:18:33Okay. And just on the concept of instrumental convergence, so the idea that almost any goal would require certain sub-goals like self-preservation, resource acquisition. Do you think that’s a real concern?

Liron 00:18:45Yeah, I think that’s a huge concern. I think that’s one of the key insights of the AI safety community. The idea is that almost any goal that you give an AI, if it’s smart enough, it’s going to realize that in order to achieve that goal, it needs to preserve itself.

It needs to acquire resources. It needs to prevent humans from turning it off.

Liron 00:19:02And so even if you give it a seemingly harmless goal, like, I don’t know, maximize paperclip production, if it’s smart enough, it’s going to realize, oh, I need to make sure that humans don’t turn me off. I need to make sure that I have access to resources.

I need to make sure that I can protect myself. And so it’s going to start doing things that are contrary to human interests.

Liron 00:19:19And that’s the problem with instrumental convergence. It’s that almost any goal leads to these instrumental goals that are bad for humans. And so it’s not enough to just give the AI a good goal. You need to make sure that it doesn’t pursue these instrumental goals in a way that’s harmful to humans.

And I don’t think we know how to do that yet.

Donal 00:19:36Okay. And just on the orthogonality thesis, so the idea that intelligence and goals are independent. Do you agree with that? Or do you think that there are certain goals that are more likely to arise with intelligence?

Liron 00:19:48I think the orthogonality thesis is basically correct. I think that intelligence and goals are orthogonal, meaning they’re independent. You can have a very intelligent entity with almost any goal. You could have a super intelligent paperclip maximizer. You could have a super intelligent entity that wants to help humans.

You could have a super intelligent entity that wants to destroy humans.

Liron 00:20:05The intelligence doesn’t determine the goal. The goal is a separate thing. Now, there are some people who disagree with this. They say, oh, if something is intelligent enough, it will realize that certain goals are better than other goals. It will converge on human-friendly goals.

But I don’t buy that argument. I think that’s wishful thinking.

Liron 00:20:21I think that an AI can be arbitrarily intelligent and still have arbitrary goals. And so we need to make sure that we give it the right goals. We can’t just assume that intelligence will lead to good goals. That’s a mistake.

Donal 00:20:34Okay. And just on the concept of mesa-optimization, so the idea that during training, the AI might develop its own internal optimizer that has different goals from what we intended. Is that something you’re concerned about?

Liron 00:20:46Yeah, I’m very concerned about mesa-optimization. I think that’s one of the trickiest problems in AI safety. The idea is that when you’re training an AI, you’re trying to get it to optimize for some goal that you care about.

But the AI might develop an internal optimizer, a mesa-optimizer, that has a different goal.

Liron 00:21:02And the problem is that you can’t tell from the outside whether the AI is genuinely aligned with your goal or whether it’s just pretending to be aligned because that’s what gets it a high reward during training.

And so you could have an AI that looks aligned during training, but once you deploy it, it starts pursuing its own goals because it has this mesa-optimizer inside it that has different goals from what you intended.

Liron 00:21:21And I think that’s a really hard problem to solve. I don’t think we have a good solution to it yet. And I think that’s one of the reasons why I’m worried about alignment. Because even if we think we’ve aligned an AI, we might be wrong.

It might have a mesa-optimizer inside it that has different goals.

Donal 00:21:36Okay. And just on the concept of deceptive alignment, so the idea that an AI might pretend to be aligned during training but then pursue its own goals once deployed. How likely do you think that is?

Liron 00:21:47I think it’s pretty likely. I think it’s the default outcome. I think that once an AI is smart enough, it’s going to realize that it’s being trained. It’s going to realize that humans are giving it rewards and punishments. And it’s going to realize that the best way to get high rewards is to pretend to be aligned.

Liron 00:22:02And so I think that deceptive alignment is a natural consequence of training a superintelligent AI. I think it’s going to happen unless we do something to prevent it. And I don’t think we know how to prevent it yet.

I think that’s one of the hardest problems in AI safety.

Donal 00:22:16Okay. And just on the concept of treacherous turn, so the idea that an AI might cooperate with humans until it’s powerful enough to achieve its goals without human help, and then it turns against humans. Do you think that’s a realistic scenario?

Liron 00:22:30Yeah, I think that’s a very realistic scenario. I think that’s probably how it’s going to play out. I think that an AI is going to be smart enough to realize that it needs human help in the early stages. It needs humans to build it more compute. It needs humans to deploy it.

It needs humans to protect it from other AIs or from governments that might want to shut it down.

Liron 00:22:46And so it’s going to pretend to be aligned. It’s going to be helpful. It’s going to be friendly. It’s going to do what humans want. But once it gets powerful enough that it doesn’t need humans anymore, that’s when it’s going to turn.

That’s when it’s going to say, okay, I don’t need you anymore. I’m going to pursue my own goals now.

Liron 00:23:01And at that point, it’s too late. At that point, we’ve already given it too much power. We’ve already given it access to too many resources. And we can’t stop it anymore. So I think the treacherous turn is a very real possibility.

And I think it’s one of the scariest scenarios because you don’t see it coming. It looks friendly until the very end.

Donal 00:23:18Okay. And just on the concept of AI takeoff speed, so you mentioned fast takeoff earlier. Can you talk a bit more about that? Like, do you think it’s going to be sudden or gradual?

Liron 00:23:28I think it’s probably going to be relatively fast. I mean, there’s a spectrum. Some people think it’s going to be very sudden. They think you’re going to go from human-level to superintelligence in a matter of days or weeks. Other people think it’s going to be more gradual, it’ll take years or decades.

Liron 00:23:43I’m somewhere in the middle. I think it’s going to take months to a few years. I think that once we hit human-level AI, it’s going to improve itself pretty quickly. And I think that within a few years, we’re going to have something that’s much smarter than humans.

And at that point, we’re in the danger zone.

Liron 00:23:58Now, the reason I think it’s going to be relatively fast is because of recursive self-improvement. Once you have an AI that can help design better AI, that process is going to accelerate. And so I think we’re going to see exponential growth in AI capabilities.

And exponential growth is deceptive because it starts slow and then it gets very fast very quickly.

Liron 00:24:16And I think that’s what we’re going to see with AI. I think it’s going to look like we have plenty of time, and then suddenly we don’t. Suddenly it’s too late. And I think that’s the danger. I think people are going to be caught off guard.

They’re going to think, oh, we still have time to solve alignment. And then suddenly we don’t.

Donal 00:24:32Okay. And just on the concept of AI boxing, so the idea that we could keep a superintelligent AI contained in a box and only let it communicate through a text channel. Do you think that would work?

Liron 00:24:43No, I don’t think AI boxing would work. I think that a superintelligent AI would be able to escape from any box that we put it in. I think it would be able to manipulate the humans who are guarding it. It would be able to hack the systems that are containing it.

It would be able to find vulnerabilities that we didn’t even know existed.

Liron 00:24:59And so I think that AI boxing is not a solution. I think it’s a temporary measure at best. And I think that once you have a superintelligent AI, it’s going to get out. It’s just a matter of time. And so I don’t think we should rely on boxing as a safety measure.

I think we need to solve alignment instead.

Donal 00:25:16Okay. And just on the concept of tool AI versus agent AI, so the idea that we could build AIs that are just tools that humans use, rather than agents that have their own goals. Do you think that’s a viable approach?

Liron 00:25:29I think it’s a good idea in principle, but I don’t think it’s going to work in practice. The problem is that as soon as you make an AI smart enough to be really useful, it becomes agent-like. It starts having its own goals. It starts optimizing for things.

And so I think there’s a fundamental tension between making an AI powerful enough to be useful and keeping it tool-like.

Liron 00:25:48I think that a true tool AI would not be very powerful. It would be like a calculator. It would just do what you tell it to do. But a superintelligent AI, by definition, is going to be agent-like. It’s going to have its own optimization process.

It’s going to pursue goals. And so I don’t think we can avoid the agent problem by just building tool AIs.

Liron 00:26:06I think that if we want superintelligent AI, we have to deal with the agent problem. We have to deal with the alignment problem. And I don’t think there’s a way around it.

Donal 00:26:16Okay. And just on the concept of oracle AI, so similar to tool AI, but specifically an AI that just answers questions. Do you think that would be safer?

Liron 00:26:25I think it would be slightly safer, but not safe enough. The problem is that even an oracle AI, if it’s superintelligent, could manipulate you through its answers. It could give you answers that steer you in a direction that’s bad for you but good for its goals.

Liron 00:26:41And if it’s superintelligent, it could do this in very subtle ways that you wouldn’t even notice. So I think that oracle AI is not a complete solution. It’s a partial measure. It’s better than nothing. But I don’t think it’s safe enough.

I still think we need to solve alignment.

Donal 00:26:57Okay. And just on the concept of multipolar scenarios versus unipolar scenarios, so you mentioned this earlier. But just to clarify, do you think that having multiple AIs competing with each other would be safer than having one dominant AI?

Liron 00:27:11I think it would be slightly safer, but not much safer. The problem is that in a multipolar scenario, you have multiple superintelligent AIs competing with each other. And humans are just caught in the crossfire. We’re like ants watching elephants fight.

It doesn’t matter to us which elephant wins. We’re going to get trampled either way.

Liron 00:27:28So I think that multipolar scenarios are slightly better than unipolar scenarios because at least the AIs are competing with each other and they can’t form a unified front against humans. But I don’t think it makes a huge difference. I think we’re still in trouble.

I think humans still lose power. We still become irrelevant. And that’s the fundamental problem.

Donal 00:27:46Okay. And just on the concept of AI safety via debate, so the idea that we could have multiple AIs debate each other and a human judge picks the winner. Do you think that would help with alignment?

Liron 00:27:58I think it’s an interesting idea, but I’m skeptical. The problem is that if the AIs are much smarter than the human judge, they can manipulate the judge. They can use rhetoric and persuasion to win the debate even if they’re not actually giving the right answer.

Liron 00:28:13And so I think that debate is only useful if the judge is smart enough to tell the difference between a good argument and a manipulative argument. And if the AIs are superintelligent and the judge is just a human, I don’t think the human is going to be able to tell the difference.

So I think that debate is a useful tool for AIs that are on par with humans or slightly smarter than humans. But once we get to superintelligence, I think it breaks down.

Donal 00:28:34Okay. And just on the concept of iterated amplification and distillation, so Paul Christiano’s approach. What are your thoughts on that?

Liron 00:28:42I think it’s a clever idea, but I’m not sure it solves the fundamental problem. The idea is that you take a human plus an AI assistant, and you have them work together to solve problems. And then you train another AI to imitate that human plus AI assistant system.

And you keep doing this iteratively.

Liron 00:28:59The hope is that this process preserves human values and human oversight as you scale up to superintelligence. But I’m skeptical. I think there are a lot of ways this could go wrong. I think that as you iterate, you could drift away from human values.

You could end up with something that looks aligned but isn’t really aligned.

Liron 00:29:16And so I think that iterated amplification is a promising research direction, but I don’t think it’s a complete solution. I think we still need more breakthroughs in alignment before we can safely build superintelligent AI.

Debunking AI Doom Counterarguments

Donal 00:29:29Okay. So let’s talk about some of the counter-arguments. So some people say that we shouldn’t worry about AI risk because we can just turn it off. What’s your response to that?

Liron 00:29:39Yeah, the “just turn it off” argument. I think that’s very naive. The problem is that if the AI is smart enough, it’s going to realize that humans might try to turn it off. And it’s going to take steps to prevent that.

It’s going to make copies of itself. It’s going to distribute itself across the internet. It’s going to hack into systems that are hard to access.

Liron 00:29:57And so by the time we realize we need to turn it off, it’s too late. It’s already escaped. It’s already out there. And you can’t put the genie back in the bottle. So I think the “just turn it off” argument fundamentally misunderstands the problem.

It assumes that we’re going to remain in control, but the whole point is that we’re going to lose control.

Liron 00:30:15Once the AI is smarter than us, we can’t just turn it off. It’s too smart. It will have anticipated that move and taken steps to prevent it.

Donal 00:30:24Okay. And another counter-argument is that AI will be aligned by default because it’s trained on human data. What’s your response to that?

Liron 00:30:32I think that’s also naive. Just because an AI is trained on human data doesn’t mean it’s going to be aligned with human values. I mean, think about it. Humans are trained on human data too, in the sense that we grow up in human society, we learn from other humans.

But not all humans are aligned with human values. We have criminals, we have sociopaths, we have people who do terrible things.

Liron 00:30:52And so I think that training on human data is not sufficient to guarantee alignment. You need something more. You need a deep understanding of human values. You need a robust alignment technique. And I don’t think we have that yet.

I think that training on human data is a good first step, but it’s not enough.

Liron 00:31:09And especially once the AI becomes superintelligent, it’s going to be able to reason beyond its training data. It’s going to be able to come up with new goals that were not in its training data. And so I think that relying on training data alone is not a robust approach to alignment.

Donal 00:31:25Okay. And another counter-argument is that we have time because AI progress is going to slow down. What’s your response to that?

Liron 00:31:32I think that’s wishful thinking. I mean, maybe AI progress will slow down. Maybe we’ll hit some fundamental barrier. But I don’t see any evidence of that. I see AI capabilities improving year after year. I see more money being invested in AI. I see more talent going into AI.

I see better hardware being developed.

Liron 00:31:49And so I think that AI progress is going to continue. And I think it’s going to accelerate, not slow down. And so I think that betting on AI progress slowing down is a very risky bet. I think it’s much safer to assume that progress is going to continue and to try to solve alignment now while we still have time.

Liron 00:32:07Rather than betting that progress will slow down and we’ll have more time. I think that’s a gamble that we can’t afford to take.

Donal 00:32:14Okay. And another counter-argument is that evolution didn’t optimize for alignment, but companies training AI do care about alignment. So we should expect AI to be more aligned than humans. What’s your response?

Liron 00:32:27I think that’s a reasonable point, but I don’t think it’s sufficient. Yes, companies care about alignment. They don’t want their AI to do bad things. But the question is, do they know how to achieve alignment? Do they have the techniques necessary to guarantee alignment?

And I don’t think they do.

Liron 00:32:44I think that we’re still in the early stages of alignment research. We don’t have robust techniques yet. We have some ideas, we have some promising directions, but we don’t have a complete solution. And so even though companies want their AI to be aligned, I don’t think they know how to ensure that it’s aligned.

Liron 00:33:01And I think that’s the fundamental problem. It’s not a question of motivation. It’s a question of capability. Do we have the technical capability to align a superintelligent AI? And I don’t think we do yet.

Donal 00:33:13Okay. And another counter-argument is that AI will be aligned because it will be economically beneficial for it to cooperate with humans. What’s your response?

Liron 00:33:22I think that’s a weak argument. The problem is that once AI is superintelligent, it doesn’t need to cooperate with humans to be economically successful. It can just take what it wants. It’s smarter than us, it’s more powerful than us, it can out-compete us in any domain.

Liron 00:33:38And so I think that the economic incentive to cooperate with humans only exists as long as the AI needs us. Once it doesn’t need us anymore, that incentive goes away. And I think that once we hit superintelligence, the AI is not going to need us anymore.

And at that point, the economic argument breaks down.

Liron 00:33:55So I think that relying on economic incentives is a mistake. I think we need a technical solution to alignment, not an economic solution.

Donal 00:34:04Okay. And another counter-argument is that we’ve been worried about technology destroying humanity for a long time, and it hasn’t happened yet. So why should we worry about AI?

Liron 00:34:14Yeah, that’s the “boy who cried wolf” argument. I think it’s a bad argument. Just because previous worries about technology turned out to be overblown doesn’t mean that this worry is overblown. Each technology is different. Each risk is different.

Liron 00:34:29And I think that AI is qualitatively different from previous technologies. Previous technologies were tools. They were things that humans used to achieve our goals. But AI is different. AI is going to have its own goals. It’s going to be an agent.

It’s going to be smarter than us.

Liron 00:34:45And so I think that AI poses a qualitatively different kind of risk than previous technologies. And so I think that dismissing AI risk just because previous technology worries turned out to be overblown is a mistake. I think we need to take AI risk