
In what sense are there coherence theorems?
Mutual Understanding · Divia Eden, Daniel Filan, and Elliott Thornley
Audio is streamed directly from the publisher (api.substack.com) as published in their RSS feed. Play Podcasts does not host this file. Rights-holders can request removal through the copyright & takedown page.
Show Notes
In this episode, Daniel Filan and I talk about Elliot Thornley’s LessWrong post There are no coherence theorems.
Some other LessWrong posts we reference include:
* A stylized dialogue on John Wentworth's claims about markets and optimization
Transcript:
Divia (00:03)
I'm here today with Elliot Thornley, who goes by EJT on less wrong and Daniel Phylin and Elliot is currently a postdoc at the global priorities Institute working on this sort of AI stuff and also some global population work. And at the end we're going to be discussing the his post on less wrong. There are no coherence theorems, which he wrote as part of the.
the case philosophy fellowship. And Daniel, you are currently doing ML, you're research manager at MATS, which is the ML alignment and theory scholars. And you also have your own AI risk research podcast. So welcome to the podcast, both of you guys.
Elliott Thornley (00:56)
All right, yeah, thanks.
Daniel (00:56)
Thanks, great to be here.
Divia (00:58)
Yeah, so I had read this post, There Are No Coherence Theorems. I missed it when it first came out and then someone linked it to me on Twitter recently and I was like, this is pretty interesting to me. It's very relevant to my interests and I also thought the discussion of it was pretty interesting. yeah, Elliot, would you mind summarizing for our audience what the post is and what it says?
Elliott Thornley (01:22)
Yeah, the sort of background context is I was reading about all this AI safety stuff, just getting into it and looking for places I felt like I could contribute. And in particular, I came across these coherence arguments. And coherence arguments are supposed to be ways in which we can predict the behavior of advanced artificial agents. you know, maybe we can't know exactly what they'll want, but maybe we can sort of
know the form in which they're wanting will take, namely that they'll be expected utility maximizers. So this means that they'll at least choose as if they assigned a real valued utility and probability to each outcome and makes choices that maximize the expectation of utility. They'll maximize expected utility in this sense. Coherence arguments are arguments for thinking that advanced artificial agents are going to be expected utility maximizers.
And the argument basically goes that if agents aren't representable as expected utility maximizers, if they don't behave in this kind of way, then they're going to be liable to pursue dominated strategies, which basically means like you present them with a series of choices or gambles or something like that. And they sort of plot away through this decision tree. They make this sequence of choices that leaves them with an outcome or lottery that they dis -prefer to some outcome or lottery that they.
could have preferred instead. And this seems like a bad consequence. If that was going to happen, then maybe it would put some pressure on you as an agent to revise your preferences. And so the thought of these coherence arguments goes that agents that are not representable as expected utility maximizers will recognize this vulnerability and so be motivated to change their preferences to the extent necessary to make them invulnerable, which will in turn make them
expected utility maximizers. And the post that I wrote is pushing it back against these arguments. So in particular, the sort of canonical version of the argument, seems to me, appeals to these so -called coherence theorems, which I define in the posters, theorems which imply sort of the result, theorems which imply that unless an agent can be represented as an expected utility maximizer,
then it's liable to pursue dominated strategies. And this seemed like kind of surprising to me. I'd sort of not heard of these coherence theorems before, like theorems that imply this particular thing before. So I went looking and I sort of read the theorems that had in various places been called coherence theorems. seems to be some like haziness or disagreement about exactly which theorems were supposed to be the coherence theorems. And I found that sort of none of the listed theorems had that particular implication.
And so this was sort of my way of pushing back against these coherence arguments and thinking like, no, actually, no, this argument isn't a good reason to think that advanced artificial agents are going to be representable as expected utility maximizers.
Divia (04:33)
Yeah, thank you so much for the summary. And is it fair in like, is it a fair summary in like somewhat less technical terms to say that the reason people think that AIs will be expected utility maximizers is because otherwise they get Dutch booked? Or is that not a not a very good summary?
Elliott Thornley (04:50)
That's a pretty good summary. like Dutch books, people kind of use it in different ways. Sometimes people use it just to mean like a particular kind of exploitation or like pursuit of a dominated strategy. But yeah, that's pretty much it.
Divia (05:05)
And then thank you for that summary. Can you also say a little bit about what your impression was about the discussion? And I think most of it's on Les Wrong, maybe a little bit of his on the alignment forum and a little bit is on the EA forum. I think this was on all in all three places. Is that right?
Elliott Thornley (05:18)
Yeah, the comments sort of all over the place, basically. Yeah, so I'm trying to remember with the comments. I think the comments were a mixed bag. So part of it was in the post, I defined coherence theorems in this particular way. So I say that coherence theorems are theorems that imply that unless an agent can be represented as an expected utility maximizer, then they're liable to pursue dominated strategies. And I sort of took this to be the
way that the term was used, at least based on my background reading. But some people said, no, that's not the way that I, like the people, use the term. They use them in some different way. And of course, if you use it in some different way, then it kind of seems silly to deny that there are no coherence theorems. if you use the term coherence theorems to refer to the von Neumann -Morgenstern theorem, like...
looks like I'm denying the existence of the von Neumann -Morgenstern term, which would be a kind of a silly thing to do. So yeah, that kind of comment, I feel, wasn't so useful. I guess maybe I sort of invited this kind of comment with the title that I gave to the post. But the sort of the main point remains like despite that comment, which is like, this coherence argument doesn't work. You for the coherence argument to work, you need this claim that
unless an agent can be represented as an expected utility maximizer, then it's liable to pursue dominated strategies. My point is that there aren't any theorems which imply that claim and sort of that remains true no matter how you define the term coherence theorems. So I feel like those lines of comments weren't so productive. There were some more productive lines of comments. I don't know if you want me to talk about those or we should hold off for a moment.
Divia (07:15)
Yeah, I think if you could briefly say something about them and then I have a few questions about why this is important.
Elliott Thornley (07:22)
Yeah, good. So, there aren't any theorems which imply that unless an agent is representable as an expected utility maximizer, then they're liable to pursue dominated strategies. That's not to say that you can't argue for that claim. So, a theorem, as I understand it and as I use the term, it's kind of like an argument with no undischarged assumptions. You sort of don't have to rely on any premises to get the conclusion.
Divia (07:38)
Mm
Elliott Thornley (07:51)
But if you are happy to bring in some premises, then you can argue for the conclusion. You can say, unless an agent is representable as an expected utility maximizer, and if that agent satisfies x and y and z, then they're liable to pursue dominated strategies. it's like, expected use
Divia (08:14)
Which is the way the VNM theory that is actually formulated, right?
Elliott Thornley (08:18)
Not quite. So, Von Neumann's Morgenstern Theorem, and this is important, makes sort of no reference to money pumps or exploitation or Dutch booking or anything like that.
Divia (08:20)
Right? Okay.
I thought it doesn't have premises about like independence and like that sort of thing.
Elliott Thornley (08:35)
Yeah, so it's purely a representation theorem. So there are four von Neumann -Morgenstern axioms, independence, continuity, completeness, and transitivity. And the theorem says that an agent is representable as an expected utility maximizer if and only if it satisfies the four axioms. And then I guess like, so you're like kind of right because often people go on to defend
one or more of the axioms by means of like money pump arguments, sort of like motivating saying, you must satisfy independence, otherwise you're liable to pursue dominated strategies. So it's like money pump arguments plus the VNM theorem can sort of make a coherence theorem or a coherence argument. Yeah, but then my point in the post is like the money pump arguments part of that equation.
Divia (09:07)
Thank
Mm
Got it.
Elliott Thornley (09:32)
rely on these premises that are sort of contentious in the normative case. And I think like very likely false in the descriptive case, the case that we're interested in, when we're talking about what artificial advanced artificial agents will be like.
Divia (09:47)
All right, thanks. And just to check, Daniel, did you read this post when it came out?
Daniel (09:50)
I think I got... So, okay, I'm going to be real. This post sort of triggered me. I read the title and I was like, I think I've seen some coherence theorems. And then I think the way you put it in the post is like, you know, a coherence argument, has to have like no substantive assumptions. Sorry, a coherence theorem has to have no substantive assumptions. And I'm like, well, that's not theorem. The way I read that, I was like, well, theorems always...
or the form A implies B or whatever, you assume A and you get B, and that's what proves A implies B. And so then I gave up on your post, which I think it was wrong. So I think like, I did read it later. I actually have not. It's been a little, it's been some time since I've read it, I'm gonna be honest. But the, yeah, maybe to say a little bit more. I think like when I originally read,
Divia (10:39)
Yeah, it is.
Daniel (10:48)
something like the introduction or first few paragraphs of Eliot's post. Like I thought it was going to be something like, you have to have some minor assumptions in order to prove these like coherence arguments. You know, like you have to assume something like, you live in a world where like you have choices or whatever. And I was going to be like, well, okay, like whatever. I'm those are good assumptions to make. So it's fine to make those assumptions and they still deserve name coherence theorems.
But I think the actual state of play is there are these von Neumann Mortgage Trader axioms, right? There are like these claims about what your preferences should be like. And the von Neumann Mortgage Trader theorem says that if your preferences are like that, then you're an expected utility maximizer. And the state of play is that if you assume like one of those von Neumann Mortgage Trader arguments, one of those axioms rather, then you can make these money pump arguments for the other axioms, right?
If you have this axiom, then if you don't obey this axiom, you're to shoot yourself in the foot or whatever. But you actually can't get there from zero of the axioms about your preferences. And so basically like,
Divia (12:00)
And sorry, for people who can't see the video, Ellie, you're nodding, right? That seems about right to you. Okay.
Elliott Thornley (12:03)
Yeah, yeah,
Daniel (12:06)
Yeah, and so the thing that seems to be true that I did not realize for a while, even after Elliot wrote his post, that the actual, the arguments, the things you need to rely on to get expected utility maximization, or the things you need to rely on to say that if you're not an expected utility maximizer, then you shoot yourself in the foot, they're just like seed -affirmatively more substantive assumptions than you might have guessed just based on the zeitgeist or just based on the way people
at least talked about these things one year ago. Yeah.
Elliott Thornley (12:40)
Yeah, I think that's great way to put it.
Divia (12:42)
Yeah, thanks. Okay, and now just to talk about another thing, we touched on this a little before we started recording, but Elliot, what was your motivation for writing the post? There's sort of, I think you say a bit about it in the post about why it seemed important, but can you lay that out for our listeners?
Elliott Thornley (13:01)
Yeah, yeah, so part of my motivation for writing the post was to sort of point out this thing that seemed like a mistake to me and a mistake about something important. Yeah, so in particular, coherence arguments appeared in this big Katya Grace post, I think it's called like counter arguments to the basic AI ex -risk case. And it sort of place in that post made me think
Divia (13:27)
Bye.
Elliott Thornley (13:31)
that people are taking coherence arguments as at least like a moderately important part of the basic case for existential risk from AI. And so it seemed important to point out this weakness as I perceived it in these coherence arguments. I say only a moderately important part because I take the point of coherence arguments to be that showing that advanced artificial agents are going to be
sort of goal -directed in a sort of concerted and concerning way. And there are other reasons to expect that besides coherence arguments. So in particular, you might think that like AI labs are in fact going to train agents to be goal -directed in some concerting way because they'll be more valuable or they'll be better able to solve problems or do things in the real world. So like only moderately important for those reasons I think coherence arguments are.
But I think, as maybe we'll get into later, I think coherence arguments are very important for another reason, which is something like this. So, you know, we've got these alignment proposals, proposals for keeping agents aligned or shut downable or corrigible or something like that. Some of them, including my own, rely on being able to create an advanced artificial agent that's not representable as an expected utility maximizer. And if you sort of...
bind coherence arguments, might think these kinds of alignment or shutdown ability proposals can't even get off the ground because as soon as your agent is sort of reflective enough to realize that it's vulnerable to pursuing all these dominated strategies, and in fact, you know, the claim is that it is vulnerable to pursuing all these dominated strategies, then it's sort of going to turn itself into an expected utility maximizer, and you're going to lose the property that kept it aligned or shutdownable. So it's like that second thing that
See, it's the main importance of pushing back on these coherence arguments, which is like making space for these proposals that rely on agents not being expected utility maximizers.
Divia (15:40)
Cool, yeah, so I think, let me try to summarize something about why this matters as I understand it. So one thing is, and I'm actually, I did look into this a little. didn't find sort of, like Daniel, you were talking about the zeitgeist. Like I think that there's sort of an impression of, yeah, these coherence theorems, of course it'll have a utility function. Maybe more than people have actually really said that anywhere.
precisely or maybe they have said it and I couldn't find it. So I'm a little bit confused about that point. But I think it's true that a lot of people, I think I had the impression like, yeah, at least people seem to think that it's gonna be some sort of expected utility maximizer. Yeah.
Daniel (16:21)
Yeah, I can maybe I say something about this. was like, so the reason that I sort of got, I don't know, developed more thoughts about this topic is that in the second half of last year, I, well, okay, first what happened is I read this comment on less wrong by someone who was complaining about like how people think that AIs are going to be expected utility maximizers. And I was like, guys, we've like proved, you know, there are proofs like in places and I'm just, I'm going to give a talk and I'm just going to like.
Divia (16:45)
Right.
Daniel (16:49)
collect all the proofs, I'm just going to make a really solid argument, and then that person will have to shut up from now on. And basically, yeah, I mean, I was under the impression that this would be really solid. And I still think you don't need that much, like the assumptions you need to get expected utility maximization.
On the one hand, they're not huge, but you only get like, you get a surprisingly weak form of expected utility maximization. And like, in order for me to make the argument that I wanted to make properly, I ended up like basically hypothesizing that, you know, this agent is going to like, try and narrow, try and like ensure that the state of the world and the future is in some like small set, which might, which you might think is the kind of assumption that you wanted to like, prove instead of the kind of thing you wanted to prove instead of assume.
But anyway, point being, I definitely assumed that the arguments here were better than they turned out to actually be.
Divia (17:51)
Yeah. Okay, so that's one thing. It seems like a bunch of people, regardless of what exactly anyone said, which I'm not sure about, a bunch of people assume that they're really solid, like, then we proved it with math type arguments for that the AI is going to be expected utility maximizer and that at least the three of us seem to agree, not true, right? And I read the comments and my impression from the comments is also like, think nobody has any real counter to that thing. And yeah, and this matters. I mean, I think it sort of matters for its own sake, part of what it means.
I know if you, I certainly identify as a rationalist. think Daniel does, Elliot, I don't know if you do, but that we care about what's true just for its own sake. And yeah, in so far as it's something that people say as part of the AI risk argument, that matters. Though again, I think all three of us are here are like, okay, but we're not trying to say like, look, we disproved it, AI is not risky. Like I think we all think probably still is risky and.
At least I will say the basic argument of like, you make something smarter than people and you have all these companies that are like, I'm going to try to connect it to the internet and have them do as powerful things as possible. I don't know. It seems kind of dangerous. I think that argument for sure is still, is still there. But then also what you're saying, Elliot is look, but if we are going to try to actually make proposals for making AI aligned and making AI safe and try to game out how this will work, then, then no, really does matter to try to pin down what we do and do not know.
about AIs so that we can have proposals that, like we can evaluate which proposals seem promising or not. Is that, is that everything I said seemed right?
Elliott Thornley (19:24)
Yeah, yeah, that's right. Yeah, so at least in so far as you mistakenly think that these coherence arguments are rock solid, you're sort of from the from the very outset ruling out this whole class of proposals, which might seem promising. And in fact, after writing this post, I sort of looked into incomplete preferences in more detail and thought that, yeah, proposal along these lines does seem promising. So in particular, like,
You don't really need to get into the whole thing here, but one reason you might think that incomplete preferences are promising is an agent with incomplete preferences can just sort of be more chill than an expected utility maximizer. So in particular, can like lack a preference between many more pairs of options. And you might think that we want this in so far as we want our artificial agents to be more sort of like chill and lacking preferences.
it's useful because like, if you lack a preference between A and B, very likely you're not going to like pay costs to shift probability mass between A and B. you know, two identical cans of, Coca -Cola, you lack a preference between them. And so you don't like, pay a dollar to get the right one with probability 0 .9 rather than like get the left one, the probability 0 .9. And so, you know, if we can create these agents with
Daniel (20:23)
Maybe.
Elliott Thornley (20:51)
incomplete preferences, we can create them such that there are sort of many more things such that they're unwilling to pay costs to shift probability mass between those things. And so sort of like, you get a sort of more chill artificial agent in this respect.
Divia (21:07)
And am I right that this, part about the agents with incomplete preferences and that motivation, that was not in the original post, right?
Elliott Thornley (21:14)
no, this is all like later thinking.
Divia (21:17)
Cool, all right. Yeah, because I do think some people, like maybe Daniel, what you're saying, saw it as more like, you're trying to nitpick and why does it even matter? Like, I think this is not a very virtuous way to, sorry Daniel, as far as I'm you, it doesn't seem like a totally virtuous way to read it. Like, well, why does it matter? Because the basic conclusion is probably right. And I think it's fine in the sense like everybody has limited time and.
Daniel (21:31)
Yeah.
Divia (21:43)
and whatever, but in terms of like actually responding to it substantively, I don't know, maybe this is just a little of my agenda. I'm like, yeah, often things matter for reasons that people hadn't even thought of at the time. And your thing about incomplete preferences seems like one of them, maybe.
Elliott Thornley (21:58)
Yeah, that's right. Although I do want to point out something that Daniel said earlier, which I think is also right, which is insofar as you're relying on these money pump arguments for the von Neumann -Morgenstern axioms and taking those two things together as like your coherence argument, it's important to note that at least the of the sort of most up -to -date, most sophisticated, weakest assumptions, money pump arguments
Divia (22:18)
Mm
Elliott Thornley (22:26)
all get found in this book by Johan Gustafsson from 2022 called Money Pump Arguments with Daniel Hens. Yeah, and the way that book is structured, he presents the money pump argument for completeness first, and then uses completeness to create the money pump for transitivity, uses completeness and transitivity, create the money pump for independence, and then uses all three, I think, maybe, maybe just like one of them to create the money pump for continuity.
Divia (22:30)
Good. That was the book here. Nice.
Elliott Thornley (22:55)
And so in that respect, like the money pumps are kind of like a house of cards where if you don't have the completeness money pump, if you're not compelled by that one, then it's hard to like get the other money pumps going as well. And this might be another reason to think that these coherence arguments are important besides like making space for these alignment proposals that depend on non -expected utility maximizers.
Divia (23:19)
Also, are you familiar with Scott Garment's work, like his geometric rationality sequence?
Elliott Thornley (23:25)
No, I haven't read that unfortunately. It's on my list.
Divia (23:26)
Okay, well, we'll leave that aside. Maybe for a different day though, as like a footnote, I found his post where he gave an example for why he's not compelled by independence to be, it certainly stuck with me and I was like, yeah, okay. And so before I had engaged with what you wrote, when people were always like, okay, but what about V I was always like, well, I'm not persuaded on the independence point. But yeah, anyway, I'll leave that aside. Okay, so yeah.
I think I'm hoping at this point for Danielle, you and Elliot to talk a little more about what in practice, what you do expect and why in terms of which in terms of the expected utility maximization stuff.
Daniel (24:10)
Yeah, it's a little bit hard for me to say this because if you take... There's this question, what does it mean to be an expected utility maximizer? Literally, what are we saying when we say that? And it means something like, for each state of the world or for each way things could be, there's some utility value attached to that. And for each thing you could do, you consider... Or there's some probabilities of...
various outcomes, and you basically reliably pick the thing that maximizes the expected value of the utility. So the expected value being like sort of probability weighted average, right? I think there are various versions of this. Like, do you have to be like thinking about the probabilities in your head or whatever? I'm a bit less concerned about that. But like one thing about this is that
The thing we're assigning utilities to is, I don't know, maybe something like states of the world or something like that. And if you're allowed to be like very, very fine grained about what counts as a state of the world, you can really just, like expected utility theory can be super expressive, right? Like if you're allowed to distinguish between tons and tons and tons of states of the world, then maybe like expected utility theory constrains you very little because you just have tons and tons and tons of utility functions that you can optimize the expectation of.
So all of that is to say, like, what do I actually expect with maximizing expected utility? I'm like, I think...
Divia (25:43)
Wait, so I can summarize that in case people didn't Maybe there's some sense that people have before they really think about it that, if you have a utility function, then you must be sort of like an act utilitarian where you're, I don't know, trying to do the good for the greatest number or at least for yourself in some kind of pretty naive seeming way. And you're like, no, it really doesn't say that at all. You could, you could formalize almost anything this way.
Daniel (25:46)
Hmm, sure.
Yeah, I think that's right. if you... Yeah, not quite everything.
Divia (26:15)
You could be like, assign super high. Yeah, but like I could assign super high utility to having this first and then that later for these reasons. And I could super low utility to having what seems like the same thing, but at a different time for different reasons. Like that type of thing.
Daniel (26:22)
Yeah.
Yeah, think the one thing that it does rule out is you've expected utility theory, if you're justifying it by money pump arguments, it does say something like, for some notion of resources in the world, it does say something like you don't give up resources for literally nothing. So you don't choose to have just less rather than more, but with literally all else held equal. But like,
What counts as all else held equal? Well, it sort of depends how strictly you want to model it. Yeah, or like, I don't know. You look at the phonon and mortgagetern axioms, like completeness, transitivity, independence, and continuity. And they're very, like they look really minimal. They look like they're not assuming all that much. And then if you read the phonon and mortgagetern theorem, you might get the sense that like,
Divia (27:23)
You say what they are, again, for people's benefit.
Daniel (27:26)
Yeah, so, yeah, let's go through them. So completeness says.
Divia (27:31)
And Elliot, I would say feel free to weigh in on these as you wish.
Elliott Thornley (27:34)
Okay.
Daniel (27:35)
Yeah. So let's start with completeness because it sounds like it's basically nothing, as I guess Elliot has written about, it's actually kind of non -trivial. So completeness says that for any two options, A and B, you either prefer A to B, or you prefer B to A, or you're indifferent between A and B, right? And so you might be thinking, wait, how is that even an assumption? Isn't that just like, isn't indifferent
doesn't indifferent just mean you don't prefer one to the other? Yeah, but it's not, but it's not because the way that you want to, or at least for the purpose of money pump arguments, guess, the way that you want to define indifferent is that if you're indifferent between like A and B and you make A slightly better, then you prefer A. Or if you make A slightly worse, then you prefer B.
Divia (28:08)
Right, so it seems almost like a tautology at first.
Daniel (28:34)
So here's an example of a way you could fail to satisfy completeness. Suppose you're asking yourself, should I become a doctor or should I become a monk? And you're like, man, I have really no idea. I just have no concrete idea of which one of those I should do. And then suppose I tell you, actually, when you were thinking about becoming a doctor versus becoming a monk,
You're using slightly out of date numbers for doctor's salaries and the salaries of doctors are actually like 3 % higher than you realized.
Divia (29:10)
Or assuming you're like, okay, well then, so if I give you this dollar to become a monk, then you will, and you'd have to see this. Yeah. Right, which it does not, that's not really how people, like that's not really how human beings behave, right?
Daniel (29:14)
Yeah, yeah, yeah. So that would...
that...
Divia (29:28)
I guess there's a question about that.
Daniel (29:28)
that's probably, yeah, there's some, it seems like it's not how human beings behave. So that's completeness. Transitivity says that if you prefer A to B and if you prefer B to C, then you also prefer A to C.
Independence basically says that suppose you prefer A to B, right? Like if you could choose between A and B, pick A. Independence says suppose that you've got some probability of C and some probability of A. That's one option. Or you can get the same probability of C and some probability of B, right? maybe like, basically it says you prefer maybe C, maybe A.
to maybe C, maybe B. So one example of this is like, suppose that you prefer eating pizza tonight than eating Thai tonight. And then like someone like basically says, hey, I'm do a coin flip, right? If the coin comes up heads, then you're definitely getting Mexican. If it comes up tails, then like maybe you're gonna get pizza and maybe you're gonna get a...
Thai, whatever the other one I said was. Yeah. But like right now you've got to decide like what's going to happen if the coin comes up tails, right? And independence says that like, if you prefer pizza to Thai, then right now you've got to say that like, if the coin comes up tails, you want the pizza world rather than the Thai world. So that's independence.
Divia (30:44)
tired.
Which through the record, I think this is often actually not true for sort of fairness reasons, given that agents I think are often best modeled as having something like internal conflict or like internally at least different preferences, which was basically the argument as I understand it that Scott Garibrand made in his post in the geometric rationality sequence.
Daniel (31:25)
Yeah, a lot of, you're not the only one to not, yeah, a lot of people don't like independence. I think it's actually pretty good, but people disagree about this. And then finally there's continuity. And continuity basically says that like small enough probabilities don't matter. So basically all these axioms are about like, you know, what you choose when your options are like having certain probabilities over just different outcomes, right? And
I forgot what continuity actually says, but it roughly says that like, if I prefer like this, you know, if I prefer this doing this thing, which has like some probabilities over various outcomes versus this other thing with some probabilities over various outcomes, there's some like tiny amount by which I can change the probabilities so that my preference remains the same. So if I'm like, I suppose someone offers me a gamble and I want to take that gamble over doing nothing.
If somebody says like, the probabilities are actually 0 .001 % different than what I said, then I'm still happy to take the gamble.
Elliott Thornley (32:32)
Yeah, think the continuity in the VNM theorem is slightly different, or at least like it's expressed slightly differently. basically, it goes, if you prefer A to B to C, then there's some combination of A and C that you prefer to be, and some combination of A and C that you dis -prefer to be. So like, yeah.
Divia (32:32)
Thanks. Appreciate it.
Daniel (32:41)
okay.
right.
have the impression that those are like, that you could do one of those assumptions or the other, like there are a different versions of continuity that can get you the same results. Is an impression.
Elliott Thornley (33:07)
Yeah, I wouldn't be surprised if they're equivalent, actually. yeah, at least in the Gustafsson book and in the V theorem, that's how it's expressed.
Daniel (33:18)
Anyway, so that's what those assumptions are, right? And, well, if you thought those assumptions were uncontroversial or something, you might think that like, the VNM theorem is like saying that if I, if I obey those assumptions, then there's actually a whole bunch of other things I have to do with my behavior or whatever to like, you know, to be consistent or whatever. But actually, like, as Elliot said earlier, the VNM theorem is a representation theorem. It says that like, if you satisfy those assumptions, then
there's some utility function for which you're maximizing it. And so it's, in some sense, utility maximization is just as weak as those four assumptions. And those four assumptions, don't tell you that much. Anyway, all this is a tangent from what do I expect in terms of utility maximization. And I'm like, yeah, there's probably going to be some sense in which AIs are going to be utility maximizers, because I think, probably because I coherence arguments are kind of good.
and partly they're persuasive to me. think that like the assumptions are stronger than one might think, but they're still like relatively, I think they're weak enough that I think they're basically right.
Divia (34:20)
Like they're persuasive, like they're...
Daniel (34:35)
Yeah, but I don't know. I'm sort of in the place where like, will AIs be expected utility maximizers? What do I expect with about that? I'm like, it feels like a weird question to ask to me.
Elliott Thornley (34:49)
Yeah, yeah, I think one thing that's worth emphasizing and that Daniel and Divya, you talked about earlier as well, which is like, if you sort of place no restriction on how you innovate, individually outcomes, no restriction on the objects of preference, that any behavior can be rationalized as expected utility maximization. So in particular, like the classic example of
Divia (35:13)
Maybe there's some trivial sense in which I'm like, the exact thing I did is worth a lot and everything else in my utility fraction is worth zero. Is that one way you can do it?
Elliott Thornley (35:24)
Yeah, that's exactly right. So yeah, if you want to get some like non -trivial prediction out of coherence arguments, then you need to like place some restriction, at least probabilistic on like how you individuate outcomes. You've got to say like, no, this artificial agent doesn't just care about like the entire history of the universe and there's sort of no more structure to their preferences than this like ordering over histories of the universe. Actually they care about
ice cream flavors or slices of pizza or something like that. And they're indifferent between any two histories of the universe that are the same with respect to slices of pizza. And if you sort of make that restriction, then you can get some predictions out of coherence arguments. Because then you can sort of think, okay, well, you the agent is trading sausage pizza for pepperoni and then they're pepperoni for mushroom and then they're trading
mushroom for sausage, and they're sort of paying to choose in this cycle. And they're doing this dominated strategy. And so now they're going to like revise their preferences and remove the cycle and things like that.
Divia (36:31)
and then they get money pumped.
Yeah, so can one of you describe like, do you think people, because again, there are these money pump arguments. So how far, or like, what do think the practical implications of the things people say about money pumping are? Or are there none? I think there's some, right?
Daniel (36:57)
mean, you utility theory is a really nice language. Like if you, if you give yourself utility functions and if you, if you're allowed to just say, yeah, my agent is going to be an expected utility maximizer and the utilities are going to be over states as defined in this way, then it becomes very, it becomes nicer to approve a variety of theorems. Right. So like, like utility functions, their functions over things, they're continuous. can like optimize things with respect to those functions. You can like vary the functions smoothly to
Divia (37:01)
Mm
Daniel (37:27)
I don't know. I mean, one -up shot is just a very nice modeling language. In terms of actual substantive risk arguments from modeling things as expected utility optimizers, I'm honestly not aware. Maybe Elliot has some in mind. I guess there's things like shutdown ability under some assumptions of
Yeah, maybe Elliot should go here.
Divia (37:56)
Okay, no, but there's something else that I'm trying to ask first, though I do want to get there that's like, okay, I, as a person, I definitely have some intuition of like, okay, yeah, I don't want to be that thing. Like it's sort of incoherent for me to be that thing that has those cyclical preferences about the pizza toppings because I don't want to lose all my money. And so there's some sort of intuition there that I think people tend to then expand. And I probably historically have sort of expanded it and I could try to speak to it, but I'm wondering like,
Do you guys know what I'm talking about with this sort of expanded intuition? Like, no, but surely I've got to kind of be like this or else.
Elliott Thornley (38:33)
Yeah, this sounds right to me. In particular, think it was the Eliezer Yudkowski post that I read. think it's called Coherent Decisions Imply Consistent Utilities or something like that, where it does this kind of thing. You give the example of the person that fails to be an expected utility maximizer by having cyclic preferences. They prefer A to B, B to C, C to A. This is really sort of like the
classic money pump, the money pump that really works if any of them do, because then you can sort of say, all right, pay me $1 and I'll switch you from A to B, pay me $1 and I'll switch you from B to C, pay me $1 and switch you from C to A. Yeah, so I think that money pump kind of basically just works. There are some technicalities, but like, it's pretty convincing to me. And so I think you shouldn't have cyclic preferences. I think the error is in extrapolating from that to the whole like
expected utility maximization, all the von Neumann, Morgenstern axioms. Because the money pump for completeness, I think in particular, is not nearly as convincing as the money pump for acyclicity, the one that says you shouldn't have cyclic preferences.
Daniel (39:46)
Yeah.
Divia (39:46)
Can you lay that out? what is, I don't think I actually know what the money pump is for completeness.
Elliott Thornley (39:52)
Yeah, okay, so I'll give the sort of non -forcing version first and then talk about the second one that Gustafsson talks about, which is supposed to be the forcing version. recall that completeness or having incomplete preferences is about having preferences that are like insensitive to some sweetening or souring, in particular, lacks of preference that are insensitive to some sweetening or souring.
Divia (40:20)
So like, I'm, yeah, I said, don't know if I want to be a doctor or a monk and then you offered to pay me a dollar to be a doctor. I'm like, yeah, I still don't know. Actually. I refuse to. Yeah.
Elliott Thornley (40:28)
Yeah, exactly. And the money pump for completeness, the non -forcing one that we'll talk about first, sort of uses this fact. So the first choice is between a doctor or monk, and you're stipulated to lack of preference. So you can choose monk at this point. And then at some later period of time, you know that you'll be offered the choice between sticking as a monk or switching to a slightly worse paid doctor than you were before, the one that you, the first choice that you had.
And then if you have incomplete preferences and in particular you lack preferences between both careers as a doctor and the careers of monk, but you prefer to be a better paid doctor to a lower paid doctor, then it seems like you could be money pumped in this situation because what you could do is you could decide to choose monk at node one instead of higher paid doctor and then later change your mind and choose
lower paid doctor at node two instead of sticking with the monk. And then you'd be money pumped in that case. And this is like.
Divia (41:30)
Okay.
Well, is it like infinitely because people could keep doing this? I get it. From my perspective, money problems don't seem that compelling unless it's infinite.
Elliott Thornley (41:44)
Yeah, good. So this is one thing to consider, which is like, yeah, the money pump for completeness is less compelling than the one for acyclicity exactly for this reason, because you couldn't sort of extract infinite money out of someone. Or at least like from the bare fact that they have incomplete preferences, you couldn't extract infinite money out of them. If they had like extremely incomplete preferences, maybe you still could extract a lot.
Divia (41:56)
Okay.
Sure, but if people are like, okay, we're gonna keep switching you from monk to even worse paid doctor, then I would think at a certain point, I'd be like, okay, well now that you're asking me to pay to become a doctor, like now I'm out. So you can't keep doing this, right?
Elliott Thornley (42:20)
Yeah, yeah, that's right. Yeah, but okay, so the main way in which I think this money pump for completeness isn't particularly compelling is it relies on this premise that Johann Gustafsson calls decision tree separability. And decision tree separability basically says that you can ignore parts of the decision tree that are no longer accessible. So in particular, like
If xNihilo you had the choice between monk and lower paid doctor, you'd lack a preference and so would maybe choose like each with some positive probability. And so since you'd do that thing xNihilo, you'll also do that thing if you previously turned down the option to be a better paid doctor.
Divia (43:12)
I see. But you're saying, no, you can just not do that. You can be like, well, I turned it down before, so I'm going to sort of stick with that.
Elliott Thornley (43:19)
Yeah, yeah. So this is the proposal for like, how you avoid this money pump for completeness is by denying decision tree separability. And actually, okay, so the sort of complication is that Gustafson is arguing about whether being representable as an expected utility maximizer is rationally required. Whereas in a, say again,
Divia (43:40)
What is it?
What does that mean, if it's rationally required?
Elliott Thornley (43:45)
Yeah, it's kind of like, it's, we're getting into like normativity and stuff. It's like a requirement of rationality, you like prudentially ought to be a von Neumann Morgenstern accent.
Divia (43:56)
Just like, prudentially, I ought to not have cyclic preferences, we might say. Is it in that sort of sense? Like, is that sort of how you mean it?
Elliott Thornley (44:00)
Yeah, yeah, it's like, I'm gonna go now. I will say that again.
Yeah, it's like arguments about what we prudentially ought to do or be. It's like rationality requires that you are.
Divia (44:13)
Okay. And so this is not like a mathematically rigorous thing. Or is it?
Elliott Thornley (44:18)
Well, it's kind of like, it's a way of interpreting this claim of decision tree separability, which might affect how compelling you find it. You know, we come into theorizing with some intuitions about what rationality requires. And maybe we think that rationality requires that we satisfy decision tree separability, that like, our decision shouldn't depend on
parts of the decision tree that we can no longer access.
Divia (44:49)
which some people have an intuition about that or some sort of moral impression about that.
Elliott Thornley (44:54)
Yeah, I think that's basically what it comes down to. And my point that I make in the post is like, know, decision tree separability interpreted as a claim about rationality is somewhat contentious. But when we're thinking about how advanced artificial agents will in fact behave, we need this like analog of the premise, which is instead claims about how artificial agents will in fact behave, namely that they will in fact
Daniel (44:58)
Yeah.
Elliott Thornley (45:24)
ignore parts of the decision tree to which they no longer have access and behave the same no matter what happened in the past. And this claim is very easy to doubt, right? So if you create an artificial agent that in fact modulates its behavior depending on what happened in the past, then you can falsify this claim. And so you could quite easily create this artificial agent that failed to satisfy decision tree separability and thereby avoid the money pump for completeness.
Divia (45:34)
Yeah.
Elliott Thornley (45:54)
yep.
Daniel (45:55)
Can I maybe use this as a jumping off point to say a slightly weird thing about this literature, which is that you might have thought that money pump arguments, or you might have thought that the way you should talk about for Norman rationality or whatever, is some theorem like, if you don't do this, then a bad thing will actually happen to you. And maybe you might think that the theorems would be something about what kinds of things agents actually do.
And if you don't actually do this thing, then a bad thing will in fact happen to you. And like, I think some arguments are kind of like this. And then some arguments are basically talking about, you know, the object of von Neumann rationality is just like your preferences, which are just things inside your head about like how you rank various outcomes that might not even be things that you actually do. And then the theorems are like, well, if you, you know, if your preferences inside your head are arranged a certain way.
then some preferences will depend on other things. But that's a crazy way for the inside of your head to be arranged. And that's like really bad. And, yeah, they'll use these words that send you to dictionary, like cognitive attitudes, which is a word for like cognitive, which I had to look up. means like relating to your desires, I believe. Yeah, so C -O -N.
Divia (47:12)
So, what attitudes?
We'll work on that later.
Hm. Can you spell that? Sorry. Yeah. cognitive. OK. Interesting.
Daniel (47:22)
A -T -I -V -E, if I recall correctly. Anyway, and so like, from my point of view, if I'm trying to think about AIs, I sort of want to think like, I want to say something like, if they aren't, or I don't know if I would like to say this, but the kind of theorem that I'm interested in hearing is if they don't actually do expected utility maximizing behavior, then a bad thing w