Total research transparency feels a bit too galaxy-brained for me. It makes non-robust assumptions that newly discovered techniques won't be usable to enhance already existing open-weights models to excessively dangerous capability levels. I also think the disincentive for research is overstated as it neglects first-mover advantage.
I like that the space governance plan supplement acknowledges that defining torture and slavery are hard problems that ASI could potentially help with:
It might take non-trivial reflection by advanced ASI to figure out the precise bounds of this.
And in the sidenote it gives examples of difficult subproblems or dependencies that would need solved in order to do this, like "What exactly are negatively-valenced experiences?" But I have a couple of problems with the rest of the scenario given this:
Thanks for the comment!
1. I think the AIs might be philosophically competent enough to solve ~all the problems, and using the AIs to solve them is basically the right move. We try to make this clear at the start of the epilogue. We wanted to be somewhat more concrete in the scenario than meta level solutions like this, and try to sketch out in concrete detail what the solutions might actually look like, which is why we didn't just stop at "handoff to the AIs", though I do think in practice we should mostly be handing off to the AIs at this point.
2. I agree that making the AIs this philosophically competent may be very difficult and not happen by default. I think this is indeed a big concern, and I wish we'd written (and thought) about this more carefully. The place we've written the most about this is here: https://ai-2040.com/supplements/alignment-roadmap#phase-4-handoff (which TBC is written from a Plan C perspective, assuming much less lead time than Plan A, and is the supplement from which we linked your post). In Plan A, the story is largely that we have a huge amount of time with ~human level AIs, and many people have access for many years, and those people will (hopefully) make progress on this.
Thanks for sharing, very interesting.
One thing that jumps out at me is that Total Research Transparency is to some extent the opposite of cybersecurity hardening to prevent hacking. The fact that there are plausible arguments for both suggests to me that we have a lot of uncertainty about what policies we should be pursuing. And this in turn would seem to suggest that we should favour corrigible plans that can be amended later, which seems like an argument against Total Research Transparency: once we have enacted it, we can no longer put that particular genie back in the bottle.
On the epilogue: I guess I'm pretty unconvinced by the idea that people who don't care much about/are mildly interested in their potential space properties will just sell off their tickets, or even be able to. You're essentially flooding the market with capital by giving everyone what is, in expectation, a 10 billionth of the lightcone. I'm not sure there'd be enough money on earth to make more than a small minority of people sell off their tickets, even if those people don't particularly care much (presumably they care somewhat, though not necessarily in any strong way) about what happens in their slice of the universe, this ignoring people who want to keep their tickets so that they can be worshipped by new life on their part of the universe, or because they want to have some sort of space harem.
Besides that, like 1/4 of the world is muslim, and many more people are religious fundamentalists of other sorts, or engage in some other, secular blend of fanaticism. I remain unconvinced that passing over large fractions of the universe to these people is a good idea.
I remain unconvinced that passing over large fractions of the universe to these people is a good idea.
the scenario has "no slavery, no torture" rules. It doesn't specify but you might enforce some kind of exit rights.
I think it's harder to define a rule about "no brainwashing people", but I think basically either you solve that, or you don't, and either way whether or not the cosmic commons is seeded with any particular culture is small potatoes compared to how memetic evolution goes over trillions of years.
Certainly laws like "no slavery, no torture" are nowhere near sufficient for this, nor are they particularly well-defined. But, even ignoring that, the loss of otherwise-possible value from these areas is still incredibly significant!
Yeah, but, the most obvious alternative is "instead of principled liberalism, all out memetic war for the future." (and betting that whoever wins is actually better than principled liberalism.)
(it'd also be pretty surprising to me if dogmatic settlers stayed the same particular flavor of dogmatic for more than a couple thousand years (really more than a few hundred)
There used to be wars all the time about religions (which, indeed makes major claims about what's good or bad that seem naively horrible to let the other guys win). And it turns that "we agree to live-and-let live about that" outperforms most other deals tried so far.
What sort of alternatives do you have in mind?
(it’d also be pretty surprising to me if dogmatic settlers stayed the same particular flavor of dogmatic for more than a couple thousand years (really more than a few hundred)
What's preventing AI-powered value lock-in in this scenario? E.g., people telling intent-aligned AIs "keep me faithful to X" as part of competitive virtue/loyalty signaling?
I guess I agree that nonzero of this will happen but I think this requires a kind of weird combination of “strategic about the singularity” and “believing in religious orthodoxy” and that specifically caching out into not evolving at all over time. I expect this to be a pretty small share of the Lightcone given the constraints on who would opt in.
And I think over a billion years people who enforcedly believe in Islam or most equivalent things would probably still find ways to create good beautiful/alien-in-the-expected-good-ways tapestry of posthuman experiences.
Between those I don’t find myself too worked up about it
There maybe can be global rules like “no permacommitments without some degree of awareness of opportunity cost
Can you explain why it requires “strategic about the singularity”? I think as soon as it's possible to tell one's AI assistant "keep me faithful to X" (and have the AI do a competent job of this), someone among the billions of religious believers is bound to do it, and then the practice will spread via imitation and competitive signaling (similar to other "tech" for enforcing faith, like "hell for non-believers"). This process does not seem to require anyone being strategic. What am I missing?
CEV, or the general category of extrapolation/idealization processes (and I am confused and dismayed how rarely I see this mentioned in these conversations nowadays).
Yeah.
Okay, I actually also have a "wtf guys why aren't we talking about CEV more?" post lined up. I think I didn't bring it up in this context because I'm treating the AI 2040 Plan A desiderata to include "it can be explained succinctly in a paragraph that the average pretty smart human will read and say 'okay I see how that would be fair/reasonable/good.'"
I think it would be great if we had such a paragraph for CEV, although I don't currently.
"Whenever you'd certainly later wish you'd been warned against your course of action, you are."
CEV is ~ optimizing the world based on what you would later wish, not just warning you (though the optimization could include just warning you about some things). Good start though!
I currently think it's a better operationalization of CEV to not "optimize based on what you'd later want", exactly.
People frequently object "but future me want all kinds of path dependent alien stuff"
To which I've replied "but, it only does the stuff that lots of different monte carlo simulations all turn out to want"
To which people "I dunno still seems like the me who's thought for 1000 years or whatever may end up alien in some way, and I don't sign up to automatically identify with that."
In my last discussion about at this, I said "I think the right way to do CEV is, you don't optimize the values of the thing-at-the-end. You optimize what future-you would do if they were specifically trying to help present you (or, by present you's values after having some kind of mediated conversation with various chains of future you's). And then, only where the values cohere across time and simulation-rolls.
(might easily change my mind about this. I'm not sure if there's more context I'm missing)
If we want to move from Plan D to Plan A, I believe the first step is to collectively agree on the problem. We are far from it, and there is a lot we can do. I wrote about this in this piece: The current bottleneck is political will, not research.
I'm excited about people commenting on this post with questions, feedback, critiques, different proposals, etc. We'll try to monitor it and respond to many of the comments.
I just noticed this part of plan S.
Chip fabs are allowed to keep scaling up, but in a slow and highly regulated fashion, and besides, there isn’t as much demand now that AI capabilities are frozen.
This seems like a mistake. Whether plan S is good or bad will depend on if the conditions for AI development are better or worse when it finally ends. If fabs are allowed to scale up and research on how to make cheaper and more efficient chips is allowed to continue, then we'll plausibly have much more compute and a much faster takeoff when plan S eventually ends, plausibly making it all net negative. Chip R&D in particular might actually be the most important thing here -- Moore's law continuing for several more decades would make takeoff much, much faster if it the deal eventually broke down.
Depending on how much political will there is (and in plan S there is presumably a lot!) I'd guess it's better to go the other way and reduce chips and fab capacity, and to ban R&D into producing more efficient chips. So that if someone ever leaves the deal, it takes maximally long for the world to resume the singularity. (Leaving enough time to coordinate on another deal, or to do research, etc.)
In fact, if you expect non-tracked compute to decrease over time (e.g. culture change makes it more and more outrageous to hide chips; replacement of leadership and rank-and-file at various places give greater probability that someone changes their mind or leaks; it's tempting to run the chips in ways that would degrade them or to sell them to buy-back programs; more total intelligence agency effort accumulating over time) then plausibly an "out" of plan S would be to restart the singularity a while later but with much less compute in the world, capping the speed of the singularity in a much more robust way than plan A.
A key part of this strategy is deterrence, "Mutually Assured Compute Destruction" which gets its own section. It doesn't mention the generalization Mutually Assured AI Malfunction (MAIM) from Schmidt, Wang, and me last year. This also spends 1/3 of the MAIM discussion on verification and how to do this in a multilateral way.
Meanwhile it cites other works like A Narrow Path. I even left this feedback to you all at AI Futures before this was released. This would constitute plagiarism in any other context. It's a bewildering unforced error--it's extremely related, it's a certainly a nontrivial idea, and I told you all this recently--I hope you all fix it.
For the past year, we at the AI Futures Project have been sinking most of our time into our next big scenario. Now it’s done!
It’s called AI 2040: Plan A.
It’s called Plan A because it’s a recommendation, not a prediction. It’s what we think should happen, not what will happen, though we think it’s plausible enough to aim for.
It’s called AI 2040 because in it, they delay the creation of superintelligence to 2040. It would have happened much sooner (in 2030, to be precise) if not for decisive action on the part of the US and Chinese governments.
As with AI 2027, summaries don’t really do it justice, since the whole point was to be detailed and comprehensive and work things out step by step rather than rely on high-level abstractions like doom or utopia.
Read the scenario at ai-2040.com. You can listen to it on audio, or view it on mobile, but the experience is significantly better on a normal computer.
What’s next for us?
Well, first we are going to respond to comments and otherwise engage with whatever conversation, responses, critiques, etc. that AI 2040: Plan A sparks. Beyond that, we aren’t sure yet. In general our mission is to help make AGI go well, and now we’ve tried out both forecasting and planning. Maybe we’ll get started on another big scenario. On the other hand, these megaprojects take so much time…