One problem is that human value might be inherently multidimensional. I think of it as the want/like/approve distinction. We seem to have separate mechanisms in our brains for 1) enjoying something in the moment, 2) wanting to do it before, and 3) approving of it afterward. It's possible to want something without enjoying it (like a person with OCD wanting to close the door exactly ten times), enjoy something without wanting it (people have said that they've reached very enjoyable meditative states but feel zero motivation to reach them again), enjoy something without approving it (porn), approve something without enjoying it (exercise), and all other combinations. This is the reason why "revealed preference" doesn't work: a person's actions are dictated disproportionally by the "want" dimension, but a good theory of value should incorporate all three. If we optimize one over the others, the tails will come apart.
A nice toy example is video games, where people are attracted to them because of the graphics, then stay because of the gameplay, and then have a warm afterglow and want to discuss afterward because of the story. Which of the three should contribute the most to the "true" quality rating of a videogame - "want", "like", or "approve"? Is this question philosophically meaningful? Will more reflection solve it?
Is this question philosophically meaningful?
Assuming you mean something like "Can philosophy answer this question?" I think "Maybe, but we probably won't know until we do a lot more philosophy." To put it another way, I think it's very plausible (but far from certain) that we can eventually answer questions like this one, given enough competent reflection, and I want to make sure we definitively find out before we make any irreversible choices based on what we think our values are.
I read your post, and I had thoughts about it. I made a vocal about it and asked fable to improve it.
I agree "Self-Correction" is a better name than "Long Reflection", though the post doesn't say why. Here is my reason: "Reflection" suggests the fix is more thinking. "Correction" admits the fix is changing what we are. That's the right framing.
But I disagree with most of the list. I think it mixes three different kinds of "flaws", and they call for very different responses.
1. The metaethics flaws are based on a framing I reject.
"Not having a workable moral framework" assumes that morality is a research problem: there is some true target out there, consequentialism and deontology are our candidate theories, and sadly they all fail. I think this picture is wrong from the start.
Here is the alternative picture. Tribes that coordinated on rules like "don't kill members of your own tribe" survived. Tribes that didn't, died out. Morality is the name we give to those rules, seen from the inside. Philosophers came much later and tried to fit general theories to this data. Utilitarians tried numbers, deontologists tried rules. Of course the theories all "have serious problems": they are rough compressions of a messy evolutionary process, not failed attempts at a real target. I wrote up this genealogy in more detail here: Dissolving moral philosophy.
To be fair about what this view doesn't give you: it's descriptive. It never crosses Hume's guillotine, and some philosophical questions stay open. But this changes what the post has to argue. "Humans lack a workable moral framework" becomes "here are the specific open questions we must answer before doing anything irreversible". That list would be much shorter, and much more debatable, than flaws 1, 2 and 7 suggest.
2. The status game flaw might be a Chesterton fence.
I see the same thing Wei sees: careful strategy and philosophy get low status in most places. Spend ten minutes on LinkedIn. But before calling it a flaw to fix, ask why the fence is there.
One possibility: society under-rewards philosophizing because, on the margin, doing things beats theorizing, and a culture that gave top status to meta-level reflection would get little done. Another: status and power are what motivates most people to do anything at all, especially now that religion doesn't. Remove that and I don't know what's left.
Same for institutions. Yes, it's annoying when a politician's mediocre report gets 200 likes and a truer analysis gets 5. But part of what holds society together is that people defer to institutions somewhat independently of the quality of their output. Legitimacy is fragile. If you "correct" deference away, you may not get a world of better epistemics. You may get a world where nothing holds, and all institutions fall apart. These fences should be moved carefully, and that cuts against listing them as simple flaws.
3.On calibration (flaw 3), the evidence is weaker than presented.
FTX looks to me like fraud plus bad incentives, not philosophical overconfidence. And competence varies a lot from person to person. Some people (Davidad comes to mind) seem to have settled enough of the philosophy to move on and build. "Humans are badly calibrated" erases exactly the variation that matters.
What I think the actual bottleneck is.
The one flaw that does real work in my model is the one the post puts in a parenthesis at the end: we are bad at large-scale, long-horizon coordination. See climate change. That one is a true precondition, both for containing the risks and for running any Long Self-Correction at all. Most of flaws 1 to 8 either dissolve (metaethics), turn out to be fences (status), or can be fixed in flight (zero-sum values). Coordination can't wait, and it's a political problem more than a reflection problem. See: The current bottleneck is political will, not research.
I'll grant flaw 8 (over-optimistic partial solutions) has real force. My own position is exposed to it too.
I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection.
Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper problem is that humans aren't safe, and can't safely serve as builders, overseers, or alignment targets for powerful AIs.
Problem with Long Reflection: It seems to imply that the main problem with humans is that we just haven't had enough time to think, that reflection is the main thing we need to do more of, and then we can get on with building powerful AIs or other technologies. Or that if we build aligned AIs that sincerely help us think a lot more, or do the thinking for us, then things will turn out fine.
So I think we need a catchy handle for a related but distinct idea, that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws.
A summary of the flaws that I have in mind:
(This list focuses on key bottlenecks that seem hard to fix even with AI assistance or intelligence enhancement, and isn't meant to be a complete list of human flaws / safety problems. It ignores e.g. that the median human is ignorant of many important issues, and that we're currently quite bad at complex large-scale coordination such as passing/implementing close-to-optimal government policies.)
My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time, so if we preserve the environment in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition.
It will probably be shortened to "The Long Correction" at some point if it catches on, similar to how "outer space" is now often just "space".
Why isn't there a version of EA that explicitly talks about how to leverage people's status motivations to do more good for the world? It's very possible that explicit talk about status is actually counterproductive at least in the short run, e.g. it heightens status motivations and makes people less altruistic, but then do we just march into the future while blindfolding ourselves to this aspect of human nature?