Should We Hand Off the Most Important Decisions to AI? (with Matthew Adelstein)
Matthew Adelstein was a Visiting Scholar at Forethought and blogs as Bentham’s Bulldog, where he writes on ethics, animal welfare, and the long-run future. He joined Fin Moorhouse to discuss: Why Matthew thinks handoff of high-stakes decision-making to AI is the most likely route to a great future Why locking in current human values would be a cosmic-scale catastrophe, why steering nowhere in particular is barely better, and why the alternative is building AIs that we trust to do reflective moral philosophy Various ways handoff might manifest politically: elected AI representatives (“President Claude”), extending political rights to digital minds, or de facto handoff through gradual disempowerment of humanity Whether handoff should be reversible (such that humans could take back control if we wanted) or if we should “tie ourselves to the mast” Whether to task AIs with “do what’s objectively good” or “do what humans would want on reflection,” why Matthew leans toward the former, and why he expects the two to ultimately converge If digital minds end up vastly outnumbering humans, and those digital minds deserve political rights, then consistently applied liberal-democratic principles may imply handoff as reducing expected disenfranchisement Whether AIs can actually do philosophy: where models currently sit, bootstrapping schemes using AI-designed evals, training AIs with different epistemic constitutions and looking for convergence, and Fin’s scepticism that recursive self-evaluation works as well as it needs to Whether we could even tell if an AI were superhuman at philosophy, given the difficulty of evaluating philosophy that’s beyond our own level of competence Whether goodness can compete under evolutionary pressure, resource races into space, and intergalactic existential risks as an argument for top-down planning Whether humans’ selfish preferences would saturate in a post-AGI world Why smart AIs won’t necessarily converge on good values, and how selection against “weird” outputs (by humans’ lights) to moral reasoning could lock in something mediocre You can read a full transcript here. To see all Forethought's published research, visit forethought.org/research. To subscribe to our newsletter, visit forethought.org/subscribe.