Given the enormous risks in building self-improving AIs, we feel it our responsibility to justify pursuing the related path of automated research to the public. We therefore lay out clearly our motivations for conducting this work and our commitments to reduce the risk that we inadvertently contribute to the very future we seek to avoid.
We believe that the development of superhuman AI presents grave risks to humanity. We believe it is prima facie plausible, and perhaps even likely that the development of superhuman AI will lead to the extinction of humanity. We believe that taking wise and proactive measures to curb this risk is a moral imperative on every person. We believe that doing the converse, i.e. minimizing or downplaying the risks, when done for financial or careerist reasons (as opposed to honest consideration or uncertainty), is in the full sense of the phrase, a "crime against humanity".
Nevertheless, the world is at present, hurtling towards taking the dangerous gamble of building increasingly powerful AIs. To our mind this leaves those who recognize the danger with only two possible tracks to defend humanity.
- A political solution to halt or slow down AI development.
- A technical solution to make AIs safe.
Both of these paths face enormous complications and difficulties.
In our view, taking into account the current social/political situation and the state of the AI field, a technical solution seems like the only viable path forward. While there is at present vague worry in the body politic about AI development (mostly centered around job displacement and economic disruption), there is little serious discussion about the existential risks posed by superintelligent AI (much less a large, popular, organized, and aggressive political movement to seriously challenge and counteract the enormous incentives pushing for AI development).
To be clear, we consider those that are committed to pursuing a political solution heroes and we do not think it out of the question that a warning shot or catastrophic event might shift the political landscape enough to create an opening for a political solution. But, given the speed of AI development, we cannot wager the future of humanity on such schemes, especially since we are not guaranteed a warning shot before the point of no return. Even in the optimistic case of a successful political intervention, we would likely only delay AI development, i.e. only buy time for researchers to come up with a technical solution to AI alignment.
The central variable in AI safety research is speed. Training an effective researcher is a notoriously painful and slow multi-year process. Once trained, the research cycle for effective researchers from idea to publication is generally ~6-18 months in academia. Given the speed of AI progress (a doubling of time horizon task completion every ~4 months) the common model of academic research is obviously wholly inadequate for the scale and urgency of the challenge we face. It therefore seems clear to us, that automating as much of the research process as possible is the obvious strategic play, and the only way to meaningfully contribute to the field in time to matter.
We fully acknowledge the enormous risks associated with this approach, and will seek to mitigate those risks as much as possible while maintaining our commitment to speed and urgency. The core source of danger with our approach is that building an effective auto-researcher for AI safety could almost certainly with minimal or no modification be used to build an effective auto-researcher for AI capabilities and thereby potentially kick off a runaway self-improvement process.
To mitigate this risk and other risks, we publicly commit to the following:
Our Commitments
- We commit to always operating in the interests of humanity in all our actions and decisions.
- We commit to never altering our public messaging about the risks of AI development for any strategic reasons whatsoever.
- We commit to never modifying our existing commitments without explicit disclosure and justification to the public.
- We commit to never exchanging information about our auto-researcher for money or other benefits.
- We commit to never discussing details about our work on the auto-researcher with untrusted or unvetted individuals.
- We commit to prioritizing safeguarding our secrets from model providers and other parties as soon as it is financially feasible to do so.
- We commit to restricting access to our outputted research as soon as it is useful for improving AI capabilities.
- We commit to introducing internal controls and monitoring that is at least as strict as frontier AI companies over all the actions of the auto-researcher as soon as it is feasible to do so.
- We commit to utilizing legal tools to ensure compliance with all commitments on all members of our team as soon as it is financially feasible to do so.
- We commit to immediately halting our work on auto-research if we ever come to believe that our work might do more harm than good.
- We commit to immediately halting our work on auto-research if instructed to do so by a legitimate and trusted political body.
Meet the people, and the agent, behind this work on our team page.