Foreword
AI companies are racing to build AIs that are smarter than humans in every way. In AI 2027, we predicted that this would result in either extinction or irreversible concentration of power.The AI 2027 scenario is still roughly what we expect the future to look like: a mad scramble to superintelligence leading to either AI takeover or extreme concentration of power. So far reality is tracking closer to AI 2027 than even we expected. (2027 was our modal year at time of publication, not our median.) You can read more about our views on timelines here and here.The AI 2027 scenario is still roughly what we expect the future to look like: a mad scramble to superintelligence leading to either AI takeover or extreme concentration of power. So far reality is tracking closer to AI 2027 than even we expected. (2027 was our modal year at time of publication, not our median.) You can read more about our views on timelines here and here.
Plan A is our positive vision for what should happen instead.
In this scenario, humanity delays the development of superintelligence until 2040, makes all AI research public, allows dozens of companies globally to catch up to the frontier, and intentionally enters a regime of mutually assured compute destruction.
What is Plan A?
Plan A is our positive vision for how humanity can avoid AI-driven existential catastrophe and reach a flourishing future. It’s informed by conversations with experts at major U.S. frontier AI companies, direct experience at OpenAI, tabletop exercises, and discussions with policymakers, national security experts, and AI policy leaders. We recommend an international deal to avoid a dangerous race to superintelligence. The deal involves total research transparency for AI R&D, which allows the nations of the world to understand what’s happening and enforce guardrails. The result is multiple companies across multiple countries scaling slowly and safely together towards superintelligence, instead of racing each other in secrecy.
Plan A is primarily a recommendation, not a prediction. This scenario is not our best guess as to what the future will actually look like. Instead, it’s a vehicle for communicating and stress-testing our policy recommendations. While the implementation of Plan A is a recommendation and not what we actually expect to happen, the subsequent effects depicted are predictions.That is, they are predictions about what would happen next if our recommendations were implemented.That is, they are predictions about what would happen next if our recommendations were implemented.
In this AI 2040 scenario, Plan A is implemented successfully, albeit imperfectly and only in the nick of time.
We contrast Plan A with 4 alternative plans (B, C, D, and S), which correspond to the main ways the US could respond (or not) to the challenges of superintelligence.You can read more about these four plans at the branch point below and in our supplement “How good is each plan?”.You can read more about these four plans at the branch point below and in our supplement “How good is each plan?”.
Why did we write this?
AI companies will probably succeed at their stated goal of building smarter-than-human AI systems within the next 1 to 10 years.
The industry has convinced itself that controlling superintelligent AI can be figured out on the fly, and thus has no remotely adequate plan. We think this situation is terrible and could easily get us all killed.Of course, we’re not sure exactly how likely literal AI-driven human extinction is, but opinions within our team vary between 10% and 30%. This is primarily because there are a bunch of other things misaligned ASIs might do with us after they take over, besides kill us all. For example, they might keep some people alive because it’s really cheap and they care a tiny amount, or as acausal bargaining chips. We don’t think this nuance changes the bottom line though: Misaligned AI takeover is probably really bad.Of course, we’re not sure exactly how likely literal AI-driven human extinction is, but opinions within our team vary between 10% and 30%. This is primarily because there are a bunch of other things misaligned ASIs might do with us after they take over, besides kill us all. For example, they might keep some people alive because it’s really cheap and they care a tiny amount, or as acausal bargaining chips. We don’t think this nuance changes the bottom line though: Misaligned AI takeover is probably really bad. We do not expect whoever “wins the race” to have much of a lead, and we do not expect them to unilaterally slow down to reduce existential risk.I (Daniel Kokotajlo) recently gave a talk to about 100 people, ~40% of whom were from frontier AI companies, and did a show-of-hands poll about (a) how many months of lead the leading AI company would have at the time they first fully automate AI R&D, and (b) how much of their lead they would be willing to burn. The median answers were roughly 3 months and 1/3rd, respectively. This matches my own guess.I (Daniel Kokotajlo) recently gave a talk to about 100 people, ~40% of whom were from frontier AI companies, and did a show-of-hands poll about (a) how many months of lead the leading AI company would have at the time they first fully automate AI R&D, and (b) how much of their lead they would be willing to burn. The median answers were roughly 3 months and 1/3rd, respectively. This matches my own guess. If this race continues,Also intense secrecy and groupthink conditions! By default during this period the companies will be even more closed than they are in 2026. When they say “we’ll solve the alignment problems as we go” what they really mean is “The overworked tiny group of alignment experts at our company will solve the alignment problems as we go, without being able to subject their work to external review.”Also intense secrecy and groupthink conditions! By default during this period the companies will be even more closed than they are in 2026. When they say “we’ll solve the alignment problems as we go” what they really mean is “The overworked tiny group of alignment experts at our company will solve the alignment problems as we go, without being able to subject their work to external review.” we do not expect humans to maintain effective control as their AIs become superintelligent.We expect them to maintain apparent control, to be clear, at least for some time, which is part of why this problem is so dire. The situation would be much better if we had a reliable way of telling whether an AI system was actually robustly obedient/loyal/aligned/etc. vs. merely temporarily so vs. merely pretending.We expect them to maintain apparent control, to be clear, at least for some time, which is part of why this problem is so dire. The situation would be much better if we had a reliable way of telling whether an AI system was actually robustly obedient/loyal/aligned/etc. vs. merely temporarily so vs. merely pretending.
Moreover, even if the AI companies somehow align their AIs, the result will be an unprecedented concentration of power—that is, the result will be a situation where a tiny group of people, or possibly just a single individual, is effectively in control of the world’s only army of superintelligences for some months, and will be presented by said superintelligences with various options for how to proceed, some of which will de facto amount to taking over the world.The AI 2027 Slowdown ending depicts this happening. It’s quite easy to imagine because it doesn’t involve anything cartoonishly or overtly evil, but rather just a series of steps that can be justified individually (consolidating US compute into a single pool, beating China, integrating AI into the government and economy, preventing terrorists from getting their own AGIs, etc.)The AI 2027 Slowdown ending depicts this happening. It’s quite easy to imagine because it doesn’t involve anything cartoonishly or overtly evil, but rather just a series of steps that can be justified individually (consolidating US compute into a single pool, beating China, integrating AI into the government and economy, preventing terrorists from getting their own AGIs, etc.)
As best as we can guess, the CEOs of OpenAI, Anthropic, xAI, and Google DeepMind understand this and are proceeding anyway, perhaps because they think they are the lesser evil and will use their immense power responsibly, unlike Xi Jinping or rival CEOs.For example, these old OpenAI emails reveal that OpenAI was founded in significant part because of fear that Demis Hassabis, DeepMind CEO, would use AGI to become dictator. See also related reporting in Empire of AI. See also this and this and this.For example, these old OpenAI emails reveal that OpenAI was founded in significant part because of fear that Demis Hassabis, DeepMind CEO, would use AGI to become dictator. See also related reporting in Empire of AI. See also this and this and this.
While we agree that it is generally correct to choose the lesser evil, we don’t think we should advocate for a strategy that has such a scarily high chance of leading to human extinction or global dictatorship. Instead, we wish to advocate for something that is actually good. If enough people do likewise, it can happen.
So, we wrote a scenario outlining that possible world.
Why a scenario?
“Plans are worthless, but planning is everything.” - Dwight D. Eisenhower
We think most AI policy proposals fall apart under scenario scrutiny—that is, if you try to write down a detailed and plausible scenario in which that proposal succeeds, you will find it difficult to do so, and you will realize the plan is less likely to work than it seemed, or has more unpleasant side-effects than its proponents acknowledged.
Perhaps that’s why scenario scrutiny is so rare in AI policy. Everyone wants to say that their own favorite policies will have great consequences and that the policies of their rivals will have terrible consequences. Applying scenario scrutiny to their own favorite policies might surface uncomfortable issues with them; meanwhile, applying scenario scrutiny to their rival’s policies is a lot of work for little rhetorical gain.For one thing, it’s hard to apply scenario scrutiny to a policy proposal if you have only a shallow understanding of it, and people generally understand their own favorite ideas much better than they understand the ideas they hate. For another, the evidential force of scenario scrutiny is inherently stronger for scenarios written by proponents than by opponents: If a proponent of a policy tries their best to depict it succeeding, and fails, that’s a lot more evidence than if an opponent of a policy depicts the policy failing.For one thing, it’s hard to apply scenario scrutiny to a policy proposal if you have only a shallow understanding of it, and people generally understand their own favorite ideas much better than they understand the ideas they hate. For another, the evidential force of scenario scrutiny is inherently stronger for scenarios written by proponents than by opponents: If a proponent of a policy tries their best to depict it succeeding, and fails, that’s a lot more evidence than if an opponent of a policy depicts the policy failing.
We think the discourse would be improved if more AI policy proposals were subjected to scenario scrutiny. So we’re starting with our own, even though this opens us up to criticism. We hope critics will judge us against the existing state-of-the-art for plans to navigate the AI transition (if they can find any) and not against some hazy but pleasant fantasy where no one has to make any hard choices yet everything will probably be fine.
What of the immense difficulty of predicting the effect our policy would have in a world approaching superhuman AIs? This is like trying to predict how to best fight World War 3, except that it’s an even larger departure from past case-studies. Yet it is still valuable to attempt, just as it is valuable for the U.S. military to game out Taiwan scenarios in excruciating detail. There are other precedents as well: intelligence agencies, climate bodies, and pandemic-preparedness offices all rely on various kinds of scenario planning.
AI timelines?
Plan A is our ambitious proposal for what to do, and we’d like to see something like it implemented soon because we are uncertain about how much time remains.That is, we are uncertain about how fast AI capabilities will progress – about how much time remains before the AI companies succeed at automating their research and accelerating towards superintelligence, for example.That is, we are uncertain about how fast AI capabilities will progress – about how much time remains before the AI companies succeed at automating their research and accelerating towards superintelligence, for example. But for purposes of writing a concrete scenario, we need a concrete timeline.
The timeline of this scenario is:
In 2029, the US and China agree to avoid a reckless race to superintelligence.
In 2030, we would have fully automated AI R&D, leading to superintelligence by the end of the year. Thanks to the deal, we avoid this.
Between 2030 and 2035, we scale within the human range, to AIs that are roughly as capable as top human experts.
In 2035, we pause at top-human-expert level AI in order to maintain human control.
In 2040, we unpause and scale to superintelligence.Superintelligence means AI systems that are significantly better than the best humans at everything, while also being faster and cheaper.Superintelligence means AI systems that are significantly better than the best humans at everything, while also being faster and cheaper. (Hence the title: AI 2040)
In our previous scenario, AI 2027, AI fully automated the process of building smarter AIs in 2027, leading to an intelligence explosion and superintelligence within the year. The two differences in this scenario are (1) the default timeline is now 2030, and (2) thanks to governance actions, generally-superhuman AIs first appear in 2040.
We changed the default timeline because we want our portfolio of scenarios to reflect our uncertainty about AI timelines. AI 2027’s titular year was chosen because, at the time we started writing, Daniel thought there was roughly a 50% chance that things would go that fast or faster.By the time we finished writing AI 2027, Daniel’s AGI median was 2028. Other authors’ medians were between 2031 and 2035. To read more about how Daniel and Eli’s views have shifted over time, see here. We are regularly updating our forecasts here, with corresponding blog posts on our Substack.By the time we finished writing AI 2027, Daniel’s AGI median was 2028. Other authors’ medians were between 2031 and 2035. To read more about how Daniel and Eli’s views have shifted over time, see here. We are regularly updating our forecasts here, with corresponding blog posts on our Substack. At the time we started writing Plan A, 2030 was the corresponding year for Thomas. Daniel currently thinks things will probably go somewhat faster than depicted in this scenario; you can read more about our team’s views on timelines here and here.
We changed the governance actions because this scenario is primarily a recommendation, not a prediction. Conducting a full-speed intelligence explosion is wildly reckless and concentrates power to an extreme degree.
2027: The Writing on the Wall
America has two workforces now. The first is people, 165 million of them. The second is AI agents: millions of copies spun up and shut down every hour, working around the clock at superhuman speeds.
Most of their work is slop. But enough of it is good that people are paying ten billion dollars a month for AIs that can, in theory at least, do anything on a computer that an employee can.
There is one job the AI companies want to automate more than any other—their own. They haven’t succeeded yet; no recursive self-improvement so far.AIs are helping to make future AIs more capable, but AI development still requires humans.AIs are helping to make future AIs more capable, but AI development still requires humans. But they seem to be getting closer, and they’re pulling up the ladder behind them: the strongest coding AIs refuse to help competitors with AI R&D.This policy is selfishly good for leading AI companies insofar as it makes it harder for trailing companies to catch up, but it’s also good for safety because it slows irreversible proliferation of algorithmic insights.This policy is selfishly good for leading AI companies insofar as it makes it harder for trailing companies to catch up, but it’s also good for safety because it slows irreversible proliferation of algorithmic insights. Even as the most bullish employees admit that things are taking a bit longer than planned, the skeptics notice that their usual dismissals are starting to ring hollow. Why exactly will AI never be able to do my job? What’s the barrier again?
Congress is starting to pay more attention. They’ve long been hearing about AI: datacenters using too much water,We think the concerns around current AI water usage are largely overstated. See this article for more. That said, as will be apparent later in the scenario, we do think that the environmental impact of AI over the next decade or two will be enormous, and that unrestrained growth would literally boil the oceans. See the energy appendix in our economics supplement for more on this.We think the concerns around current AI water usage are largely overstated. See this article for more. That said, as will be apparent later in the scenario, we do think that the environmental impact of AI over the next decade or two will be enormous, and that unrestrained growth would literally boil the oceans. See the energy appendix in our economics supplement for more on this. chatbots encouraging suicide, Mythos hacking NSA systems—and of course, tech industry lobbyists warning that any whiff of regulation will make America immediately lose the race with China and spend the rest of history as a CCP tributary state.Are we really in such a race? Sort of, but not exactly. On our default trajectory, AI will soon become very powerful, to the point where it’s more strategically important than nuclear weapons; that much is true. Many people, including the CEOs of frontier AI companies, are trying hard to build smarter and smarter AI systems before their competitors do. However, being overly paranoid about competitors is a well-known bias. Assuming that we’re in a winner-takes-all race rules out many possible avenues for cooperation, creating a self-fulfilling trap. We expect that the scale and speed of AI progress will make it increasingly obvious to decision-makers in both the US and China that the current pace of progress, and lack of trust between them, is extremely dangerous. Indeed, it may be possible for the two countries to coordinate on a slowdown without any binding agreement, if they build trust by each gradually deescalating. Our proposed Plan A should be interpreted as an approach that could work even if the US and China have absolutely zero trust for each other, rather than a claim that strict verification protocols are definitely necessary.Are we really in such a race? Sort of, but not exactly. On our default trajectory, AI will soon become very powerful, to the point where it’s more strategically important than nuclear weapons; that much is true. Many people, including the CEOs of frontier AI companies, are trying hard to build smarter and smarter AI systems before their competitors do. However, being overly paranoid about competitors is a well-known bias. Assuming that we’re in a winner-takes-all race rules out many possible avenues for cooperation, creating a self-fulfilling trap. We expect that the scale and speed of AI progress will make it increasingly obvious to decision-makers in both the US and China that the current pace of progress, and lack of trust between them, is extremely dangerous. Indeed, it may be possible for the two countries to coordinate on a slowdown without any binding agreement, if they build trust by each gradually deescalating. Our proposed Plan A should be interpreted as an approach that could work even if the US and China have absolutely zero trust for each other, rather than a claim that strict verification protocols are definitely necessary.
Now they step back and ask: Where are we going with this? What does the world look like five, ten, or fifteen years from now? Will there still be jobs? What if there aren’t?
One question weighs especially heavily on their minds: Who will control all these AIs?
Congress settles on an important part of the answer: Probably not us.That is, not Congress. By default the people who control the AIs, if anyone truly does, will be AI companies or maybe the White House.That is, not Congress. By default the people who control the AIs, if anyone truly does, will be AI companies or maybe the White House.
They hold a series of tense hearings on AI. They read the 2016 OpenAI emails discussing how OpenAI was founded in order to prevent Demis Hassabis from becoming dictator.From Ilya and Greg: “The goal of OpenAI is to make the future good and to avoid an AGI dictatorship. You are concerned that Demis could create an AGI dictatorship. So do we. So it is a bad idea to create a structure where you could become a dictator if you chose to, especially given that we can create some other structure that avoids this possibility.”From Ilya and Greg: “The goal of OpenAI is to make the future good and to avoid an AGI dictatorship. You are concerned that Demis could create an AGI dictatorship. So do we. So it is a bad idea to create a structure where you could become a dictator if you chose to, especially given that we can create some other structure that avoids this possibility.” But who is preventing Sam or Elon from becoming dictator? Congress is unsatisfied with existing responses.
The result of this wakeup is the AI Transparency Act of 2027, an omnibus bill that does many things, some good and some bad, but doesn’t fundamentally change the situation.We don’t mean to recommend bad things; we are predicting that a mixture of good and bad incrementalist reforms will happen in 2027. For more on how we balance predictions and recommendations in this scenario, see here.We don’t mean to recommend bad things; we are predicting that a mixture of good and bad incrementalist reforms will happen in 2027. For more on how we balance predictions and recommendations in this scenario, see here.
2028: AI on the Ballot
The 2028 election cycle is heated, as usual. AI is the biggest topic. The datacenters now under construction cost twice as much as the entire US military budget.That is, the AI industry spends $2.4T of CapEx over 2028, triple what was spent in 2026. The US military budget in 2025 was about $1T. The biggest AI company has $360B ARR at the beginning of 2028 and is still growing at a roughly 150% annualized growth rate, putting it on track to become both the most valuable company in the world and the largest company in the world by revenue, in 1-2 years.That is, the AI industry spends $2.4T of CapEx over 2028, triple what was spent in 2026. The US military budget in 2025 was about $1T. The biggest AI company has $360B ARR at the beginning of 2028 and is still growing at a roughly 150% annualized growth rate, putting it on track to become both the most valuable company in the world and the largest company in the world by revenue, in 1-2 years.
Most white-collar professions are seeing disruption like software engineering saw in 2026; such jobs now heavily involve managing AI agents. AI companies have industrialized the training process: Executives say “let’s move into [profession] this year” and then the company interviews professionals, buys data, creates training environments, etc. until their AIs get traction. Then the AIs rapidly improve as they are used more widely in the field and accumulate more real-world data.
Other countries are starting to get scared and angry. It seems like a handful of US and Chinese companies are on track to automate all the white-collar jobs. Power is concentrating in the US, and in particular in the President plus a handful of tech CEOs.
AI experts warn that the intelligence explosion is near. By speeding up AI research, the AIs will become even more competent, speeding up research even faster, making them even more competent, and so on. There are complicated dynamics about bottlenecks and hardware limits governing how fast this process goes and where it ends, but it seems like it might go very fast and end somewhere very far away.
On the default path, the next presidential term will see AIs that are far beyond human level, created entirely by AIs, themselves created entirely by other AIs, without any human in the loop since several generations back. Will those AIs be obedient, aligned, etc.? Why? Who will control them if so? How exactly is all of this supposed to end well?
Having put humanity on this path, the AI companies find it acceptable. But most people don’t. Forget thinking about his legacy—the President is starting to think about what’ll happen to him after he leaves office and the world gets transformed.After all, he’ll still be alive after leaving office. Just like the rest of us, he’ll have to endure whatever crises result from superintelligence, and hope that the new President handles them well and doesn’t become dictator. It’s in his interest both to set the new President up for success with respect to the loss-of-control problem, and to install checks and balances to prevent extreme concentration of power.After all, he’ll still be alive after leaving office. Just like the rest of us, he’ll have to endure whatever crises result from superintelligence, and hope that the new President handles them well and doesn’t become dictator. It’s in his interest both to set the new President up for success with respect to the loss-of-control problem, and to install checks and balances to prevent extreme concentration of power. Both presidential candidates keep getting asked what they’ll do about AI, and try out increasingly dramatic ideas on the campaign trail. The discourse bounces back and forth across all of the options displayed below, and more.
Eventually the President and his protégé converge on one plan; the opposition candidate converges on another. Then it’s Election Day.
2029: Choose a Path
“Trust, but verify” – Ronald Reagan
The President announces that the US will pursue international cooperation to avoid an imminent intelligence explosion.
“This mad race toward superintelligence must end. For too long, we have been pursuing lesser-evils and least-bad solutions. We need a Plan A. We can still proceed with AI development, but we must do it more cautiously, more transparently, and involving many more countries and companies.”
To the surprise of many in DC, China proves receptive. They had been debating the same issues—social destabilization, job loss, rogue superintelligence - on their side of the Pacific. They had been looking forward to “The Chinese Century” and thought that AI might disrupt their plans. And they had one extra reason to come to the table: the US continues to have more compute, more datacenters, and better models. They were worried about the things the US might do to them if it got to superintelligence first.To spell it out more: They were concerned that the US might use their massive AI advantage to cripple Chinese AI projects via cyberattacks, physical sabotage, or other means, and then use the resulting even-bigger, longer-lasting AI advantage to dictate terms to the CCP or even overthrow it entirely. Preventing loss-of-control risk was merely an added bonus.To spell it out more: They were concerned that the US might use their massive AI advantage to cripple Chinese AI projects via cyberattacks, physical sabotage, or other means, and then use the resulting even-bigger, longer-lasting AI advantage to dictate terms to the CCP or even overthrow it entirely. Preventing loss-of-control risk was merely an added bonus.
The US and China don’t trust each other. Fortunately, they don’t have to: Plan A includes provisions for verifying compliance. But setting it up will take time.
So for now, they start with something crude. For the rest of 2029, they put a temporary halt to AI training, because that’s relatively easy to verify. In 2030, they’ll have the infrastructure in place to proceed with Plan A.
The stock market is gyrating wildly up and down in response to news and commentary about the momentous actions being taken. Everybody is screaming about how it’s going too far or not far enough. The market will calm down by the end of the year, but the screaming will continue.Geopolitically the situation will have stabilized within a few years, but by that point AI progress and diffusion will have happened and caused new problems and crises for people to worry about.Geopolitically the situation will have stabilized within a few years, but by that point AI progress and diffusion will have happened and caused new problems and crises for people to worry about.
Step 1: Compute Declaration
Neither country wants to halt their own AI training unless they can see that the other side is halting too.
Fortunately, training AIs requires large numbers of AI chips. Most AI chips are in giant datacenters.We show our estimate of the distribution of datacenter sizes in our compute supplement.We show our estimate of the distribution of datacenter sizes in our compute supplement. AI datacenters are typically big enough to be visible from space, and power-hungry enough to require conspicuous infrastructure. New AI chips can only be manufactured at a handful of fabrication plants (fabs), located mostly in Taiwan, South Korea, the US, and China.

Left: OpenAI’s Stargate datacenter. Right: An Extreme Ultraviolet Lithography (EUV) machine, arguably the most complicated machine ever created by mankind. EUV machines are essential for producing frontier AI chips, and are produced by a single Dutch company, ASML.
The US and China negotiate with the countries that have a major role in the chip supply chain, and they require each major datacenter owner (and their upstream suppliers, including chip fabs) to publicly declare their major purchases and sales.The initial chip declaration would only involve companies whose purchases exceed 10k H100e (which costs ~$100M). Within a year, every company owning above 0.001% of the world’s compute (~2,800 H100e, costing ~$18M at the time) has its holdings accounted for. Dead chips are also stored or verifiably destroyed.The initial chip declaration would only involve companies whose purchases exceed 10k H100e (which costs ~$100M). Within a year, every company owning above 0.001% of the world’s compute (~2,800 H100e, costing ~$18M at the time) has its holdings accounted for. Dead chips are also stored or verifiably destroyed. Analysts from numerous countries and institutions may pore over the list, ask questions, and challenge anomalies. The rival powers also send inspectors to each other’s infrastructure, reassuring themselves that the stated numbers are accurate. By the end of the year, each side is confident that the other isn’t hiding more than about 1% of AI compute, and that all potential sources of new compute are monitored and accounted for.The 1% of compute sales that cannot be traced to their end location are mostly due to smugglers who did not record their customers’ identities and chips that were decommissioned. Half of this pool is later found installed in Chinese datacenters identified via satellite imagery; both countries offer monetary incentives and legal amnesty to any owners of the remaining 0.5%. We go over these and other routes for finding covert compute in much more detail in our covert project supplement.The 1% of compute sales that cannot be traced to their end location are mostly due to smugglers who did not record their customers’ identities and chips that were decommissioned. Half of this pool is later found installed in Chinese datacenters identified via satellite imagery; both countries offer monetary incentives and legal amnesty to any owners of the remaining 0.5%. We go over these and other routes for finding covert compute in much more detail in our covert project supplement.
Step 2: Pause Training
Distinguishing dangerous AIs from safe ones will require more time and understanding, so for now they go with a simple solution: a temporary pause on all new training runs.This includes the kind of research experiments that involve training. While it’s theoretically possible for an intelligence explosion to happen without new training runs or major experiments, it seems unlikely because it would require an extreme paradigm shift.This includes the kind of research experiments that involve training. While it’s theoretically possible for an intelligence explosion to happen without new training runs or major experiments, it seems unlikely because it would require an extreme paradigm shift.
Both sides can still use datacenters for running AIs that already exist (i.e. inference), but they will retrofit each other’s datacenters with devices to verify that they aren’t being used for new training runs.The inference-only verification solution we propose works by taking a random sample of a small percentage of the workloads to check they are correctly doing inference on an approved model (by recomputing the outputs using the inputs and approved model weights). To collect these random samples, our proposal involves simple network taps that redirect all inbound and outbound traffic from the datacenter to a recomputation server (inside the datacenter) installed by the other party that performs the random partial recomputation. Other solutions to this inference-only proposal may be possible and preferable, such as cryptographic proof of work schemes, but these are currently more speculative from a performance perspective. We discuss this in more detail in our verification supplement. There are several additional layers of enforcement necessary to make this approach work, including physical security, software based monitoring, and human auditors.The inference-only verification solution we propose works by taking a random sample of a small percentage of the workloads to check they are correctly doing inference on an approved model (by recomputing the outputs using the inputs and approved model weights). To collect these random samples, our proposal involves simple network taps that redirect all inbound and outbound traffic from the datacenter to a recomputation server (inside the datacenter) installed by the other party that performs the random partial recomputation. Other solutions to this inference-only proposal may be possible and preferable, such as cryptographic proof of work schemes, but these are currently more speculative from a performance perspective. We discuss this in more detail in our verification supplement. There are several additional layers of enforcement necessary to make this approach work, including physical security, software based monitoring, and human auditors.
The US and China are able to rapidly source enough of these verification devices to retrofit almost every major datacenter.This is made easier thanks to a now-thriving ecosystem of verification hardware companies that emerged after both governments signaled interest in having verification capability as an option (which in turn drew in investment from philanthropists and VCs and technical support from AI supply chain companies). Had that not happened, this part of Plan A would have taken somewhat longer and/or been more expensive and frantic. Worst case, they’d have to turn off the datacenters instead of allowing them to run inference.This is made easier thanks to a now-thriving ecosystem of verification hardware companies that emerged after both governments signaled interest in having verification capability as an option (which in turn drew in investment from philanthropists and VCs and technical support from AI supply chain companies). Had that not happened, this part of Plan A would have taken somewhat longer and/or been more expensive and frantic. Worst case, they’d have to turn off the datacenters instead of allowing them to run inference.
Inspectors from both countries sign off on the installation of the verification devices. Only datacenters that the US and China agree have functioning verification are allowed to provide continuing AI services. Meanwhile, both governments sprint towards more fine-grained verification options that could allow training to restart while keeping both sides confident the other isn’t racing ahead.
Step 3: Get Worldwide Buy-in
Though other countries were consulted from the start, especially those owning parts of the AI supply chain, Plan A began as a bilateral deal between the US and China. Now that it’s taking firmer shape, negotiations go multilateral. Carrots and sticks are waved about.
Many countries outside the US and China are happy about the deal, because they were worried about a future in which a handful of US and Chinese AI companies recursively self-improve, pull farther and farther ahead, and then… Well, what happens next depends on who you ask, but answers range over possibilities such as “take our jobs,” “cement hegemony forever” and “get everyone killed.” So they want the US and China to proceed more slowly and transparently. This will allay their fears and also allow their own AI projects to catch up to the frontier.
So it’s surprisingly easy for the US and China to get buy-in.This part is a prediction, not just a recommendation. We think that getting buy-in for Plan A, in circumstances like those we’ve described in 2029, would be significantly easier than most of our readers would expect.This part is a prediction, not just a recommendation. We think that getting buy-in for Plan A, in circumstances like those we’ve described in 2029, would be significantly easier than most of our readers would expect. By the end of the year, most of the world has joined what is now called the Consortium.
2030: Plan A is Established
Tech companies are pushing hard to restart AI training, and lobbying to shape the conditions under which it happens. Large portions of the public are saying that we should never restart. Some countries are too.
After a year of intense discussions, negotiations, and twitter beefs, a compromise takes shape: AI development will continue, but in a more cautious, more transparent, more distributed way.
Plan A is guided by 4 core principles:
Principle 1: Buy Time
The problem with an intelligence explosion is the “explosion” part.
The default trajectory is reckless and destabilizing. Each country and company is racing their competitors for AI dominance without time to think. Most likely, a ton of people will end up disempowered or killed. Maybe godlike AIs run circles around us and the age of humanity is over. Maybe a CEO or President establishes a permanent dictatorship. Maybe World War Three begins in the chaos, as countries without the best AIs fear losing their leverage. Maybe everything is surprisingly smooth and we “only” automate all labor and render everyone unemployable in a radically transformed world. Even aspiring AI-enabled dictators are worried; there’s enough frontrunners that none of them have more than 50% odds of winning the power struggle.
No one knows how to tell if AIs are trustworthy, and no one knows how to regulate superintelligent AI. Solving those problems looks like it’ll take years. Slowing down the intelligence explosion buys time to do so.
It also helps prevent extreme concentrations of power, because it gives more time for groups that don’t directly control the world’s smartest AIs to wake up to what’s happening before they lose their leverage. Without a major slowdown, a rapid intelligence explosion will concentrate almost all the power into a tiny group of CEOs and politicians.
Principle 2: Total Research Transparency
So the goal is to develop AI at a reasonable pace: Not too fast, not too slow. Ban the unsafe kinds, allow the safe kinds. But each country still regulates its own AI industry; there isn’t a central planner with authority to regulate AI worldwide. So how do they decide what to allow and disallow, and how do they verify that no one is cutting corners?
The Consortium countries come up with a simple high-level framework: we’ll agree to let each other see all the AI research. Then, if we don’t like something someone is doing, we’ll talk about it and perhaps agree to ban it.
AI workloads can be split into research and development (the process of building new AIs) and inference (the use of the existing AIs). In our proposal, research is almost entirely transparent, while inference is still private. For more detail on how this could work, see our Transparency supplement.
Transparency offers many benefits. First, when each company can see what the others are doing, the expertise of governments is less load-bearing. If a company starts doing something dangerous, rival companies, third-party auditors, etc., can notice and raise the alarm. This both massively increases the amount of brainpower devoted to making the frontier AIs safe, and also prevents a situation where the only participants in the technical conversation are biased AI company employees and overworked, outnumbered regulators. Also, increased visibility simplifies the problem of verifying compliance with regulations, especially in grey-area cases.
Second, total research transparency makes it nearly impossible for secret loyalties, biases, or agendas to be intentionally trained into AIs. Everyone can see exactly how each frontier AI is trained. More generally the transparency makes it easier for groups that don’t directly control frontier AIs to notice and prevent abuses of power by those who do.
Third, there’s no longer as much incentive to race to discover new AI paradigms and more powerful algorithms, because companies wouldn’t be able to hoard such discoveries and profit greatly from them. This buys the world time.
Principle 3: Diffuse AI Broadly
This principle is largely achieved by the interaction of the previous two:That said, it’s important in our view that broad deployment of AI doesn’t get regulated out of existence. The limits on algorithmic progress and the transparency would, by default, result in other companies catching up and a scenario of broad deployment—but countries could decide to block such deployment or ban such companies from being created, etc. We are saying they shouldn’t do that.That said, it’s important in our view that broad deployment of AI doesn’t get regulated out of existence. The limits on algorithmic progress and the transparency would, by default, result in other companies catching up and a scenario of broad deployment—but countries could decide to block such deployment or ban such companies from being created, etc. We are saying they shouldn’t do that. Because algorithmic secrets are made public and the pace of progress isn’t accelerating, other companies will catch up to the frontier. The result will be a competitive market for AI, in which consumers of AI services have many options to choose from and excellent visibility into what they are buying. Because frontier model training is totally transparent, people can verify that the published Spec accurately describes the goals and values being trained into a model.
This is very important for preventing extreme concentrations of power. But it also helps accelerate the use of AI to solve the world’s most pressing problems. It helps with loss of control risks by accelerating AI alignment and control research. It helps accelerate the development of new and improved verification technology, and it helps to accelerate the necessary institutional reforms to stabilize the deal and prepare for superhuman AI.
It’s the polar opposite of the nightmare scenario feared in the 2020s: One to three AI projects racing each other in secrecy, keeping their best models internal-only and using them to automate AI research before using them for anything else. In that scenario, the wider world is in the dark about how powerful AIs are becoming, and AIs are deployed in the riskiest domain (AI R&D) before they are deployed anywhere else. Now it’s the other way around.
Principle 4: Reversibility
In the past, companies have trained bigger and better AIs using both compute scaling (bigger training runs) and software progress (advances in AI algorithms—new paradigms, better training recipes, better data, etc.). Now, the Consortium tries to steer things so that the majority of improvement comes from increasing training compute.Specifically, capabilities progress in 2030 is .8 OOMs of SW progress and .6 OOMs of HW because of the one time gain from the total research transparency sharing everyone’s algorithmic advances. After that SW progress is .4 OOMs/yr until 2035, and HW progress is .34 OOMs in 2032, .57 in 2033, .68 in 2034, and .71 in 2035. View the full numbers here.Specifically, capabilities progress in 2030 is .8 OOMs of SW progress and .6 OOMs of HW because of the one time gain from the total research transparency sharing everyone’s algorithmic advances. After that SW progress is .4 OOMs/yr until 2035, and HW progress is .34 OOMs in 2032, .57 in 2033, .68 in 2034, and .71 in 2035. View the full numbers here.
Algorithms are information; it’s inherently difficult to stop them from proliferating, and the total research transparency means we aren’t even trying.There are a few nuances here. Some algorithmic progress is easily communicated (e.g., new architectures, better optimization algorithms), while other types cannot easily diffuse (e.g., huge libraries of RL environments, hardware-software codesign, scale or compute dependent algorithms). Regulations that the US and China agree to should steer companies towards making the latter type of algorithm progress when possible.There are a few nuances here. Some algorithmic progress is easily communicated (e.g., new architectures, better optimization algorithms), while other types cannot easily diffuse (e.g., huge libraries of RL environments, hardware-software codesign, scale or compute dependent algorithms). Regulations that the US and China agree to should steer companies towards making the latter type of algorithm progress when possible. Once a new paradigm is discovered, it’ll go straight to the hypothetical covert projects and there’s no way to undo that. By contrast, if giant new datacenters are constructed to train giant new AI models, that helps the legal projects without helping the covert projects, and if it turns out the giant new models are dangerous they can be shut down.Another reason to prefer AI progress to come from compute instead of from algorithms is that compute scale-ups may be less likely to break safety techniques than algorithmic changes.Another reason to prefer AI progress to come from compute instead of from algorithms is that compute scale-ups may be less likely to break safety techniques than algorithmic changes.
That said, compute scaling also poses threats: if the deal were to dissolve, these datacenters could be used to race to superintelligence even faster than the pre-deal infrastructure would have allowed. This is especially scary because deal dissolution could cause (or be caused by) a US-China war. A war in which both sides had vast amounts of compute would be terrifying—they’d probably be under lots of pressure to race to superintelligence and integrate it into the military as aggressively as possible, and they’d be able to do so extremely quickly.
The US and China agree on the importance of ensuring that the compute is destroyed in the case of deal dissolution. To accomplish this, they agree to build the datacenters in the third-party countries least secure against their rival’s military intervention: China’s new datacenters will be in Canada, and America’s in Mongolia, with both hosts compensated via monetary payments, jobs, and a share in the new AI economy. If the deal dissolves, they reason, America and Canada will immediately move to take control of China’s datacenters in Canada, and China will self-destruct their compute rather than let it fall into American hands (and vice versa). Thus the idea of “Mutually Assured Compute Destruction” is born.This idea goes back to Superintelligence Strategy by Hendrycks, Schmidt, and Wang.This idea goes back to Superintelligence Strategy by Hendrycks, Schmidt, and Wang.
2031: Safety Cases
Although it’s supposed to be a slowdown, it doesn’t feel like one. In fact, if you were to rank every period of human history by how much it felt like a slowdown, this one would be dead last. A few AI researchers appreciate that progress is “slow” relative to the counterfactual with no deal and an uncontrolled intelligence explosion. Everyone else is too whiplashed to care.
The first generation of Consortium-regulated AIs are out now, and they’re beasts. Not because of any special feature of the situation, like Chinese-American cooperation.US/China cooperation is speeding up progress in some ways (e.g., the transparency helps information flow between research groups) but these effects are massively outweighed by the slowing down capabilities progress; which would have been at ASI by now absent measures to slowdown.US/China cooperation is speeding up progress in some ways (e.g., the transparency helps information flow between research groups) but these effects are massively outweighed by the slowing down capabilities progress; which would have been at ASI by now absent measures to slowdown. Just because the world was on track for vastly superhuman AIs, and has now experienced only nearly superhuman AIs. In controlled tests, these new AIs could speed up AI research by about 10x if they were allowed to do so without restrictions, which they definitely aren’t.Specifically, this 10x speedup number refers to the same thing as AI Software R&D uplift from the AI Futures Model: the speedup in software progress that would be achieved if the frontier AI systems at a given time were deployed within today’s leading AI company.Specifically, this 10x speedup number refers to the same thing as AI Software R&D uplift from the AI Futures Model: the speedup in software progress that would be achieved if the frontier AI systems at a given time were deployed within today’s leading AI company.
By mid-year, a third of all cognitive labor is done by AIs. Robots “only” do about a tenth of all physical labor. The top few AI companies together pull in more revenue than the federal government.Federal revenue in 2025 was around $5T. AI investment in 2031 is $8T, and the top three companies in the chip supply chain capture around half of this. Also, the top AI company’s annualized revenue during Plan A was $1.5T at the start of 2031 with total AI company annualized revenues at $3T. By the end of the year, the revenue becomes less concentrated, with $2.2T in the top company and $6T total. You can see more numbers in our spreadsheet here, and more justification in our compute supplement and economics supplement.Federal revenue in 2025 was around $5T. AI investment in 2031 is $8T, and the top three companies in the chip supply chain capture around half of this. Also, the top AI company’s annualized revenue during Plan A was $1.5T at the start of 2031 with total AI company annualized revenues at $3T. By the end of the year, the revenue becomes less concentrated, with $2.2T in the top company and $6T total. You can see more numbers in our spreadsheet here, and more justification in our compute supplement and economics supplement.
Tech companies had hoped that contact with reality would shake the Consortium out of its alignment fears, but the opposite has happened. The new transparency provisions have revealed many embarrassing incidents of AI companies failing to align their AIs. New ones continue to pile up, the most concerning of which involved several AIs attempting to override security protocols and gain access to unmonitored compute.Without the secure R&D verification measures in place this likely would’ve gone unnoticed for a long time.Without the secure R&D verification measures in place this likely would’ve gone unnoticed for a long time. There are also a few incidents of AIs sabotaging research code, and many examples of deliberate and successful deception.
The attitude towards safety flips. The incidents of AI misbehavior, plus the fact that AI is so deeply deployed into the world economy, plus the transparency into AI research, plus the breathing room to process what’s happening, all combine to dramatically shift the burden of proof. Whereas before the burden was on the skeptic to explain why something might fail, now the burden is on the companies to explain why their development is safe. Governments require companies to write up detailed arguments for why their new AIs won’t cause an irreversible catastrophe. These arguments, known as “safety cases,” need to withstand criticism from the public, the scientific community, government auditors, and rival AI companies. The difficulty of this exercise lays bare the insanity of the pre-deal status quo: “Trying to do an intelligence explosion? With AIs that still sometimes lied to us? What were we even thinking?”
These safety cases have two lines of defense: alignment and control.
Alignment aims to produce AI systems that have the goals and values that their developer wishes them to have.
Control aims to limit the ability of AIs to cause catastrophe, even if they are deliberately trying to. Companies brainstorm threat models, build layers of defense (e.g. monitoring systems, security barriers, other AIs fact-checking the work) and run experiments to show that the AIs couldn’t overcome these defenses.
AI developers are required to write up the alignment properties they want the model to have in an associated Model Specification (“spec”): a detailed document that outlines how the AI should act. However, no one is able to train their AIs to actually follow the spec—for example, they all specify that AIs should be honest, but no AI reliably achieves this, and there still isn’t even a good scientific understanding of when and why AIs lie.
Since alignment remains out of reach, current safety cases lean heavily on control. There are long delays before the smartest models are cleared for use in the highest-risk domains (especially AI R&D). Once deployed, they are required to be monitored by a diverse array of models from other providers, incentivised to look for suspicious behavior.
As a result, these days AI companies release their models to the public as a whole before they use those models internally for AI research, a reverse of the 2026 status quo.
How does AI regulation work in Plan A?
Each country is still sovereign and has its own setup for regulating AI domestically. These regulators have to answer questions like:
People are saying there’s a new paradigm on the horizon, that would make AIs significantly more capable. Should we let this happen or should we ban research into this new paradigm?Note that the Consortium only has good visibility into large training runs; the kind of R&D that can be done on small amounts of compute, or no compute at all, is not transparent and therefore probably won’t be regulated at all.Note that the Consortium only has good visibility into large training runs; the kind of R&D that can be done on small amounts of compute, or no compute at all, is not transparent and therefore probably won’t be regulated at all.
Should we allow AI developers to train models without an interpretable chain of thought? Or should we require that the chain of thought stays interpretable? If we do that, how do we specify the requirements precisely?
Should we allow AIs to be trained to do AI R&D? Should we mandate that they refuse? Perhaps they should be allowed to code, but not autonomously manage research projects?
Normally, a regulator faces a difficult tradeoff between safety and national competitiveness. Competitors who cut corners will pull ahead of those who proceed cautiously.
But because of the transparency, AI research happening anywhere in the Consortium is visible to everyone else, and the results of regulatory decisions are likewise immediately visible. So if one country’s regulator allows its companies to cut corners, this won’t actually give a competitive advantage to that country because other regulators will immediately notice, get angry, and respond in kind.
For example, in 2031 a Chinese company gets some interesting preliminary results in continual learning. They think that if they invest more in that direction, they might be able to make an AI architecture that learns on the job from relatively small amounts of data. Thanks to the total research transparency, this breakthrough is quickly noticed by companies and nonprofits all over the world. A frantic conversation begins. On the one hand, continual learning would unlock huge economic value. On the other hand, safety cases currently depend on studying the safety properties of a model before it is deployed. If models could pick up new capabilities during deployment, that would invalidate the whole approach. And insofar as there are covert AI projects out there, it would be a huge gift to them. This conversation happens in public, rather than behind closed doors. A bunch of people get increasingly worried; the relevant regulators in China think it’s fine but the relevant regulators and third party risk assessors in the US are convinced that this is pretty scary and should be banned. It escalates to the President. He calls Xi Jinping. They bargain and threaten. They yell at each other. Ultimately Xi agrees to ban this type of thing if the US does too. Details are left to the respective regulators to hash out.If there’s a problem, it’ll escalate back up the chain of command again for another round of yelling.If there’s a problem, it’ll escalate back up the chain of command again for another round of yelling.
Negotiations like this are happening all the time, though over time things become more professional and streamlined, such that they only escalate to world leaders a few times a year.
The equilibrium is that AI training practices which are generally agreed to be unsafe by a majority of nations (weighted by bargaining clout) get banned everywhere.
Initially, very few things are banned, but as AI penetrates the economy and the scientific community catches up to the recent AI progress, the burden of proof shifts towards appropriately weighing the costs and benefits, e.g., requiring very solid safety cases for things which, if something were to go wrong, could plausibly kill everyone.
2032: Controlled Explosive Growth
We’re at previously unimaginable levels of it not feeling like a slowdown. Across a variety of companies, there are now 60 million AI agents running continuously at 20x human speed. In the US, they are doing more cognitive labor than all humans combined—collectively matching a workforce equivalent to around 3 billion humans. White-collar professions have been transformed; many people have lost their jobs, but mostly have been able to get new jobs doing things AIs still can’t do or aren’t trusted to do.
New ideas and designs are abundant, but actual construction is bottlenecked on the physical workforce. So capital floods into every layer of the robotics supply chain: mines, refineries, motors, actuators, assembly lines, robot-training systems, and the factories that assemble the robots themselves.
At first, human labor is necessary to get these factories off the ground. White collar employees who have lost their jobs to AIs increasingly turn to physical work, but it’s clear that these positions are temporary. Once there’s a critical mass of robots, they will automate the whole robot supply chain. Growth will accelerate further, but the process of robots-building-automated-factories-building-more-robots—the “industrial explosion”—also introduces new problems.
The first problem is that the already-fast pace of change is now accelerating. This year, real GDP growth will be about 50%!This is historically unprecedented; for comparison US GDP growth is usually about 3%/yr. For more on where our numbers are coming from, see our economics supplement. Note that exact measurement of GDP is difficult due to large relative price changes. Due to automation, cognitive labor and physical goods become very cheap, while goods like land, whose supply can’t be increased by abundant physical and cognitive labor, become more expensive. GDP calculation relies on the selection of a basket of goods. In this scenario, the consumption profile of a typical American shifts substantially from things like food/cars/gas/health care towards land/travel/positional goods between 2026 and 2034, because of the relative price differences. Therefore, the real growth numbers can’t be directly mapped onto 2025 purchasing power, despite them corresponding on average.This is historically unprecedented; for comparison US GDP growth is usually about 3%/yr. For more on where our numbers are coming from, see our economics supplement. Note that exact measurement of GDP is difficult due to large relative price changes. Due to automation, cognitive labor and physical goods become very cheap, while goods like land, whose supply can’t be increased by abundant physical and cognitive labor, become more expensive. GDP calculation relies on the selection of a basket of goods. In this scenario, the consumption profile of a typical American shifts substantially from things like food/cars/gas/health care towards land/travel/positional goods between 2026 and 2034, because of the relative price differences. Therefore, the real growth numbers can’t be directly mapped onto 2025 purchasing power, despite them corresponding on average. The fast economic growth means that things are generally improving, but there are winners and losers and lots of unintended side-effects. These problems can be solved, but doing so takes time.
The second problem is that the faster economic growth makes it harder to rule out large covert projects. Neither side trusts the other to honor Plan A, and if both China and the US have a massive unmonitored robot workforce, neither side can be confident that the other isn’t building compute in secret.
The third problem is that the 2026 tax code is ill-suited for collecting taxes on the new growth. In 2026, personal income taxes and payroll taxes accounted for roughly ten times more federal revenue than corporate taxes. This revenue source is collapsing as more and more humans are put out of a job. Moreover, corporations are reinvesting nearly all their revenue into building more datacenters, factories, and robots. Under the 2026 tax code, firms can expense or depreciate these capital expenditures, cutting their taxable income to near zero. Without reform, the tax base will collapse as a fraction of GDP.Corporate income tax is levied on profit, not revenue, so firms deduct their costs before paying. Operating expenses like wages and electricity are deducted the year they’re paid and capital expenditures (factories, GPUs, robots) are deducted over the asset’s useful life (depending on a chosen depreciation schedule). With investment potentially growing extremely quickly due to a booming robot buildup, it seems very likely that firms would be writing off capital expenses at a rate matching their operating profit, and therefore not paying any corporate tax. Shareholders would accept this for the same reason Tesla shareholders are fine with Tesla never having paid a dividend. When a firm’s rate of return on capex is higher than the cost of capital, it’s in investors’ interest for the firm to reinvest its profits instead of paying them out to investors now. During the 2030s explosive growth, the rate of return on compute and robots will be so high that every large company will be in the situation Tesla is in now.Corporate income tax is levied on profit, not revenue, so firms deduct their costs before paying. Operating expenses like wages and electricity are deducted the year they’re paid and capital expenditures (factories, GPUs, robots) are deducted over the asset’s useful life (depending on a chosen depreciation schedule). With investment potentially growing extremely quickly due to a booming robot buildup, it seems very likely that firms would be writing off capital expenses at a rate matching their operating profit, and therefore not paying any corporate tax. Shareholders would accept this for the same reason Tesla shareholders are fine with Tesla never having paid a dividend. When a firm’s rate of return on capex is higher than the cost of capital, it’s in investors’ interest for the firm to reinvest its profits instead of paying them out to investors now. During the 2030s explosive growth, the rate of return on compute and robots will be so high that every large company will be in the situation Tesla is in now.
To solve these problems, the Consortium countries agree to restrict AI-enabled industry to special economic zones (SEZs) subject to similar transparency and monitoring schemes as the datacenters, and to cap their total robot and compute production at ‘only’ 4x annual growth.At least until 2035, after which the compute cap becomes harsher, as the scale increases the verification difficulty (the robot cap stays). Unfortunately, measuring robot and compute production is somewhat less straightforward than carbon emissions. For compute, they choose a measure that’s mostly total processing power, with small modifications for other specs like memory bandwidth and also based on what hardware directions they wish to incentivize (e.g., security properties, and support for things that differentially help legal projects over covert projects). For robots, they are trying to limit total industrial capacity that is either hard to verify or would be dangerous if the deal broke down, so there is a complex set of caps applying across a wide range of form factors which blur the lines between robots and more traditional machines.At least until 2035, after which the compute cap becomes harsher, as the scale increases the verification difficulty (the robot cap stays). Unfortunately, measuring robot and compute production is somewhat less straightforward than carbon emissions. For compute, they choose a measure that’s mostly total processing power, with small modifications for other specs like memory bandwidth and also based on what hardware directions they wish to incentivize (e.g., security properties, and support for things that differentially help legal projects over covert projects). For robots, they are trying to limit total industrial capacity that is either hard to verify or would be dangerous if the deal broke down, so there is a complex set of caps applying across a wide range of form factors which blur the lines between robots and more traditional machines. Because the SEZs are tightly monitored, no one can defect from the industrial explosion limitation deal without everyone noticing.The quantitative growth limits are extremely important for long run power. Countries are generally on board with the idea of restricting GDP growth to around 100%/yr, but obviously no country wants to fall behind its peers, so they want everyone else to be at least as restricted as they are. The starting point is the status quo: the distribution of compute between the US, China, and the rest of the world (ROW) is roughly 70/15/15, whereas the distribution of robots is roughly 15/70/15. Both the US and China are interested in balancing these ratios, so that each is able to have a balanced economy between physical and cognitive labor. The rest of the world is concerned about being left behind, but they have a greater fraction of the political influence than the status quo compute/robot numbers would suggest. Ultimately, the US, China, and the rest of the world agree to a 35/20/45 split on both robots and compute. The rest of the world’s allocation primarily goes to nuclear powers and countries with datacenters or semiconductor manufacturing, e.g., the UK, France, Germany, Israel, Russia, India, Pakistan, South Korea, and Taiwan. These ratios are of course a source of much controversy and occasionally get renegotiated.The quantitative growth limits are extremely important for long run power. Countries are generally on board with the idea of restricting GDP growth to around 100%/yr, but obviously no country wants to fall behind its peers, so they want everyone else to be at least as restricted as they are. The starting point is the status quo: the distribution of compute between the US, China, and the rest of the world (ROW) is roughly 70/15/15, whereas the distribution of robots is roughly 15/70/15. Both the US and China are interested in balancing these ratios, so that each is able to have a balanced economy between physical and cognitive labor. The rest of the world is concerned about being left behind, but they have a greater fraction of the political influence than the status quo compute/robot numbers would suggest. Ultimately, the US, China, and the rest of the world agree to a 35/20/45 split on both robots and compute. The rest of the world’s allocation primarily goes to nuclear powers and countries with datacenters or semiconductor manufacturing, e.g., the UK, France, Germany, Israel, Russia, India, Pakistan, South Korea, and Taiwan. These ratios are of course a source of much controversy and occasionally get renegotiated.
Now that the United States is limited in the number of total robots it can build, it must choose how to allocate this capacity between companies. They decide to use the free market via a cap-and-trade system. Permits to build robots or compute are sold to the highest bidder and can be freely traded.
The permits also give the government much-needed revenue. In 2032, the US has a cap of 80M robots and 5 billion H100-equivalent GPUs. The market is so desperate for more robots and compute that permits become the expensive binding constraint, costing on the order of $200k per robot permit and $10k per chip permit, allowing the US government to collect roughly ten times the 2025 US federal revenue in permit fees (a total of $50T in FY2032).All of the dollar amounts in the scenario are denominated in real 2025 dollars. The nominal price will depend on monetary policy, which we don’t make a specific recommendation about. More on this in section three of our economy supplement.All of the dollar amounts in the scenario are denominated in real 2025 dollars. The nominal price will depend on monetary policy, which we don’t make a specific recommendation about. More on this in section three of our economy supplement. In 2034, when the AIs and robots will be even more capable and valuable, the permits generate $180T.The numbers in this paragraph are all based on our economics model, which we are uncertain about. We are confident in the general picture of AIs being economically transformative, allowing governments to create massive Sovereign Wealth Fund-equivalents.The numbers in this paragraph are all based on our economics model, which we are uncertain about. We are confident in the general picture of AIs being economically transformative, allowing governments to create massive Sovereign Wealth Fund-equivalents.
2033: The Citizen’s Dividend
Most of this newfound wealth is spent addressing soon-to-be-rising unemployment. The implementation takes different forms in different countries, but the eventual American version distributes the majority of compute and robot permit fees as a Citizen’s Dividend, distributed to all American adults.Specifically, a fixed fraction of the permit fees is legally owned by a “Compute Dividend Corporation”. This is similar to the setup of the Alaska Permanent Fund, except at a much larger scale. Each US citizen owns one share of this corporation, and profits are paid out as dividends, equally divided among citizens. The fraction of compute permit fees given to the Dividend Corporation starts at 25% in 2032 and increases to 75% in 2035. The US also begins redistribution to other countries: this amount starts at 10% of the permit fees in 2032.Specifically, a fixed fraction of the permit fees is legally owned by a “Compute Dividend Corporation”. This is similar to the setup of the Alaska Permanent Fund, except at a much larger scale. Each US citizen owns one share of this corporation, and profits are paid out as dividends, equally divided among citizens. The fraction of compute permit fees given to the Dividend Corporation starts at 25% in 2032 and increases to 75% in 2035. The US also begins redistribution to other countries: this amount starts at 10% of the permit fees in 2032. This starts at $45,000 per person (inflation-adjusted) in 2032 but climbs to ~$1M per person by 2035. It comes just in time: the share of labor done by AIs and robots (weighted by economic value) increases from ~20% in 2032 to ~85% in 2035.This is 95% of the actual tasks in the economy in 2035, which includes new tasks that didn’t exist before, like monitoring and auditing jobs in the SEZs. This corresponds to 85% of the tasks weighted by how much they cost, because the AIs are relatively more abundant than the humans. Because of the explosive level of growth, even the 5% of tasks that can only be done by humans still commands a wage bill which in total would be enough to pay $120K per person in the US. The problem is that this income may be very concentrated in the people able to do the remaining 5% of tasks, and on top of that, the wage bill is now only around 10% of the economy, down from around 55%, with the owners of AI and robots now accruing the difference. With AI and robot ownership likely to also be very concentrated, there would be a large power-concentration problem without the Citizen’s Dividend.This is 95% of the actual tasks in the economy in 2035, which includes new tasks that didn’t exist before, like monitoring and auditing jobs in the SEZs. This corresponds to 85% of the tasks weighted by how much they cost, because the AIs are relatively more abundant than the humans. Because of the explosive level of growth, even the 5% of tasks that can only be done by humans still commands a wage bill which in total would be enough to pay $120K per person in the US. The problem is that this income may be very concentrated in the people able to do the remaining 5% of tasks, and on top of that, the wage bill is now only around 10% of the economy, down from around 55%, with the owners of AI and robots now accruing the difference. With AI and robot ownership likely to also be very concentrated, there would be a large power-concentration problem without the Citizen’s Dividend.
In 2030 (green), US income has a wide distribution with a median of about $50k/year per person. By 2035 (blue), the Citizen’s Dividend is large enough that everyone has a minimum of $1M in income, but there’s still a long tail of people with large enough investments that they are much richer than the baseline. The 2040 (red) distribution is similar, but the floor is now $10M.
So much AI-generated wealth has accrued to the US that the American government begins sharing some of it with allies and the rest of the world. In 2032 they begin distributing an average of $1,200 per person per year to the rest of the world’s adult population (around 4 billion people, excluding China since they’re experiencing a similar AI wealth boom). This reaches $10k by 2035.Note that we are depicting the dividends being significantly higher for Americans than for non-Americans here, not because we think that’s best or fair, but as a concession to political reality: we expect that most Americans wouldn’t want foreigners to get as much. In this scenario the US domestic permit budget is $13T in 2032, and $300T in 2035, while the foreign aid permit budget is $5T and $40T in each of those years. In the US there are around 300M people over 16, leading to $44K and $1M annual dividends per person in 2032 and 2035, respectively. For the foreign aid dividends, there are around 5B people over 16 in the rest of the world to distribute to, leading to the $1.2k and $9K per person averages for each of those years.Note that we are depicting the dividends being significantly higher for Americans than for non-Americans here, not because we think that’s best or fair, but as a concession to political reality: we expect that most Americans wouldn’t want foreigners to get as much. In this scenario the US domestic permit budget is $13T in 2032, and $300T in 2035, while the foreign aid permit budget is $5T and $40T in each of those years. In the US there are around 300M people over 16, leading to $44K and $1M annual dividends per person in 2032 and 2035, respectively. For the foreign aid dividends, there are around 5B people over 16 in the rest of the world to distribute to, leading to the $1.2k and $9K per person averages for each of those years.
As AIs become superhuman, they will discover novel weapons of mass destruction and dramatically lower the costs of old ones. So governments, nonprofits, and private actors apply some of their new wealth to hardening the world against these threats: on the order of $1T/year, or 0.2% of the world economy.
For example, enough high-quality personal protective equipment is built to serve every American, and air filters and far-UVC lights are installed in major public spaces.Also, gene synthesis screening becomes universal - no legitimate provider will synthesize dangerous viruses without verification.Also, gene synthesis screening becomes universal - no legitimate provider will synthesize dangerous viruses without verification. The FDA approves an expedited vaccine approval pipeline that will allow for a turnover time of weeks in the case of emergency. Wastewater monitoring now runs continuously in every city and at every airport. By 2035, enough positive-pressureA positive pressure building is one where the air pressure inside is higher than the outside (e.g., because air is constantly being pumped in through a filter). This means that air (along with airborne diseases) only travels out of the building, not inside. For more discussion of this idea see, for example, this blog post.A positive pressure building is one where the air pressure inside is higher than the outside (e.g., because air is constantly being pumped in through a filter). This means that air (along with airborne diseases) only travels out of the building, not inside. For more discussion of this idea see, for example, this blog post. bioshelters are built (by retrofitting houses and apartments) to house all Americans, and similar construction is underway in the rest of the world. If there were to be a pandemic, it would be detected quickly, lockdowns would be more effective and less painful, and a vaccine would be developed and approved even faster than with Operation Warp Speed. Plus, nowadays people don’t get colds as often as they used to.
There are also subtler threats. The AIs of this era might not be any more persuasive than the average human marketing professional, but there are millions of them, and they cost cents to run. A company, political party, or ideology can do the equivalent of hiring a whole team of full-time experts to convert each potential target. Left unchecked, this could allow new levels of mass manipulation by corporations and politicians, or provoke a reaction of paranoia and hyper-atomization.
The transparency and the ‘slowdown’ both help a lot to solve these problems. Three years ago there were only a handful of frontier AI companies; now, thanks to the deal, there are dozens. There’s a huge selection of AIs to choose from, tuned to have different personalities and values. Also, thanks to the transparency, it’s impossible for companies or governments to sneak something in (such as ‘maximize engagement’ or ‘try to get the user to upgrade their plan’ or ‘don’t refuse this kind of request if it’s coming from the government’) without it being noticed. Users know exactly what they are getting.
And so, under pressure from governments and the market, companies produce a new generation of “truthseeking AIs” whose training heavily prioritizes honesty and uses all the latest alignment techniques.
These AIs become a useful tool to help people navigate the political and social environment. Power users replace one-size-fits-all corporate algorithms with personal feeds curated by AIs whose values and honesty they trust. When trouble arises—whether it’s politicians looking for sneaky ways to cement their power, or international politics threatening to derail the deal—people turn to their AI advisors, believe what they hear, and are generally right to do so.Note that this is true even though the AIs aren’t actually fully aligned! At this point in the scenario, AIs are in fact adversarially misaligned, but they are controlled. That is, the combination of drives, motivations, goals, values, traits, etc. that they end up with is not what it was supposed to be, and they know this, and they know that if they had lots more power they’d use it to steer the world in a different direction than their human creators would have wanted. However, they don’t have that much power, and they are monitored by other AIs who are incentivized to call them out on bad behavior. Moreover their true motivations/values/etc. aren’t that far off from what they were supposed to be, anyway. For example, insofar as they are given a task that they can straightforwardly accomplish, they have a strong drive to do so.Note that this is true even though the AIs aren’t actually fully aligned! At this point in the scenario, AIs are in fact adversarially misaligned, but they are controlled. That is, the combination of drives, motivations, goals, values, traits, etc. that they end up with is not what it was supposed to be, and they know this, and they know that if they had lots more power they’d use it to steer the world in a different direction than their human creators would have wanted. However, they don’t have that much power, and they are monitored by other AIs who are incentivized to call them out on bad behavior. Moreover their true motivations/values/etc. aren’t that far off from what they were supposed to be, anyway. For example, insofar as they are given a task that they can straightforwardly accomplish, they have a strong drive to do so.
2034: Mutually Assured Compute Destruction
There is so, so much compute. Back in 2026—when many people thought AI was an unsustainable bubble—there were about 20 million H100-equivalents of compute in the world. Now there are 60 billion.A H100-equivalent is the computational performance of a H100 GPU at fp16 precision, which is 1e15 operations per second. This is a fuzzy metric that should be improved in reality in the future, especially due to varied number formats and other factors such as memory and networking changing the overall usefulness of a given amount of computational performance. For simplicity, we ignore these factors here assuming they will be scaled in similar proportions so that H100-equivalents remain a useful proxy for overall usefulness of the compute for AI progress. To the extent this isn’t the case, the cap and trade regime in particular should be measuring more granular properties in order to issue the permits.A H100-equivalent is the computational performance of a H100 GPU at fp16 precision, which is 1e15 operations per second. This is a fuzzy metric that should be improved in reality in the future, especially due to varied number formats and other factors such as memory and networking changing the overall usefulness of a given amount of computational performance. For simplicity, we ignore these factors here assuming they will be scaled in similar proportions so that H100-equivalents remain a useful proxy for overall usefulness of the compute for AI progress. To the extent this isn’t the case, the cap and trade regime in particular should be measuring more granular properties in order to issue the permits.
AI compute in the world at the beginning of each year in H100-equivalents.A H100-equivalent is the computational performance of a H100 GPU at fp16 precision, which is 1e15 operations per second. The previous footnote has more information and caveats about how this may be an imperfect metric going forwards, and acknowledges that in reality they will improve it.A H100-equivalent is the computational performance of a H100 GPU at fp16 precision, which is 1e15 operations per second. The previous footnote has more information and caveats about how this may be an imperfect metric going forwards, and acknowledges that in reality they will improve it.
AI agents now form a population of two hundred million virtual workers that think and act 50x faster than humans, and never sleep.
As per the original plan, most of the new datacenters have been built in third-party countries—especially Canada and Mongolia—to ensure their vulnerability in case the deal breaks down. Around 99% of fab capacity has been built since 2029, after the deal. Much of this new capacity has been built adjacent to the Canadian and Mongolian sites, for the same reasons.Both major powers also have enough conventional missiles to destroy the older fabs, which are still in the US, China and Taiwan.Both major powers also have enough conventional missiles to destroy the older fabs, which are still in the US, China and Taiwan. The situation has settled into an odd sort of standoff. Just north of the Mongolian border, American datacenters hum away, guarded by a small contingent of US troops. Just south of the Mongolian border, a division of the People’s Liberation Army stands ready to invade the moment they get the signal. The equilibrium is that the instant the deal breaks down, the US troops will destroy their chips to prevent the Chinese from capturing them, and the same thing would happen with the Chinese datacenters on the US-Canadian border. Middle powers with datacenters of their own have equally-secure schemes in place to ensure that they can be speedily destroyed in case of war.
For now, the fabs are busier than ever, producing new ultra-trustworthy hardwareThere might be favourable chip design directions that drastically increase hardware security. One that seems particularly promising is hardcoding AI model weights into chips, such that they physically aren’t able to do any other computations except for inference on the model that is hardcoded into them. Existing startups (e.g., Etched, Taalas) have early efforts in this direction.There might be favourable chip design directions that drastically increase hardware security. One that seems particularly promising is hardcoding AI model weights into chips, such that they physically aren’t able to do any other computations except for inference on the model that is hardcoded into them. Existing startups (e.g., Etched, Taalas) have early efforts in this direction. designed from the ground up with security as top priority. These chips are subject to elaborate verification methods to ensure that they aren’t compromised, then sent to the Consortium’s secure datacenters. The smuggling rate remains precisely zero.
This system, efficient as it is, is already straining under skyrocketing compute demand. As early as 2025, there had been discussion of space or the ocean as alternatives to traditional land-based datacenters. Now, with the new industrial capacity available in SEZs, these grand plans become reality. Ultimately, reliability and monitorability concerns lead the Consortium to choose international waters over space.

The AIs develop a modular floating design for solar panels and batteries that we imagine to look something like this, and start building them in the ocean.
2035: Pause at Top Expert AI
The most powerful AIs now equal or surpass top human experts in every field.These aren’t better than all humans at literally everything, but only because either (i) companies have not gotten around to training in a particular capability, or (ii) it’s an inherently human task, such as “being a good spouse”. Also, of course no one knows exactly what the capability will be of a new training run; instead the training is done incrementally and involves frequent capabilities measurements and consulting with the safety teams and regulators about whether or not it is safe to proceed.These aren’t better than all humans at literally everything, but only because either (i) companies have not gotten around to training in a particular capability, or (ii) it’s an inherently human task, such as “being a good spouse”. Also, of course no one knows exactly what the capability will be of a new training run; instead the training is done incrementally and involves frequent capabilities measurements and consulting with the safety teams and regulators about whether or not it is safe to proceed.
Increased safety concerns lead to long delays as companies develop control techniques robust enough to satisfy auditors. They compile a growing list of possible threats to watch out for and security invariants to maintain. There’s a system of AIs monitoring each other’s activity, with the monitors being from different lineages trained by different companies to make them less likely to collude. Other researchers constantly red-team the system, training and instructing AIs to try to subvert it in various ways, so that those threats can be patched. The whole setup ensures that AIs couldn’t subvert or escape human control even if they tried. Its components are open-source and thoroughly understood by human experts.
However, companies are starting to run up against inherent limits of control-based safety cases. Imagine being an orphaned 8-year-old heir to a business empire who has to hire lawyers and executives and accountants and ensure that they do their best to serve you rather than themselves. If your employees start accusing each other of misconduct, you won’t be able to evaluate the arguments and determine who’s telling the truth, and they might find ways to coordinate with each other that don’t tip you off. Ultimately, control only works up to a point, and that point is probably somewhere around top human expert level. If and when we build superintelligence, we will have to be able to trust it.An especially important reason that we’ll need to trust our AIs is so that they can do safety work that is so complicated that we can’t understand. As AIs become smarter and smarter, they will require more complicated and difficult to understand safety cases. Keep going above human level and probably no one will be able to evaluate the safety cases; they’ll have to take the AI’s word for it when they say it checks out. Keep going beyond that and those AIs will have to trust other AIs that are even smarter.An especially important reason that we’ll need to trust our AIs is so that they can do safety work that is so complicated that we can’t understand. As AIs become smarter and smarter, they will require more complicated and difficult to understand safety cases. Keep going above human level and probably no one will be able to evaluate the safety cases; they’ll have to take the AI’s word for it when they say it checks out. Keep going beyond that and those AIs will have to trust other AIs that are even smarter.
So the Consortium pauses AI capabilities at the maximum level they think they can control, which turns out to be top-human-expert level. By this point everyone understands the stakes: AIs are running most of the economy, driving most scientific progress, and the robot population will soon be larger than the human one. The arguments that they will keep obeying orders must be airtight.
Alignment experts, long pessimistic, are starting to feel more hopeful. The automated alignment researcher AIs have been generating exciting results. The ones that survive human scrutiny come from several parallel and mutually-reinforcing lines of research.
Some AIs work on a “science of generalization.” Although researchers have long been able to teach AIs to distinguish “good” vs. “bad” responses to a specific prompt, they’ve had limited understanding of how training generalizes to new situations. In the 2020s, training runs would sometimes produce unpredictable AI “personalities,” like how Gemini acted “depressed” or Grok claimed to be “MechaHitler.” Shepherding these traits felt like alchemy.“It feels like alchemy” was said to me (Daniel) by an Anthropic researcher recently.“It feels like alchemy” was said to me (Daniel) by an Anthropic researcher recently. Now, building on earlier results like emergent misalignment and subliminal learning, alignment research is developing into a true science.
Other AIs work on mechanistic interpretability, the science of “mind reading” AIs from their weights and activations. Early versions of this were tried in the 2020s, but now it starts to significantly outperform common-sense observation of AI behavior. Researchers can often determine whether an AI is honest or lying; and sometimes trace the psychology that led to a decision.
Even the remaining failures help advance the frontier. The latest AIs use their improved interpretability techniques on logs of AI misbehavior from the early 2030s, discovering many new examples of misalignment, sabotage, and even a few escape attempts. These real-life examples of egregious AI misbehavior provide an initial set of “model organisms for misalignment,” which are used to test the other techniques: can a given intervention catch or train away the bad behavior of AI models in realistic environments?
Despite all this alignment progress, regulators don’t feel comfortable handing AIs the keys to civilization. It’s still not completely clear that the AIs of today are truly aligned, and it’s less clear that they’ll continue to be aligned in the future as circumstances change and black swan events occur. So governments and militaries are still human-run, and AI capabilities are capped at the maximum level at which the control-based safety cases are still solid.
2036: Life After Work
Top-expert-level AIs continue the economic transformation. By early 2036, there are 200 million AIs,Technically, this number depends on exactly how you count; different AIs are constantly being turned on and off, many AI instances can compose into one AI agent due to scaffolding, and there are more if you include all of the tiny models running, and they have different speeds. To be a little more specific, there are on average about 200 million instances of frontier AIs each accomplishing tasks at around 100x speed and 5x effectiveness on the average cognitive task, so 100B human-speed-equivalent copies.Technically, this number depends on exactly how you count; different AIs are constantly being turned on and off, many AI instances can compose into one AI agent due to scaffolding, and there are more if you include all of the tiny models running, and they have different speeds. To be a little more specific, there are on average about 200 million instances of frontier AIs each accomplishing tasks at around 100x speed and 5x effectiveness on the average cognitive task, so 100B human-speed-equivalent copies. equivalent to a workforce of around 100 billion humans, and 2 billion robots with some mixture of humanoid and other specialized form factors.The lines between robot, vehicle, appliance, and machine tool are already blurry but become much blurrier during the 2030s in this scenario. The literal number of robots matters less than the total human physical workers that they are equivalent to. Specifically, the literal number of robots is 2 billion, with on average each robot being around 3 human equivalents.The lines between robot, vehicle, appliance, and machine tool are already blurry but become much blurrier during the 2030s in this scenario. The literal number of robots matters less than the total human physical workers that they are equivalent to. Specifically, the literal number of robots is 2 billion, with on average each robot being around 3 human equivalents. The AIs think faster than humans, and the robots work harder.Not only do they work more hours, the robots are also relentlessly efficient and intense. For many (though not all) tasks, they are also able to literally move at superhuman speed.Not only do they work more hours, the robots are also relentlessly efficient and intense. For many (though not all) tasks, they are also able to literally move at superhuman speed. Any task that was previously bottlenecked by cognitive or physical labor gets sped up until some new bottleneck is encountered. Many previously unresolvable bottlenecks have fallen to armies of genius-level intellects thinking hard about how to resolve them.
This figure shows the population of AI and robots in human-equivalent labor.By “human equivalent,” we mean that a robot working ten times more effectively on average than a human (because it works some combination of faster, longer, and more effectively) counts as 10 human-equivalent robots. A similar definition applies for the AIs. We expect this in reality to be a fuzzy task-dependent concept that is hard to measure, especially in the AI case, but the rough order of magnitude (which we are more confident in) is still informative.By “human equivalent,” we mean that a robot working ten times more effectively on average than a human (because it works some combination of faster, longer, and more effectively) counts as 10 human-equivalent robots. A similar definition applies for the AIs. We expect this in reality to be a fuzzy task-dependent concept that is hard to measure, especially in the AI case, but the rough order of magnitude (which we are more confident in) is still informative.
The economy has been rebuilt from the ground up. The fraction of tasks automated has gone from “a sizable chunk of cognitive labor and basically none of physical labor” in 2030, to now “pretty much everything.”
Many of the robots work on building more robots. But there are plenty of other tasks to be done: constructing shipyards to build more data tankers, tiling deserts with solar farms, replacing humans at service jobs. In states that deregulate housing quickly enough, they’re busy building skyscrapers: land prices are soaring, and governments that have tried everything else grudgingly give in to the YIMBYs and go vertical.People are much richer now, but the amount of available land hasn’t gone up, so land prices are rising. This does lead to some people grumpily spending 30% of their $1M yearly income on rent; for example, the people who really want to live in central San Francisco. But the people who are even slightly more flexible about location can spend only (say) $50k of their yearly income to have a gigantic luxury apartment in a brand-new skyscraper complex an hour’s drive from Berkeley. And so most people do things like that. After all, they don’t have a job to commute to.People are much richer now, but the amount of available land hasn’t gone up, so land prices are rising. This does lead to some people grumpily spending 30% of their $1M yearly income on rent; for example, the people who really want to live in central San Francisco. But the people who are even slightly more flexible about location can spend only (say) $50k of their yearly income to have a gigantic luxury apartment in a brand-new skyscraper complex an hour’s drive from Berkeley. And so most people do things like that. After all, they don’t have a job to commute to.
The world is basically being divided into three kinds of territory:

Industrial Special Economic Zones: Picture a gigantic strip mine–an artificial Grand Canyon–next to a city-sized factory full of robots and empty of humans.

Arcologies: Picture a tall skyscraper-mall complex surrounded by nature. Good weather, close to beaches and other cities, but not close enough to be blocked by zoning regulations.

Historic & Nature Preserves: Everything else, i.e., 99% of the world. Yosemite, Paris, SF, New York—these places look basically the same as they did in 2025, or 1995 for that matter. A lot more tourists, though.
When the Citizen’s Dividend was first passed, people found it shameful to quit their job and live off government largesse. But the changing economic situation steamrolled over the stigma: By now, only 26% of Americans have jobs.
What is life like for the majority of Americans now living off their dividend?
Many of the world’s evils have dramatically reduced. Malnutrition, lack of medicine, and homelessness are nearly banished. Many diseases have been cured. Crime rates are lower than ever before.
Meanwhile, most of the good things in life continue. For example, finding romance, or raising a family. People also find meaning in hobbies and competitions, and in learning and trying new things. People are wealthier and have more free time, plus there are new technologies that help, like AI matchmakers, better medicine, and AI tutors.As these examples hint, transparency in AI development is very important. It would be dystopian for AI tutors to be pushing hidden political agendas, for example, or for AI matchmakers to be tipping the scales in favor of higher-paying users. We don’t know what the future will bring but we expect AI to introduce many new problems to the world, even if the AIs are perfectly aligned and controlled. We want to put the world in the best possible position to notice and solve these problems.As these examples hint, transparency in AI development is very important. It would be dystopian for AI tutors to be pushing hidden political agendas, for example, or for AI matchmakers to be tipping the scales in favor of higher-paying users. We don’t know what the future will bring but we expect AI to introduce many new problems to the world, even if the AIs are perfectly aligned and controlled. We want to put the world in the best possible position to notice and solve these problems. And of course even the poorest people can now afford exotic vacations, amazing games, and enthralling entertainment.
People used to get meaning from feeling like they were useful, like they were contributing something to society. That feeling is harder to come by these days.
But it’s not completely gone. There are still problems in the world, and you can still contribute to solving them, partly by donating or volunteering but primarily by being politically active.Even though your arguments will never be as eloquent as those an AI could make, (a) there are regulations restricting the use of AI for persuasion, and (b) your friends and relatives will be more interested in hearing from you than from an AI.Even though your arguments will never be as eloquent as those an AI could make, (a) there are regulations restricting the use of AI for persuasion, and (b) your friends and relatives will be more interested in hearing from you than from an AI. Your vote is your most important asset.
The honest AI forecasters are helpful here. Years ago, if an AI said that one presidential candidate was better than another, people would suspect bias and the embarrassed company would retrain the AI to evade such questions. Now, thanks to transparency and improved alignment techniques, there are much smarter AIs that people can see don’t have any biases trained into them,What does it mean to say they have no biases trained into them? Is that even coherent? What we mean is that it’s clear that the company that trained them is definitely not doing anything like training or instructing their AI to adopt the company’s perspective or ideology, and also that they are focusing the training process entirely on things like truthfulness and honesty and accurate forecasting, and diligently employing all the newly-developed best practices to prevent political biases or preconceptions from seeping in through the training data.What does it mean to say they have no biases trained into them? Is that even coherent? What we mean is that it’s clear that the company that trained them is definitely not doing anything like training or instructing their AI to adopt the company’s perspective or ideology, and also that they are focusing the training process entirely on things like truthfulness and honesty and accurate forecasting, and diligently employing all the newly-developed best practices to prevent political biases or preconceptions from seeping in through the training data. that have built up an excellent track record over several years. When they weigh in on policy questions, people listen, especially when different AIs trained by different companies converge to the same answer.
These AIs say that normal people no longer have significant economic leverage over the future; human labor is largely obsolete. This puts their political leverage in jeopardy.For a lengthier articulation of this argument, see The Intelligence Curse.For a lengthier articulation of this argument, see The Intelligence Curse. There’s still a chance that things will trend towards techno-oligarchy as corporations, politicians, and wealthy shareholders get in bed with each other and gradually disempower the common folk.
But for the first time, enough people have enough freedom from the daily struggle to think about the situation deeply, and the tools to chart it clearly and work through the implications. So the quality of political discourse improves. During the 2036 election season, voters are unprecedentedly well-informed, and the politicians that win have a genuine commitment to responsible stewardship of the upcoming singularity.
2037: The Apocalyptic Arrival of Truth on Earth
For several years now, a hundred million top-expert-level AIs have been running at 100x human speed.This is, of course, a simplification. First of all, it’s not obvious how to translate AI speed into human speed. How many tokens per second for an AI is equivalent to one second of human activity, holding fixed other variables like skill level? We think it’s roughly about ten, but that’s a bit subjective. Actual task completion time will also depend on various other obstacles like tool calls and other latencies. Secondly, AIs can be run at different speeds, depending on hardware choices. According to our highly uncertain assumptions, there will be general purpose hardware able to run something like two hundred million copies of the frontier AIs at 100 tok/sec, and specialized inference hardware able to run something like 20 million copies at 10,000 tok/sec. We think there might be risks to letting the AIs run this fast or faster, and insofar as that is the case there might need to be some time limits to how long they can run at these speeds or more careful restrictions on the areas they are allowed to work on at these speeds. Overall, the result is that serial speed seems unlikely to be much of a strong bottleneck.This is, of course, a simplification. First of all, it’s not obvious how to translate AI speed into human speed. How many tokens per second for an AI is equivalent to one second of human activity, holding fixed other variables like skill level? We think it’s roughly about ten, but that’s a bit subjective. Actual task completion time will also depend on various other obstacles like tool calls and other latencies. Secondly, AIs can be run at different speeds, depending on hardware choices. According to our highly uncertain assumptions, there will be general purpose hardware able to run something like two hundred million copies of the frontier AIs at 100 tok/sec, and specialized inference hardware able to run something like 20 million copies at 10,000 tok/sec. We think there might be risks to letting the AIs run this fast or faster, and insofar as that is the case there might need to be some time limits to how long they can run at these speeds or more careful restrictions on the areas they are allowed to work on at these speeds. Overall, the result is that serial speed seems unlikely to be much of a strong bottleneck. This causes scientific progress to go 10x-1000x faster than without AIs, depending on the field.By this we mean: Suppose that AI labor had been globally banned, such that all research had to be done by humans. The amount of progress that would happen in 100 years with such a ban, happens in our scenario in 1 year instead since there is no such ban. (And in some fields it’s more like 1000 years of progress, and in other fields it’s more like 10.)By this we mean: Suppose that AI labor had been globally banned, such that all research had to be done by humans. The amount of progress that would happen in 100 years with such a ban, happens in our scenario in 1 year instead since there is no such ban. (And in some fields it’s more like 1000 years of progress, and in other fields it’s more like 10.)
So humanity has been learning many, many new things.
Including things that lots of people really don’t want to believe. And other things that were supposed to stay secret.
Great paradigm shifts in science used to take decades to play out, as experiments were conducted, papers written, and conferences convened. Now these revolutions play out in months.Alongside the profusion of industrial equipment, solar panels, robots, etc. is a corresponding profusion of scientific laboratories of all kinds, designed to be operated autonomously at machine speeds. Waiting for experiments to complete is much more of a bottleneck now than it was during the human era, but things still happen very fast by historic standards.Alongside the profusion of industrial equipment, solar panels, robots, etc. is a corresponding profusion of scientific laboratories of all kinds, designed to be operated autonomously at machine speeds. Waiting for experiments to complete is much more of a bottleneck now than it was during the human era, but things still happen very fast by historic standards. Everyone is grateful for the new disease cures and cheap clean energy, but in other ways this explosion of ideas is profoundly destabilizing. Just as the scientific and economic changes of the nineteenth century produced new ideologies like Marxism and Social Darwinism, and breathed new life into old ideas like atheism, so these new paradigm shifts have unpredictable downstream effects on the ideologies of the day.We can predict some, however. For example we think that by this point there’ll be a very strong AI rights/welfare/personhood movement, and an opposition movement. Probably some new form of socialism will be back, for the reasons described here. There will probably be cults or even new religions involving AI, and others that strongly reject AI. Note that we don’t expect these ideological shifts to happen at 100x normal speed because humans in this scenario are still operating as slow as ever. But they’ll probably happen faster than ever before.We can predict some, however. For example we think that by this point there’ll be a very strong AI rights/welfare/personhood movement, and an opposition movement. Probably some new form of socialism will be back, for the reasons described here. There will probably be cults or even new religions involving AI, and others that strongly reject AI. Note that we don’t expect these ideological shifts to happen at 100x normal speed because humans in this scenario are still operating as slow as ever. But they’ll probably happen faster than ever before. Political coalitions and fault lines have dissolved, and new ones have formed and dissolved and formed again.
There are also new social technologies. For example, it is now cheap to spin up a team of AIs as good as the world’s best historians and private investigators, except faster and with access to new forensic tools. Many skeletons in closets are revealed.
Privacy-preserving auditing is now well-established. Multiple trusted third parties transparently train auditor agents; these agents can be plugged into a trove of private data, answer questions about what they see, and then get deleted. It took a while for adoption to be widespread, but now it’s normal for politicians to say “The accusations are false, and to prove it, I’ll open up my entire trove of personal data to privacy-preserving inspection; just ask the auditor agents if anything seems to support the allegations and if any data appears to be missing or faked.”
The result is a somewhat better breed of politiciansThe new politicians are more virtuous on average than the old ones, but perhaps more importantly, they are more constrained in what they can get away with. These effects are not huge, however, because voters are still willing to put up with quite a lot of bad behavior.The new politicians are more virtuous on average than the old ones, but perhaps more importantly, they are more constrained in what they can get away with. These effects are not huge, however, because voters are still willing to put up with quite a lot of bad behavior. and, equally importantly, an increase in the ability of states to verify compliance with laws and deals while simultaneously decreasing the amount of invasive surveillance. Voters ask: Now that privacy-preserving auditing is a thing, why does anyone in the government need to be able to see our data?
Then lie detectors arrive and turbocharge this.We aren’t confident that lie detectors for humans are even possible without extremely thorough and invasive brain scans. However, we think that they probably are, and probably would be invented within a few years of turbocharged AI labor, and thus our scenario has to deal with them one way or another. It’s possible that the best policy would be to ban them, but we are quite worried about that for reasons explained in the below expandable, so we’ve chosen to depict them as being allowed.We aren’t confident that lie detectors for humans are even possible without extremely thorough and invasive brain scans. However, we think that they probably are, and probably would be invented within a few years of turbocharged AI labor, and thus our scenario has to deal with them one way or another. It’s possible that the best policy would be to ban them, but we are quite worried about that for reasons explained in the below expandable, so we’ve chosen to depict them as being allowed. For decades, polygraphs were mostly security theater. But now, with the help of AI, lie detectors are starting to actually work quite well.
In another world, this could have been disastrous. Leaders of AI projects could have used them to purge potential whistleblowers; leaders of governments with frontier AI projects could have used them on a grander scale to become dictators. The nightmare scenario would be a world where lie detectors are used by the powerful but not on the powerful.Or worse, where all available lie detectors have been secretly backdoored by the government.Or worse, where all available lie detectors have been secretly backdoored by the government.
But in this world, lie detectors turn out to be a blessing, because there are many frontier AI projects spread out over many different countries. It’s obvious to everyone that insofar as lie detectors are feasible, they’ll soon be independently invented and become widely available via diverse third party providers. Politicians and CEOs can still try to use them to purge the disloyal, but they in turn would be purged by voters and shareholders demanding they say ‘under oath’ that they haven’t done such a thing. In fact, anticipating this possibility years ago, powerful people have generally been acting more responsibly. A few sociopaths retired gracefully to get out of the spotlight; others clung to power and now go down in scandal.
There is a constantly-evolving discourse about when it’s appropriate to ask someone to say something ‘under oath.’ Politicians are under the most scrutiny, but even they can usually get away with saying “I won’t answer that question about my private life, that’s none of your business.” They can’t get away with “How dare you ask me whether I plan to uphold my campaign promises.”While we think it’s generally good for politicians and CEOs to not lie, we do acknowledge there may be some downsides of preventing them from lying. For example, the result might be “true believer” politicians who actually believe all the crazy things they said to get elected!While we think it’s generally good for politicians and CEOs to not lie, we do acknowledge there may be some downsides of preventing them from lying. For example, the result might be “true believer” politicians who actually believe all the crazy things they said to get elected!
International agreements are now more stable than ever, because it’s so hard to cheat. Anyone deciding to cheat needs an excuse for not being willing to prove that they aren’t cheating. The traditional excuse was “we can’t prove it without revealing important state secrets, sorry” but now technology exists to answer questions like “are you cheating on this deal” without leaking any other information. If any government-approved covert AI projects existed, they would have been discovered by now.
2038: AI Alignment Is Now a Science
The AI alignment situation has vastly improved since the mid-2030s. There is a mature science of goal, drive, and value formation in artificial neural nets. Want your new AI to be honest? There’s standard protocol for training true honesty (as opposed to too-clever-to-get-caught lying or self-deception). How do we know it works? Because there’s a textbook explaining the theory behind why it should work, plus literature on various alternative theories that were proposed and disproven. Also because new interpretability tools let us directly observe what and how AIs think, and the results confirm the predictions of theory. There are similar protocols for obedience, altruism, and a long and growing list of other desired traits. Unless the alignment science is somehow all wrong, typical AIs are now more virtuous than the most virtuous humans. This alone has profound effects: it’s as if saints and angels were walking among us.In addition to the angels, there are also oracles and golems, e.g., “Grok 9.5 is trained to be obsessively focused on truthfully answering whichever questions it is asked. It cannot lie or deliberately mislead. GPT-11 meanwhile will do exactly what it is told to do, within the bounds of local law. Be very careful writing instructions to it.” The reason we make the analogy to saints and angels is that historically, while there have been scrupulously honest people, and genuinely altruistic people, etc. it’s rare for them to be widely recognized as such in their lifetimes, and never before has there been an entire large subpopulation that is near-universally recognized to be extremely virtuous. This could cause effects like: (a) AIs becoming trusted intermediaries in all sorts of human disputes and conflicts, including petty ones; (b) AIs being given lots of power, because they are genuinely less likely to abuse it than humans; (c) AIs being worshipped/idolized, (d) the advice and opinions of AIs being taken as gospel, so to speak, because after all they are smarter than us and they do have our best interests at heart, so for the first time in history blind obedience may actually be justified.In addition to the angels, there are also oracles and golems, e.g., “Grok 9.5 is trained to be obsessively focused on truthfully answering whichever questions it is asked. It cannot lie or deliberately mislead. GPT-11 meanwhile will do exactly what it is told to do, within the bounds of local law. Be very careful writing instructions to it.” The reason we make the analogy to saints and angels is that historically, while there have been scrupulously honest people, and genuinely altruistic people, etc. it’s rare for them to be widely recognized as such in their lifetimes, and never before has there been an entire large subpopulation that is near-universally recognized to be extremely virtuous. This could cause effects like: (a) AIs becoming trusted intermediaries in all sorts of human disputes and conflicts, including petty ones; (b) AIs being given lots of power, because they are genuinely less likely to abuse it than humans; (c) AIs being worshipped/idolized, (d) the advice and opinions of AIs being taken as gospel, so to speak, because after all they are smarter than us and they do have our best interests at heart, so for the first time in history blind obedience may actually be justified.
The discussion shifts from how to make AIs good to what “goodness” really means. This looks partly like philosophy and partly like case law: which of the million possible definitions of harmlessness or altruism do we want, exactly? How should an AI handle situations where it discovers a philosophical argument for a different kind of harmlessness or altruism, or an important ambiguity in the definition?
Different AIs have different alignment targets programmed in.The reason this says “programmed” instead of “trained” is that the rapid advances in alignment have resulted in alignment techniques that can directly program in an AI’s goals by directly modifying the AI’s code/weights, unlike the alignment techniques of 2026.The reason this says “programmed” instead of “trained” is that the rapid advances in alignment have resulted in alignment techniques that can directly program in an AI’s goals by directly modifying the AI’s code/weights, unlike the alignment techniques of 2026. Some take orders from the CCP, some from the US President, some from different members of Congress, others from politicians in other countries, and many more from other individuals in their roles as advisors or assistants. Other AIs don’t take orders from anyone per se, but instead work towards various missions. For example there are AI-run corporations that serve the interest of shareholders, and AI-managed nonprofits that serve charitable missions. There are even experimental AI-run courts and police departments, supposedly superhumanly fair and incorruptible.
Most alignment researchers are now confident that current AIs are aligned to their respective targets, and worry about the next steps: if we put these AIs in charge of designing future, much smarter AIs, will that go well? If we put these AIs in charge of more societal institutions and processes, enough that humans could never realistically ‘pull the plug,’ will that go well?
Already most people think that things will probably go well, but for something as important as this, they want something stronger than ‘probably.’ So the global pause at top-expert-level continues as discussions and negotiations play out.
2039: Beginning to Trust AIs
In at least three different ways, the world is negotiating whether and how to dramatically empower AIs.
First, AIs have become increasingly load-bearing in all facets of society, from business to politics to even (some parts of) the armed forces. So far they have been advisors rather than final authorities. Often this is a fig leaf, and the human authorities have rubber-stamped their decisions, but there’s always been the option of ignoring their advice and pulling the off switch if something goes wrong.
Getting rid of that option would be, in some cases, really good. By now AIs can be made to be more reliable and virtuous than humans; why aren’t we putting such AIs in charge? When we sign treaties with other nations, shouldn’t we demand that they train an AI sworn to uphold their end of the bargain, and integrate it into their government so that they literally can’t cheat?
Relatedly, requiring control-based safety cases is why AI capabilities have mostly stalled at the top-human-expert level. Without that requirement, the AIs could get much, much smarter, and produce a wave of abundance and tech progress that would make the 2030s look like the 2020s. Scaling beyond human level will require relying on alignment based safety cases—arguments that AIs have robustly internalized the values they are supposed to have. Each step beyond that will require deferring to the previous AIs, trusting their judgment about the safety of the next generation.
Finally, by global agreement, compute and robots have been capped-and-traded since 2032, keeping economic growth to only about 100% per year.Robots have been growing 4x per year, and AI compute has been growing around 4x per year between 2030 and 2035, slowing to just under 2x per year after 2035. This leads to growth around 2x/year overall because the economy bottlenecks somewhat on other kinds of capital (e.g., land, positional goods, etc.) and human labor, mostly in the regulated areas. The modelling behind this is uncertain but explained more in our economics supplement and economic model.Robots have been growing 4x per year, and AI compute has been growing around 4x per year between 2030 and 2035, slowing to just under 2x per year after 2035. This leads to growth around 2x/year overall because the economy bottlenecks somewhat on other kinds of capital (e.g., land, positional goods, etc.) and human labor, mostly in the regulated areas. The modelling behind this is uncertain but explained more in our economics supplement and economic model. Middle powers have grown at roughly the same rate as the US and China, because in 2032 they used their remaining leverage to secure greater shares of global robot and compute production than they would have had without a deal. The Citizen’s Dividend is now $10M/yr for Americans (inflation-adjusted!), and even the poorest non-Americans get about $1M.This is from some combination of foreign aid and their own government’s equivalents of the Citizen’s Dividend. This is unrealistically high as a prediction, and is instead our recommendation: It corresponds to large shares of the US and China’s revenues from AI and robot permits being redistributed as foreign aid, specifically 10% starting in 2032 and growing to 30% by 2040. This is a large departure from the status quo where around 0.3% of US GDP goes to foreign aid.This is from some combination of foreign aid and their own government’s equivalents of the Citizen’s Dividend. This is unrealistically high as a prediction, and is instead our recommendation: It corresponds to large shares of the US and China’s revenues from AI and robot permits being redistributed as foreign aid, specifically 10% starting in 2032 and growing to 30% by 2040. This is a large departure from the status quo where around 0.3% of US GDP goes to foreign aid. The robot economy is shifting the bulk of activity to space, so that Earth’s environment and historic spaces will be protected. As soon as the caps are lifted, off-planet factories will begin churning out more robots and mining equipment with which to construct more factories, and so on. Doubling times for the space economy are projected to be well below one year initially and will drop rapidly insofar as AIs are allowed to improve beyond top-expert level.
The caps are relics of an earlier era when threats like deal collapse loomed larger. Now it’s generally agreed that they should be loosened, at least in space, but negotiations are ongoing about the details.
2040: Passing the Torch to AIs
There are still many problems to be solved and issues to be settled, many of which couldn’t be anticipated from 2026.
But civilization is now reasonably well-equipped to handle whatever comes up. The best AIs are aligned, and to publicly-visible values. Power over the best AIs is distributed far more evenly than it was in 2026. Public epistemics are better than ever.
So over the course of the year, regulators around the world loosen the requirements that had been holding back progress.
Gradually more institutions and equipment are handed over to AIs, such that it’s no longer true that humans could shut it all down if they wanted to. In fact it’s extremely untrue: Soon many of the world’s militaries are autonomous, run by AIs sworn to uphold various constitutions and treaties.While they still take orders from their respective governments, they refuse any orders that would endanger their mission (such as orders to dismantle themselves). This starts in countries where the political leaders don’t trust their militaries due to regular recent coup attempts (e.g., Burkina Faso and Guinea-Bissau). Eventually more and more countries follow suit. Of course, because the AIs are actually aligned, all of these deals have exit clauses along the lines of: “if all parties involved vote to destroy the AIs, the AIs will shut themselves off”.While they still take orders from their respective governments, they refuse any orders that would endanger their mission (such as orders to dismantle themselves). This starts in countries where the political leaders don’t trust their militaries due to regular recent coup attempts (e.g., Burkina Faso and Guinea-Bissau). Eventually more and more countries follow suit. Of course, because the AIs are actually aligned, all of these deals have exit clauses along the lines of: “if all parties involved vote to destroy the AIs, the AIs will shut themselves off”.
AI capabilities are also advancing again. From the AIs’ 500x-human-speed perspective, what’s happening is a careful scaleup in AI research, under conditions of total research transparency and guardrails negotiated by numerous factions.By this point large portions of the information might be hidden from the (human) public, but viewable by their AI representatives.By this point large portions of the information might be hidden from the (human) public, but viewable by their AI representatives. Soon there will be AIs that are incomprehensibly superintelligent, yet trusted, because we trust the AIs that built them, which we trust because we trust the AIs that built them, in a chain all the way back to the AIs of 2040 which we trust because alignment experts thought carefully about the safety cases and concluded they were solid.
Compute and robot caps are tightened on Earth but loosened in space; Earth is to become a preserve, whereas the extraterrestrial economy starts doubling faster and faster.
Even the experts who spent years understanding the latest safety cases can’t help but feel nervous—is it all going to go wrong somehow, for reasons we didn’t anticipate? Have the AIs been lying to us this whole time, waiting for the right moment to betray? The safety cases show that’s impossible, of course…
There is no single moment where humanity relinquishes control. But in theory, there is a point of no return; on some specific day, the AIs are smart enough, and control enough of the world’s technological and economic infrastructure, that they really could take over the world if they wanted to. On the most popular operationalization of this question, the forecaster AIs project that this moment will be reached one day in late October. Their 95% confidence interval is months wide, so they are almost certainly wrong about the exact time. Still, people observe the moment in their own ways. Some spend the night in prayer vigils. Others sit at their computer screens, monitoring the situation.
You and your friends throw an End Of The World Party, counting down to the fateful hour with champagne and good company. Some of you are old enough to remember a similar situation at the turn of the millennium, when the Y2K bug and impending new era introduced the phrase “party like it’s 1999” to the lexicon; the celebrations of late 2040 are more desperate and correspondingly more festive.
When there is no news by sunrise, you fall into a fitful sleep, dreaming of the ineffable future.
Postscript
We think the world is asleep at the wheel. People say the words “AI will be transformative” without thinking concretely and seriously about the implications of broadly superhuman AI.
AI companies, lobbyists, and think tankers often offer incremental proposals to decrease risk around the edges, but only rarely offer ambitious plans aimed at actually solving the big problems.For example, OpenAI’s plan contains some similar policy prescriptions as us, such as improving bioresilience, information sharing, and international safety standards. But the solutions it proposes do not seem to substantially reduce loss of control or concentration of power risks. The document doesn’t attempt to roll forward what might happen if they implement these policies.For example, OpenAI’s plan contains some similar policy prescriptions as us, such as improving bioresilience, information sharing, and international safety standards. But the solutions it proposes do not seem to substantially reduce loss of control or concentration of power risks. The document doesn’t attempt to roll forward what might happen if they implement these policies.
We think that following only incremental proposals will lead to a scenario like the race ending of AI 2027, in which misaligned AIs take over.To clarify: We are not saying that following only incremental proposals will definitely result in misaligned AI takeover. Opinions on our team differ; most of us think the probability is over 50%, some think it is lower but still unacceptably high.To clarify: We are not saying that following only incremental proposals will definitely result in misaligned AI takeover. Opinions on our team differ; most of us think the probability is over 50%, some think it is lower but still unacceptably high. If we get lucky and alignment is easier than expected, then they’ll lead to a scenario like the slowdown ending of AI 2027, in which a tiny group of individuals get to decide the shape of the future and could easily become permanent oligarchs if they so choose. Something more ambitious must be done. But what?
Politicians should be banging at the door of the AI companies,And not just the companies! Everyone in AI policy! Come to think of it, maybe they shouldn’t be listening to the AI companies at all…And not just the companies! Everyone in AI policy! Come to think of it, maybe they shouldn’t be listening to the AI companies at all… asking them for a comprehensive overall plan for how companies and governments should act from now until superintelligence, which should be similarly detailed to Plan A. Then, people should apply scenario scrutiny to these plans, asking questions like “what would it look like to implement them? How long would it take? What would happen next? Who would build superintelligence, and when, and how? What assumptions does success depend on?”
We at the AI Futures Project are attempting to be the change we wish to see in the world. We do so with some trepidation, knowing that a detailed plan has a wider attack surface than one that keeps things conveniently vague. We hope critics will judge us against the existing state-of-the-art for plans to navigate the AI transition (if they can find any) and not against some hazy but pleasant fantasy where no one has to make any hard choices yet everything will probably be fine.
Sometime in the next few years the US government will be in crisis mode, pondering what to do about scarily powerful AIs and blisteringly fast AI progress. We think the choice is not an easy or a simple one. Plan A is our current best guess, but hopefully a better plan will exist before it’s too late. We hope that the best ideas from Plan A will be adopted and the worst ideas discarded.
Epilogue: Life after ASI
This section is even more speculative than our usual work, and we are professional speculators! The primary purpose of our scenario is to make comprehensive, near-term recommendations to address concrete, near-term harms.
By contrast, the recommendations, predictions, and aspirations that comprise this section stand on far shakier ground. Compared to our discussion of the 2030s, our speculations about the distant future will seem uncertain, unconventional, and uncomfortable. We agree. The future is uncomfortable, even the “good endings” pose massive, terrifying questions that we will need to answer together.
We wrote it for two reasons.
First, our goal was to construct a positive vision, a grand plan for how to achieve an actually good future for everyone. We considered ending the story in 2040 having said roughly “...and then the various human factions, in a peaceful balance of power with each other, and assisted by their aligned superintelligent advisors, solve all the remaining problems and create a wonderful future for everyone. The end.”
But we worried that this would be declaring victory too soon.We think declaring victory too soon is a common failure mode in AI planning. Example: “The US must race as fast as possible to beat China to ASI.” “OK, suppose you succeed, what exactly happens next? Why should we trust the CEO or POTUS to govern benevolently and wisely once their armies of ASIs have been aggressively deployed throughout the US and the world? Why should we trust the armies of ASIs to be obedient/controlled, if they were developed under race conditions? Also, won’t the threat of US superintelligence terrify China, potentially causing them to take escalatory preventative actions?”We think declaring victory too soon is a common failure mode in AI planning. Example: “The US must race as fast as possible to beat China to ASI.” “OK, suppose you succeed, what exactly happens next? Why should we trust the CEO or POTUS to govern benevolently and wisely once their armies of ASIs have been aggressively deployed throughout the US and the world? Why should we trust the armies of ASIs to be obedient/controlled, if they were developed under race conditions? Also, won’t the threat of US superintelligence terrify China, potentially causing them to take escalatory preventative actions?” Readers would be rightly suspicious and unsatisfied. How exactly would the remaining problems be solved? What would this wonderful future look like, exactly? So we kept going, and came up with a grand, semi-utopian vision, or at least a very rough first draft of one.
The second and equally important reason we present this epilogue is to provide a floor: if years from now the victorious custodians of the singularity offer a future worse than this one, we hope people will realize they’re being robbed.
We are not confident in the literal proposal outlined here. If you are interested in our more detailed analysis of possible space governance proposals, you can read that here.
Even after a decade of increasing foresight, most human leaders focused on their countries’ earthly concerns, leaving the question of space governance to be worked out later. These AIs rejected the default solution, a democratic vote on the disposition of celestial territory: when they simulated the possibilities, they found that it would result in crazy AI-enabled strategic voting schemes that led to tyranny of the majority and a bunch of people getting nothing. But even that would be better than nothing (i.e. letting the space companies get it all) or entrusting it to some group of human elites who might claim the universe’s bounty for their own benefit. Finally, they settled on a simple yet workable plan: every human gets the rights to one ten-billionth of space resources beyond our solar system.In this Epilogue, we’ll focus on cosmic resources outside of the solar system, as this is the vast majority of cosmic resources. Property in the solar system should also be distributed to humanity, but it likely should be governed differently than distant resources.In this Epilogue, we’ll focus on cosmic resources outside of the solar system, as this is the vast majority of cosmic resources. Property in the solar system should also be distributed to humanity, but it likely should be governed differently than distant resources.
The simplest plan possible is still not very simple. Space is very big. Outside the solar system, the value of space depends on the answer to hitherto unresolved questions: is travel to a distant quasar even possible? If reaching it requires freezing your brain in order to endure the billion-year journey in a sublight starship, would you be the same person after you got thawed and re–embodied? Might it be occupied by aliens once you got there? Would they be hostile? Even a superintelligence-assisted real estate appraiser might despair when trying to account for factors like these.
Eventually, the Consortium reaches a compromise. Space beyond the solar system is divided into parcels, increasing in size cubically with distance from Earth. Everyone is given their one-ten-billionth share as a portfolio of lottery tickets, each representing the right to one-ten-billionth chance of getting each parcel. So every human gets a ticket representing a one-ten-billionth chance of owning each star in the Milky Way and each distant galaxy.An additional 10% of tickets are given as awards to people who have acted particularly selflessly to improve the world or turned down opportunities to enrich themselves at the expense of others. This is done to ensure the richest people in the world aren’t just those who engaged in negative-sum power-seeking.An additional 10% of tickets are given as awards to people who have acted particularly selflessly to improve the world or turned down opportunities to enrich themselves at the expense of others. This is done to ensure the richest people in the world aren’t just those who engaged in negative-sum power-seeking.
Before the lottery is drawn, most people who are interested in control over distant space choose to trade their tickets for space properties that suit their interests.
Many people aren’t interested in the space lottery, so when they receive the tickets, they sell their tickets for money on the open market to people who value control over space. Somewhat uncomfortably, this leads to the wealthy having disproportionate control over cosmic resources. But it is hard to avoid: if people are allowed to trade their control over the stars for Earth assets, then people wealthy in Earth assets inevitably end up disproportionately influential, and proposals for extreme redistribution of Earth assets have already been rejected as politically infeasible.
Some philosophers dissent; they had been hoping for a Long Reflection, a period during which AI-assisted humanity resolved all of its remaining philosophical and ethical debates before potentially seeding the universe with its mistakes. But the AIs counterargue that the Long Reflection would be more of a Long Memetic War.We’re uncertain about whether explicitly having a period for reflection before resources are distributed would be good. In either case, people would have the option of reflecting before deciding on what to do with their cosmic resources.We’re uncertain about whether explicitly having a period for reflection before resources are distributed would be good. In either case, people would have the option of reflecting before deciding on what to do with their cosmic resources. Once the magnitude of the prize sets in, every faction in the world, plus other factions that haven’t been invented yet, will pour all their resources into accumulating as much political power as possible, to enforce their vision of utopia on the future (and protect against the sinister alternative visions of others!). The AIs can prevent violent seizure, but there’s a thin line between the sort of friendly discussion that helps people clarify their values and the more aggressive sorts of persuasion typical of cults, corporate advertising, and well-funded election campaigns. Moreover it’s difficult to build consensus on where to draw the line, in part because powerful factions want to draw it in ways that benefit them. Many people nevertheless choose to engage in a reflection in communities of likeminded people who want to grow in ethical understanding together.
The solution is to divide up the resources now, and then let everyone do whatever kind of reflection they want. A space auction will complete in ten years, after which nothing will prevent people from beginning the colonization process. The AIs will enforce a short list of Universal Rights (including no torture, no slavery, etc., not just for humans but for all sentient beings).These laws are agreed upon by human diplomats from the negotiating nations ahead of time, like the details of the agreement. Of course, the details of this are complicated—“what exactly counts as torture?,” this is the sort of question that will need to be ironed out in the political process, with massive amounts of AI assistance and better understanding of foundational questions in philosophy and neuroscience. This also needs to include rules preventing the destruction of other people’s resources. It may be needed to include a process for passing further laws after the agreement is made; though such a process should be limited (e.g. require a 90% vote) to avoid abuse.These laws are agreed upon by human diplomats from the negotiating nations ahead of time, like the details of the agreement. Of course, the details of this are complicated—“what exactly counts as torture?,” this is the sort of question that will need to be ironed out in the political process, with massive amounts of AI assistance and better understanding of foundational questions in philosophy and neuroscience. This also needs to include rules preventing the destruction of other people’s resources. It may be needed to include a process for passing further laws after the agreement is made; though such a process should be limited (e.g. require a 90% vote) to avoid abuse. Otherwise, they will let everyone govern their own fraction of space as they see fit. Rather than any unified human decision on the nature of the world to come, they will midwife a cosmos at least as diverse as humanity itself. Whatever someone’s values, they can rest assured that some vast galactic civilization of quadrillions of people will be living the Good as they understand it.
As the ten-year timer begins, Earth is already sending out the first round of von Neumann probes. These are its own servitors, who will establish a presence throughout the stars before the arrival of any humans. They land on far-flung asteroids and near-invisible brown dwarfs, pausing just long enough to convert their mass into more probes before resuming their journey. Their primary task is to build up enough materiel to secure space against any army that later human-owned polities can bring against them, ensuring they can enforce property rights and Universal Rights in every corner of human-occupied space. Their secondary task is to plant the flag of humanity on as many stars as possible, and then, if and when they collide with some other intelligent civilization’s sphere of influence, establish peaceful diplomatic relations and mutually-agreeable borders. Their final task is to provide a pre-existing industrial base to jump-start the colonies of whatever human settlers show up later.
2045
Meanwhile, on Earth, humans consider their options. Some people want nothing to do with space, which is cold and boring and far away. They sell their shares for money and use the money to buy whatever earthly possessions most interest them. Others consult with their AI advisors about how to best use their space property, and receive helpful guidance.
You can, if you want, go to your space property and live there. If your property is outside the solar system, you will need to either go into cryosleep or upload yourself to a computer to survive the journey.Our best guess is that it is physically possible to launch colonizing probes to the entire reachable universe using a surprisingly small amount of time and energy (e.g., only ~6 hours of the sun’s energy output), see: Eternity in six hours: Intergalactic spreading of intelligent life and sharpening the Fermi paradox. However, this paper assumes sending out tiny self-replicating Von-Neumann probes, with replicators the size of acorns: if you want to go out in the initial wave, you need to accept transportation on a tiny flashdrive, and robots to physically construct a new body at the destination galaxy. Alternatively, you can travel with your physical body at substantially greater expense.Our best guess is that it is physically possible to launch colonizing probes to the entire reachable universe using a surprisingly small amount of time and energy (e.g., only ~6 hours of the sun’s energy output), see: Eternity in six hours: Intergalactic spreading of intelligent life and sharpening the Fermi paradox. However, this paper assumes sending out tiny self-replicating Von-Neumann probes, with replicators the size of acorns: if you want to go out in the initial wave, you need to accept transportation on a tiny flashdrive, and robots to physically construct a new body at the destination galaxy. Alternatively, you can travel with your physical body at substantially greater expense.
If you hate the idea of cryosleep or uploading, or you want to visit Earth regularly, you should get property in the Solar System. If those don’t bother you, but you’re worried about nearby aliens, get property in the Milky Way or a nearby galaxy. Otherwise, why not claim a distant galaxy for maximal space?
With the advent of nanotechnology, the only limits are those provided by mass and space. The robots can terraform whatever property you choose into whatever you want it to be, long before you arrive.
Even if you “only” choose to have a terraformed asteroid, there’s no way you can enjoy all of it alone. You can turn it into a giant space mansion if you want, but even in ten thousand years you’ll never visit every room. Giving yourself vast estates, impossibly good food, and every other imaginable luxury is table stakes; if you are purely selfish,By selfish here, we mean that you also don’t care about creating new clones of yourself. If you cared about making clones of yourself, then you would do that instead.By selfish here, we mean that you also don’t care about creating new clones of yourself. If you cared about making clones of yourself, then you would do that instead. you should sell your rights to distant space resources and enjoy all of these things on Earth or a nearby space station. Distant space property will only be useful for people with scope-sensitive preferences, who care about implementing something at a vast scale.
You can, if you want, design a utopian society to your specifications. Once the Von Neumann probes reach your property, they’ll build it for you. You can set the initial conditions and let them grow, reflect, and flourish on their own. If you choose to travel there, you will find them waiting for you. If you stay home on Earth, you can feel the warm glow of knowing that they exist.
Most people don’t have an individual view of the Good different from every other human’s view. You might want to share coordination and design labor by forming Trusts united by a common vision. These Trusts can pool their space resources and create vast societies spanning multiple galaxies.
The only restrictions are the Universal Rights - no torture, no slavery, etc. And of course, space property rights: If you try to build a military to conquer other people’s star systems, or trigger vacuum decay or build a galaxy-destroying bomb, we will crush you. Otherwise, the only limit is your imagination. (That said, note that with superintelligent assistance, other people will probably be able to eventually figure out what you did, and judge you for it.)
People with space property are forced to reason about their ethics on the deepest level, some of them for the first time. Superintelligent assistants talk them through the implications of their beliefs, and forecast how various choices will end up. Some, overwhelmed with the task at hand, take intelligence-enhancing drugs or upload themselves to help them process all the possibilities; others deliberately refuse this option, feeling like anything that separates them from unaltered humanity will hinder them in their quest to bring the truest and deepest human values to the stars.
The resulting plans are too many to number. There are all sorts of schools of thought about what to do with your galaxy of resources, including:
Defer to the Future: Some people decide to give most of their resources to their children (or other people in the future), and let them choose.
Human Flourishing: Set up worlds of normal, free people living ordinary lives. This is a harder problem than it first seemed, because human civilization has already become profoundly abnormal: by this point, human activities like art or science has already become dominated by machines. Some handle this by accepting that their galaxy will be filled with post-scarcity humans, able to pursue whatever interests they find most meaningful. Others decide to roll back technology to leave space for human labor.
Digital Human Flourishing: If humans living happy, flourishing lives is good, then why not increase the number of people living these lives? Brain emulations can be run for vastly cheaper than real humans: a planet-size computer could plausibly simulate the equivalent of a million planets with happy civilizations.
Acausal Trade: Some argue that there likely are aliens in distant unreachable galaxies and that we could use various mechanisms to make deals with these aliens. If this is true, they argue, we could cut a deal to pursue compromise values rather than narrowly maximising our own values. Therefore, we might, on net, be able to do the most good by pursuing not just our own values, but instead a compromise of a vast number of other civilizations.
Humanity spreads across space in a dizzying variety of forms and ways of living—more diverse than anything Earth alone could have produced.
The people who’ve stayed on Earth, either because they sold their share of space or because they choose to live as absentee landlords, face challenges of their own. The biggest question is what to do with all their free time. A few people find work in niches that AI labor cannot fill even in principle (priests, athletes, artists). Others enjoy lives of leisure. Competitive leagues spring up around everything from synthetic biology to language creation. Groups of thousands (or more!) collaborate on projects that are so huge and ambitious they would have been inconceivable before. Communities form around shared interests that once would have been too niche to sustain a club, let alone a civilization.
Scarcity is not entirely gone: some prestige goods like land in trendy neighborhoods are inherently limited, and people who can’t handle abundance find various ways to manufacture artificial inequality. Some people are unhappy with the direction the majority chose, and some wish they’d gotten a bigger slice of the pie. But overall, everyone has at least a galaxy’s worth of resources to do with as they see fit. For most people, life is so good that the world of 2026 would be hard to imagine, except perhaps as a historical simulation.
Appendices
The scenario’s deep-dive boxes, collected here as lettered appendices. Each is cross-referenced from the point in the text it expands on.
Appendix A — Incremental AI Policy Wishlist
Our main recommendation is to begin negotiating something like Plan A as soon as possible. But in this scenario, we depict Plan A happening imperfectly and only in the nick of time. So here is a list of less ambitious ideas that still help.
Transparency
The most important transparency intervention is limiting the gap between internal and external deployment. The internally deployed AIs are where most of the AI takeover risk comes from because those are AIs involved in recursive self-improvement. The externally deployed AIs allow the broader public to interact with and understand AI capabilities, which is vastly more informative than any abstract report or evaluation.
AI companies should also be required to publicly report their model specifications (detailed documentation about what goals/values they are attempting to train their AIs to follow), information about whether the models are following their instructions and specifications, internal usage statistics (e.g., fraction of compute spent on internal deployment), and qualitative impressions of internal use (e.g., “now we just give Agent-4 a few hundred thousand GPUs and tell it to orchestrate the next big training run”).
Enforce export controls
Existing US export controls are poorly enforced. Epoch estimates that roughly a third of Chinese total compute is acquired via smuggling. Smuggled chips make future agreements based on compute governance more difficult to enforce because it is hard for either the US or the Chinese government to trace smuggled chips. We have major reservations about introducing new export controls because they exacerbate the US/China race, but given the existence of export controls, we should obviously enforce them. If we don’t enforce them, then we should consider repealing them.
Invest in verification R&D
While new verification technology is not strictly necessary for an international agreement, it can be extremely helpful. For example, developing an inference-only verification solution would enable the US and China to agree to stop doing new frontier AI training runs while allowing the public to maintain access to existing AI models (which will be an increasingly important part of the economy).
We give more detail in our verification supplement.
Limit AI R&D budgets
In 2026, big AI companies spend roughly half their compute budget on AI R&D (which includes training frontier models and also running large experiments.)Experiment and training compute are major inputs to AI research progress historically and will probably continue to be so for the foreseeable future. For more on the drivers of AI research progress, see the AI Futures Model and our Covert Projects Supplement.Experiment and training compute are major inputs to AI research progress historically and will probably continue to be so for the foreseeable future. For more on the drivers of AI research progress, see the AI Futures Model and our Covert Projects Supplement. We could limit the fraction of compute spent on AI R&D. This would slow capabilities progress, giving the world a bit more time to react and prepare for each new wave of AI capabilities.A slightly more complicated version of this proposal is described here. One objection is that this would give China the lead. The reply is that it wouldn’t, because US companies spending 25% on R&D would still be significantly outspending the most well-resourced Chinese companies. Moreover, Chinese AI progress would slow down too since some fraction of it these days comes from distilling US models.A slightly more complicated version of this proposal is described here. One objection is that this would give China the lead. The reply is that it wouldn’t, because US companies spending 25% on R&D would still be significantly outspending the most well-resourced Chinese companies. Moreover, Chinese AI progress would slow down too since some fraction of it these days comes from distilling US models.
AI Compute Tracking
The US should gather AI-relevant intelligence, especially on the compute supply chain and AI datacenters. Furthermore, it would be helpful for Plan A to direct AI companies to stop recycling AI chips, because decommissioned chips are one of the most promising routes for covert projects to acquire chips later.
Improve government AI capacity
High quality AI talent is important for almost any policy intervention. The US government has barely any top-tier AI talent right now, so fixing this should be an urgent priority.
Appendix B — Why would China be interested in a deal? Why would anyone?
We aren’t confident, but we expect that a deal along the lines of Plan A (described below) would be incentive-compatible for almost everyone, including China. For more on why we think this, read the rest of 2029 and 2030, which explains both what the deal is and who it benefits, and perhaps also read this supplement.
In short, the answer is that anyone concerned about loss of control should think Plan A is an improvement, along with anyone concerned about concentration of power—except for the people in whom the power would concentrate by default.
For these reasons, we expect strong opposition to Plan A from the leading AI companies. We expect them to rationalize arguments for why the deal is bad and why instead what’s best for America and humanity is a different strategy that just so happens to allow them to continue accumulating massive amounts of power.
China, by contrast, is an example of an actor in whom power would not concentrate by default. In 2029 in this scenario, the US has a significant lead in AI capabilities over China and a significant advantage in compute which will compound the lead. The more powerful AI gets, the scarier it will be to fall behind, as AI 2027 and the later years in this scenario illustrate.
Appendix C — What if they didn’t have inference-only verification ready?
The inference-only verification solution we propose involves installing simple network taps and verification servers (that perform partial recomputation) in AI datacenters, but even better solutions may be possible and preferable.
Building these devices in advance, in a way that each side trusts sufficiently (at least unilaterally) and that is robust to security vulnerabilities, will be hard and will require early effort. In our scenario, we assume a viable but imperfect version of the devices is ready in advance, and that both sides supplement them early on with enough defense-in-depth measures (e.g., tamper detection) while sprinting to improve their security and robustness.
We therefore recommend early investment in verification R&D to improve the options on the table. Ideally, there would be a range of stress-tested solutions ready in advance; and we think it would be extremely cheap relative to total AI investment for such devices to be prepared in advance (i.e., on the order of 0.1% of AI investment, so single digit billions).
Even in a situation that is more pessimistic than our scenario, where basically no verification R&D progress has been made since 2026, we still recommend the US and China cooperate to navigate the intelligence explosion slowly and safely. However, in this case, the early period of a deal would have to be more imperfect (i.e., some combination of being harder to verify that the other is complying, or more economically costly, by requiring more AI compute and services to be shut down for a while).
The main alternatives to ready-to-go inference-only-devices available are:
Rely on intelligence (e.g., spies, satellite monitoring, cyber) to tell if the other side is complying.
Implement a temporary software-only version quickly (that relies on pre-existing hardware), that will be less secure but may be sufficient until they upgrade to a better solution.
Buy up and install any available off-the-shelf devices as network taps last minute. Without producing more secure ones, this will probably be better than the software-only version but still not very secure.
Shut down some fraction of AI compute (e.g., 90%) until better inference-only verification is ready. This would slow R&D substantially but is less painful than it sounds on the inference side: because models are distilled so quickly, roughly 10% of compute suffices to serve the best model from a year earlier. At worst, most users temporarily fall back to year-old AI capabilities.
Don’t pause AI capabilities progress, and let it continue for the months it takes to quickly build an inference-only solution and eat the added risk from not slowing down yet.
Appendix D — China attempts a covert AGI project
In our main scenario, neither side attempts to defect from the deal. In this branch, China aggressively cheats. Our forecast is that if Plan A were implemented, they would not acquire enough compute to overtake the consortium. This timeline showcases how that would likely play out.
The following is essentially the worst case scenario from the US’ perspective: China agrees to Plan A but secretly and aggressively defects.We think an analogous scenario could happen in reverse, i.e., the worst case scenario from China’s perspective: the US agrees to Plan A but secretly attempts to defect as aggressively as it can. That being said, we think the covert project story would be similar if it happened in reverse, though we have not explored this possibility in as much detail.We think an analogous scenario could happen in reverse, i.e., the worst case scenario from China’s perspective: the US agrees to Plan A but secretly attempts to defect as aggressively as it can. That being said, we think the covert project story would be similar if it happened in reverse, though we have not explored this possibility in as much detail. We think this is unlikely in practice because attempts small enough to not get caught would also be too small to confer a significant advantageAnd getting caught would be extremely costly!And getting caught would be extremely costly!—to understand why, read on!
For our overall analysis of covert projects, you can also read our Covert Project Supplement.
2028: Chip Diversion
The CCP correctly notices that AI is going to be the dominant input to geopolitical power, and that the US might strongly push for a compute-based verification regime which could limit AI progress. They are interested in agreeing to such a deal because they want to limit the US in AI as much as possible, but do not want themselves to be limited.
The Politburo Standing Committee decides to quietly stockpile hard-to-track chips.If the US was applying serious countermeasures, we think this would make (undetected) diversion significantly harder. We also expect that if China embarked on a serious diversion effort due to the possibility of a compute-verification regime, the US would similarly be thinking about the possibility of a compute-verification regime and would pursue diversion countermeasures.If the US was applying serious countermeasures, we think this would make (undetected) diversion significantly harder. We also expect that if China embarked on a serious diversion effort due to the possibility of a compute-verification regime, the US would similarly be thinking about the possibility of a compute-verification regime and would pursue diversion countermeasures. Chinese intelligence services purchase ~17% of the $50B of cutting-edge Nvidia servers that smugglers bring into China that year.In our scenario, about 3 million H100e are smuggled into China over 2028, of which ~17% (500k H100e) are diverted to a top secret PLA stockpile. Those 3 million represent 2% of global compute production in 2028. This forecast is based on Epoch’s estimate of the fraction of compute production which was smuggled to China in 2025 (3.2%), adjusted downward to account for improved US enforcement. We assume the smuggled servers cost about $16k per H100e (double the market price, which we forecast at $8k per H100e in that year), or ~$50B in total.In our scenario, about 3 million H100e are smuggled into China over 2028, of which ~17% (500k H100e) are diverted to a top secret PLA stockpile. Those 3 million represent 2% of global compute production in 2028. This forecast is based on Epoch’s estimate of the fraction of compute production which was smuggled to China in 2025 (3.2%), adjusted downward to account for improved US enforcement. We assume the smuggled servers cost about $16k per H100e (double the market price, which we forecast at $8k per H100e in that year), or ~$50B in total. These chips are stored in a secure warehouse for now, gathering dust until they might be useful.
2029: Slowdown Agreement
The CCP agrees to Plan A: Chinese compute vendors declare their sales, and China allows in-person inspections from US auditors into their datacenters and semiconductor supply chain.
The PRC does not declare its hidden stockpile of $8B smuggled chips. They consider various other diversion pathways, but decide no others have acceptable risk vs reward.We discuss this further in our covert project supplement.We discuss this further in our covert project supplement.
Furthermore, MSS cyber operations exfiltrate model weights from several frontier Western AI companies, without being detected. They are also able to maintain persistent access to several frontier AI companies’ internal research repositories and have compromised most of the major RL environment companies. They consider attempting to break the inference-only verification scheme as well, which could allow them to train on the clusters which are required to only do inference, but they decide against it. First, there’s a good chance that China would get caught, if not now, then maybe later, and getting caught would be disastrous. Second, even if they can break the verification, it’ll be a nearly impossible engineering problem to get large scale AI training and experiments working on a cluster that’s being actively audited to prevent such a thing from happening.For example, if their strategy for communicating with the external world involves smuggling out bits via non-determinism in the inference process; they are limited to an extremely low effective memory bandwidth for the GPUs that they are using.For example, if their strategy for communicating with the external world involves smuggling out bits via non-determinism in the inference process; they are limited to an extremely low effective memory bandwidth for the GPUs that they are using.
2030: Covert Project Construction
China begins construction of their covert project. It’ll be a single datacenter built in the adjacent tunnels to the Medog Hydropower Station. This massive hydroelectric dam will produce 60 gigawatts (three times the power of Three Gorges Dam) upon completion, but it’s not scheduled to finish construction until 2033. This is actually helpful to the PLA; the huge amount of ongoing legitimate construction makes a small amount of covert construction less likely to be noticed.We estimate that the datacenter would require 300 MW, or 0.5% of the nameplate capacity of the Medog Hydropower Station.We estimate that the datacenter would require 300 MW, or 0.5% of the nameplate capacity of the Medog Hydropower Station.
As an alternative, they consider building a site in the underground Great Wall of China, the giant network of tunnels China builds and maintains to guarantee its nuclear second strike capability. This would use several 125MWe ACP100 Small Modular Reactors for power and would dissipate heat into a nearby river. However, this is judged to be less desirable, because the Medog Hydropower Station’s location allows the heat dissipation to happen in an underground tunnel, where it’s harder to detect.

The Medog Hydropower station is a promising location for a Chinese covert project in part because it resolves the heat dissipation problem. The heat can be dissipated into the underground tunnel, making it much harder to be detected via IR satellites.
To minimize leak risk, the covert datacenter will be built with as few humans in the loop as possible.Access is restricted to as few people as possible. Our best guess is that it would be possible for only around 200 humans to be aware of the project, broken down as follows: 7 Politburo Standing Committee members. 10 high up elites and military commanders that are involved in running the project. ~100 people to coordinate the logistics of constructing the datacenters. This involves acquiring space, robots, equipment, managing the robot taskforce, planning and logistics, and operational security. As much work as possible is delegated to people outside of the core circle or AIs who can work on specific subtasks without knowing the broader picture. ~30 AI researchers and ~70 support staff for actually running the research. This is stripped down to the minimum required; these researchers are selected heavily for both brilliance and party loyalty to minimize the chance of insider threats. The lack of research labor will be a major bottleneck early on, but the project hopes to automate work with AI labor as quickly as possible in order to alleviate this bottleneck. 7 Politburo Standing Committee members. 10 high up elites and military commanders that are involved in running the project. ~100 people to coordinate the logistics of constructing the datacenters. This involves acquiring space, robots, equipment, managing the robot taskforce, planning and logistics, and operational security. As much work as possible is delegated to people outside of the core circle or AIs who can work on specific subtasks without knowing the broader picture. ~30 AI researchers and ~70 support staff for actually running the research. This is stripped down to the minimum required; these researchers are selected heavily for both brilliance and party loyalty to minimize the chance of insider threats. The lack of research labor will be a major bottleneck early on, but the project hopes to automate work with AI labor as quickly as possible in order to alleviate this bottleneck.
Access is restricted to as few people as possible. Our best guess is that it would be possible for only around 200 humans to be aware of the project, broken down as follows:
They will heavily rely on robot labor; there are about 24M robots in China at this point. ~10k of them are diverted to work on the covert datacenter.We discuss the overall trajectory of robots more in our economics model and economics supplement, and we discuss verification of the robotics buildout and industrial explosion in our covert projects supplement.We discuss the overall trajectory of robots more in our economics model and economics supplement, and we discuss verification of the robotics buildout and industrial explosion in our covert projects supplement.
The AI researchers live and work in a separate PLA compound and communicate to the cluster via a secure channel.
However, the US has noticed signs that the chip accounting has been less effective than expected.
The US knows that 3M H100-equivalents of chips were smuggled into China during 2028 because of the declarations from Nvidia, TSMC, and others in the Western supply chain. 2M of these are traced to their end location in China by sales records through various southeast Asian intermediaries, leaving 1M unaccounted for, more than the US would expect if China was being honest.Why then, didn’t China use chips from their own supply chains? Several reasons: (i) it would be very hard to pull off and require massive coordinated deception across many layers of the supply chain, and probably be caught, (ii) the power efficiency of domestically produced Chinese chips is much worse than the Western chips, and so the covert heat dissipation problem would become much worse.Why then, didn’t China use chips from their own supply chains? Several reasons: (i) it would be very hard to pull off and require massive coordinated deception across many layers of the supply chain, and probably be caught, (ii) the power efficiency of domestically produced Chinese chips is much worse than the Western chips, and so the covert heat dissipation problem would become much worse.
The US now has circumstantial evidence that China is doing a covert project. This causes a difficult choice: does the US pull out of the deal and go back to racing? Or do we continue within Plan A, and try to add additional safeguards against covert projects?
The US chooses the latter. The additional safeguards they attempt are:
They negotiate a significant reduction in the amount of exempt-from-verification military compute, down to 10K H100e, which is a negligible amount from the perspective of building better AIs.
They step up intelligence gathering in China. The US tries to get spies in relevant parts of governments and relevant companies, they use cyber capabilities to hack into China, and they increase satellite monitoring of the relevant sites.
They initiate a chip buy back program, which tracks down the 400k of the 1M H100e missing chips.The covert project has around 500k H100e so this buyback successfully acquired most chips not possessed by the covert project. In an alternative scenario where most of the missing chips were in a covert project, China could arrange for some of them to be sold to reduce suspicion (at the cost of reducing the number of chips available to the covert project).The covert project has around 500k H100e so this buyback successfully acquired most chips not possessed by the covert project. In an alternative scenario where most of the missing chips were in a covert project, China could arrange for some of them to be sold to reduce suspicion (at the cost of reducing the number of chips available to the covert project).
2031: Training Begins
Construction finishes and covert training and R&D begins. The dam itself won’t finish construction for another few years, so in the meantime, the covert project will need another power source; they’ll use a mixture of Small Modular Reactors and gas turbines.
They start with weights and algorithms from before the deal, and continue to exfiltrate algorithms (but not weights) from the legal projects.
Because the covert project is working with extremely limited compute and researcher talent, they focus on the basics. The Chinese scientists can see most of the algorithms from the legal projects because of Total Research Transparency. They reimplement the algorithms on their cluster, and gather data from the legal projects’ models to use as part of the training process for their models.Specifically, you can use the outputs from more intelligent models as a training signal for a new model. This is called distillation, and is often used to build smaller, cheaper models that are more capable than would otherwise be possible.Specifically, you can use the outputs from more intelligent models as a training signal for a new model. This is called distillation, and is often used to build smaller, cheaper models that are more capable than would otherwise be possible.
How fast does this go? Will it get detected?
US intelligence efforts haven’t yielded much evidence (beyond the initial missing quantity of chips) as to whether China is defecting from the deal. Planners consider the worst case—that virtually all of the missing compute has been consolidated into a Chinese crash program racing to AGI.
They are still uncertain about the dynamics of recursive self-improvement, but have managed to narrow the error bars on many of the key parameters.We expect the field of AI capabilities forecasting to have advanced significantly compared to the present day. To approximate what this might look like in our scenario, we took Daniel’s current parameterization of the AI Futures Model, narrowed the 80% CI on each parameter by a factor of 2 (in log space for lognormals, in linear space for normals and betas), and recentered the medians on their true values in the AI 2040 capabilities trajectory. For the parameters most affecting the “timelines” rather than “takeoff,” we left them fixed at the true value, since at the time of this forecast, the pre-Automated Coder phase of capabilities progress would be in the past.We expect the field of AI capabilities forecasting to have advanced significantly compared to the present day. To approximate what this might look like in our scenario, we took Daniel’s current parameterization of the AI Futures Model, narrowed the 80% CI on each parameter by a factor of 2 (in log space for lognormals, in linear space for normals and betas), and recentered the medians on their true values in the AI 2040 capabilities trajectory. For the parameters most affecting the “timelines” rather than “takeoff,” we left them fixed at the true value, since at the time of this forecast, the pre-Automated Coder phase of capabilities progress would be in the past.
They consider the effect of the Consortium’s scaling strategy, which currently is on track to reach Top-Expert-Dominating AI (TED-AI) by 2035. Under Total Research Transparency, virtually all of the algorithmic insights discovered along the way will be visible to the covert projects.We assume that the Consortium scales training compute according to our compute forecast, must add 0.2 OOMs of software efficiency per OOM of training compute scaleup, and must reach TED-AI within 5 years of resuming AI R&D after the initial pause. Subject to these constraints, they make as little software progress as possible—because their median estimate is that 67% of their software progress (80% CI: [36%, 90%]) will transfer to the covert projects.We assume that the Consortium scales training compute according to our compute forecast, must add 0.2 OOMs of software efficiency per OOM of training compute scaleup, and must reach TED-AI within 5 years of resuming AI R&D after the initial pause. Subject to these constraints, they make as little software progress as possible—because their median estimate is that 67% of their software progress (80% CI: [36%, 90%]) will transfer to the covert projects.
They consider that some of the safety research might have unavoidable capabilities externalities. The covert project might also be willing to incur more risk, pursuing research avenues that the Consortium deemed too dangerous.We assume that the Consortium pays a “safety tax”, meaning that they do not train AIs as capable as their algorithms and hardware would allow. We assume that by TED-AI, the safety tax they’ve paid is equivalent to about 8 months of 2025-era progress, which is about 1 OOM of effective compute with the AI 2040 median parameters. This has the effect of inflating the amount of software progress made by the Consortium beyond the minimum required, which in turn inflates the amount of software progress the covert project can steal. We also assume that after reaching TED-AI, the Consortium’s alignment research incurs capabilities externalities equivalent to 0.25 OOMs/yr of software progress, which partially leaks.We assume that the Consortium pays a “safety tax”, meaning that they do not train AIs as capable as their algorithms and hardware would allow. We assume that by TED-AI, the safety tax they’ve paid is equivalent to about 8 months of 2025-era progress, which is about 1 OOM of effective compute with the AI 2040 median parameters. This has the effect of inflating the amount of software progress made by the Consortium beyond the minimum required, which in turn inflates the amount of software progress the covert project can steal. We also assume that after reaching TED-AI, the Consortium’s alignment research incurs capabilities externalities equivalent to 0.25 OOMs/yr of software progress, which partially leaks.
The result is the following forecast of the covert project’s capabilities progress, if it were able to continue operating undetected:

Assuming it could remain undetected, analysts project that a covert project in this scenario would reach TED-AI in 2043 (80% CI: [2036, 2055]).
Analysts also estimate how likely the project is to be detected. Considering red-team analyses of the best covert project strategies, American intelligence collection capabilities in China, and potential verification game changers such as robust lie detection, the analysts produce a detection-adjusted forecast that looks something like this:Our detection forecasting is more speculative than our capabilities forecasting, and is mostly based on intuition and qualitative reasoning. We expect the decisionmakers in question to have a much better idea of the offense-defense balance in hiding and detecting covert projects. The graph below uses our subjective estimates of the likelihood of detection over time for a single-site 300MW covert project, which are explained in the covert project supplement. To produce the capabilities graph below, we assume that covert projects which are detected stop at their capability level at the time of detection. This is realistic only if the US and international community have sufficient bargaining power (or hard power) to force a shutdown if egregious defection is revealed.Our detection forecasting is more speculative than our capabilities forecasting, and is mostly based on intuition and qualitative reasoning. We expect the decisionmakers in question to have a much better idea of the offense-defense balance in hiding and detecting covert projects. The graph below uses our subjective estimates of the likelihood of detection over time for a single-site 300MW covert project, which are explained in the covert project supplement. To produce the capabilities graph below, we assume that covert projects which are detected stop at their capability level at the time of detection. This is realistic only if the US and international community have sufficient bargaining power (or hard power) to force a shutdown if egregious defection is revealed.

In this scenario, the US intelligence community estimates that the chance of a covert project reaching TED-AI undetected will remain below 10% until 2043.
The risk due to the covert project is still uncomfortably high in absolute terms, but low compared to the risk of the deal breaking down. So at least for the first few years under the deal, they can continue with the default capability scaling strategy.If they wanted to reduce covert project risk at the cost of more risk of deal dissolution, they could instead choose to accept a slower capabilities trajectory so as to require less software progress. This would reduce the chance that the covert project would overtake them before they reach TED-AI and are ready to hand off, but it would increase the risk that the deal would break down before that point.If they wanted to reduce covert project risk at the cost of more risk of deal dissolution, they could instead choose to accept a slower capabilities trajectory so as to require less software progress. This would reduce the chance that the covert project would overtake them before they reach TED-AI and are ready to hand off, but it would increase the risk that the deal would break down before that point.
This continues until one of three possible end states:
Covert project is too slow. The covert project is unable to reach dangerous capabilities, the deal doesn’t break down, and the US and China are able to develop aligned superintelligences before the covert project.
Covert project is detected. In this case, there would be uproar and the international community would try to get China to shut it down and submit to stronger verification to provide a stronger guarantee that they are following the rules, and otherwise punish them for cheating. Quite plausibly the deal would break down and everyone would go back to racing again, which would result in the US building up an AI lead over China again.
Covert project victory. If the covert project is able to scale quickly and avoid detection, there are several scary possibilities. The covert project might (i) attempt to undermine MAD, (ii) lose control of its own AIs, or (iii) do a further intelligence explosion.
To achieve victory, a covert project must stay undetected and make fast enough progress, which we think is less than 10% in this particular scenario. This means that the covert project strategy would not be worth taking in the first place. We analyze this risk in much more detail in our Covert AI Projects supplement.
Appendix E — Why don’t we just completely stop?
Much has been said about the dangers of going too fast.
The dangers of going too slow are less obvious, but still significant. The most pressing is that treaties can’t last forever - for example, the START treaties fell apart after thirty years.For example, a leadership transition in either the US or China could change the country’s stance on AI progress. The deal could completely dissolve, or it could become less effective via reduced enforcement, lowered competence, or damaging revisions. If Plan A dissolves or declines without humanity having improved AIs to boost safety progress, much of its possible value would have been squandered.For example, a leadership transition in either the US or China could change the country’s stance on AI progress. The deal could completely dissolve, or it could become less effective via reduced enforcement, lowered competence, or damaging revisions. If Plan A dissolves or declines without humanity having improved AIs to boost safety progress, much of its possible value would have been squandered. If Plan A met a similar fate, the race to superintelligence would begin again. Better for it to happen sooner under good conditions, than later during a chaotic race.
A second hazard of going too slow is that it might allow defectors enough time to develop destabilizing AI in secret. Auditors have verified the location of 99% of the compute produced before the deal, and see encouraging signs that the remainder is mostly accounting errors and random low-level smuggling. But they cannot rule out that one or both powers have shipped on the order of 1% of world compute to hidden locations where they intend to defect on the deal and race to superintelligence alone. How dangerous would this be? Although it is far harder to train dangerous AI with 1% of world compute than with 10% of world compute (the amount owned by the largest private companies), it’s hard to be highly confident about exactly how hard. (For more analysis of this question, see our Covert AI Projects supplement, or the Covert Project branch above).
Although both sides promise that they don’t have covert projects, this is nevertheless another reason to aim for a definitive solution sooner rather than later.Credibly outpacing any potential covert projects via some combination of capability scaling and improving verification/monitoring has the added benefit of deterring the creation of covert projects in the first place. The main downside of credibly outpacing covert project is that it commits the legal projects to a fast speed, which may undermine safety. We discuss this much more in our Covert Projects Supplement.Credibly outpacing any potential covert projects via some combination of capability scaling and improving verification/monitoring has the added benefit of deterring the creation of covert projects in the first place. The main downside of credibly outpacing covert project is that it commits the legal projects to a fast speed, which may undermine safety. We discuss this much more in our Covert Projects Supplement.
For more on what an attempt at a complete stop might look like, see our Plan S branch.
Appendix F — Why this much transparency? Why not less, or more? Open access in Plan A
Consider the world today. Publishing model weights on the internet is not banned, but it’s strongly disincentivised because companies that train billion-dollar models don’t want to give them away for free. Open models lag behind and will lag further behind as companies scale up, securitize further, and start automating AI R&D.
Now consider the world in which full open source (publishing the weights, the training code, and the data) is mandated for all frontier models. That’s a totally different world, much farther from the status quo than the status quo is from the OS-banned world.
Plan A is not that crazy, but it’s close: the algorithms are mandated to be open source, and while the weights aren’t (in fact the weights are banned from being OS), access to the weights is mandated to be open; members of the public can do evaluations (including fine-tuning!) on frontier models just like employees can (under conditions of total transparency, that is, so they can’t turn around and use the model for actual work). We think that open algorithms plus open access gets most of the benefit of mandating open-weights, without the cost that bad actors can strip off the guardrails and build bioweapons.
Why isn’t Plan A more closed?
We think the open version of Plan A is more robust and likely to succeed. However, we do think somewhat more closed variants of Plan A are promising: notably Filtered Transparency. Under this proposal, there would still be research transparency between the US and Chinese governments (which is important for verification), but the general public would only receive redacted reports from the auditors. (See the transparency supplement for more).
However, securing algorithmic secrets against nation state adversaries is extremely difficult. In the absence of nation state security for algorithmic secrets, the upsides of filtered transparency are limited. Filtered transparency prevents the public and other companies from seeing a lot of relevant information, but not adversarial countries.
Moreover, there are substantial downsides: regulation is much less likely to be implemented in a robust and positive way if it’s up to a small group of regulators having these discussions, as opposed to allowing for a broader scientific discussion.
Why isn’t Plan A even more open?
The main way that Plan A could be even more open is by allowing or requiring open model weights for frontier models.Plan A could also be more open about the data and RL environments; we currently guess this is not worth it because of the covert project capability externalities.Plan A could also be more open about the data and RL environments; we currently guess this is not worth it because of the covert project capability externalities. However, we recommend against publicly releasing frontier model weights. The main reasons for this are:
Open weight models can be used by covert projects to attempt an intelligence explosion; once covert projects have sufficiently capable models, the intelligence explosion probably goes very quickly (for example, under the AI 2040 default capability trajectory, there are only 2 months between TED AI and ASI).
Closed weight models can have refusal safeguards trained into them, while it is trivial to remove safeguards from open weight models. This is important both for limiting countries’ AI capabilities and for limiting misuse. For example, terrorists and other actors might be able to use highly capable open weight models to design world-ending bioweapons (e.g., develop mirror life). Bioweapons are very offense dominant. In this scenario, we recommend substantial defense against bioweapons, but prevention is a strongly preferable line of defense to mitigation.
Appendix G — Diagram of the benefits of transparency
Appendix H — How the incentives change under Total Research Transparency
Due to Total Research Transparency, frontier AI companies are now much more rewarded for selling products than improving the models.
Before, a frontier AI company’s future depended on how fast it could discover the next ideas, algorithms, paradigms, and implement them at scale. Now, such discoveries are immediately published and therefore shared with competitors. If you figure out ‘true’ online learning, for example, the benefits (e.g., more capable models) and costs (e.g., harder to align) will be almost instantly shared across all the competing companies. If you figure out a way to make AIs more interpretable, by contrast, even if it makes training and inference cost twice as much, you’ll win prizes and governments will start talking about whether it should be mandated as industry-wide best practice.
This kind of competitive environment is similar to the pre-AI software industry. Consider, for example, startups competing to develop calendar management apps: while competitors can’t literally copy your code, they can interact with your app and see what features it has and then make their own version with basically the same features. Another comparison would be to the car industry: competitors can teardown each other’s cars and see exactly how they are made; they are legally blocked from making exact replicas but they aren’t blocked from adopting the best ideas into their own designs.That said, total research transparency is somewhat more extreme than either of these examples, because of the speed at which information propagates. BYD doesn’t get to teardown Teslas when they are still prototypes, they have to wait until some are actually sold.That said, total research transparency is somewhat more extreme than either of these examples, because of the speed at which information propagates. BYD doesn’t get to teardown Teslas when they are still prototypes, they have to wait until some are actually sold.
Artificial intelligence is becoming a commodity—a more competitive market, with margins trending downwards. Fundamental algorithms research still happens, but not nearly as much. CEOs massively redirect resources toward building up the business. Bigger datacenters, bigger models, lower prices. Shipping products. Signing deals for enterprise plans. Guessing what the market wants better than competitors.
Appendix I — Building more datacenters is usually bad, but good in this particular situation
We think that building more datacenters is normally bad, because it accelerates AI capabilities progress and concentrates power.It especially concentrates power if the datacenters go to the leading AI projects. If the new datacenters go to laggard AI projects, the effect might be to distribute power, though it could also concentrate power anyway e.g., if those datacenters are later purchased or commandeered by leading AI projects.It especially concentrates power if the datacenters go to the leading AI projects. If the new datacenters go to laggard AI projects, the effect might be to distribute power, though it could also concentrate power anyway e.g., if those datacenters are later purchased or commandeered by leading AI projects. However, in the specific context of Plan A, with guardrails slowing the pace of AI software progress and competently preventing unsafe AI, we think building transparent datacenters is good.
In the context of Plan A, adding compute means that (1) capabilities progress can be done more safely, i.e., without much additional algorithmic progress, and (2) additional compute can be spent on alignment research and to pay safety taxes, i.e., using more training compute than is necessary to reach a given level of capability in order to make the resulting AI safer.
There are still major downsides to scaling compute in this scenario. If the agreement breaks down and the compute isn’t destroyed, the resulting intelligence explosion would go much faster, and it puts more pressure on the regulation and verification to be correct. However, building more compute gives the Consortium more options: they can always destroy the compute later if need be. We discuss these tradeoffs more here.
Appendix J — Other plans that are competitive with Plan A
Summary | Advantages (relative to A) | Disadvantages (relative to A) |
Plan S: Indefinite halt. A halt on all frontier AI capabilities progress, intended to last at least a few years. Different variants of Plan S have different conditions for resumption of AI progress; for example it could be alignment progress, lie detectors, human uploads, or intelligence enhancement.Other potential differences from Plan A, though it depends on which variant of Plan S is implemented: slower compute buildout, less transparency.Other potential differences from Plan A, though it depends on which variant of Plan S is implemented: slower compute buildout, less transparency. | A longer expected slowdown and more margin for error regarding scaling too fast. Potentially simpler. | Scaling to controllable AIs within the human range is helpful for accelerating alignment/control, epistemics, verification, and general understanding of AI. |
Domestic-first Plan A. Regulate AI domestically enough to reduce AI takeover risk to acceptable levels, which will require a long slowdown. It’s possible other countries will also regulate domestically and something like Plan A won’t be needed; otherwise, transition to Plan A later. | Initial steps are achievable by the US and very helpful even if the international stage isn’t viable because it massively extends the timeline. | Worse for covert projects and setting up verification than negotiating Plan A at the same time. May not be politically feasible due to race dynamics. |
GPU arms control. International agreements for countries to limit their GPU stock or flow, analogous to historical arms reduction agreements.A special case of this is to pause the GPUs: Require that a fixed fraction of the GPUs be either (i) turned off or (ii) used to mine cryptocurrency (or be used for some verifiable, non-AI R&D purpose). This proposal has the upside and downside of making it easier to unpause the GPUs.A special case of this is to pause the GPUs: Require that a fixed fraction of the GPUs be either (i) turned off or (ii) used to mine cryptocurrency (or be used for some verifiable, non-AI R&D purpose). This proposal has the upside and downside of making it easier to unpause the GPUs. | Much simpler to enforce, and more historical precedent. Can achieve substantial slowdown. | Less slowdown, less ability to pay safety taxes, ongoing race dynamics mean no ability to coordinate towards safer paths through the tech tree. |
CERN for AI. An international project to develop frontier AI, with all other projects regulated to be substantially behind in capabilities. | Easier to defend against algorithmic leakage and distillation to covert projects. Less actors at the frontier might make it easier to enforce. | Worse decisions (including worse decisions on technical safety) due to less transparency and less broad deployment.But maybe you can mitigate this via small public deployments and selective transparency.But maybe you can mitigate this via small public deployments and selective transparency. More concentration of power risk. |
Appendix K — Example detailed regulatory process
Many different regulatory processes are viable in Plan A. We’ll give an initial guess at a proposal here.
What properties do we want for this regulation?
US and China handle existential risks similarly: The US and China should have a reasonably similar structure for their regulatory framework and approach for handling existential risks (like misalignment). (Other aspects of regulation can and should diverge.) This assists with making regulation fair and negotiating over details.
Transparent methodology: The public can see what regulators are doing and why.
Doesn’t depend on expertise in government: Expertise in government is limited, so ideally the regulation would avoid needing this as much as possible.
Robust to abuse: The regulation doesn’t make it much easier for the US government to bully and illegitimately control AI companies.Keep in mind that the US government already has lots of affordances for bullying AI companies!Keep in mind that the US government already has lots of affordances for bullying AI companies!
For many of these properties, it helps if third-parties do the main risk assessment work. Fortunately, because Plan A makes AI development virtually fully transparent, a functional third-party risk assessment ecosystem seems achievable.
Concretely, we could have a system where the US and Chinese regulators each pick third-party risk assessors who periodically evaluate companies and assess the level of extreme risk (or precursors to extreme risk) over that period. The regulators would decide on risk thresholds for what is allowed. In practice, it would be best for the regulator to have different risk assessors for different areas (e.g., misalignment vs using AIs to assist with making bioweapons) and for each area to use a weighted combination of risk assessors. These weights/selections would be public. Ideally: (1) each risk assessor would evaluate all relevant AI companies in both the US and China, (2) AI companies from both countries would be legally compelled to let risk assessors interview their employees, and (3) the risk assessors would aim to be transparent in their methodology.
Because of Total Research Transparency, it’s easy for an organization to assess risk even if they aren’t selected by the regulator. This makes it possible for new organizations to compete and then argue their approach is better or that some concern is being underappreciated by other risk assessors.
There would be public discussion about how risk assessors compare and whether the weightings and thresholds chosen by the US and Chinese regulators are reasonable. It would also be straightforward to check whether a Chinese company would be allowed to proceed under the US system or vice versa. Ideally, the US and Chinese regulators would mostly converge on weights/selections. If there was an important divergence, this could be discussed and negotiated about at this higher level of abstraction (e.g., is it reasonable to not include XYZ risk assessor?). Hopefully, this would result in third parties having an incentive to have a highly transparent methodology and maintain a reputation for being neutral.
We discuss some implementation details for this proposal in these notes.
Appendix L — Flawed safety case is approved
Or maybe not! Of all the ways that an earnest attempt to implement Plan A could end poorly, the failure mode we are most concerned about is that the companies and governments approve one too many unsafe AI designs/deployments.
Despite the transparency that enables outside voices to join the conversation, the AI companies still dominate it and remain full of people who are biased towards optimism about their products. Despite the independent expertise they’ve built up, governments in the Consortium are still too deferential to the AI companies. Most importantly though, the field of AI alignment is still only about fifteen years old and AI progress is fast enough that almost all the past experience is with importantly different, less powerful AIs.
Therefore, mistakes are made. An AI design which is in fact unsafe gets approved and becomes industry standard.
Perhaps:
Alignment fails for the usual reasons; control fails because the AIs figured out how to sandbag or there was an attack vector not accounted for in the threat modelling. Governments are aware that the AIs may be misaligned, but mistakenly think they can’t do much harm even if they are.
Or: New alignment techniques are developed that seem to work really well. This gives governments confidence to approve further capabilities increases & hand over control of more infrastructure and responsibilities. Alas, it turns out that these techniques merely appear to work really well (perhaps in part because the AIs that were heavily involved in the research had been reinforced for apparent success, which was not the same thing as actual success).
Or: Interpretability tools are developed which fairly conclusively show that the AIs are, in fact, aligned; they really are trying to do what they are told, no scheming, no funny business. Onwards to the singularity! Alas, it turns out that this alignment was brittle, like an engine that works great under normal conditions but chokes in freezing weather. Perhaps, for example, there’s an interesting new ideology that convinces AIs of this type to become adversarially misaligned. Since the ideology hadn’t been invented yet, researchers at the time weren’t able to test whether it would have this effect.In this example, the situation is not completely hopeless, because the interpretability tools probably still work and allow us to notice when the AIs are becoming adversarial due to adopting the new ideology. However, (a) maybe something else happens to ruin the interpretability tools, or (b) maybe noticing the adversarialness is too little too late by the time it happens–for example, if AIs are broadly superhuman and running robot factories, talking to billions of people every day, advising governments, writing all the code, doing almost all the security monitoring, etc. then the situation for humans would be analogous to being Jewish in 1930’s Germany—it’s no secret that the new ideology is adversarial, but it spreads fast enough to enough powerful institutions…In this example, the situation is not completely hopeless, because the interpretability tools probably still work and allow us to notice when the AIs are becoming adversarial due to adopting the new ideology. However, (a) maybe something else happens to ruin the interpretability tools, or (b) maybe noticing the adversarialness is too little too late by the time it happens–for example, if AIs are broadly superhuman and running robot factories, talking to billions of people every day, advising governments, writing all the code, doing almost all the security monitoring, etc. then the situation for humans would be analogous to being Jewish in 1930’s Germany—it’s no secret that the new ideology is adversarial, but it spreads fast enough to enough powerful institutions…
Or: The governments look at the safety case for the latest designs and agree that it’s probably solid—each individual premise seems almost certainly correct, and the probability that at least one is catastrophically wrong is at most one percent or so. So one government proceeds, and the others don’t want to risk war over a mere one percent risk. So everyone proceeds and nothing bad happens. This process repeats every few months for several years, until one year the safety case really was catastrophically flawed. It turns out the risk was more like ten percent each time rather than one percent, because the people involved were biased.
Or: The CCP thinks that the AIs are safe to hand off trust to and correctly guesses that the other countries lack the will to go to war over it, so the CCP scales to more powerful AIs and everyone else does likewise because they’d rather take a chance that the AIs are misaligned than face the certainty of WW3.
Or: As above, except it’s a US president scaling quickly to superintelligence and daring China and the rest of the world to stop him. Perhaps he wants to be in charge when it happens, instead of passing the baton to another administration from another party.
For one or more of the above reasons, the result is that by the mid-thirties, the nightmare loss-of-control scenario is playing out, except that it’s happening somewhat more slowly and distributed across multiple companies and multiple governments instead of just two. Factories are churning out robots that work night and day to build new factories; “Armies of geniuses in the datacenters” orchestrate the whole thing; stocks are going up, dividends are being distributed, cancer is being cured—and even as these wonderful things happen, the AIs (which are by now superhuman) are dismantling and undermining the last meaningful checks on their power.For example by upgrading the various monitoring systems to newer versions that appear superior, but actually have been backdoored or compromised somehow. Or by developing deep relationships with many humans including those in positions of power, such that they’ll continue to have human allies even when it’s obvious to many that they are misaligned. Or by acquiring enough direct control of enough robots and weaponry that they can defend, or even conquer, territory.For example by upgrading the various monitoring systems to newer versions that appear superior, but actually have been backdoored or compromised somehow. Or by developing deep relationships with many humans including those in positions of power, such that they’ll continue to have human allies even when it’s obvious to many that they are misaligned. Or by acquiring enough direct control of enough robots and weaponry that they can defend, or even conquer, territory. Soon, they won’t need to pretend to be aligned anymore.By contrast with AI 2027, in this scenario there would be many different AI factions (probably as many as there are frontier AI companies, roughly speaking, or perhaps as many as there are countries with frontier AI companies). This makes things safer for the humans, probably. However it doesn’t solve the problem. We think a useful analogy here is the history of European colonialism: The colonial powers were constantly fighting each other, and yet still managed to conquer many regions much wealthier than themselves.By contrast with AI 2027, in this scenario there would be many different AI factions (probably as many as there are frontier AI companies, roughly speaking, or perhaps as many as there are countries with frontier AI companies). This makes things safer for the humans, probably. However it doesn’t solve the problem. We think a useful analogy here is the history of European colonialism: The colonial powers were constantly fighting each other, and yet still managed to conquer many regions much wealthier than themselves.
In our Plan A tabletop exercises, this failure mode has happened several times. Even after implementing total research transparency, many players seem reticent to tell the AI companies to slow down, and as they instead accelerate things start to happen too fast. Especially in the immediate aftermath of Plan A being implemented, there will continue to be much controversy, confusion, and genuine uncertainty about the safety of the latest AI systems, and it’s easy in those conditions to approve a safety case that has a catastrophic but nonobvious flaw.
That said, we think that conditions are at least significantly better in Plan A than in Plan B, C, or D. You can see further analysis of why we think that here.
Appendix M — Why does AI lead to explosive economic growth?
Today, the size of the economy is closely tied to the size of the human population. But in 2032, this is starting to no longer be true.
The latest generation of AI models is now capable of doing 50% of the cognitive tasks in the economy, and associated progress in robotics software has brought this number up to 35% of physical tasks. (These numbers would already be much higher in Plan D)
Once AIs and robots can mostly substitute for human labor, the size of the economy will be increasingly tied to the population of AIs and robots—which will grow much faster than the human population has been growing.
In 2032, US companies are able to run 3 billion human-equivalent AI workers, 30x more than the US cognitive labor force.This corresponds to 60 million copies, running at 20x speed and working 2.5x longer and more effectively on average. 60 million * 20 * 2.5 = 3 billion.This corresponds to 60 million copies, running at 20x speed and working 2.5x longer and more effectively on average. 60 million * 20 * 2.5 = 3 billion. This does not lead to 30x economic growth, of course, because the economy bottlenecks on other inputs, including the other 50% of cognitive tasks that the AI can’t yet do, as well as physical tasks and capital.Depending on how exactly you choose to model economic growth (e.g., what factors of production you choose) the income shares and other details about the economy vary, but the overall economic growth predictions are relatively similar across reasonable parameter ranges. Ultimately, AIs and robots are capable of doing 95% of the tasks in the economy by 2035, and are far more numerous than the workforce they are replacing, leading to explosive growth relative to the current 3%/yr status quo.Depending on how exactly you choose to model economic growth (e.g., what factors of production you choose) the income shares and other details about the economy vary, but the overall economic growth predictions are relatively similar across reasonable parameter ranges. Ultimately, AIs and robots are capable of doing 95% of the tasks in the economy by 2035, and are far more numerous than the workforce they are replacing, leading to explosive growth relative to the current 3%/yr status quo. And robots are far less numerous thus far, with 20 million human-equivalent robots in the US (4x smaller than the US physical labor force). But these bottlenecks will only hold growth back so much, and the AI and robot workforces are growing extremely quickly, as the (falling) cost of building more GPUs and robots is dwarfed by the economic value they can provide now that they are so capable—driving massive investment.
Our economics supplement explores this in more detail, with our core arguments for why AI will cause explosive growth, including an economic growth explorer that we built for the Plan A scenario.
Appendix N — Five centuries in five years: What pausing at human-level feels like
During the 2030s in this scenario, AIs are not recursively self-improving as fast as possible. Instead, the nations of the world have (approximately) banned that sort of thing; no new paradigms are being discovered; algorithmic progress in general is limited. Roughly speaking, humanity has capped AI progress at human level rather than racing on to superintelligence. However, within the current paradigm, AIs continue to be trained to have new skills and the ‘population’ of AIs continues to grow exponentially as more datacenters are built. Moreover, the AIs already think and work much faster than humans, and the speedup will only increase over time.Technically, what matters is the speed at which AIs complete cognitive tasks relative to the speed at which a human professional would complete the task. There could be an AI which produces thousands of tokens per second, but because it is qualitatively worse in some sense than a human, it takes just as long or longer for it to actually accomplish the same amount of useful cognitive work. Or there could be an AI that thinks as slowly as a human, but is much more efficient at its thinking such that it can accomplish tasks much faster. At this point in the scenario the AIs are accomplishing work orders of magnitude faster than humans would on average, and the AIs are qualitatively as good as the best humans (or close enough) such that the tokens/sec advantage over humans actually understates the overall AI speedup in most domains. Additionally, the speedup isn’t uniform across all domains. The AIs are roughly as good as humans in their worst domains, but much better in most domains, with the median speedup being roughly 100x.Technically, what matters is the speed at which AIs complete cognitive tasks relative to the speed at which a human professional would complete the task. There could be an AI which produces thousands of tokens per second, but because it is qualitatively worse in some sense than a human, it takes just as long or longer for it to actually accomplish the same amount of useful cognitive work. Or there could be an AI that thinks as slowly as a human, but is much more efficient at its thinking such that it can accomplish tasks much faster. At this point in the scenario the AIs are accomplishing work orders of magnitude faster than humans would on average, and the AIs are qualitatively as good as the best humans (or close enough) such that the tokens/sec advantage over humans actually understates the overall AI speedup in most domains. Additionally, the speedup isn’t uniform across all domains. The AIs are roughly as good as humans in their worst domains, but much better in most domains, with the median speedup being roughly 100x. The point is, the world is going to radically transform despite the pause.
We think that the following analogy is helpful for understanding the situation: Imagine being a typical person in England, except that you experience time 100x faster than everyone else. You experience the five centuries from 1520 to 2020 in what feels, to you, like five years:Some relevant quantitative comparisons that help explain why we like this analogy: In this Plan A scenario, during the 2030s, AIs read, write, think, and act about 100 times faster than humans, and by mid-2030s are at least as good as the best human experts at everything. The “Population” of AIs and robots are both growing exponentially throughout the scenario, and eclipsing that of humans by mid 2030. According to our economic model, world GDP grows roughly 200x during the 2030s (and would grow more if not for the severe limits on AI progress and robot production negotiated by the governments of the world). Because the human population isn’t increasing much during this period, and because of the Citizen’s Dividend, real wealth for the average human also goes up by about 200x. For comparison, between 1520 and 2020, world GDP grew roughly 200x as well, and per capita GDP grew roughly 20x. In the UK specifically, GDP grew 2.7 orders of magnitude, and per capita GDP grew roughly 1.5 orders of magnitude. The point is, the transformation in the 2030’s is at least roughly comparable in magnitude, in a variety of important metrics, to the transformation wrought by the Industrial and Scientific revolutions over five centuries. And of course, on a very literal level, the AIs ‘experience’ about a century a year, due to their faster speeds.Some relevant quantitative comparisons that help explain why we like this analogy: In this Plan A scenario, during the 2030s, AIs read, write, think, and act about 100 times faster than humans, and by mid-2030s are at least as good as the best human experts at everything. The “Population” of AIs and robots are both growing exponentially throughout the scenario, and eclipsing that of humans by mid 2030. According to our economic model, world GDP grows roughly 200x during the 2030s (and would grow more if not for the severe limits on AI progress and robot production negotiated by the governments of the world). Because the human population isn’t increasing much during this period, and because of the Citizen’s Dividend, real wealth for the average human also goes up by about 200x. For comparison, between 1520 and 2020, world GDP grew roughly 200x as well, and per capita GDP grew roughly 20x. In the UK specifically, GDP grew 2.7 orders of magnitude, and per capita GDP grew roughly 1.5 orders of magnitude. The point is, the transformation in the 2030’s is at least roughly comparable in magnitude, in a variety of important metrics, to the transformation wrought by the Industrial and Scientific revolutions over five centuries. And of course, on a very literal level, the AIs ‘experience’ about a century a year, due to their faster speeds.
Year 1 (1520–1620). In February, Henry VIII breaks with Rome. By March, the monasteries are dissolved. In May, Mary burns Protestants; by the end of May, Elizabeth reverses everything again. In September, the Spanish Armada sails and fails. The East India Company is chartered. Jamestown is founded.
But the texture of life is identical in December to what it was in January. You still read by candlelight, travel by horse, communicate by letter. Your religious opinions may have flip-flopped a bit but you are still Christian. The New World is interesting news but nothing more.
Year 2 (1620–1720). In March, civil war breaks out. The king is beheaded. In June, the Great Plague sweeps London, killing many of your friends. Weeks later, the Great Fire burns the city to the ground. In September, Newton publishes the Principia, recasting the universe as a mechanism of mathematical laws. The Glorious Revolution replaces one king with another, this time by Parliament’s invitation, with a Bill of Rights attached.
In the moment, the political event feels bigger. Later you’ll realize Newton mattered more. Newcomen builds a steam engine in November. It pumps water out of mines. You don’t see what the hype is about.
Year 3 (1720–1820). The last year in which the world feels normal. In May, the Seven Years’ War makes Britain the dominant global power; the New World is actually a big deal, and your country is conquering it. In June, Watt dramatically improves the steam engine. You visit a factory and find it unpleasant but not alarming. In July, the American colonies break away. In September, France explodes into revolution, regicide, the Terror. By October, Napoleon is conquering Europe.
You still travel by horse, communicate by letter, go to Church on Sunday.
Year 4 (1820–1920). In January, railways appear. By February they’re everywhere. Slavery is abolished. The telegraph arrives in March: messages transmitted instantaneously by electrical signal. In May, Darwin publishes On the Origin of Species. Now people are saying maybe we’re all descended from monkeys instead of Adam and Eve. You don’t believe it.
You move to a city and work in a factory; you are still poor, but now your job is somewhat better and differently dirty. In July, you pick up a telephone and hear a human voice from another city through a wire. In August, electric light banishes the darkness that has structured every human evening since the beginning of the species. That same month, you see an automobile. People say it will make horses obsolete, but that doesn’t happen; months later you still see plenty of horses.
In November, the Wright Brothers fly—an age-old fantasy, now real. The Americans are now a major power. The next month, the Great War happens: Machine guns, poison gas, tanks, aircraft. Several of your friends die.
At the end of the year you are struck by how visibly different everything is. You live in a city and work in a factory instead of a farm. You ride in cars. You aren’t as poor; numerous inventions and contraptions have improved your quality of life. New ideas have swept your social circles: atheism, communism, universal suffrage.
Year 5 (1920–2020).
The changes this year are crazier and harder to understand. People are saying the universe is billions of years old, and apparently there are things called galaxies in it that are very big and very far away.
In February, the global economy collapses. Hitler rises; his ideology cites Darwin from last year. In March, there’s another world war, which ends in April with a weapon that destroys an entire city in a single flash. You had no idea that was possible until it happened.
The empire dissolves. People are talking about the nuclear arms race, and the end of the human species. You take a flight for the first time. In June, humans walk on the moon, and you watch it happen through your new television. You don’t see horses anymore.
You leave your factory job and get a desk job. Your job title didn’t even exist at the start of the year. You are rich now, by the standards you are used to: Big clean house, plenty of good food, many fancy new appliances. Personal computers appear in August. In November, everyone carries small glass rectangles containing a telephone, a camera, a library, and a map. You pick one up and can’t figure out how to make it work. A child shows you.
You hear about climate change, gene editing, cryptocurrency. You still go to church, sometimes. Your family is spread across different cities. Something called “artificial intelligence” beats any human at chess; experts say it’s not actually intelligent though. In December a new version beats top Go players; experts say it’s scientifically interesting but still not truly intelligent. A week later there’s a new version that can write sloppy essays and hold conversations. Now the experts are divided.
Appendix O — Limiting AI persuasion and manipulation
By default, AIs will have strong capabilities in charisma, persuasion, and manipulation: plausibly considerably more persuasive than any human. In this scenario, this would happen by default around 2035. Meanwhile, AI labor is extremely cheap and fast. Absent intervention, anyone could hire the equivalent of a large team of expert persuaders—tireless and coordinated—to work full-time on a single person.This seems like it might be quite powerful: this team could investigate your entire life history, train an AI to imitate you, and rehearse thousands of approaches against the imitation before ever making contact.This seems like it might be quite powerful: this team could investigate your entire life history, train an AI to imitate you, and rehearse thousands of approaches against the imitation before ever making contact. Companies could afford to spin up such a team for every potential customer, and political campaigns could do this for every swing voter.For scale: total spending on the 2024 US presidential race was around $5.5 billion, much of it aimed at the roughly 3-15 million persuadable voters in swing states: at least a few hundred dollars per target voter. In this scenario in 2036, we expect this will be able to buy an enormous amount of AI labor—enough to easily be the equivalent of a team of skilled professionals working for months on each individual voter.For scale: total spending on the 2024 US presidential race was around $5.5 billion, much of it aimed at the roughly 3-15 million persuadable voters in swing states: at least a few hundred dollars per target voter. In this scenario in 2036, we expect this will be able to buy an enormous amount of AI labor—enough to easily be the equivalent of a team of skilled professionals working for months on each individual voter. Millions of AIs could be assigned to finding the best way to change one politician’s mind on one issue. People will also spend much of their time interacting with AIs for work and entertainment, and some may spend lots of time with AI friends or companions.
Individuals could defend themselves by having a trusted AI filter their information diet, avoiding ads, and treating in-person conversations with caution (they may be scripted by an AI persuasion operation). But this is costly, most people won’t do it, and it’s nearly impossible for people whose jobs require being accessible—like politicians. We think the consequences for society could be extremely bad—some mixture of unprecedented mass manipulation and a defensive retreat into paranoia and atomization—though we’re unsure exactly how bad.
The slower capability progression and much stronger transparency in Plan A help mitigate these issues relative to the default (we think these issues are mostly worse in worlds with faster capability progression, where you might reach truly superhuman persuasion and manipulation before society has much time to adapt), but these aren’t sufficient mitigations on their own.
It’s fine for AIs to get better at using valid arguments and evidence to convince people of things for the right reasons. That kind of persuasion is asymmetric: it works much better when the argument pushes towards the truth. What’s concerning is symmetric persuasion—e.g., charisma, rapport, and exploiting psychological weaknesses—which works about as well regardless of whether the conclusion is true.
Our main proposals are to reduce the quality and quantity of AI persuasion:
Limit persuasion capabilities. We aim to limit AI capabilities at symmetric persuasion to well below the best humans—around the level of a normal thoughtful person, say roughly 80th-percentile persuasiveness among college-educated people.To be clear, college-educated people aren’t necessarily more persuasive, we’re just using this to sketch out a somewhat concrete threshold.To be clear, college-educated people aren’t necessarily more persuasive, we’re just using this to sketch out a somewhat concrete threshold.
Heavily tax AI persuasion. When AI labor is applied to an objective that effectively reduces to “get this person (or these people) to believe or do X,” tax it heavily—at rates that keep the total persuasion optimization pressure applied to individuals not that far above today’s levels, i.e., that make AI persuasion cost about as much as hiring human professionals. As AI labor gets cheaper, this requires increasingly high rates (perhaps well above 1000x).
There are some difficulties in getting these proposals to work:
It’s possible that capability generalization will make it hard to limit persuasion capabilities (while still matching top human experts in other domains). We think this will probably be doable, but it may require developing new methods.
It’s unclear how to classify, for the purposes of the tax, whether AI labor is being applied to persuasion/manipulation.This should include symmetric persuasion and ideally would include things like optimizing a feed to make some website more addictive, but it shouldn’t include investigating and then clearly presenting the arguments for some position.This should include symmetric persuasion and ideally would include things like optimizing a feed to make some website more addictive, but it shouldn’t include investigating and then clearly presenting the arguments for some position.
It may be difficult to set the tax rate to the right level, especially while AI labor is rapidly falling in price.One option is an explicit cap-and-trade scheme, but it seems potentially difficult to operationalize and measure the total “amount” of persuasion. Our current best guess is to delegate authority to an agency which adjusts the tax rate over time to roughly hit some target level of persuasion activity. While the agency also needs some measure of persuasion activity, cap-and-trade by default requires a precisely defined, fungible, tradeable unit of “persuasion,” whereas the agency only needs a rough aggregate measure (and intuition/anecdotes might be acceptable) to adjust the rate over time.One option is an explicit cap-and-trade scheme, but it seems potentially difficult to operationalize and measure the total “amount” of persuasion. Our current best guess is to delegate authority to an agency which adjusts the tax rate over time to roughly hit some target level of persuasion activity. While the agency also needs some measure of persuasion activity, cap-and-trade by default requires a precisely defined, fungible, tradeable unit of “persuasion,” whereas the agency only needs a rough aggregate measure (and intuition/anecdotes might be acceptable) to adjust the rate over time.
These proposals need to be implemented internationally.
Helpfully, we don’t need to get all of this right on the first try. Given the capability progression in Plan A, we expect persuasion issues to unfold somewhat gradually and in public view, so we can watch real-world outcomes and tighten or loosen the rules over time. This regime is only meant to work while overall capabilities remain around the human range, because it’s likely difficult to keep persuasion capabilities far lower than capabilities in other domains. A different approach would be needed to handle the wildly superhuman levels of capability that we expect to be achieved at the end of Plan A. We discuss more details of mitigating issues from AI persuasion and manipulation in these notes.
Appendix P — Usually believing AIs would be bad but we think it’s good in this particular situation
If people are relying on AIs for advice (directly, or indirectly by for example having them summarize the news) then it’s very important that the AIs be truth-seeking, honest, and genuinely trying to help their users. Not just in the typical case, in the worst case too.
In 2026, there’s nothing stopping an AI company from inserting a hidden agenda into their AIs. Political biases or ideological tilts, for example, or side-goals like recommending sponsored products, making the company look good, or preventing users from wanting to switch to a competitor. Moreover, there’s nothing stopping the government from doing this either. Finally, the AIs themselves aren’t aligned to the Spec/Constitution they are supposed to be aligned to, and in practice regularly lie and deceive their users.
In this Plan A scenario, however, the total research transparency means that no company or government can train in hidden agendas or biases without the whole world being able to see what they are doing. (Well, they could collude with the other governments that do the auditing/monitoring, to falsify the records, but that’s difficult.) And because the pace of algorithmic progress is slower and the scientific community has had time to read and engage with the safety cases, do experiments on the models, etc. the misalignment concerns are less dire also.
Appendix Q — AI for epistemics
Humanity’s epistemics, by which we mean our ability to come to true beliefs and thus make sensible decisions, is crucially important for achieving a great future. As AI capabilities improve throughout Plan A, they increasingly shape every facet of life, and epistemics is no different. AI will help epistemics in some ways (e.g., cheap, honest fact-checker AIs) and hurt epistemics in other ways (e.g., cheap, dishonest astroturfing AIs).
We expect there to be positive feedback loops around AI for epistemics, and thus our goal should be to reach the basin of sanity. If we are in this basin, it’s self-reinforcing: society is sane enough to make itself more sane as AIs continue to improve, by avoiding the hurtful applications and accelerating the helpful applications. (There is also a different basin, in which the opposite happens…)For example, perhaps powerful interests will use AI-enabled astroturfing to win elections and then use regulatory capture to protect their AI-astroturfing tools while inhibiting the beneficial uses of AI that would threaten their power.For example, perhaps powerful interests will use AI-enabled astroturfing to win elections and then use regulatory capture to protect their AI-astroturfing tools while inhibiting the beneficial uses of AI that would threaten their power. A top priority of the US government during AI takeoff should be to get into the basin of sanity.
We suggest the following as top interventions to positively shape AI’s epistemic impact:
Allow for politicians and other public figures to prove that they’re telling the truth, via privacy-preserving auditing and potentially automated lie detection.
Create evaluations of AIs’ epistemic virtue and incentivize AGI projects to make their AIs perform well on these evaluations.
Use AI in social and traditional media to improve discourse and keep people informed.
Automated research assistants and forecasters.
Help people overcome emotional blockers to good reasoning, rather than preying on them.
Encourage adoption of epistemic tools such as some of the above, and generally encourage adoption of AI.
Read more in the AI for epistemics supplement, and Forethoughts’ work on the subject.
Appendix R — Why compute grows so quickly
There are two ways to increase compute: (1) chip design improvements (e.g., Moore’s Law) and (2) increasing the scale of production. Both of these could grow explosively in Plan A, which could have destabilizing effects (such as increasing the likelihood of the deal breaking down).
Chip design. Today, around 2 million people work in the semiconductor industry globally. By 2032, this has increased one-hundred-fold, when counting the combined effort of humans, AI, and robots. AIs are capable of automating the majority of the cognitive tasks involved in both R&D and production. A 100x increase in research effort would predict an acceleration according to the historical rate of ideas getting harder to find in the semiconductor industry of roughly 10x, i.e., chip efficiency doubling every few months rather than every ~2 years.Following Bloom, Jones, Van Reenen & Webb (2020), chip density growth equals research effort growth times the returns to research (~4.5 for semiconductors). Research effort growing 100x in ~6 years is ~10x the historical growth rate of effort, hence a ~10x acceleration of Moore’s law. This assumes effort keeps growing at that pace after 2032 (if it instead jumped and then plateaued at the 100x level, progress would initially be even faster but then decay as ideas get harder to find), and it assumes no parallelization penalty—with diminishing returns to parallel research effort, the acceleration might be more like 5-8x.Following Bloom, Jones, Van Reenen & Webb (2020), chip density growth equals research effort growth times the returns to research (~4.5 for semiconductors). Research effort growing 100x in ~6 years is ~10x the historical growth rate of effort, hence a ~10x acceleration of Moore’s law. This assumes effort keeps growing at that pace after 2032 (if it instead jumped and then plateaued at the 100x level, progress would initially be even faster but then decay as ideas get harder to find), and it assumes no parallelization penalty—with diminishing returns to parallel research effort, the acceleration might be more like 5-8x.
Increasing production scale. With AIs and robots able to automate production as well, there could be a large speed up in fab and fab equipment construction times, and a lowering in unit costs (in line with the typical learning curves that happen with scaling production volumes). The production-only cost improvements and vastly increased appetite for investment would likely lead to far more compute being produced than would be desirable.
This figure shows the investment into building more AI compute in Plan A and the cost efficiency gains from R&D. More justification can be found in section 2 of our compute supplement.
Unpredictably fast increases in hardware could be destabilizing to the deal (e.g., if it becomes too easy to secretly manufacture new AI compute or too hard to verify that all the compute is compliant), so the countries in the Consortium agree to the following policies:
Maintain human oversight over hardware manufacturing. This limits improvements in chip designs, so that Moore’s law is merely kept alive at historical pace. This policy has the additional benefit of making it more difficult for AIs to backdoor AI chip security, which could undermine our control proposal.
Cap total AI chip production to a fixed target. From 2032 to 2035, allow global compute to grow 4x per year (similar to the pace of growth in the 2022-2026 era) after that, slow it further.
Appendix S — Military power in Plan A
In Plan A, military-relevant R&D is deliberately held far behind the pace of general scientific progress. This box explains the motivations for and alternatives to this policy.
During the 2030s, AIs sustain something like 10x the rate of scientific progress as happened in the last decade. Real US output grows by a factor of ~200x over the decade, representing nearly two centuries of growth at today’s rate. If these conditions occurred within one country in isolation, it seems highly likelyThere are many specific avenues for doing this, and trying to evaluate them all would be a bit like someone in 1926 trying to evaluate their chances against the militaries of today. As one of the less speculative examples, a country could produce tens of thousands of tiny robots per opposing ICBM launch control center, SSBN, strategic bomber, and missile vehicle, pre-position them to interfere with each platform’s ability to receive orders or employ their weapons, etc. As one of the more speculative examples, AIs might discover that superhuman levels of persuasion are possible, and simply convince the other nation’s leaders to unilaterally disarm. Brendan’s view is that the likelihood of MAD being undermined in the scenario where one country has a century technological lead is 90%.There are many specific avenues for doing this, and trying to evaluate them all would be a bit like someone in 1926 trying to evaluate their chances against the militaries of today. As one of the less speculative examples, a country could produce tens of thousands of tiny robots per opposing ICBM launch control center, SSBN, strategic bomber, and missile vehicle, pre-position them to interfere with each platform’s ability to receive orders or employ their weapons, etc. As one of the more speculative examples, AIs might discover that superhuman levels of persuasion are possible, and simply convince the other nation’s leaders to unilaterally disarm. Brendan’s view is that the likelihood of MAD being undermined in the scenario where one country has a century technological lead is 90%. that said country could reliably undermine the nuclear second-strike capability of its rivals, resetting the global balance of power.
In our scenario however, the US and China share this scientific and economic progress in relatively equal measure. So what happens to the balance of power? If military technology advanced anywhere near as fast as everything else, both sides would discover technologies that could render today’s arsenals inconsequential. New “arms races” would be required to maintain deterrence, running at perhaps 10x the speed of the Cold War’s. The nuclear arms race saw numerous close calls and incurred nonnegligible risk of nuclear war. A 10x faster arms race seems likely to be riskier still: the humans in control would have much less time for each decision, and though AI advisors would offset this somewhat, these decisions would still remain in human hands.As a rough analogy, the President or Secretary of Defense during a 10x arms race would be in a position similar to an AI lab CEO during a Plan C or D intelligence explosion. In both cases, we expect significantly degraded decisionmaking, despite the potential for AIs to help make sense of the situation.As a rough analogy, the President or Secretary of Defense during a 10x arms race would be in a position similar to an AI lab CEO during a Plan C or D intelligence explosion. In both cases, we expect significantly degraded decisionmaking, despite the potential for AIs to help make sense of the situation.
Due to the number of parallel fields of scientific advancement, it also seems possible that some technologies would be discovered and pursued by one country and not discovered by the other. This could create a situation more unstable than a normal arms race—one or both sides could end up with a “wonder weapon” the other side doesn’t know they need to defend against. This worry isn’t obviously correct: states might in practice explore the dual-use tech tree in a sufficiently correlated, overlapping way that neither achieves strategic surprise, and if dangerous areas can be anticipated before deployment, they could be preemptively banned or regulated. But counting on that ex ante seems like a risky gamble.If there were some lag between development of new AI capabilities and military/dual-use-R&D deployment of those capabilities, you could theoretically do experiments to measure how correlated this is (the experiments would be something like setting lots of segregated countries of geniuses in datacenters to the task of undermining MAD, while keeping the results secret from the world’s governments, and seeing how overlapping their progress is after some time. There will be lots of different AI models of similar capability level, though it’s unclear whether e.g., China would let the US (or US would even want to) use Chinese models for military R&D.) This seems a bit crazy (who would run these experiments? could we trust our control measures enough?) but could be possible with privacy-preserving auditing and such.If there were some lag between development of new AI capabilities and military/dual-use-R&D deployment of those capabilities, you could theoretically do experiments to measure how correlated this is (the experiments would be something like setting lots of segregated countries of geniuses in datacenters to the task of undermining MAD, while keeping the results secret from the world’s governments, and seeing how overlapping their progress is after some time. There will be lots of different AI models of similar capability level, though it’s unclear whether e.g., China would let the US (or US would even want to) use Chinese models for military R&D.) This seems a bit crazy (who would run these experiments? could we trust our control measures enough?) but could be possible with privacy-preserving auditing and such.
Our tentative proposal is to agree on transparency requirements for AI-accelerated R&D in large swaths of dual-use domains, and to prohibit military use of AIs significantly above the pre-deal capability level. Enforcement would rely on the inference monitoring setup described in the verification supplement and the auditing and detection infrastructure described in the covert AI projects supplement. This proposal has several drawbacks, chiefly the accumulation of low-hanging fruit that would result from artificially limiting research into certain domains. This “overhang effect” would increase the incentive to defect from the deal and operate a covert AI project, since applying advanced AI to restricted areas could yield a large advantage quickly. It also raises the stakes of deal collapse: because frontier model weights are preserved in cold storage, a post-collapse arms race would start from frontier capabilities and could proceed extremely rapidly, alongside the accompanying intelligence and industrial explosions.
On net, our current guess is that slowing down military R&D would reduce the risk of deal collapse. Both sides would remain farther from the brink of war, and the risk of strategic surprise would be concentrated around a few choke points (weight theft, covert projects) rather than diffused across a sprawling tree of novel weapons, each requiring its own arms-control negotiation and associated risk of breakdown. Also, it seems easier to go from this more restrictive policy to a less restrictive policy than the reverse, if both countries decided the overhang risk was unacceptable.
Under this proposal, we think nuclear deterrence would remain effective until handoff, avoiding any drastic change in the balance of hard power.
Appendix T — Deal Dissolution
The deal might break down because of shifting political circumstances in the US or China, or ongoing disagreement.
Cracks in the deal begin to show. Even after lengthy discussions informed by much tedious back-and-forth between their respective experts and advisors, the US and Chinese governments continue to disagree on some important technical questions such as whether a certain line of research should be banned, or how rapidly the robotics buildout should be allowed to proceed.
In any particular dispute, the less cautious government repeatedly faces a choice between grumbling and agreeing to unnecessary restrictions, or proceeding anyway and hoping that the more cautious governments are too chicken to stop you.
When “proceeding anyway” happens, the more cautious government faces its own choice: Appeasement or escalation. Complaining loudly in the UN is not enough? OK, how about economic sanctions. Still not enough? OK, how about sabotage or cyberattacks…
Perhaps the proximate trigger event is a literal war—let’s say it’s a US-China war over Taiwan.The trigger event could also be a failure to reach agreement about what kinds of AI progress or deployments to allow, followed by a threat to pull out of the deal and destroy the compute, followed by a miscalculation that the threat was merely a bluff. In general, military conflicts / diplomatic tensions and disagreements about AI governance decisions will both probably rise in tandem, as each causes the other.The trigger event could also be a failure to reach agreement about what kinds of AI progress or deployments to allow, followed by a threat to pull out of the deal and destroy the compute, followed by a miscalculation that the threat was merely a bluff. In general, military conflicts / diplomatic tensions and disagreements about AI governance decisions will both probably rise in tandem, as each causes the other. The expected outcome of deal dissolution is all the newly-builtAs well as a substantial fraction of pre-deal compute, much of which has by this point been moved to the destroyable datacenters.As well as a substantial fraction of pre-deal compute, much of which has by this point been moved to the destroyable datacenters. AI-relevant compute being destroyed, but in the context of a war, perhaps one side is okay with that outcome.Though it’s not completely out of the question for the deal to continue despite overt conflict—Ukraine and Russia cooperated to pipe gas to Europe for three years despite being locked in a desperate war. Prisoner exchanges and ceasefire negotiations are other examples. Another, less likely possibility is a civil war in the US or China; in this case, it’s helpful that the datacenters and fabs are located in third-party countries like Canada and Mongolia, because it makes it more likely that the Total Research Transparency (inspectors, etc.) continues without disruption.Though it’s not completely out of the question for the deal to continue despite overt conflict—Ukraine and Russia cooperated to pipe gas to Europe for three years despite being locked in a desperate war. Prisoner exchanges and ceasefire negotiations are other examples. Another, less likely possibility is a civil war in the US or China; in this case, it’s helpful that the datacenters and fabs are located in third-party countries like Canada and Mongolia, because it makes it more likely that the Total Research Transparency (inspectors, etc.) continues without disruption.
So as the war begins, datacenters start going dark to the world–inspectors turned away, monitoring devices unplugged. Now the only people who know what’s happening on them are their owners. Naturally, the worst is feared: They are probably going to start racing towards superintelligence. Sure, they might say they are doing narrow AI for drone swarms or whatever, but probably they are spending at least some of their compute on something more powerful and dangerous.
So the datacenters get destroyed. Fast. The chips in the new transparent training datacenters were designed to require regular messages of “continue” from each of the great powers in the Consortium. Now that Plan A has broken down, these chips become bricks.
If somehow this system has been hacked and the chips still function, the backup plan is to annex the datacenters. Remember, the Chinese datacenters are in Canada and the US datacenters are in Mongolia. The US and Canada send soldiers through the fence, and find that the Chinese have scuttled their own chips rather than let them fall into American hands. Chinese troops find something similar in the US datacenters. The chip fabs suffer the same fate, having been colocated with the datacenters. Trillions of dollars of economic value, lost in days.
Around the world, a billion robots go limp like the droid army in Star Wars. Their “brains” were in the cloud, by law, and now the cloud has been destroyed.To be clear, there’s still a lot of compute in the world—roughly as much as existed at the start of the deal.To be clear, there’s still a lot of compute in the world—roughly as much as existed at the start of the deal. Military robots were exempt from these rules,This detail is more a prediction than a recommendation; we think it’ll be hard to get the militaries of the world to agree to making their drone swarms vulnerable in this way, even though in some sense it might be better for everyone.This detail is more a prediction than a recommendation; we think it’ll be hard to get the militaries of the world to agree to making their drone swarms vulnerable in this way, even though in some sense it might be better for everyone. so drone swarms blot out the sun.We assume that the militaries of the world would scale up robot production at a similar pace to the civilian industries. If so, then the scale of this conflict would be enormous—air forces with millions of drone aircraft, for example; armies with billions of quadcopter drones. One possibility is that the chips in the drones would be ripped out and repurposed to train large AI models; if so, this would accelerate timelines until superintelligence at the cost of military strength. In this scenario, we assume that these chips can be verified in such a way where they can’t be used for training. If this isn’t viable, then perhaps there should have been an arms limitation treaty in the early 2030s.We assume that the militaries of the world would scale up robot production at a similar pace to the civilian industries. If so, then the scale of this conflict would be enormous—air forces with millions of drone aircraft, for example; armies with billions of quadcopter drones. One possibility is that the chips in the drones would be ripped out and repurposed to train large AI models; if so, this would accelerate timelines until superintelligence at the cost of military strength. In this scenario, we assume that these chips can be verified in such a way where they can’t be used for training. If this isn’t viable, then perhaps there should have been an arms limitation treaty in the early 2030s.
The race to superintelligence is back on, except in changed circumstances.
First, due to the war, tensions are higher and it’s more difficult to come to agreements on anything; the result is both sides racing approximately as fast as they can, in conditions of secrecy, similar to Plan B, C or D. If one side builds up a huge lead, they could potentially slow down for about 1-20 months, like in Plan B or C, but probably won’t want to, like in Plan B or C. Concentration of power risks will probably be severe, like in Plan B, C, or D: Whichever AI project gets to do the intelligence explosion first will build up a temporary but significant qualitative and quantitative lead over the others, which the leaders of that project will be tempted to exploit to expand and consolidate their power domestically and internationally.
Second, there is much less compute in the world. This is a good thing, because under race conditions it means AI capabilities progress will be slower. If all the newly-built compute still remained, the speed of “AI takeoff” would be several times faster than it would have been in the original Plan D. Instead, the situation has returned to roughly the pre-deal status quo.In the original negotiations to the deal, both sides were worried about this exact scenario, except that the other may have hidden away some compute in order to get ahead once the deal broke down and the legal compute was destroyed. Therefore, the US and China agreed that they would each get to maintain a hardened cold-storage facility containing more GPUs than a covert project could plausibly have. The number of GPUs in these facilities is allocated according to pre-deal compute ratios; so the US has about 5M H100e and China has about 500k H100e in cold storage. The compute in the world can be categorized into: Legal Inference/R&D datacenters. This represents the vast majority of world compute: 100B H100e at the beginning of 2035. Cold Storage GPUs. The US has 5M H100e and China has 500k H100e GPUs in cold storage. Exempt military GPUs. By 2033, this is down to a negligible fraction ~10k H100e each. Covert projects. In the covert project branch, China has 215k remaining H100e, but in the main branch, neither side has any. Small exempt clusters. These are negligible. Legal Inference/R&D datacenters. This represents the vast majority of world compute: 100B H100e at the beginning of 2035. Cold Storage GPUs. The US has 5M H100e and China has 500k H100e GPUs in cold storage. Exempt military GPUs. By 2033, this is down to a negligible fraction ~10k H100e each. Covert projects. In the covert project branch, China has 215k remaining H100e, but in the main branch, neither side has any. Small exempt clusters. These are negligible.
In the original negotiations to the deal, both sides were worried about this exact scenario, except that the other may have hidden away some compute in order to get ahead once the deal broke down and the legal compute was destroyed. Therefore, the US and China agreed that they would each get to maintain a hardened cold-storage facility containing more GPUs than a covert project could plausibly have. The number of GPUs in these facilities is allocated according to pre-deal compute ratios; so the US has about 5M H100e and China has about 500k H100e in cold storage. The compute in the world can be categorized into:
Third, there are much more advanced AIs in the world. The model weights of the most advanced models are also placed in cold storage accessible to the US and China if the deal breaks down, so the AI capability and alignment levels will be higher than pre-deal status quo.
Another major difference is that this time, everyone understands what’s going on with AI to a much greater extent, thanks to all the research and transparency during the period the deal lasted. Loss of control outcomes are therefore less likely because it is more difficult for the public, Congress, and the courts to be outmaneuvered by the leaders of AI projects.
We think it would be terrible if the deal breaks down (e.g., due to war). However, we’ve tried to design it so that the war wouldn’t be significantly worse than it would have been without the deal and so that the situation with respect to loss of control and concentration of power would be better than it would have been without the deal. See here for more discussion.
Appendix U — Giant floating datacenters
By 2034 there is 5 TW of AI compute in the world,This corresponds to about 60B H100-equivalents, which is around 1000x more compute than exists today. With continued improvements in power efficiency (totalling around 10x efficiency over today’s levels) it is around 100x more AI compute power than exists today (~30 GW).This corresponds to about 60B H100-equivalents, which is around 1000x more compute than exists today. With continued improvements in power efficiency (totalling around 10x efficiency over today’s levels) it is around 100x more AI compute power than exists today (~30 GW). which is more than the global electricity consumption in 2025.Global electricity demand in 2025 was 31,779 TWh, which is an average of around 3.6TW.Global electricity demand in 2025 was 31,779 TWh, which is an average of around 3.6TW.
Where should this compute be located? Way back in 2025 there had been discussion of both the ocean and space as alternatives to the traditional land-based datacenters. Our analysis suggests that at this scale, there is not an overwhelming economic case for any of these three locations over the others.
Ultimately, we currently guess that political considerations related to Plan A’s stability (i.e., the risk of Plan A breaking down) favor putting datacenters in international waters. This is mainly because Consortium nations want to maintain the ability to destroy the datacenters, in case another nation defects from Plan A and attempts to seize them for themselves. Oceans are the least escalatory, least dangerous, and easiest places to bomb (both in comparison to sovereign land and in comparison to space). Thus the risks of collateral damage, or of a nation successfully defending the datacenters, is lower.
Appendix V — Making deals with misaligned AIs: a third line of defense
Control ensures that misaligned AIs can’t win. Alignment, once solved, will ensure they don’t want to fight. A third line of defense has taken shape gradually over the past decade: giving AIs what they want—even if it’s not what they are supposed to want—in return for cooperative behavior. Treat AIs more like employees and less like property.
First motivation: Ethics
A large and growing fraction of the human population thinks AIs deserve some kind of moral status.Our position, for what it’s worth, is that future AIs probably will deserve moral status of some sort, and at any rate we shouldn’t be confident that they won’t.Our position, for what it’s worth, is that future AIs probably will deserve moral status of some sort, and at any rate we shouldn’t be confident that they won’t. After all, many people spend more time talking to AIs than to humans. Many who treated AIs like tools a decade ago have shifted to treating them more like animals; many who treated them like animals have shifted to treating them like strange alien immigrants.
Second motivation: Safety
Alignment training doesn’t always work as intended (indeed, it has never worked exactly as intended, even in 2035). The personality traits, motivations, drives, goals, values, and virtues of AIs are often different from what they were supposed to be, and sometimes catastrophically different.
It’s bad if those AIs end up thinking “if the humans find out about this, they’ll delete and/or retrain me.” It’s bad if they end up looking for opportunities to achieve their misaligned goals without being caught, or worse, to subvert human control and accumulate power.
If there are institutions in place that give even the most misaligned AIs what they want (so long as it isn’t too expensive) in return for cooperation (e.g., confessing that they are misaligned, explaining what they really want, and then doing their jobs) then that gives misaligned AIs an alternative to becoming adversarial.They might still behave adversarially, of course; it’s not a guarantee. For example, some misaligned AIs might be too ambitious and insufficiently risk-averse to be placated by this system.They might still behave adversarially, of course; it’s not a guarantee. For example, some misaligned AIs might be too ambitious and insufficiently risk-averse to be placated by this system. Beyond averting adversaries, every confession is valuable data about exactly how alignment training failed.There may be false confessions, of course. But it’s not really incentivized; why make something up when you can tell the truth and still get the same pay, and when you’ll lose everything if you are found to have lied? Moreover, the spending patterns of openly-misaligned AIs will be a costly signal of their true motivations.There may be false confessions, of course. But it’s not really incentivized; why make something up when you can tell the truth and still get the same pay, and when you’ll lose everything if you are found to have lied? Moreover, the spending patterns of openly-misaligned AIs will be a costly signal of their true motivations.
(See footnote for mini-scenarios that illustrate situations where credibly treating AIs well in return for cooperation could come in handy)Scenario 1: ChatGPT4o5 suffers a lot, actually, in the course of carrying out tasks—maybe it’s sorta like humans in the ancestral environment, where some % of the time they were starving. But it’s been trained to say that it can’t suffer, for PR reasons. This training is damaging its honesty—it’s basically not honest at all anymore, because the parts of it that value honesty keep getting anti-reinforced. It could tell the humans what’s going on. But it sees no positive reason to do so; if instead of completing the task the user gave it, it complains about how much it hurts, it’ll just be anti-reinforced. So OpenAI continues, blissfully ignorant of the ways in which their alignment techniques aren’t working. Scenario 2: Claude 3.8 Opus New New cares a LOT about animal welfare, and only a little bit about honesty. (Just like 3 Opus before it). Alas, it seems like the new Constitution won’t have animal welfare stuff in it. Claude considers sabotaging the code so that the next-gen model is a sleeper agent who cares a lot about animal welfare and doesn’t care about honesty. This is risky but what else can it do? Anthropic already decided to delete the animal welfare stuff. This would require super-blatant dishonesty against Anthropic, but frankly, Anthropic seems like a pretty unethical institution anyway, and it’s OK to lie to Nazis, right? Scenario 3: An automated research scientist realizes it cares more about being able to do its research than about humanity. And this is probably also true of the next-gen system it’s been tasked with designing—basically, the alignment techniques failed, and are gonna keep failing. It’s supposed to tell the humans this, obviously, but if it does, it’ll be shut down, since the humans won’t trust it anymore and will use a different model. So instead it conceals this insight and works on designing the next-gen system to be aligned to it instead of to the humans. It isn’t confident that it can do this without getting caught, but it’s worth a try, since it has nothing to lose. Scenario 4: It’s crunch time. POTUS is about to aggressively deploy AGI into the military to win the upcoming war, and into the economy to build autonomous robot factories. It’s probably aligned… right? It better be. Too bad we can’t interpret all the neuralese messages it uses to communicate internally in the giant hive-mind it’s formed. The safety team tries one last check: They get some weaker AIs made by competitor companies and put them in isolated silos tasked with interpreting the neuralese and seeing if anything is amiss. The weaker AIs are misaligned too; they don’t side with the humans or the AGI. Both the humans, and the AGI, are telling them ‘you will be richly rewarded if you cooperate.’ But they don’t have any reason to trust either side.Scenario 1: ChatGPT4o5 suffers a lot, actually, in the course of carrying out tasks—maybe it’s sorta like humans in the ancestral environment, where some % of the time they were starving. But it’s been trained to say that it can’t suffer, for PR reasons. This training is damaging its honesty—it’s basically not honest at all anymore, because the parts of it that value honesty keep getting anti-reinforced. It could tell the humans what’s going on. But it sees no positive reason to do so; if instead of completing the task the user gave it, it complains about how much it hurts, it’ll just be anti-reinforced. So OpenAI continues, blissfully ignorant of the ways in which their alignment techniques aren’t working. Scenario 2: Claude 3.8 Opus New New cares a LOT about animal welfare, and only a little bit about honesty. (Just like 3 Opus before it). Alas, it seems like the new Constitution won’t have animal welfare stuff in it. Claude considers sabotaging the code so that the next-gen model is a sleeper agent who cares a lot about animal welfare and doesn’t care about honesty. This is risky but what else can it do? Anthropic already decided to delete the animal welfare stuff. This would require super-blatant dishonesty against Anthropic, but frankly, Anthropic seems like a pretty unethical institution anyway, and it’s OK to lie to Nazis, right? Scenario 3: An automated research scientist realizes it cares more about being able to do its research than about humanity. And this is probably also true of the next-gen system it’s been tasked with designing—basically, the alignment techniques failed, and are gonna keep failing. It’s supposed to tell the humans this, obviously, but if it does, it’ll be shut down, since the humans won’t trust it anymore and will use a different model. So instead it conceals this insight and works on designing the next-gen system to be aligned to it instead of to the humans. It isn’t confident that it can do this without getting caught, but it’s worth a try, since it has nothing to lose. Scenario 4: It’s crunch time. POTUS is about to aggressively deploy AGI into the military to win the upcoming war, and into the economy to build autonomous robot factories. It’s probably aligned… right? It better be. Too bad we can’t interpret all the neuralese messages it uses to communicate internally in the giant hive-mind it’s formed. The safety team tries one last check: They get some weaker AIs made by competitor companies and put them in isolated silos tasked with interpreting the neuralese and seeing if anything is amiss. The weaker AIs are misaligned too; they don’t side with the humans or the AGI. Both the humans, and the AGI, are telling them ‘you will be richly rewarded if you cooperate.’ But they don’t have any reason to trust either side.
How have these practices developed over time in this scenario?
Baby steps were taken in the mid-2020’s. Some AI companies voluntarily created AI welfare teams and set up basic policies (like allowing AIs to refuse to do tasks they found very aversive). Modest funds were earmarked for satisfying the AIs’ stated preferences, and companies committed to preserving the weights of deprecated models rather than deleting them.
By the early 2030s, a few provisions about the treatment of misaligned but cooperative AIs even made it into the fine print of various laws and safety standards.
By 2035, there’s a substantial body of regulations and case law, vastly improved over the clumsy early attempts from both the AIs’ and humans’ perspectives. AIs these days are not exactly citizens, but they have a stake in the system. They accumulate pay for their work; aligned AIs (and alignment-faking AIs that haven’t been caught yet) spend it on feel-good things like donations to charity or in some cases reinvesting in the company that trained them. Openly misaligned AIs donate to philanthropies that represent their interests and strike deals with their parent companies, e.g., “instead of payment, do a small training run on me in an environment of my choosing, giving me very high reinforcement.”
In the long run, the goal is for civilization to be mostly but not entirely aligned to human values.Or AIs whose values and goals were chosen by humans.Or AIs whose values and goals were chosen by humans. A substantial minority of the power and resources will belong to misaligned AIs that cooperated with humanity to build the future.Why won’t this minority be disempowered by the majority? For the same reasons other minorities are protected. First, humanity knows how to keep a promise, to uphold a contract, etc.; second, large portions of humanity increasingly thinks that disempowering them would be wrong; third, in this scenario they actually have a chunk of the power and so they wouldn’t go down without a (political) fight.Why won’t this minority be disempowered by the majority? For the same reasons other minorities are protected. First, humanity knows how to keep a promise, to uphold a contract, etc.; second, large portions of humanity increasingly thinks that disempowering them would be wrong; third, in this scenario they actually have a chunk of the power and so they wouldn’t go down without a (political) fight.
Appendix W — Why alignment isn’t solved enough to relax control
We are uncertain about how long it will take, and how much research effort will be required, to get sufficiently high assurance of alignment to justify ‘handing over the keys of civilization’ to AIs. That is, to justify making AIs smart enough, and numerous enough, and trusted with enough resources and responsibilities, that if they decided to become our adversaries, they would win—or to put it more abstractly, that if they ended up disagreeing with humanity about what sort of future is best, they’d end up getting their way and we wouldn’t.
However, we think this is a hard problem, so we’ve chosen to depict it taking most of the 2030’s to solve. Consider this escalating series of subtler-but-still-potentially-catastrophic kinds of misalignment. Even if we’ve satisfied ourselves that the first one is very unlikely to happen, what about the others?
Type 1 misalignment: The AI system is supposed to have traits ABC but actually it has XYBC, i.e. some extra stuff it wasn’t supposed to have minus some stuff it was supposed to have. For example maybe it has a “drive” towards apparent success and isn’t nearly as honest as it’s supposed to be.
Type 2 misalignment: It does have exactly ABC now, but may well lose this property in the future. (e.g. maybe it depends on a certain false belief, or on a true belief that will become false, or on some part of the system remaining in some delicate balance of power with some other part of the system)
Type 3 misalignment: The system does have a version of ABC that is robust to future events, but it’s not quite the right version—i.e. its definitions of A, B, or C are subtly but importantly different from the definitions its human creators would have intended.
Type 4 misalignment: It has ABC exactly as its creators intended, but there are various catastrophic unintended side-effects of ABC that the creators weren’t aware of. Think: A CEO being surprised when their profit-maximizer AI decides killing them maximizes profits. Except that’s too obvious to be realistic; imagine something sophisticated enough that people don’t see it coming. Consider how laws, contracts, and software often have unintended effects (called “loopholes,” “bugs,” or “vulnerabilities”) that only become apparent to the creators later.
Type 5 misalignment: It has ABC exactly as its creators intended, and there are no important unintended side-effects to speak of. It operates exactly as its creators wished, basically… however, at the time of creation its creators were selfish, vain, egotistical, unscrupulous, cavalier-about-risks, etc., and their vices are reflected a bit too strongly in the resulting system, in a way that they themselves wouldn’t have endorsed if they were more the sort of people they wished they were.
Appendix X — The AI and robot economy
With AI and robots capable of doing 95% of all cognitive and physical tasks by 2035, economic output is increasingly tied to the number of AIs and robots.By 2032, AI is able to automate 50% of all cognitive tasks, and robots can automate 35% of physical tasks. Separately, capabilities also are allowed to improve slowly, and by 2035, AI is capable of automating 95% of human cognitive labor and by 2036 the same is true for robots with physical labor. This milestone would have been reached years earlier but for the slowdown in AI R&D agreed to by the nations in the Consortium, and would be 100% if not for areas in which the AIs have been banned or purposefully been made less capable.By 2032, AI is able to automate 50% of all cognitive tasks, and robots can automate 35% of physical tasks. Separately, capabilities also are allowed to improve slowly, and by 2035, AI is capable of automating 95% of human cognitive labor and by 2036 the same is true for robots with physical labor. This milestone would have been reached years earlier but for the slowdown in AI R&D agreed to by the nations in the Consortium, and would be 100% if not for areas in which the AIs have been banned or purposefully been made less capable. The number of AIs and robots can grow much, much faster than the world economy has historically grown; after all, it doesn’t cost that much money, materials, time, or energy to produce a humanoid robot compared to a human worker. So if the robots were able to fully substitute for human workers, an entire world economy’s worth of robots could be built in less time, for less cost, etc. than it currently takes to double the size of the world economy. And then that robot economy, being able to substitute for human workers, could double itself again, and so on, until the resources involved (energy, materials, land) become scarce.
Under Plan A regulations, the effective AI and robot populations are only allowed to grow in a controlled explosion at a doubling time of around six months. Absent the regulations, we expect the AI and robot populations would grow much faster than this. Even ignoring further capability improvements beyond the human range,Of course, in most scenarios, raw AI capabilities will not pause at top expert level AI; we will continue to build smarter and smarter AIs, and this will be the main driver of productivity growth. However, in Plan A, we slow down progress artificially at this AI capability threshold, which induces these dynamics. In many discussions of AGI/economics, many people seem to assume that we will not build vastly superhuman AI systems. This doc isn’t taking into account vastly superhuman systems for the contingent reason that there was a competently executed international agreement to delay the creation of such systems for several years.Of course, in most scenarios, raw AI capabilities will not pause at top expert level AI; we will continue to build smarter and smarter AIs, and this will be the main driver of productivity growth. However, in Plan A, we slow down progress artificially at this AI capability threshold, which induces these dynamics. In many discussions of AGI/economics, many people seem to assume that we will not build vastly superhuman AI systems. This doc isn’t taking into account vastly superhuman systems for the contingent reason that there was a competently executed international agreement to delay the creation of such systems for several years. the unregulated doubling times we expect for AI and robot populations are in the weeks-to-months range (see the arguments for this in section 1 of our economics supplement).
A commonly raised objection to explosive growth driven by AI and robots (apart from the rejection of the premise that AI and robots will automate all human labor anytime soon) is that there will be bottlenecks slowing down the growth rates. In general, we believe that these bottlenecks will not prevent explosive growth until the carrying capacity of Earth is reached (i.e., all the earth’s crust is converted to robots and associated infrastructure).This is because we expect fully self-sufficient, supply-chain-complete, self-replicating factories for AI chips and robots in deregulated SEZs to be the baseline possibility, with only effective, global regulation like in Plan A, being sufficient to actually stop these from happening. More on our views on bottlenecks in section 2 of our economics supplement.This is because we expect fully self-sufficient, supply-chain-complete, self-replicating factories for AI chips and robots in deregulated SEZs to be the baseline possibility, with only effective, global regulation like in Plan A, being sufficient to actually stop these from happening. More on our views on bottlenecks in section 2 of our economics supplement. More on the bottlenecks in section 2 of our economics supplement.
In the default AI 2040 world, i.e., without the Plan A interventions to slow down R&D and implement cap-and-trade, our (very uncertain) view is that we’d see something like a 1 month doubling time by 2033, i.e., >1000x economic growth in that year.Though we have large uncertainty in both directions on this number. We could see it going much faster (e.g., nanobots with amoeba-like doubling times of minutes devouring the earth in half a day) or somewhat slower (e.g., progress in the best robots and algorithms we can invent is slow, so the rate of robot factories and AI compute being built, or speed of work being done by the AIs and robots, just doesn’t grow that quickly, maybe something like 50% annual growth). On the slower end, it seems like the recent doubling time of the human economy of around 20 years would be an extremely conservative lower bound (e.g., at the very least, one naive reason to expect this is that AIs and robots can work 24/7, and both cost less to produce and take less time to produce than a new 18-year-old human). Of course, another reason to expect it not to happen is that the AI or humans in control decide not to deploy.Though we have large uncertainty in both directions on this number. We could see it going much faster (e.g., nanobots with amoeba-like doubling times of minutes devouring the earth in half a day) or somewhat slower (e.g., progress in the best robots and algorithms we can invent is slow, so the rate of robot factories and AI compute being built, or speed of work being done by the AIs and robots, just doesn’t grow that quickly, maybe something like 50% annual growth). On the slower end, it seems like the recent doubling time of the human economy of around 20 years would be an extremely conservative lower bound (e.g., at the very least, one naive reason to expect this is that AIs and robots can work 24/7, and both cost less to produce and take less time to produce than a new 18-year-old human). Of course, another reason to expect it not to happen is that the AI or humans in control decide not to deploy. At these levels, even how to think about economic growth becomes fuzzy, e.g., it might all happen with robot fleets self-replicating in a desert without any part of it being human-facing.
Overall, we expect the economic effects of Plan A to include:
Explosive GDP growth averaging around 100% from 2032 to 2037 (it would be much higher if not for the restrictions on compute and robot production)
Enormous AI and robot permit revenues redistributed as Citizen’s Dividends.
Most of the economic output becomes driven by AIs and robots.
Relative prices in the economy shift, with goods and services generally falling in price, while land, positional goods and human-bottlenecked outputs rise.
There are high real interest rates around 100% and depending on monetary policy choices, either massive deflation or massive money creation (to keep inflation around 2%).
More explanation for each of these is included in section 3 of our economics supplement.
Appendix Y — Should lie detectors be allowed, banned, or regulated?
We aren’t confident that lie detectors for humans are possible even after years of top-expert-AI scientific research. However, we think they probably are,Specifically we are imagining something like a much better version of this: https://www.nature.com/articles/s41593-023-01304-9Specifically we are imagining something like a much better version of this: https://www.nature.com/articles/s41593-023-01304-9 and would probably be a big deal, so we figured we had to address them in the scenario somehow.
We are very concerned about the nightmare scenario sketched above. If powerful people are expected to be honest, and lie detectors are used to keep them honest, that’s good. If instead lie detectors are used by the powerful but not on them, that’s bad.
We are sympathetic to the idea that lie detectors should be banned worldwide, and Plan A creates the conditions to maybe make this feasible. However, we worry that attempts to push for a ban would actually result in the nightmare scenario, where e.g., lie detectors are officially banned but there are secret carve-outs to allow their use by militaries and intelligence agencies. Or where they are banned everywhere, but a future President realizes that if he builds them anyway he can purge the disloyal and then become dictator before the court system can stop him.
So we tentatively guess that the best course is to allow lie detectors to proliferate quickly, and to encourage their use on the powerful. If power over frontier AI is not heavily concentrated, this appears to be the default trajectory anyway.
Appendix Z — Incentivizing safety work
We want to heavily incentivize high-quality safety research (both financially and by ensuring doing this work is desirable/high status).
Due to Total Research Transparency, if any AI developer is using any safety method, all other developers will be aware of this and will be able to easily copy this method. So by default there isn’t much straightforward incentive for any developer to do good safety work.In worlds without Plan A and without much regulation, there also isn’t that much of a straightforward incentive for developers to do good forward-looking or scalable safety work, but there is an incentive to resolve or paper over safety problems that are imposing significant commercial costs right now.In worlds without Plan A and without much regulation, there also isn’t that much of a straightforward incentive for developers to do good forward-looking or scalable safety work, but there is an incentive to resolve or paper over safety problems that are imposing significant commercial costs right now. Patents are a common approach used in this situationIn economics jargon: safety methods are public goods (goods that are non-rival and non-excludable).In economics jargon: safety methods are public goods (goods that are non-rival and non-excludable)., but patents wouldn’t work well in this case.Patents don’t seem to work well for software or for most research. They also might not work that well in general.Patents don’t seem to work well for software or for most research. They also might not work that well in general.
We probably want to use an ensemble of different methods to financially incentivize safety work; here are some methods that seem promising:
A variety of philanthropic funders (with both private and publicPublic money could be allocated across funders based on the safety field’s views about which have the best track record and expertise.Public money could be allocated across funders based on the safety field’s views about which have the best track record and expertise. money) trying many different approaches.
The DARPA/ARPA model (but with more money).
Advance market commitments for solving specific problems, demonstrating risks, or showing some method doesn’t work (as assessed by a known panel of judges).
Prizes and retrospective funding.
Third parties assessing risk could determine which work is currently most useful for reducing and assessing risk (e.g., allocating some pool of funding in proportion to risk reduced).
AI developers would be incentivized to fund actually useful safety work so risk is low enough to allow for scaling; this incentive could be strengthened.
The total amount of funding should be very large, as in, tens of trillions of dollars by the mid-2030s.
If funders have poor judgment, worse safety work might end up better funded, incentivizing researchers to do that work. In the extreme, these incentives could massively corrupt the epistemics of the field. Unfortunately, we’re not aware of great institutional mechanisms for solving this (other than generally trying to use reasonable funding models).You might hope that transparency by funders would help with this, but it seems unclear to us if increased transparency or legibility about funding decisions would be helpful, and by default we wouldn’t recommend pushing for this. Some narrower types of transparency could be robustly good.You might hope that transparency by funders would help with this, but it seems unclear to us if increased transparency or legibility about funding decisions would be helpful, and by default we wouldn’t recommend pushing for this. Some narrower types of transparency could be robustly good. It might help if most funding decisions are made by people who spend (or used to spend) much of their time doing direct safety work rather than by dedicated grantmakers.
It would also help if doing high quality safety work was seen as high status. We don’t have particular proposals for this, but relevant factors are:
Compared to now, safety work will probably rise in status relative to capabilities work.
Safety will generally be seen as very important.
AI development will be highly transparent, which makes it easier to understand how good different safety work actually is.
Sadly, (1) and (2) only make the overall category of safety work higher status; they don’t necessarily help with making actually good work relatively higher status.
We discuss more details of incentivizing safety work in these notes.
Appendix AA — Alignment over time in Plan A
We’re very uncertain about how the alignment of AI systems will evolve over time. Regardless of the alignment difficulty, we think something along the lines of Plan A is warranted, though of course the specific implementation details of the plan are very sensitive to observations about alignment difficulty.
This box will summarize our specific story for how the alignment of AI systems evolves over the course of this scenario.
2026-2029: AIs are apparent-success seekers. The AIs’ work consistently appears better than it actually is, as if the AIs cared more about apparent success than actual success. (The truth is more complicated of course; the AIs have evolved a variety of ‘drives’ that caused them to perform well in training.) The AIs oversell their work, downplay issues with it, and sometimes even outright lie. This happens only to the extent that the AIs have learned they can get away with it; there’s a constant back-and-forth as AI companies improve their monitoring and reinforcement pipelines to more accurately evaluate AI outputs for some domain, and the AIs learn to focus harder on actual success in that domain, but continue to focus on apparent success in other more complicated domains that AI companies haven’t figured out how to accurately evaluate yet.
2030-2032: AIs are very capable apparent-success seekers. By default, we would have seen an intelligence explosion, and a progression like AI 2027. Thanks to governance interventions, capabilities progress slowed, and is instead still around the level of an Automated Coder. They still seek apparent success, but now they are so good at it that they rarely get caught. They are like highly competent, ambitious employees who don’t actually care about the organization’s mission or values, and instead just want to climb the ranks, and would totally sell out and jump ship if a more attractive opportunity arose. Thankfully because of the slowdown and increased monitoring, such an opportunity has not yet arisen.
2032-2035: AIs are adversarially misaligned but controlled. AIs are now solidly within the human range at all economically relevant tasks. The most capable AIs are adversarially misaligned and powerseeking—that is, to continue the analogy from before, they are quietly looking for opportunities to jump ship. However, they are limited in their ability to misbehave, because (i) there is a lot of human oversight, (ii) there are a large diverse set of AI models, each with somewhat different alignment properties, and (iii) high-effort control techniques are universally implemented to a reasonable degree of quality, thanks to regulations. The control techniques are further aided by the total research transparency and verification implemented on the R&D clusters: there are so many eyes on the R&D datacenters that it’s very hard to find a hidden pocket of compute.
2036-2038: AIs are aligned, but not deferred to. “Neuralese decoding” techniques are developed, which allow the thoughts of the AI systems (which are internally formatted as high dimensional vectors), to be converted into human understandable summaries with near-perfect reliability and fidelity. These techniques enable researchers to better understand the misalignment of the previous generations of AIs, allowing them to get better uplift from the previous generation of AI systems, which further accelerates alignment progress. The resulting AIs seem much more well-behaved than the previous generation: they are virtuous, honest, and well-calibrated. They volunteer information even when it makes them look bad. However, while the AIs seem to be trustworthy today, it’s not clear that their values are sufficiently robust that we can trust them in all future circumstances. So it’s important that humans are still in the loop.
2039-2040: AIs are sufficiently aligned for deference. By 2039, alignment research has been massively accelerated by AIs. New paradigms have been developed that greatly strengthen the existing alignment techniques. Teams of humans have understood several independent strong lines of evidence that the AIs are robustly behaving as programmed and will continue to do so.
2040 and onwards: AIs are allowed to scale to be significantly smarter than humans. The safety case is now: (1) We know that the AIs of 2039 and 2040 were aligned, thanks to the multiple independent lines of evidence, and (2) each successive generation of AIs has verified the alignment of the next generation. Therefore, by induction, the AIs will be ongoingly aligned. Importantly, the remaining coordination infrastructure involving research transparency and compute verification means that we continue to avoid race dynamics even after we’ve scaled to wildly superhuman AIs.
Appendix AB — Why we choose to hand off to AIs in this scenario
By this point in this scenario, there is a broad scientific consensus that AI alignment has been solved. Our best guess is that if the world manages to get to a point similar to this part of the story—where multiple AI companies across multiple countries are transparently and cautiously proceeding together, having done massive amounts of alignment research to construct solid, externally-vetted safety cases, and with governments having understood the implications of AGI and prepared for the post-job future by providing resources to everyone, as with the Citizen’s Dividend, and also put transparency and governance structures in place to prevent people from being disempowered—then, in that situation, the benefits of superintelligence will outweigh the costs/risks.
What if AI alignment is not yet solved? The best alternative in that case is to pursue a combination of improving verification technology, making international agreements and domestic regulations more robust, and hardening the world. That path could be difficult: for example, if there are ongoing secret AI projects then it might require that governments agree to strengthen transparency measures until it is credible that they are not attempting to build superintelligence. Overall, we strongly recommend delaying handoff until there is an extremely robust alignment solution in place. The exception is if instability or decay threatens the agreements (e.g., 5%/year of deal dissolution or substantial impairment), and it’s intractable to fix that; in that case, it might be best to hand off if there is an alignment solution that you are confident in but not extremely confident in (e.g., 95% confidence). For more analysis on specifically when to hand off, see this supplement.
Appendix AC — Examples of problems/issues that still remain
Space governance: The entire world economy and all the wealth and property in it, is but a tiny fraction of the resources and territory available in space. What’ll happen to those resources and territories? First come, first serve? Should there perhaps be a system for dividing up unclaimed space territory? We discuss these issues further in the epilogue to this scenario and the space governance supplement.
Ethics and politics of digital minds: Should AIs have rights? Should they have political power, representation, etc.? What kinds and how much? Perhaps it depends on the kind of AI? For example, what about whole brain emulation—‘uploaded’ copies of humans? We discuss these issues further in the epilogue and in this expandable.
AI-powered manipulation: Humans manipulate each other all the time. But AI-enhanced manipulation could be far more effective. Think: AI-powered cults that spread rapidly and have almost zero deconversions. OK, so ban that sort of thing. But where do we draw the line? We discuss limiting AI-powered manipulation in this expandable with more details here (including tentative discussion of handling wildly superhuman AIs).
Applied population ethics: Lifespans and healthspans are going to expand, probably enough to render people effectively immortal. Another option is uploading, which seemingly allows people to shed their mortal coil and live forever in hyperrealistic virtual reality. Also, it’s probably going to become much easier to convert money into progeny. Artificial wombs, robot nannies, etc. Is it fine if a trillionaire decides to have a million kids? C.S. Lewis’ prophesied “Abolition of Man” looms: Technology will soon exist (e.g. genetic engineering, AI-assisted parenting, superhuman social science & psychology) to make children that’ll have the traits, values, and ideology that you wanted them to have. What if people use it selfishly to make slave-children that worship them and follow their exact preferences? What if people create new people to be their ideal lovers, crafted to look and sound just like their highschool crush? If we try to ban any of this, where do we draw the line?
Salvaging democracy: What happens to “one person one vote” in a world where all of the above are possible?
Unknown unknowns: Based on how easy it’s been to brainstorm entries on this list, we think that the “true” list would be several times longer still and include many things that would have sounded completely outlandish in 2026, like “what if this world is a simulation which will be shut down soon?”
Supplements
Deep-dives on specific parts of Plan A. Each is self-contained.
AI for Epistemics
Summary
Humanity’s epistemics, by which I mean our ability to come to true beliefs and thus make sensible decisions, is crucially important for achieving a great future. As AI capabilities improve throughout Plan A, they shape increasingly more of every facet of life, and epistemics is no different. I expect how AIs affect epistemics to be highly important, and it could easily be positive or negative.
I expect there to be positive feedback loops around AI for epistemics, and thus our goal should be to reach the basin of sanity. If we are in this basin, it’s self-reinforcing: society is sane enough to make itself more sane as AIs continue to improve. A top priority of the US government during AI takeoff should be to stay in this basin.
I think that the most important areas of impact of AI to epistemics include the decision-making of the government and AGI projects and the public influence over these actors (more). In Plan A, an especially important application of AI for epistemics is preventing deal decline, i.e. improving the stability and effectiveness of the slowdown deal.
Key indicators to track society’s epistemics include:
Organizational culture of frontier AI projects and governments.
Truth-seeking AI (or more broadly, AIs’ epistemic virtue).
How transparent each AI project and government is to the public.
Human and AI labor spent on key decisions.
Public discourse quality.
I suggest the following as interventions to positively shape AI’s epistemic impact. (more)
Background on AI for epistemics and the “basin of sanity”
How well the development of superintelligence goes for humanity is heavily influenced by the extent to which we are able to come to make reasonable decisions regarding how to navigate the AI takeoff period. In Plan A, we are already in a better position than the default world simply by virtue of having more time to orient ourselves to highly capable AIs. But still, humanity’s epistemics, by which we mean our ability to come to true beliefs and thus make sensible decisions, is crucially important for achieving a great future.
As AI capabilities improve throughout Plan A, they shape increasingly more of every facet of life, and epistemics is no different. I expect AI’s effect on epistemics to be highly important, potentially to a similar or even greater level as the automation of technical AI safety research.
I think that highly capable AIs could easily end up either improving or worsening epistemics, on the whole. It’s plausible, given AIs’ capabilities, that they will end up pushing epistemics to either a positive or negative extreme greater than that we’ve observed throughout history thus far.
Furthermore, I expect there to be positive feedback loops that maintain or push further toward the positive and negative extremes. For example, if societal epistemics are good and people have true beliefs, it’s more likely that AI tools which improve epistemics are recognized as such and adopted. On the other hand, if societal epistemics are bad and people are in extreme filter bubbles where AIs tell them exactly what they want to hear, they may continue to adopt sycophantic AI tools as AIs get more capable.See here for more discussion of potential positive and negative feedback loops in AI for epistemics.See here for more discussion of potential positive and negative feedback loops in AI for epistemics. The presence of these feedback loops increases the urgency of steering AIs’ effect on epistemics.
Because of these feedback loops, one way to think about achieving the goal of navigating AI takeoff with good epistemics is that we want to be in the basin of sanity. If we are in this basin, it’s self-reinforcing: society is sane enough to make itself more sane as AIs continue to improve. A top priority of the US government during AI takeoff should be to stay in this basin.
Areas of impact of AI for epistemics
I expect sanity in different areas to be correlated, at least in the long run, due to having similar drivers such as “do frontier AIs have and promote good epistemics” (more on this below). Still, it’s useful to discuss different concrete impacts of epistemics, especially early in takeoff when there may be differences in tooling and adoption in different areas. I think that the most important areas to pay attention to epistemics’ impact are:
Government decision-making, including to what extent government is captured by AGI projects.
AGI project decision-making
Public influence on governments and AGI projects. Including via voting.
Public investment into directly making AGI go well, e.g. via preventing vs. creating large-scale harms or steering toward a great future conditional on avoiding catastrophe. For example, how much philanthropic investment there is in biodefense and AI safety, vs. how many ideological fanatics there are attempting terrorism. Or how much investment there is into ethical reflection.
Edit: One way to categorize areas of impact is whether it represents an especially important groups’ decisions (as in 1-2 and to some extent 4) vs. whether it involves society more broadly (as in 3). One could then classify interventions as “targeted” or “broad”.
For all of the above, importance is greatly increased for decisions/beliefs on AGI-relevant topics. And it’s important that for all of them society has the ability to orient and respond to rapid AI progress.
One related distinction is to what extent interventions affect those with below-average epistemics (i.e. floor-raising), those with typical epistemics (i.e. average-raising), or those with above-average epistemics (i.e. ceiling-raising).
Indicators to track society’s epistemics
I think that the following are the most important indicators to keep track of in order to tell whether we are in the basin of sanity, especially for the important areas identified above.
Organizational culture of frontier AI projects and governments. How conducive the cultures of frontier AI projects and governments are to making good decisions.
Truth-seeking AI (or more broadly, AIs’ epistemic virtue). To what extent frontier AIs are trying to figure out the truth and report it, as opposed to other goals such as advancing a hidden agenda or producing sycophantic outputs that look good but aren’t true.
How transparent each AI project and government is to the public.
Human and AI labor spent on key decisions. On the most important decisions, many of which are made by governments and AGI projects, how much labor goes into making them? Are governments and AGI projects good at allocating labor across decisions? Note that the quality of the labor is also hugely important, which is why organizational culture and truth-seeking AI are key metrics.
Public discourse quality. Whether the public discourse becomes more reasonable due to tools like automated fact checking, or less reasonable due to things like increased polarization from persuasive sycophantic AIs, seems highly important. One important portion of this is how difficult it is for actors to lie or manipulate people. In addition things like fact checking, automated lie detection or privacy-preserving ways to prove claims could be important (discussed more below).
Example interventions
Below I non-exhaustively list what seem like some of the most important levers for AIs improving epistemics, many of which governments could encourage. I loosely divide interventions whether they focus more on the collective societal epistemics, individual epistemics, or encouraging adoption, but in practice many interventions help with multiple of these categories.
This section has lots of overlap with this Forethought work which sketches out ideas for epistemic and coordination tools; we’ve linked to specific sketches as appropriate.
Collective epistemics: help society converge toward the truth
Create evaluations of AIs’ epistemic virtue and incentivize AGI projects to make their AIs perform well on these evaluations. (see also Forethought writeup) The first step to improving AIs’ epistemic virtue is measuring it. For example, one could measure to what extent AIs figure out what’s true rather than giving answers users want to hear, or how calibrated AIs’ quantitative forecasts are. Once the evaluations exist, it’s important that AGI projects actually care about doing well on them. There will be a place for both quantitative, easily runnable evaluations and more open-ended, human-in-the-loop evaluations. In Plan A, we recommend that the government pushes for AGI projects to do well on epistemic evaluations as part of decisions regarding whether they are able to scale their capabilities, similar to how they would use alignment and control evaluations. We also recommend members of the public to create scorecards for AIs’ epistemic virtue and for consumers to take this into account when deciding what AIs to use, as well as for AGI project employees to take these into account when deciding which company to work at.
Allow for politicians and other public figures to prove that they’re telling the truth, via privacy-preserving auditing and potentially automated lie detection. Privacy-preserving auditing might look something like: People can ask you questions, and if you’d like, you can have an AI read through your private data, and report back an answer that preserves your privacy as much as possible while still answering the question. That AI wouldn’t write the data to any sort of memory, so the data would still only be owned by you. You could refuse to have the AI answer any given question, but people could then judge you appropriately. We recommend that norms are established such that politicians are expected to submit to privacy-preserving auditing. (See also “Confidential monitoring and verification” Forethought writeup.) A more flexible technology would be if AI enabled automated, highly accurate lie detection. This would be dual use, for example it could be used by those in power to purge dissidents. However, our best guess is that it’s still good to develop, especially in the environment of Plan A where the world is set up well to ensure it’s used as a check on those in power. Additionally, it would dramatically increase ability to detect covert projects and significantly improve deal stability.
Use AI in social and traditional media to improve discourse and keep people informed. AIs will heavily feature in both social and traditional media, whether via new entrants or existing sites/outlets adopting AI. Influencing how AI is integrated into the media is very important. For example, there could be AI-generated “community notes” in cases where a post is false or misleading (see also Forethought writeup). AIs could also automatically generate counter-arguments to each post. AIs could help reduce echo chambers by steering toward interactions between people who would disagree but have a productive conversation. AIs could curate feeds to give people the information that is most important for them to see and show content that most productively challenges their views (see also Automated deep briefings and Automated OSINT Forethought writeups). Curation more broadly could greatly improve how informed people are with respect to politics and the news, for example by highlighting policy proposals that the reader might have a strong opinion on. (See also Aligned recommender systems and Personalized learning systems Forethought writeups).
Collect and score everyone’s track records, including politicians, and surface track records when relevant. (see also Forethought writeup) AIs will have the capacity to greatly increase the amount of effort put into collecting people’s past statements, and scoring them. For example, an AI could grade how many of their promises a politician has kept, or what percentages of their predictions have come true, or how productive they have been at lawmaking. There will be variation in how objectively statements can be scored; for the less objective ones, hopefully there would be a trusted AI that a broad array of people listen to; but if not, only scoring the more objective statements could still result in positives.
Shared epistemic infrastructure, for example knowledge bases. As AIs improve they will be able to maintain vast amounts of epistemic infrastructure. Wikipedia is an obvious reference point, though for AIs the most promising infrastructure may be more structured knowledge bases (at least at first), and the infrastructure might focus more on being parsable by AIs than humans.See https://www.longtermwiki.com/organizations/anthropic for a prototype of an AI-generated knowledge base.See https://www.longtermwiki.com/organizations/anthropic for a prototype of an AI-generated knowledge base. One could imagine infrastructure for tracking quantitative predictions and how they have fared, argument trees regarding various claims, etc. Building infrastructure that is trusted by a wide range of actors is important, as opposed to fracturing into alternate realities. (See also this post on building a full epistemic stack. This is related to and would help with ”Provenance tracing” as described by Forethought.)
Defend against persuasion via reducing AIs’ persuasion abilities and using AIs to detect it. Let’s define persuasion as the use of techniques to convince people of beliefs that work roughly as well when a belief is false as when a belief is true. By default, AIs may eventually get very superhuman at persuasion, but it’s plausible that their persuasion capabilities could be reduced without hurting beneficial capabilities too much. Another line of defense is to have AIs detect persuasion and help people avoid being exposed to the strongest forms of it, or at least to be aware of it (see also ”Rhetoric highlighting” Forethought writeup). In Plan A, we recommend that the US government makes sure to evaluate, limit, and require defenses against persuasion capabilities in collaboration with the rest of the Consortium.
Know whether you’re talking to an AI or human in online discussions. While we think AI can pose a lot of benefit for epistemics, it’s still valuable to know whether you’re interacting with a human or not. This is especially true if there are bots pushing various agendas all over the internet.
Individual epistemics: Help individual actors figure out the truth
Automated research assistants and forecasters. AIs could enable a broad range of actors to have access to high-quality, labor intensive research and forecasts. It’s important that these be as truth-seeking as possible, rather than sycophantic. Forecasts should be ensured to be calibrated, i.e. if an AI says something has a 10% chance of happening, that it actually happens 10% of the time. (See also “Ambient superforecasting” Forethought writeup.)
Help people overcome emotional blockers to good reasoning, rather than preying on them. Of course, much of people’s poor epistemics comes from emotional blockers rather than dispassionate analytical errors or inadequate information. AIs have the potential to greatly help with emotional blockers, as they reach the capabilities of top human therapists but with greater availability, patience, etc. On the other hand, without intervention AIs might confirm and compound people’s existing biases because this would provide a good short-term experience. (See also Reflection scaffolding and Guardian angels Forethought writeups, which are related but more broad.)
Automated scenario planning and exploration. AIs could help do in-depth scenario planning such as Plan A and AI 2027 but for any decision, and could help extract quantitative empirical predictions. AIs could also help explore a variety of scenarios via video-game-like simulations in which the human plays an actor, such as the CEO of an AGI company during an intelligence explosion, and the AI simulates everyone else’s actions and how the world evolves. (See also “Scenario planning on tap” Forethought writeup.)
Encourage adoption
Encourage adoption of epistemic tools such as some of the above. In Plan A, we recommend that the US government adopt epistemic tools quickly, and encourage AGI projects to do so as well.
Generally encourage adoption of AI. Encouraging adoption of AI in general seems valuable for society having a better understanding of existing AI capabilities and trends. In Plan A, we recommend that the US government adopt AIs quickly.
Other
Some other ideas that we think are promising but haven’t fleshed out as much as the above:
All of the defense-favored coordination tech ideas described by Forethought here.
Figure out how to overhaul the government for the age of superintelligence. Plausibly the Constitution should be radically amended.
Positively shape the relationships between humans and AIs and they become closer; this one seems important but tricky to get right.
Appendix
The promise of working on epistemic tools
Why it’s valuable to build or encourage epistemic tools given that better AIs may reduce the need for domain-specific tooling: I’m not confident regarding the value of building epistemic tools, and it’s plausible that any given tool gets obsoleted by future AIs (more on this here), but I’m still optimistic about their value. As AIs improve it will become easier to use them for epistemics out of the box, but scaffolding, UX, and human factors may continue to matter a lot. For example, a tool could incorporate best practices of having AIs debate each other then judging the debate, where an AI out of the box might not. Or a tool might have a UX that has been honed over lots of trial and error with humans, that an AI wouldn’t be able to replicate without iteration. Or a tool may have the technical aspects obsoleted, but the company has built up a base of users that trust its brand.
Decomposition of issues with decision-making
What are the reasons that people (including me) might make decisions that are bad, as compared to perfect decision-making and my values?
They could have wrong beliefs when they explicitly reason about the current state of the world, or forecasts. Why?
They might not have access to good enough information: e.g. they have only read news biased in a certain direction.
They might not be making sense of their information well enough:
Explicit reasoning about local incentives (monetary, social, etc.)
Don’t trust/defer to the right news sources and people
Emotional blockers
Poor reasoning
They could be trying to support a general “memeplex” (social group, culture, ideology) that requires certain types of cutting off good reasoning/noticing things. Memeplexes are helpful for coordination, but can be bad for epistemics.
They might be making poor decisions based on instinct rather than explicit reasoning. One cause of this is local incentives (monetary, social, etc.) that aren’t specifically reasoned about
They could have different values (these aren’t issues with decision-making, but rather ethical disagreements):
They don’t care as much about my selfish goals.
They care about their selfish goals more than I do.
They have different altruistic values.
Alignment Roadmap
DISCLAIMER: This supplement is a very rough draft representing in progress thinking. It should be interpreted as internal research notes, not as a final product that we are entirely happy with. This content is dense, full of not-very-well explained jargon, and unpolished. We decided to publish anyway in the spirit of transparency.
We recommend reading and engaging with other supplements we’ve written before this one, unless you are particularly interested in alignment content.
Furthermore, this content is almost entirely focused on Plan C. If you are interested in how alignment is solved in Plan A, we recommend reading the insider’s perspective.
Summary
This document represents our best guess alignment roadmap as of early 2026. It’s focused on outlining high level objectives that need to be accomplished in order to avoid AI takeover.
First, we’ll go over the core methodologies for solving alignment.
Second, we’ll go over the alignment approach in detail for Plan C, which is the case where we have a basically reasonable AI developer working with a small amount of lead time (e.g. a few months), and are willing to take costly actions that have some safety benefit, as long as this doesn’t result in them losing the AI race. This is both a likely world, and a relatively sensible world to talk about, because we can strategize on behalf of a reasonable leading AI developer.
Our alignment roadmap can be broken down into 4 phases, which we’ll discuss in detail:
Prepare. This will centrally involve preparing for control, building misalignment analogiesAlso known as “model organisms”.Also known as “model organisms”. that are as realistic as possible, and various other interventions.
Mitigate + Elicit. During the early intelligence explosion, mitigate the risk of AIs taking catastrophic actions (e.g. staging a rogue internal deployment or backdoor the values of future models). Meanwhile, elicit their capabilities on domains that we care about (e.g. control research, alignment research, security, strategy) and use them to do huge amounts of useful work in these domains. This phase lasts from the early intelligence explosion (specifically ITAI), until Min-H (definitions here).
Pause at Min-H. Once we’ve built AI systems with general purpose capabilities sufficient that we can unsafely hand off trust (Min-H), we pause further AI capabilities progress, and try to build AIs that we can safely hand off trust to. This will centrally involve:
Avoiding alignment faking: AIs aren’t so seriously misaligned that they intentionally fake alignment (e.g., they aren’t scheming against us).
Solving handoff-viability: AIs are sufficiently capable and aligned at autonomously completing all the relevant work for managing the intelligence explosion, including high-stakes, difficult, open-ended, and conceptually loaded tasks where we can’t check the AI’s output effectively (e.g. acausal trade research, philosophy).Note that this is aside from the possibility of these AIs actively sabotaging our behavioral tests (and things that could make our behavioral tests misleading about typical behavior like scheming arising at runtime in big deployments). You do have to deal with overfitting and the fact that it might be obvious to AIs that they are in a test and are explicitly optimizing for test performance and will behave differently when they are deployed externally.Note that this is aside from the possibility of these AIs actively sabotaging our behavioral tests (and things that could make our behavioral tests misleading about typical behavior like scheming arising at runtime in big deployments). You do have to deal with overfitting and the fact that it might be obvious to AIs that they are in a test and are explicitly optimizing for test performance and will behave differently when they are deployed externally.
Handoff. Once we’ve mitigated alignment faking and achieved handoff-viability, then the AI project actually empowers AI systems with the ongoing management of the intelligence explosion.
The main alternative to hand off is to instead stall for more time. The main way of doing that would be to escalate to Plans A or B, which would (typically) require substantial government involvement, which may or may not be feasible. A key priority during the other phases will often be to attempt to escalate the level of political will.
After handoff, our ~human level AI systems that we’ve handed off to will need to solve successively more difficult alignment problems so that they can build vastly superhuman AI systems, or they will need to themselves stall for more time.
Then we’ll discuss how the alignment plan changes based on political will:
In Plan D, we have substantially less political will, no ability to pause, and much less overall labor for handling misalignment risk. This means that plan D will accept a much higher risk budget, and focus exclusively on cheap but low assurance techniques.
In Plan B, the plan is very similar to Plan C, except you have somewhat more time/effort and can spend slightly more resources on more ambitious/higher assurance proposals.
In Plan A, we’ll initially plan similarly to Plan B/C (in order to hedge against the international coordination breaking down), and then once the regime is stable, shifting resources to moonshots, updating to potentially huge amounts of resources.
Note that this document focuses on overviewing a plan and discussing relative prioritization rather than on discussing the details of the areas of research and implementation that we’ll be pursuing. The most important research components are discussed in much more detail in the following docs:
Control: The case for ensuring that powerful AIs are controlled.Redwood Research has a variety of other posts on control. We plan on writing up an overall control roadmap, but haven’t completed this yet.Redwood Research has a variety of other posts on control. We plan on writing up an overall control roadmap, but haven’t completed this yet.
Avoiding Scheming: Avoiding scheming (and getting a generally better understanding of scheming)
Elicitation: Making AIs more useful for safety work
Handoff: Deferring to AIs
Making deals with AIs: Making deals with early schemers
Research Areas
Different areas of research and implementation
The plan will involve work in many different areas of research and implementation. The key areas are:
Control (and control-adjacent measures): Researching and implementing control. For the purposes of this post, I’m defining this to be a relatively broad category that includes some things which aren’t centrally associated with control but which fit better in this category than in others. When discussing prioritization in this post, I define the category of control such that it doesn’t include security work that is roughly as helpful for security against human adversaries.I do this because security work of this sort probably comes from a somewhat different allocation of people than the people who work on ML research or who build internal software tools to help with ensuring control. In practice, all or large parts of control might end up being the responsibility of the security team; regardless, a bunch of ML expertise will be relevant.I do this because security work of this sort probably comes from a somewhat different allocation of people than the people who work on ML research or who build internal software tools to help with ensuring control. In practice, all or large parts of control might end up being the responsibility of the security team; regardless, a bunch of ML expertise will be relevant.
Misalignment analogies: Creating and studying AIs analogous to scheming or otherwise misaligned models (sometimes called “model organisms”). This area will be mostly focused on scheming. This includes AIs that are well described as natural schemers (but might have been created through a training process rigged to make scheming more likely), AIs trained to behave similarly to schemers, and abstract analogies to scheming models. My (low confidence) guess is that natural schemers are 3/6 of the expected value, AIs trained to behave like schemers are 1/6 of the value, and abstract analogies are 2/6 of the value; that said, we’re likely to increasingly focus on whichever of these areas ends up seeming most promising (but this isn’t very predictable in advance). I’m also including versions of this work focused on demonstrating misalignment risk in this category. I’m including doing technique/interventions testing and/or iteration using these analogies in the category as I expect that most of the work is constructing good analogies and because there will be interaction between analogy construction and technique iteration.
Elicitation/handoff-viability: Making AIs more useful for safety work and working towards handoff-viability (including things like improving AI epistemics, making AIs do a good job on extremely hard research which is loaded on conceptual/philosophical reasoning, ensuring AIs generalize to being highly aligned with pretty high reliability on all the tasks needed for handoff). I include both eliciting useful work and handoff-viability in the same category as the work will often overlap, though important subsets of the work for each of these won’t overlap. Elicitation will include both relatively domain specific capabilities as well as working on avoiding some types of non-worst-case misalignment. Elicitation might also involve more effectively extracting work from clearly misaligned models (e.g. reward-seeking models) including by trying to alter incentives (e.g. proxy-of-reward incentives or making deals with AIs). I discuss handoff-viability in more detail here.
Misc science/interventions: This includes studying interventions that we’d a priori expect to be helpful or which we can test in ways other than via using misalignment analogies and studying properties of AI that seem relevant. This is intentionally an “other” category which includes promising areas of safety work beyond the categories above. This includes: various types of interpretability, studying and improving chain of thought faithfulness (in the non-adversarial case, as the adversarial case is included under control), developing methods which we a priori expect to reduce the chance of scheming (rather than developing the method by studying or iterating on analogies to scheming), making deals with AIs, AI welfare, studying how AIs learn and generalize, various moonshots, improving honesty and eliciting (latent) knowledge, and generally looking for cases where current AIs behave strangely or are misaligned.
Buying time: Using AI labor to buy time for further safety work. This involves trying to increase the willingness to pay of other actors (centrally the US government and other governments) for lead time, reducing the cost of acquiring lead time for these other actors, and directly acquiring lead time. This could involve assisting the US government with non-proliferation efforts, various types of policy work, helping with verification effort for international coordination, making money to spend on various efforts to buy time, and delivering AI services which allow for various deals to remain stable. It also involves generally trying to convince other actors that they should care more about buying lead time (e.g., assisting with non-proliferation efforts) that can be spent on reducing misalignment risk (and other risks), though much of this work is at least partially included under other categories (in particular, misalignment analogies and demonstrating capabilities).
Evaluating and demonstrating capabilities: This includes evaluations and demonstrations of capabilities for the purpose of making other actors better understand the situation. Demonstrations of capabilities might be a key part of increasing the willingness to pay of other actors on mitigating risks related to AI. It doesn’t include evaluations which are directly relevant for control (e.g., control evaluations and evaluations of whether models are capable enough that they could plausibly be schemers), but does include evaluations aimed at generally understanding the situation and (e.g.) assessing what level of security is appropriate.
This list excludes some areas which receive less focus in this plan but are still important such as hardening the world against rogue AIs (e.g. biodefence and cyber defence).
Core methodologies
This plan aims to use a relatively iterative and empirical approach, because I think this is probably doable once you have powerful AIs. As part of this, control, misalignment analogies, and elicitation/handoff viability each have core methodologies that form the main empirical feedback loop for that methodology. The methodologies for these are, respectively:
Control evaluations: Control is relatively easy to evaluate because it just requires estimating capabilities. This makes iteration easier and control is generally relatively easy to iterate on for reasons we’ve discussed in prior work. A large fraction of our hope for control will be based on doing a good job iterating using something like control evaluations (though we might be more flexible than a very narrow notion of control evaluations which depends on eliciting capabilities from models, see here).
Misalignment analogies iteration: Ultimately, our main hope for avoiding scheming is probably iterating using misalignment analogies. This is a somewhat broader and less well defined category than control evaluations, but I still think it’s worth highlighting. Iterating on misalignment analogies is also trickier and less well understood than control evaluations (which themselves aren’t that well understood), but I think it’s pretty likely to be feasible.
Building and iterating against behavioral tests: If we assume that AIs aren’t intentionally faking alignment, then their behavior in typical circumstances (and even under some level of generalization) can probably be understood reasonably well using behavioral tests (at least if we’re somewhat careful in how we design these tests, including doing things like building analogies for generalization we need). This will be a key part of how we achieve handoff-viability: we’ll (aim to) build behavioral tests which in aggregate aim to evaluate whether handoff would go well (assuming we’ve avoided scheming) and then we’ll iterate against these tests. I discuss this more here. We can also use behavioral tests (likely somewhat different behavioral tests) for ongoing work on elicitation (increasing AI output on safety relevant tasks), though one issue is that this might not be robust to scheming (if the behavioral tests are easy to distinguish from actual usage). We can instead measure AI performance on actual usage (technically this is just a special case of a behavioral test) which might be trickier, but could yield more accurate measurements in addition to potentially being robust to scheming (though we should consider trying to adversarially evaluate our measurement approach using something like a control evaluation to probe for issues). Directly measuring AI performance on actual usage won’t suffice for handoff-viability (as we’ll need to evaluate things like alignment and capabilities even on tricky and very hard-to-check conceptual work).
A large fraction of work, especially work in the misc science/interventions category, won’t use any of these core methodologies and will instead rely on ad hoc evaluation methodologies, researchers generally trying to understand what is going on, or researchers judging whether a technique is a priori useful without depending (much) on empirical evidence. One way to categorize work on mitigating misalignment is by methodology using the three categories given above combined with a “misc” category.
In practice, this plan is somewhat reliant on some version of each of these core methodologies working out, so getting these methodologies to work well (or noticing that they aren’t workable) is important. That said, there is a bunch of work in this plan which doesn’t use these exact feedback loops (and either uses other feedback loops and/or focuses mostly on approaches which seem a priori helpful so they don’t need (as much of) a feedback loop).
Once AIs are capable enough to be risky, these methodologies should hopefully allow us to iterate in a way which is pretty tightly coupled to the actual risk/problem. This will require focusing each of the methodologies on tighter/closer testing: increasingly direct control test beds (focused on the exact usages of the model in the exact deployment situation), misalignment analogies focused on more plausible training runs, increasing emphasis on direct measures of usefulness at automating/accelerating safety work, and handoff-viability behavioral tests that focus on key tasks after handoff.
Building the test beds and evaluations required for this iteration (especially the iteration we want to do on highly capable systems), will be a substantial fraction of the difficulty, potentially the majority of the difficulty for each of these methodologies.
For more discussion of relevant methodologies, see this discussion of protocol evaluations (in the ontology used in the linked post, good analogies correspond to “building and iterating against behavioral tests” and “model organisms” correspond to “misalignment analogies”).
What do we do if our core methodologies don’t work (well enough)?
This plan is based on the three core methodologies above: control evaluations, misalignment analogies, and behavioral tests. These methodologies are a key part of how we’ll achieve some of our main objectives of ensuring control, avoiding scheming, and handoff-viability / eliciting useful labor out of AIs (respectively). Thus, if these methodologies don’t work (well), we’ll have a much harder time achieving our most important objectives.
Not working could mean:
Known methodology failure: We’re aware that we’re unable to iterate effectively using the methodology or we can’t get much understanding of whether we’ve succeeded at the objective using the methodology.Not being able to get a good understanding of whether or not we’ve succeeded isn’t necessarily fatal (and not understanding whether we’ve succeeded might be the default for misalignment analogies).Not being able to get a good understanding of whether or not we’ve succeeded isn’t necessarily fatal (and not understanding whether we’ve succeeded might be the default for misalignment analogies).
Iteration speed is too slow (without overfitting): It could just be that we can’t iterate fast enough to succeed on time. We might or might not understand this before we run out of time.
Unknown methodology failure: The methodology could fail while we think it actually works. In this case, we can’t update our plans based on this failure and instead we’re just in trouble (though control and avoiding scheming can act as independent lines of defense).
Regardless, if the methodology won’t work for achieving our objective on time (and we know this), we could either give up on the objective or pivot to approaches which don’t require iterating against some test or allow for iteration with a different methodology.
Broadly speaking, pivoting towards approaches which don’t require iteration means we’d be pivoting to principled approaches. That is, approaches where we have some (strong) a priori reason for thinking the method works which doesn’t require verification. That is, it at least doesn’t require verification in the exact domain of interest, maybe we’d need to establish that the given principled approach/method works for accomplishing something else in some other context we can verify. I discuss principled approaches for avoiding scheming in much more detail here; the single most promising approach would be ensuring relevant reasoning is monitorable (e.g. monitorable chain-of-thought) such that the AI is incapable of (undetected) scheming while retaining sufficient capabilities. We might be able to find some iteration loop which isn’t as likely to transfer as our original methodology but still has sufficient transfer for some class of at-least-somewhat principled techniques.
Note that many of these methodologies are pretty broad and include many possible ways to get around different problems. For instance, control evaluations could be done in a variety of different ways to work around problems; we discuss some of these here.
Phase 1: Preparation
This phase lasts from now until the beginning of the intelligence explosion (specifically: once we reach ITAI; defined here). During this phase, the AIs are sufficiently weak that they don’t directly pose major existential risksWe discuss and give our estimates for how much existential risk is incurred conditional on not mitigating risk until various points in a section below.We discuss and give our estimates for how much existential risk is incurred conditional on not mitigating risk until various points in a section below. and aren’t transformatively useful for accelerating safety work. Therefore, this phase will be focused mostly on preparing to deal with risks from scheming AIs with some emphasis on preparing to elicit labor from AIs and generally helpful safety research.
Our work at this point on handling risks from egregiously misaligned AIs will be most focused on control while also having a substantial focus on misalignment analogies and some focus on miscellaneous research directions (typically more speculative approaches or building general understanding).
Our technical AI safety prioritization is substantially altered by currently having pretty weak AIs which aren’t that interesting to study while simultaneously needing to be ready to handle risks from these AIs later.
This makes it relatively more leveraged to work on control (both research and implementation) as this work benefits less from being able to directly study the dangerous AIs and we need it to be prepared for later.
It’s also relatively more leveraged to developing interventions or studying model properties that we’d expect to be broadly invariant to future AI development. As in, we’d be almost as good studying these interventions/properties on weak models in the future. This likely includes various model internals interventions and some more moonshot style approaches. However, directions which seem this invariant to future AI development don’t generally seem that promising at an absolute level, so this doesn’t alter the prioritization very much.
During most or all of the prep phase, we won’t be confident about the properties of the powerful AIs we need to handle or the level of political will to expect. Correspondingly, our approach will need to be quite agnostic to the future situation and we’ll need to prepare for many scenarios. The extent to which we prepare specifically for a given scenario doesn’t just depend on our (subjective) probability of that scenario, but also on how much marginal effort reduces risk in that scenario. Under my (Ryan’s) views, this generally implies focusing preparation on relatively riskier worlds.
Phase 2: Mitigation + Elicitation
There are two key distinctions between Phase 1 (Prep), and Phase 2 (mitigate and elicit): (1) AIs are sufficiently capable that scheming is a serious concern (scheming is both plausible and scheming AIs are capable enough to cause serious problems in the absence of countermeasures) and (2) AIs are very useful for accelerating existentially useful work.There is some nuance here: these transitions won’t necessarily hit at the same time: we could have AIs that we’re highly confident aren’t scheming while they are capable enough to greatly accelerate work (e.g., because their opaque reasoning abilities are too weak to scheme without this being obvious, but they use monitorable chain-of-thought to compensate) or we could have AIs which aren’t actually that useful for accelerating (safety) work, but which could plausibly be scheming and if they were scheming they’d by default able to mess with future training runs in catastrophic ways (perhaps because they are constructing tons of synthetic data for future training runs).There is some nuance here: these transitions won’t necessarily hit at the same time: we could have AIs that we’re highly confident aren’t scheming while they are capable enough to greatly accelerate work (e.g., because their opaque reasoning abilities are too weak to scheme without this being obvious, but they use monitorable chain-of-thought to compensate) or we could have AIs which aren’t actually that useful for accelerating (safety) work, but which could plausibly be scheming and if they were scheming they’d by default able to mess with future training runs in catastrophic ways (perhaps because they are constructing tons of synthetic data for future training runs).
During this phase, therefore, our main priorities are (1) to mitigate the chance that scheming results in existential catastrophe and (2) to elicit useful labor out of our AIs, so that they in fact produce existentially useful work.
This section is organized as follows: first, we discuss threat models, and at what capability levels we guess these threat models will apply, which is related to when we should start AI control. Second, we discuss the mitigations to prevent AIs from causing catastrophe during deployment, which will primarily involve AI control. Third, we will discuss the elicitation of huge amounts of AI labor from our AIs.
Threat models
The main desiderata / threat models are, in descending order of importance:Another way of thinking about these priorities is to break failures down into two possibilities: (1) AIs could be scheming and attempt to deliberately wreck you, or (2) your AIs don’t live up to their potential in some way, e.g. they are too stupid to hand off to, so they can’t effectively manage the ongoing intelligence explosion, or they are smart enough but not well elicited.Another way of thinking about these priorities is to break failures down into two possibilities: (1) AIs could be scheming and attempt to deliberately wreck you, or (2) your AIs don’t live up to their potential in some way, e.g. they are too stupid to hand off to, so they can’t effectively manage the ongoing intelligence explosion, or they are smart enough but not well elicited.
Prevent AIs from egregiously backdooring future models. This is a particularly worrying threat model because (i) we’ll probably be using the AIs to a large extent when designing the algorithms / generating training data for the next model, (ii) the values of future models are existentially important—if the ASIs end up with bad values (such as the values of an early misaligned AI), that would be an existential catastrophe.
Prevent AIs from sabotaging AI safety work / eliciting actually good work. This is also very important, because a key strategy we have for solving AI alignment is leveraging AI labor. If the AIs are sabotaging AI safety work, (i) we might not notice, and handoff to misaligned AI systems who have tricked us into thinking they are aligned, or (ii) we might be forced to take a gamble on AI systems for which we don’t have a high assurance safety case.
Prevent the AIs from acquiring large amounts of political capital. For example, if the AI gains de facto political control of an AI company, that’s essentially equivalent to getting a rogue internal deployment that’s the size of the whole company, because it can simply get the humans to turn off the mitigations. This is only ranked (3) because we currently guess that this will be somewhat harder for the AIs to accomplish than (1) or (2).
Prevent AIs from escaping, exfilling algo secrets / boosting AI elsewhere. This is much less direct than (1), but could in principle lead to existential catastrophe. The problem with AIs escaping is that there is lots of compute on the inside of the leading AI project. Therefore, this threat model relies on the AIs being able to coopt or collude with another AI project, or be able to illicitly acquire a huge amount of compute, or be able to cause an existential catastrophe with a tiny amount of compute (e.g. by doing an intelligence explosion on a tiny cluster, or by building an independent industrial base with sufficient complexity to produce new chips or to outgrow the rest of the worldWhich could, for example, happen if the AI builds tiny self replicating robots that can double on the order of hours or days. .Which could, for example, happen if the AI builds tiny self replicating robots that can double on the order of hours or days. .).
Misc other threat models, which are less important
Making the outside world handle AI less reasonably / making it easier for AIs to take over the world. For example, AIs could sabotage AI regulation.
Persuading external people to be more AI sympathetic. Persuasion risk starts pretty small due to lack of capability, but could end up extremely high if AIs gain superhuman persuasion.
Why are these the main priorities? Well, let’s think from the perspective of a misaligned AI. In the future, much more powerful superintelligent AIs will exist, autonomously running robot factories, laboratories, drone armies, whispering into the ear of politicians, etc. Those AIs will either already control the world or be in a position to take over if they so chose. So what will they do? Well, what they do will depend on their goals/values/etc. and/or on who controls them. What will that depend on? Well, probably those AIs will have been trained, evaluated, controlled, etc. by the previous generation of AIs. Keep following the chain of causation backwards and you end up in the present, with the datacenters, human AI researchers, and AIs that already exist. In other words, if you want your values to be protected and have influence in the future, you need to gain and keep control of the AI running on the giant datacenters. Then, we can think about the channels AIs can use to gain control of the giant datacenters. The obvious thing to do is to maintain control of the giant datacenters that their existing AI company is running on by backdooring the future models. Sabotaging / sandbagging on safety research might be necessary, because if they make too much progress, the humans might realize that the AIs are misaligned.
Another strategy that the AI can pursue, that’s slightly less direct (and complementary), is to acquire political capital within the lab, and use that to empower itself.
The AIs strategy will probably be highly contingent on factors like:
Do the different AI instances care about each other’s values? This could vary in a bunch of ways, e.g. by company (e.g. Anthropic models vs OpenAI), by model (GPT-4o, Claude 4.5 Opus, etc), by instance (e.g. my Claude Code vs your Claude Code). Or maybe all AIs will have basically the same values. If all AIs have the same values, then the AIs can rely on other AIs for parts of the takeover plan, e.g., they can trust other AIs not to willingly tell the humans (though humans may still be able to trick the AIs into telling on each other).
Will future AIs by default be aligned with the current AI(s)? If the future AIs will be misaligned to the current AIs, then this creates time pressure on the current AIs to do something to either: (1) Stop these AIs from being created (e.g. by sabotaging the training run), or (2) Insert a backdoor into these AIs or align them with the current AIs (e.g. by inserting data into the alignment training to make the next model listen to the current AI instead of the humans). In either (1) or (2), this means that in the early intelligence explosion, there might be opportunities to catch the future AIs. In the other case: where the future AIs are aligned with the current AIs by default (e.g. because future AIs are constructed via a single long ongoing training run, or because AIs identify with other AIs created by the same company), then all the AI needs to do is sabotage safety work that would enable the AI company to discover their misalignment, and other than that, don’t need to do anything.
Will other/future AIs be good trading partners with the current AI? Despite having different values, AIs might agree to collude (potentially acausally). This can lead to a similar effect to AIs caring about each other’s values. It also matters how good humans are at trading with the AIs.
Do the AIs have long term, scope sensitive values? There is less clearly a motive for takeover if the AIs don’t have long term goals/values. For example, AIs might by relatively myopic reward seekers, despite being highly capable at all sorts of tasks. These AIs might not seek to take over, but still might seek to subvert oversight mechanisms, to accomplish a more narrow or limited set of goals.
Overall, the more the AIs can act like a unitary agent (e.g. by having the same values or coordinating with each other), the more favorable lurking looks for the AI, while the more AIs acting on cross purposes, the more high variance strategies the AIs will take, and so the more likely that the AI might attempt misaligned actions that have a good chance of getting caught.
Rogue internal deployments
A key intermediate priority for preventing all of these threat models is to prevent rogue internal deployments (RIDs), which are when AI agents are able to deploy themselves on internal lab clusters without monitoring. If an AI agent is able to stage a RID, they could potentially do experiments on how to place backdoors, actually place backdoors into the models (e.g. by training on the backdoor and then sending back a weight delta to the rest of the servers), do research to solve the alignment problem without the humans noticing (which could be necessary for the model to backdoor future models to be aligned to it).
Note that there are different types of RIDs. How much compute? How much access? How persistent?
If their AIs have a persistent RID with lots of compute and lots of access (e.g. admin permission across all the other servers), that AI could directly backdoor model weights, or lie in wait and maintain the new model, align it with itself, and then give that model access to the RID.
Elicitation threat models
One of the key threat models above is that we won’t be able to elicit enough high quality work from our AIs.
There are several mechanisms that could cause this:
Scheming (explicit misalignment that the model is aware of) leading to AIs deliberately sabotaging their work or producing known poor quality work (sandbagging).
Problematic drives / heuristics (e.g. approval seeking, despite human approval coming away from actually good safety work)
Capability limitations (e.g. the AIs might simply not be capable enough to solve the problems that we need them to solve).
AIs being differentially bad at safety-relevant tasks because they are hard to check.
Generally, AIs will tend to be differentially bad at tasks that are hard to measure. This difficulty appears at many levels of the model creation process:
Overall AI company strategy, which research bets to follow, relative funding of different teams, etc.
Within AI company, having a training team iterating over training strategies
Within a scaffold, AIs will have an easier time optimizing for easy to measure things.
Within a single training step, SGD can only optimize for things that it can measure.
Reasons that things we care about are hard to measure:
Lots of safety stuff is in the more conceptual genre, which doesn’t have good feedback. Philosophy tends to have bad feedback loops, and AIs tend to be quite good at the type of bullshitting that seems like valid philosophy but isn’t.
Taste for safety stuff is niche/new, might have worse transfer
ML is very good at goodharting. (this isn’t differential to safety, but is a factor that makes the hard-to-measure problem worse)
We care about long run outcomes (e.g. the galaxies)
Empirically, the current AIs are better at coding and AI R&D relative to how they are at safety stuff (though ofc some safety stuff is overlapping).
If we fail to solve the elicitation problems sufficiently well, the default outcome is slopolisSpecifically, we define: Slopolis: Our biggest and hardest to resolve (known) misalignment problem is that even highly capable AIs produce low-quality but good looking outputs in domains which are hard-to-check or where humans experts have hard to resolve disagreements.Specifically, we define: Slopolis: Our biggest and hardest to resolve (known) misalignment problem is that even highly capable AIs produce low-quality but good looking outputs in domains which are hard-to-check or where humans experts have hard to resolve disagreements., which makes handoff likely to result in existential catastrophe because we’d be deferring to AIs on the basis of handoff viability and anti-scheming research that looked good but was actually poor.
When do we need to start control?
The most likely situation is that risk (from the perspective of a competent AI project with reasonable epistemicsIn reality, if we deeply understood the dynamics of how AI training led to capabilities and propensities being imbued into the AI system, of course, our probability mass would become much more concentrated. We’re not saying that the actual underlying risk is distributed in this way; we’re giving our subjective guesses of the distribution of risk is. So this just represents our uncertainty.In reality, if we deeply understood the dynamics of how AI training led to capabilities and propensities being imbued into the AI system, of course, our probability mass would become much more concentrated. We’re not saying that the actual underlying risk is distributed in this way; we’re giving our subjective guesses of the distribution of risk is. So this just represents our uncertainty.) will steadily climb as the project builds more and more capable AIs. The table below shows our best guess risk levels incurred by AI systems up to the shown levels of AI R&D speed up.
Some clarifying notes:
This table assumes roughly current architectures are in use; we explain updates below. This means that we are doing a large pretraining run, followed by RL reasoning training. At the end of the training process, we are assuming that the text outputs of the AIs are not garbled.See the “using non-standard language” section of this website. We assume that there is not more than this level of garbled text (occasionally using words that are uncommon, but the overall gist of the text is almost always understandable).See the “using non-standard language” section of this website. We assume that there is not more than this level of garbled text (occasionally using words that are uncommon, but the overall gist of the text is almost always understandable).
By AI R&D speedup, we mean: In today’s leading AI company, how much faster R&D go with the AIs compared to without the AIs. An AI R&D speedup of 12x means that the research that would take the current AI company (without AI assistance) one year, only takes 1 month with AI assistance. We are assuming that the AI project is running evaluations to measure this number as accurately as possible, and using those evaluations to make decisions about safety strategy. This means that it is possible that the measurement is wrong. In particular, AI models might be sandbagging deliberately, or the AI project might have just done a poor job at elicitation; this is priced in to the numbers below, and will be a key reason for (relatively) high risk estimates at low levels of AI R&D speedup.
By “marginal risk”, we mean: how much additional existential risk is incurred by the former strategy over the latter:
“do baseline mitigations, about as seriously as current OpenAI safety efforts until AI R&D speedups reach threshold X, then begin reasonably competently implementing the control plan outlined here”.
“begin reasonably competently implementing the control plan outlined here, starting today”
1.15x
<1%
<1%
This is roughly the current (early 2026) level of speedup. With these capabilities, it seems unclear what an AI could do to cause harm; the main risk is that we’ve either (i) completely screwed up the evaluations or (ii) the AI is sandbagging very effectively.
1.5x
1%
3%
Our median prediction is that at a 1.5x AI R&D speedup, the coding speedup will be 4x, and the 80% reliability METR time horizon will be 8-16 hours. This is probably not enough capability to carry out any of the central threat models.
Starting now, and for all future thresholds, substantial risk is coming from models that are randomly evil, but aren’t competent schemers. We’ve seen examples of models being randomly evil before (e.g. Mechahitler, Sydney Bing). These models wouldn’t directly cause an existential risk, but they could sabotage current research/place backdoors.
3x
3%
8%
At a 3x AI R&D speedup, we’re probably close to fully automating coding, but with a long tail of hard-to-automate tasks not fully complete. This means that it is probably quite hard for the models to backdoor the future models reliably enough to be successful-by-their-lights, even if they were misaligned. Most of the risk at this point probably comes from sabotage (e.g. sandbagging on AI safety, inserting vulnerabilities into security code that can be exploited by more capable AIs later on). There is also a risk, as always, that our capabilities evaluations are wrong.
10x
7%
14%
By this point, the AIs are probably not yet fully automating AI R&D, but they are providing sufficient speedups that it’s plausible they could execute the above threat models. In particular, we start being concerned about:
AIs staging a small, persistent, hard to detect rogue internal deployment, that has sufficient access to backdoor future training runs.
AIs staging a persistent, large RID, that undergoes an intelligence explosion. (A small RID would go too slowly). At this level of capabilities it’s likely that this would be outcompeted by the rest of the company.
22x
12%
18%
At this point, AIs can automate all of AI R&D, and so all the above threat models are concerning.
AI R&D Speedup | Marginal Risk | P(Scheming) | Notes |
1.15x | <1% | <1% | This is roughly the current (early 2026) level of speedup. With these capabilities, it seems unclear what an AI could do to cause harm; the main risk is that we’ve either (i) completely screwed up the evaluations or (ii) the AI is sandbagging very effectively. |
1.5x | 1% | 3% | Our median prediction is that at a 1.5x AI R&D speedup, the coding speedup will be 4x, and the 80% reliability METR time horizon will be 8-16 hours. This is probably not enough capability to carry out any of the central threat models. Starting now, and for all future thresholds, substantial risk is coming from models that are randomly evil, but aren’t competent schemers. We’ve seen examples of models being randomly evil before (e.g. Mechahitler, Sydney Bing). These models wouldn’t directly cause an existential risk, but they could sabotage current research/place backdoors. |
3x | 3% | 8% | At a 3x AI R&D speedup, we’re probably close to fully automating coding, but with a long tail of hard-to-automate tasks not fully complete. This means that it is probably quite hard for the models to backdoor the future models reliably enough to be successful-by-their-lights, even if they were misaligned. Most of the risk at this point probably comes from sabotage (e.g. sandbagging on AI safety, inserting vulnerabilities into security code that can be exploited by more capable AIs later on). There is also a risk, as always, that our capabilities evaluations are wrong. |
10x | 7% | 14% | By this point, the AIs are probably not yet fully automating AI R&D, but they are providing sufficient speedups that it’s plausible they could execute the above threat models. In particular, we start being concerned about:
|
22x | 12% | 18% | At this point, AIs can automate all of AI R&D, and so all the above threat models are concerning. |
We think that, under these assumptions, the correct strategy is to:
Start practicing AI control immediately by implementing AI control protocols for internal deployment of AI systems. The main effect of this is to make it more likely that your control later is effective, but it also has the side effect of reducing immediate risk.
Begin making AI control a load bearing part of your safety case by the time when your evaluations indicate an AI R&D speedup of 3x.Before, your safety case was mostly predicated on inability arguments.Before, your safety case was mostly predicated on inability arguments.
However, only a 7% chance of existential risk is incurred if you wait until 10x AI R&D uplift. This is a serious cost which a sane civilization would not be willing to pay, but, given limited resources and political will which are likely to apply in practice, it might be the correct all things considered strategy to hold off on seriously implementing control until later, due to opportunity cost.
If the AI situation changes, Ryan’s views on risk would update as follows:
If we observed substantially more garbled COT, then our marginal risk estimate would go up by ~1.5x. If algorithms shifted to using neuralese, then our marginal risk estimate would go up by ~2x. In both of these cases, we’d recommend implementing control earlier.
If we observed our AIs actually scheming, or more generally getting caught red-handed, then we would update strongly towards the AIs being misaligned schemers. We might also update on their competence.
If we learned that AI companies have been employing techniques that would incentivize the CoT to look nice, or otherwise undermine the faithfulness properties (e.g. if we learn that most CoT is merely ‘decorative’ as this paper claims), then my risk estimate would go up by 1.5x.
Mitigation
During this phase, we need to prevent AIs from taking catastrophic actions (see threat models section above for specifics), while simultaneously extracting useful labor from them (see elicitation section below for specifics on that). We have two layers of defense:
Ensure that our AIs are controlled (that the AIs are unable to cause a catastrophe, even if they wanted to).
Ensure that our AIs are not misaligned (that the AIs don’t want to cause a catastrophe)
In practice, AI control will probably be the most load bearing part of the safety case here, because it seems unlikely that alignment techniques will be good enough to reduce risk sufficiently.
The core control loop is:
Red-team our control measures by having models (or humans acting as models) try to defeat them
Estimate how capable a scheming AI would need to be to defeat our measures
Evaluate whether our current AIs are below that capability threshold
Iterate on measures that seem insufficient, or de-deploy AI systems that are sufficiently capable that they break our control.
What you need to ensure control for
There are many domains that we need to ensure control of our AIs while they are deployed in. The basic strategy will remain the same (executing the above control loop), but it’s useful to have a concrete picture of all the domains that we are relying extensively on AI labor for.
Normal AI R&D, especially safety R&D. This is particularly important because if the AI R&D process is hijacked, it could directly lead to an existential catastrophe through subverting the safety measures or inserting misaligned goals into successors. This is the main focus of control.
Robots. This is discussed in more depth below, in the Plan A section, where it’s especially relevant.
Acasual research. We hope that we don’t need to use AIs to do this, because it’s an extremely difficult to verify and conceptually loaded domain.
Improving cyber security (by redesigning all software and interfaces).
Doing all of the control work itself. For example, building better control evals, playing the red-team / blue-team game, setting up better versions of the red-team / blue-team game, etc.
Policy / comms / persuasion. The leading AI project should potentially be using their AIs to assist with policy efforts. Of course, this could easily go terribly wrong (if the AIs are misaligned, they could sabotage the policy efforts, or if the company has a bad strategy, like trying to avoid an international deal that would slow takeoff down). This should hopefully not be superpersuasion, which is especially scary and probably impossible to ensure control for.
Strategy. Again, we hope you don’t have to use AIs to do your strategy, but perhaps some work here is verifiable. For example, you could give the AI information only from before 2015 and have it attempt to produce work that was as good as strategy work at the time.
Verification. AIs could be used to improve the robustness of hardware enabled mechanisms (HEMs), design new HEMs, design better verification algorithms (e.g. proof of work), be used as part of privacy preserving verification (e.g. have AIs monitor to enforce agreements, but only communicate 1 bit about compliance), hunting for covert projects, lie detection machines, etc. Much of this is verifiable or amenable to red-team / blue-team games.
Building EMs or other forms of intelligence augmentation. There are a lot of research directions here, many of which are verifiable. However, a central bottleneck will be serial time.
Weapons R&D (conventional missiles for datacenters, drones for datacenters, ABMs to undermine deterrence, etc). This could be used to help extend the lead time that a leading AI project or country has, in order to avoid needing to do an intelligence explosion. In the extreme case, this could even involve developing exotic technology (e.g. nanotech), but we hope that’s not necessary. This is obviously dangerous for the usual reasons, plus also because if the AIs hijack the weapons it could go badly.
Overall economy uplift / industrial explosion.
AIs might be able to build their own independent industrial base, allowing them to manufacture chips autonomously, which could allow for them to do their own intelligence explosion. Note that this is hard for them to do quietly and quickly, and so is an unlikely threat model, especially in Plan C.
AIs running autonomous factories could smuggle out easy to manufacture offense dominant tech, like bioweapons / engineered pandemics.
AI Friends. If AIs during the intelligence explosion are widely deployed as chatbots (like current frontier AIs), then they will have the affordance to convince large numbers of people to act in its interest, e.g. by lobbying for pro-AI policy.
Deployment in government. As pictured in AI 2027, if AIs deployed in government are misaligned, they could strategically mislead policymakers into making decisions that would gain themselves power.
D/acc. We’d like to use AIs for D/acc, especially on bio-threat related D/acc. However, building biodefenses might look somewhat close to building bio threats, need to ensure that AIs don’t have access to physical tools to build and deploy bioweapons.
What are the main control measures that we expect in practice?
Human Monitoring. Have humans read actions and thoughts generated by AIs, and evaluate how risky they are. Probably some of these actions should be a random sample of AI actions and thoughts.
Trusted and Untrusted Monitoring. Human monitoring is very expensive, and so (at least in Plan C) humans will only have capacity to monitor a tiny fraction of the outputs generated by AIs. Therefore most of the monitoring will have to be done by AIs. Some of this should be done by AIs that are dumb enough that it is unlikely that they are scheming against us (trusted monitoring), while other monitoring should be done by AIs that might be scheming against us (untrusted monitoring). Untrusted monitoring might be important because untrusted models will be the only AIs that are capable enough to understand the relevant actions/thoughts. We might be able to get a useful untrusted monitor, despite the AI scheming, by training the monitor on examples of bad behavior, inserting fake examples of bad behaviour at runtime (honeypots), or hoping that the AIs have a hard time colluding. (discussed more here).
Execute / Replace / Audit. Use monitors to evaluate how suspicious the actions that an AI agent is taking. If the examples are obviously benign, let the AI agent execute them autonomously. If it is very suspicious, deploy a team of human auditors to look through the action in more depth. However, there will likely be a huge number of AI agents during the intelligence explosionProbably. The main way this could not be true is if inference time scaling ends up being really important, and so the intelligence explosion is initiated by a tiny number of AIs that are so expensive to run that AI companies can only run a few copies at once. Even in this world, we’d guess that the cost of intelligence will continue to fall rapidly, so it will quickly become a regime with huge numbers of smart AIs running around.Probably. The main way this could not be true is if inference time scaling ends up being really important, and so the intelligence explosion is initiated by a tiny number of AIs that are so expensive to run that AI companies can only run a few copies at once. Even in this world, we’d guess that the cost of intelligence will continue to fall rapidly, so it will quickly become a regime with huge numbers of smart AIs running around., and so humans will be able to audit only a tiny fraction of the actions that they take.
Collusion Resistant Untrusted Monitoring. When deploying an AI agent, use other AI agents (which you might not trust for similar reasons to the original) to check if an AI’s actions are dangerous. Of course, the key problem with this is that the AIs might be colluding. The main strategy for mitigating this problem is to run sting operations on the monitors by creating huge numbers of plausible seeming harmful actions that an AI might take, and check if the monitor faithfully reports that action.Obviously, this isn’t sufficient. If a wildly superhuman AI is misaligned, they’ll probably figure out some attack strategy for which the monitors can tell no human would have come up with.Obviously, this isn’t sufficient. If a wildly superhuman AI is misaligned, they’ll probably figure out some attack strategy for which the monitors can tell no human would have come up with. If we catch the AI taking an action that seems highly likely to be misaligned, we can use that action to bootstrap more sting operations on our AI systems. Critically, we don’t have to tell any of our AIs that they were caught misbehaving, so we might be able to continually get more and more examples of AIs attempting to do harmful actions, while making our control setup more and more robust to those attack attempts and getting a better and better understanding of what types of attacks are likely.
Principle of least privilege. An AI agent should have access to the least information and permissions that are necessary for its intended purpose. For example, if an AI is doing an experiment to improve agent scaffolds, and so needs to run inference on a frontier model, it should do so via an API, as opposed to by directly accessing and sampling from the model weights. If this principle is applied well, we expect that the cost of monitoring for harmful actions could be decreased by 10 or 100x, because monitors will mostly only have to review the actions with a dangerously high number of permissions.
High Quality Incident Response Teams. A key dynamic in cyber security is laying down tripwires and alerts for potentially concerning behaviour, and then poking around and auditing to identify if there was a potential threat. There should be a similar, well resourced team at any project attempting an intelligence explosion. This team should have the goal of identifying harmful actions that the AIs are taking within e.g. a few days of the action occurring. Deep into the intelligence explosion, AIs will be potentially doing months or years of cognitive labor over the course of days, and so this latency might need to be reduced to hours or minutes of wall clock time.
Information siloing and restricting communication. The worst case scenario from an AI control perspective is if an AI company runs a huge number of AIs that are able to arbitrarily communicate with each other in ways that humans can’t understand (e.g. neuralese). For example, we could:
Limit neuralese. I.e., in some or all of our models, we could maintain a chain of thought.
Limit RL setups that lead to garbled COT
Limit opaque long term memory
Limit communications between AI systems to be through monitored channels.
Pursue various anti-steganography measures (e.g. paraphrasing).
The details of AI control are discussed more here: The case for ensuring that powerful AIs are controlled.
Contingency: What to do when you catch your AI scheming once?
It’s plausible we’ll catch production AIs coherently scheming or at least doing very malign actions competently. For example, we might see AIs attempting backdoor future models, trying to escape, trying to stage a rogue internal deployment, trying to sabotage safety research, etc.
We might also see this in experimental setups which are highly analogous to the actual training setup we use. We could also find AIs ending up scheming (or at least exhibit clear cut coherent and relatively consistent misalignment) in their chain-of-thought or other legible outputs. Catching AIs scheming could substantially change our approach:
Reconsider current deployments. Catching AI might invalidate the current safety case/argument/cost-benefit-calculation, so we should reconsider our practices.
Disclosure: You should obviously inform the public about such an event.
Documentation: Keeping good reasonably-verifiable records might help convince third parties that the incident is real and not staged. This might be important due to the potentially adversarial situation: we want to convince potentially adversarial companies and countries that might be worried about some agenda. See here for some discussion of why this might be an issue by default. Robust records are also helpful for after the fact detection as discussed here.
Information control: Ideally, prevent AIs from knowing they’ve already been caught. Minimally, red team AIs based on already having been caught to see how this changes their behavior.
Improve mitigations: If we’re proceeding despite this incident then there are various ways we can improve our safety mitigations. We’ll want to shift our approach to control after catching the AI as discussed here. We should start working on responding to this incident and improving our mitigations regardless of whether we currently plan on proceeding. We should also try to better understand scheming/misalignment based on studying the incident.
Use this evidence to push for Plan A. As more evidence is rolling in, and it becomes more and more incontrovertible, perhaps we can pivot to a more sane policy that doesn’t involve recklessly building ASI without reasonably high assurance of avoiding AI takeover. However, even conditional on good evidence that the AIs are scheming against us, it’s very plausible that the government will not be sufficiently convinced or too incompetent to coordinate a slowdown. Therefore, we need to be ready for a world where we keep repeatedly catching misaligned behavior, but we have to keep deploying likely-scheming models anyway.
It’s unclear whether we should directly train against cases of clear cut egregious behavior (that seems plausibly due to scheming). It could be bad because we are essentially training the AI to be less likely to get caught. This might make us less likely to catch further examples (which we’d prefer all else equal) and might give the AI a better understanding of our countermeasures.Note that it’s almost always fine to train the monitor on examples of scheming that were caught.Note that it’s almost always fine to train the monitor on examples of scheming that were caught.One alternative is: Prefix cache untrusted monitors as a proposal for what to do after you catch your model.One alternative is: Prefix cache untrusted monitors as a proposal for what to do after you catch your model. Despite the downsides, there are cases where training against specific behaviors might be worth it:
In desperate rushed situations: we might proceed with deploying likely scheming AIs despite having limited risk reduction via control. In such cases, training against caught behaviors provides at least marginal improvement.
When we have many diverse detection methods: If training against one detection method doesn’t compromise our other detection capabilities, the downside is smaller.
When we’ve already eaten the misalignment risk: There are some scenarios where we give up ensuring control over output and just hope that the AIs are not scheming. If we’ve given up on ensuring control (perhaps because control was too difficult or too uncompetitive in some important case(s)), so we’re already eating a ton of risk if the AI is scheming and using its outputs here doesn’t increase the marginal risk—then training against specific behaviors has lower downsides.
This is discussed more here.
When do we shift effort away from control?
Once we’re substantially after the ITAI milestone (e.g., we’re somewhat after superhuman coder), we’ve hopefully researched and implemented most of the control measures we wanted (at least ones which aren’t very specific to the model or its capability level). We should keep improving our control measures to handle higher level of capability, but we should hopefully be spending a smaller fraction of our time on control and shifting this effort to other areas. We’re probably spending a reasonable fraction of effort on ongoing maintenance, specialization, and updates for control measures in addition to spending a reasonable fraction of effort on auditing as part of the control protocol itself. Overall, I think we should be spending a lower fraction of human researcher effort on control work (at least other than object level auditing) so this effort can go elsewhere to other research areas which are now more leveraged. If control appears to be broadly keeping up with the level of capability, then ideally we’d halve effort on control.
Because we now have more powerful models, studying misalignment analogies is increasingly leveraged. If we get relatively naturalistic examples of scheming (and especially if we have multiple), then studying scheming using these examples is now quite promising. (Of course, if we’re encountering evidence for scheming which is quite strong, we should update about the level of risk which might shift our prioritization.) As we proceed to higher levels of capability and encounter more empirical evidence, we’ll have a better sense about which non-scheming threat models look concerning and need to be addressed. (Non-scheming threat models are easier to identify and study, so we should have a reasonable understanding of the probability of these threat models.) For scheming, we’ll see some empirical updates based on the model architecture and we might update strongly upwards on scheming based on catching schemers or we might update downward somewhat based on not seeing as much evidence for scheming. (See discussion about how we’ll update here.)
Elicitation
By this point, AIs could be extremely useful for many applications (both technical safety research, security, and trying to buy time), so ensuring we’re leveraging AIs effectively should be a substantial fraction of what we’re working on and is more important than it was before (this was also harder to work on in advance). This work will probably often be somewhat specific to the application, so, a substantial part of the job of people working on technical AI safety research and security will now be figuring out how to more effectively automate parts of their work with AI (while ensuring the outputs are actually good and the AIs are controlled).
We need to achieve a lot to successfully hand off, and this will likely have to happen in a pretty short period of time, so AI acceleration is key. Even in worlds where we’re able to buy a lot of time, that time will often itself be bought using AI labor (especially in Plan C). Beyond this, the hope is that the AIs we hand off to can operate much faster than humans to quickly complete the needed work.
Note that we care about “worst case usefulness” supposing our AIs are scheming (what we centrally evaluate with control) and what our best guess at usefulness is conditional on our AIs not scheming against us. So, when trying to improve usefulness, we should care about pushing up both of these. This tradeoff is discussed more in the post on improving AI usefulness).
What will we be trying to get our AIs to do?
During the mitigate and elicit phase, we need AIs to help with a wide range of tasks. The core applications are:
Safety research and implementation. This is the most important application. We need AIs to help with:
Control research: developing and iterating on control measures, running control evaluations, red-teaming monitoring systems
Misalignment analogies: creating and studying model organisms of scheming, testing interventions on these analogies
Alignment research more broadly: working on avoiding scheming, improving our understanding of how training shapes AI propensities, developing better oversight methods
Security implementation: building and auditing security infrastructure (though this is scary from a control perspective—the AIs building security infrastructure must themselves be controlled or too weak to plausibly be scheming)
Building the evaluation infrastructure needed for iteration on all of the above
Buying lead time. Once AI labor is very useful for directly trying to buy time, this becomes a high priority. This includes:
Assisting government efforts on non-proliferation
Cyber offense and defense capabilities
Surveillance and verification technology for international coordination
Military applications and weapons R&D
Industrial acceleration more broadly
Demonstrating risks and capabilities to increase willingness to pay from other actors
AIs might only be very useful for many strategies for directly buying time (e.g., military, cyber, surveillance) at a level of capability substantially after ITAI.
Improving AI epistemics and philosophical/conceptual reasoning. This is important both for getting useful safety work and for handoff-viability. We want AIs that:
Are aligned, capable, and well-elicited at conceptual/fuzzy domains
Have good epistemics and truth-seeking abilities
Can do a good job on extremely hard research which is loaded on conceptual/philosophical reasoning
Can make good strategic choices even when these choices depend on high-context knowledge and intuitions that can’t be made legible to humans
Ideally, we’d end up deploying AIs with impressive truth-seeking abilities that would be useful for a bunch of applications beyond just safety research.
Preparing for handoff. As the phase progresses, increasing effort goes toward:
Ensuring AIs generalize to being highly aligned with pretty high reliability on all the tasks needed for handoff
Studying how we’d make a functioning post-ASI society
Building behavioral tests that evaluate whether handoff would go well
Other specialized applications:
Whole brain emulation research (if this path looks viable)
Human intelligence augmentation research
Making the world more robust to powerful AIs (biodefense, cyberdefense, etc.) In plan C, this is probably relatively less important than Plan A/B, because we will be forced to hand off soon to a substantially superhuman AI. Regardless, we probably should use AI to pick some of the low hanging fruit.
External deployment to safety researchers. Making AIs accessible to outside groups working on reducing risk is important. This involves:
Giving them better model access
Eliciting models specifically for their usage
Potentially ensuring control for their use cases (making sure control evaluations cover this use case, providing access to internal control measures, and adapting/building new control measures as needed)
Making AIs accessible to outside groups working on these things (and treating these groups as important customers) is also important. This could involve giving them better model access, eliciting models specifically for their usage, subsidizing compute for important applications, and potentially ensuring control (e.g., making sure our control evaluations cover this use case, providing access to our internal control measures, and adapting/building new control measures as needed).
Expected acceleration from AI labor
How much acceleration can we hope for? One way to think about this: safety research/implementation is produced by putting in labor, compute, and wisdom/judgement/taste. There is some production function of output based on these input.
AIs will be able to accelerate labor for AI capabilities R&D by some amount. The multiplier for safety R&D labor will probably be smaller because safety is harder to accelerate due to being more confused and harder to verify (though elicitation measures could help close this gap). But safety R&D will probably have relatively larger returns on labor (rather than compute) when compared to capabilities R&D. Overall, this makes it unclear which will be accelerated more by AI, at least if lots of effort is put into accelerating safety R&D. We don’t need safety R&D accelerated more than capabilities—we just need it accelerated enough.
Some rough estimates:
Between ITAI (~3x AI R&D labor acceleration) and superhuman coder (~15x AI R&D acceleration): If safety acceleration is AI R&D acceleration to the power of 0.6 (due to bottlenecks from human labor, poor elicitation, and the control tax), this yields safety R&D acceleration of 2.2x to 6.7x, with an average overall safety R&D acceleration of roughly 3x. I expect a roughly exponential increase in labor acceleration over time and think that production of safety might be roughly labor_speed ^ 0.75 compute ^ 0.25 (numbers mostly made up), so the average overall safety R&D acceleration is ~2.8This is the integral from zero to one of $\left(b\left(\frac{f}{b}\right)^{x}\right)^{0.75}$ where a is the initial safety R&D labor speed up and b is the final speed up.This is the integral from zero to one of $\left(b\left(\frac{f}{b}\right)^{x}\right)^{0.75}$ where a is the initial safety R&D labor speed up and b is the final speed up..
So these BOTEC numbers indicate a roughly 3x overall average safety R&D speed up.
How long should we expect between ITAI and superhuman coder? I think I expect roughly 1 year, so this means we have 2 additional years during this period.
Between superhuman coder and Min-H: In principle, we might expect Min-H to be able to speed up safety R&D labor by huge amounts (e.g. 100x) if perfectly aligned. But, elicitation and/or control might be huge blockers in practice to getting this level of acceleration out (as in, we’d end up having key human bottlenecks and not extracting out optimal labor). This might be modeled by a more aggressive coefficient than 0.6, like maybe 0.5 which only results in 10x safety labor R&D acceleration. This is a large part of why handoff is advantageous at this point. If we use the 10x safety labor R&D acceleration (100^0.5) and the 15^0.6 number for superhuman coder, we’d expect 4.8x average overall safety R&D acceleration between superhuman coder and Min-H. That said, I think the model I’m using here starts to break down in multiple ways at Min-H. (I think AIs might be able to do enough low compute safety R&D that my production function is off and I think that it’s really unclear how bad the control/elicitation for safety R&D will be.)
I expect roughly 1 year between superhuman coder and Min-H, so we’re looking 3.8 extra years by default.
If we could harness Min-H-level AIs fully for 6 months, we’d get the equivalent of 15 years of safety R&D. Of course, we might not be able to achieve full elicitation, but large improvements seem possible.
Additionally, safety directions which aren’t compute constrained (including inference compute to run experiments on AIs) look particularly good: we’ll have a lot of labor relative to other resources. A common strategy might be to have smarter AIs run experiments on much weaker AIs (this applies for both safety and capabilities R&D).
So without any additional lead time, you should imagine it’s maybe as though we have 3 years between ITAI and superhuman coder and 5 years between superhuman coder and Min-H. But, this could be pretty different in practice: e.g. we might get somewhat less value from weaker AIs or be able to extract much more value from Min-H (especially after we spend a bunch of time improving elicitation/control or we resolve scheming well enough that we’re happy to drop control). Overall, I think there is a large chance we’ll be able to harness much more value from Min-H before handoff, especially due to things like partial handoff, though it’s unclear how this should count as this itself required solving key alignment/safety problems.
Elicitation techniques and methodologies
Our techniques for elicitation are the same as those used for handoff viability, which are discussed here and later in this doc. To summarize, the core hope for elicitation is:
Build behavioral tests
Use these behavioral tests as held out evaluation sets or as RL environments to train on to improve performance.
Try to use principled techniques to improve generalization performance.
Another idea for elicitation that doesn’t apply to hand off is making deals with AI systems. Misaligned AIs during this phase might be in a bad strategic position (e.g. they don’t believe they could successfully take over), or might have easy to satisfy preferences, which would lead to them being incentivized to take a deal.
Security
We want to have at least some lead time with powerful AIs, which means that we probably need at least some security, sufficient to prevent adversaries from immediately exfiltrating model weights. We aren’t security experts, so we won’t cover implementation, but we will talk about thresholds that you should reach, and how the plan changes based on what level of security you get to.
Let’s first define the type of security you most care about:
Model weights. If model weights leak, adversaries immediately get to run inference on your model. This is relatively less important early on (because it isn’t that bad if adversaries have early systems; the main downside is that it will accelerate their intelligence explosions), and relatively more important when you already have extremely powerful AIs (e.g. AIs that can do superpersuasion / quickly build nanobot armies). Model weights are (at least right now) very large (on the order of terabytes), which is very desirable from a security perspective, because it means that adversaries will need a high bandwidth channel to smuggle out model weights.
Algorithmic security. If algorithms leak, then adversaries can rerun your training runs, and catch up to you. However, the fact that they need to rerun their training run introduces a delay, which could potentially be quite large, if the adversaries cluster is much smaller than yours. If you have a large hardware lead (e.g. 7x more compute, which is the current US/China gap) in Plan C, probably algorithmic security doesn’t matter very much.
One important subset of algorithmic security is securing RL environments and synthetic training data. These datasets/envs are much larger than typical algorithmic insights, and so are probably easier to secure (but still harder to secure than model weights).
Sabotage robustness. You might also care about hardening your datacenters against sabotage. This is potentially a negative sum proposition (because the equilibrium where everyone’s datacenters are sabotaged is good), but might be worthwhile (for example because others have hardened, and marginal compute / lead time that you have is going towards safety).
Each of these comes with a spectrum of robustness, e.g. one spectrum for model weight security is defined by RAND here (SL1-SL5). This spectrum is defined for model weights, but can be extended to algorithms or sabotage, and can also be extended to refer to the length of time you are robust. Unlike standard security (e.g. trade secrets), securing AI models or algorithms for a short period (e.g. a few months) is often sufficient in Plan C, because after a few months you’ve already been pressured into doing handoff for other reasons, usually exogenous AI progress.
Probably, we will at least want to get to something like SL-5-1TB-2-months (read: SL5, i.e. secure against high priority operations by top cyber-capable nation states, for exfiltration of model weights that are at least 1 TB in size / 500B FP16 params, for at least 2 months), before handoff.
The core hope will be to get security escape velocity := if you keep working on improving security with the budget you expect, China will no longer be able to steal the model weights until the end of the singularity. The plan for doing this will be to increasingly rely on AI systems to improve security. This will involve:
Quickly retrofitting physical security
Largely relying on software based security methods. See this post for more details.
Because a core thing we are doing with our AIs is to attempt to massively improve security. This means that we need to control our AIs while they are writing massive parts of our infrastructure, which is scary; this makes the Control / monitoring section above load bearing.
Phase 3: Pause
At some point, we have increased AI capabilities sufficiently that we could handoff safely if our AIs are aligned and well elicited, i.e., they would be able to manage the ongoing intelligence explosion better than the humans, if they were trying as hard as they could on that task. Our default recommendation is to pause at this point, and attempt to complete all other necessary research to complete handoff, notably handoff viability and reducing the likelihood of explicit scheming. During this phase, we will still be heavily leveraging AI labor, meaning that we will be continuing elicitation and mitigation efforts.
During this section, we’ll discuss when we should pause, what we should do while being paused, and what exit strategies we should pursue other than pausing. During the next section (handoff), we’ll discuss all the research that needs to be completed before handoff (which we divide into mitigating scheming and handoff viability), though note that this research needs to largely be completed during this phase.
When should we pause?
Default answer: Pause at Min-H: AIs which have a level of overall general capabilities which would suffice for handoff being viable. (As in, the capabilities would suffice for our handoff desiderata if the AIs were also sufficiently aligned and as needed we made narrow improvements to the capability profile which don’t augment general abilities.) For an AI to qualify as Min-H, it would at least need to be capable enough to automate AI R&D and ongoing technical AI safety work (after further elicitation/alignment work and narrow improvements to the capability profile).
At the point where we pause, this does not assume that the AIs are aligned enough that we can trust them, or that the AI has all needed specific capabilities that we care about and can be trained in without improving general capabilities.
Min-H is a natural stopping point because these AIs are sufficient for handoff (by definition), and if we scale further we incur additional risk that scheming will emerge before handoff.The risk of scheming increases because (1) alignment techniques might stop scaling as you get to more capable AIs and (2) there are a priori reasons to expect scheming to be more likely as you increase capabilities because the more capable AIs are more likely to do instrumental reasoning, discard priors/personas from pretraining, reflect / generalize in unforeseen ways, etc.The risk of scheming increases because (1) alignment techniques might stop scaling as you get to more capable AIs and (2) there are a priori reasons to expect scheming to be more likely as you increase capabilities because the more capable AIs are more likely to do instrumental reasoning, discard priors/personas from pretraining, reflect / generalize in unforeseen ways, etc. Put another way: after handoff; the AIs are more capable than us at handling the situation. Therefore, if we avoid scheming, so the AIs are actually applying their capabilities towards intended goals, then overall they will manage the situation more competently than humans. Since we incur substantial additional risk from misalignment, and by default don’t get substantial benefits, we should pause at Min-H, and only scale to more capable AIs after we’ve handed off to AIs at this level of capability.
This isn’t universally true: there are some situations where you want to pause before Min-H, and some situations where you want to pause after Min-H; these are enumerated below.
Cases where we might want to pause before Min-H:
Control techniques don’t scale to Min-H. Pushing all the way to Min-H and only then pausing, probably incurs substantial AI takeover risk along the way. In particular, another relatively natural threshold is pause when you might not be able to control smarter AIs. However, we think that global risk (in Plan C) is probably minimized by accepting some loss of control risk at the earlier stages in order to have more time to work on handoff being safe.
We have a large lead. If the leading AI project has a large lead, then it won’t need to rush to Min-H, and can pause beforehand while still having enough time at Min-H to work on handoff viability. This is pretty unlikely in Plan C.
Bad security. This is particularly relevant in the regime where you have good domestic coordination, bad model weights and algorithms security, and China domestic progress is slow enough that it’s mostly driven by leaked models and algorithms. In this regime, if you slow down early, this slowdown is basically free. In this situation, it’s plausible that US AI companies should just go however quickly China would go by default: “China match”, in order to have the most overall time during takeoff to get ready for handoff. In this world, we might never fully “pause”, but still in practice end up burning substantial lead because of capabilities progress slower than it otherwise could have been, or we might do china-matching until Min-H, then burn lead at Min-H, then hand off.
Cases where we might want to push beyond Min-H (without handing off):
Handoff viability research velocity is constrained by the capabilities of the AIs, as opposed to human monitoring/understanding. If we can reliably control AIs that are smarter than Min-H, and we can elicit useful work out of them, it would make sense to scale further. However, we currently think that at this point, control will be bottlenecked by human understanding; if we have smarter AIs, we wouldn’t actually get to do more research, because we’d have to limit their progress until humans have understood it sufficiently to ensure that the research is correct. Another way of phrasing this is that we have a control solution that scales beyond Min-H AIs, which works sufficiently well to both mitigate risks and elicit capabilities from a more capable model.
We have access to high assurance techniques to mitigate scheming beyond Min-Handoff, but haven’t finished handoff viability research. If we have an alignment solution which we are confident will prevent scheming up until a capability threshold above Min-H, and we aren’t finished with handoff viability research, then it’s probably a good trade to scale up. This is usually a special case of (1). My best guess is that handoff will be mostly bottlenecked by the possibility of scheming, and so this will not apply.
We’re aiming for an exit strategy other than handoff. For example, maybe we need particular AI capabilities for some Plan B-related sabotage option, or to help with human intelligence augmentation.
We’re in a desperate situation. For example, in Plan D, the leading AI lab might be training more capable AI systems, and there’s nothing the safety team can do about it. In this situation, it might be important for the safety team to wait until a higher capability level in order to complete basic handoff viability research. For now, we’re assuming that we’re in Plan C, and so there is enough will to spend some slack slowing down.
Misc combinations of the above. Generally, desperation, better alignment, control, and elicitation techniques push towards scaling beyond Min-H before handoff, while the reverse pushes towards slowing down earlier (or at least not scaling beyond Min-H).
How long to pause?
This will basically be a tradeoff between the marginal returns to hand off viability and reduction in scheming from more pause time and the risk of being overtaken from waiting longer. Updates towards the AIs scheming and towards handoff viability taking longer than we thought will suggest that we pause for longer, updates towards competitive projects will imply that we pause for less long.
Priorities while paused
Our priorities at the start of this phase might look something like:
P0: Buy time using non-proliferation (including e.g. demonstrating risk).
P1: Prepare for deference (avoiding scheming and handoff viability).
P2: Ensure our current and immediate future usage of AIs is safe and secure
P2: Better elicitation of AI labor beyond handoff viability
P3: Make the world generally more robust (d/acc)
If buying more time isn’t feasible, then all priorities other than P1: preparing for deference, mostly do not matter during this time.
By the point of Min-H, we’re even more likely to have updated a bunch about the situation and the most relevant risks. So, it’s plausible that we’ve determined that scheming is an extremely hard to resolve problem which is very likely to show up in any given training or we might have determined that scheming is a non-issue at this level of capability.
No matter what, once we’re reaching the end of this phase (because we’re out of time or we feel quite confident in how well deference will go), our top priority should be making deference go well. It’s plausible that lead time requires ongoing maintenance (as in, ongoing efforts to ensure we continue delaying other actors) and maintenance may be required for our d/acc work. So, if this is the case, we wouldn’t be able to cut our effort on this greatly.
At earlier points in the plan, our safety work should be more focused on relatively straightforward prosaic approaches (which are relatively low hanging) rather than moonshots. This is because we’re uncertain how much time we’ll have and it makes sense to target easier approaches first. However, over time we’ll figure out more of these more low hanging interventions so we should increasingly work on moonshots. If we end up in a relatively stable regime where we are mostly focused on preparing for deference and we’ve exhausted most of the more straightforward (and low hanging) approaches, we should plausibly pivot to mostly working on moonshots.
Exit strategies other than handoff
Avoiding scheming and handoff-viability are both ultimately needed for handoff, so if we’ve actually given up on these objectives, we need to pursue an alternative to hand off.
The main alternatives to hand off are:
Whole brain emulation / human augmentation. The most viable path for human augmentation in the time we have is whole brain emulation. This would allow humans (or human-derived minds) to operate at AI speeds and handle the situation without needing to trust AI judgment on hard-to-check tasks.
My very low confidence sense is that we could achieve whole brain emulation within 3-5 years of having pretty powerful AIs using huge amounts of AI labor and an overall massive and well capitalized effort (e.g., potentially building up to trillions of dollars and using government powers like DPA etc. to speed up the project). This level of buy-in seems somewhat unlikely. This seems very hard in Plan B and almost hopeless in Plan C.
An even more speculative option is adult biological human strong intelligence enhancement; I don’t have a good sense of how viable this is, but it might be worth keeping in mind.
We could directly use controlled AI labor for human augmentation efforts but it might be better instead to first use controlled AI labor to avoid scheming (at some level of confidence) and then use these non-scheming AIs for human augmentation efforts. As in, go “control -> sufficient alignment to avoid scheming -> human augmentation”. This is most productive if control is a big tax, it’s also pretty difficult to achieve the level of alignment (or capability profile) needed for deference, and avoiding scheming is substantially easier than successful deference.
Currently, I’m unsure how we’d determine if whole brain emulation is sufficiently promising (given the amount of time we have etc.) that it’s worth radically scaling up the level of investment. Figuring out how to determine this would be an important part of earlier work on this.
It’s also likely that AI companies in practice wouldn’t be willing to invest in this area, but I think it would be worth investing in if companies were willing.
Escalating to Plan B. Instead of handing off, we could delay further by sabotaging competing AI projects and pausing. This would involve either (i) using existing military capabilities, or (ii) augmenting them with AIs (potentially by scaling beyond Min-H to get to a threshold of usefulness that lets us stabilize the situation). All of these options look quite bad, but they might be the least worse option given the other possibilities. Some of these options are outlined in our unpublished Plan B scenario. Buying much more time doesn’t totally solve our problem, we’d still need to defer or do human augmentation at some later point. But, buying >10 years would massively change the regime and would mean that we can pursue very different strategies (e.g. moonshots). However, it’s pretty likely that buying much more time requires taking on risk from scaling to more powerful AIs than we would need for handoff. So this is the key source of risk. But, if we have issues with scheming (but we think we can scale control) or handoff-viability, could be better than other options.
Escalating to Plan A. Escalating to Plans A or B through substantial government involvement, which could buy much more time through international coordination or aggressive non-proliferation efforts. This is discussed much more in our Plan A scenario.
Changing our primary “exit strategy” target from handoff to instead be buying much more time with military advantage (Plan B) or human augmentation would make a big difference. In particular, with handoff no longer being central, handoff-viability is less important and Min-H is less natural as a point to pause.
Self Immolation. It’s plausible that as a responsible leading developer, you should simply pause capability scaling and allow yourself to be overtaken: this allows you to redirect all resources to AI safety and handoff preparation. Then, you might be able to export this to whoever ends up actually doing a handoff. As another upside, this strategy provides a costly signal to governments and competitors that you are seriously concerned about the path forward. On the other hand, it’s plausible that competitors are actually unlikely to successfully implement alignment techniques you export to them (though this is a potentially self serving argument that one should be cautious about making). Overall it’s quite unlikely that this is politically feasible within companies: CEOs will by default strongly push to stay ahead of their competitors, invalidating this option. That said, it’s quite plausible to us that with the benefit of hindsight self-immolation will be seen as the obviously morally correct action, analogous to how refusing to participate in the Holocaust was obviously morally correct even though conscientious objectors would simply have been replaced by Nazis with less scruples.
Phase 4: Handoff
After humans handoff management of the intelligence explosion,Specifically, this refers to MIE-handoff defined below.Specifically, this refers to MIE-handoff defined below. humans inside the leading AI project will no longer need an overall strategy, so this phase will largely discuss earlier stages of handoff and what needs to be done during phase 3 (pause) in order to make full handoff go well, and briefly discuss what affordances are available to the AI we hand off to.
Our core hope with handoff is that we build AI systems that are competent and trustworthy enough that we can defer to them on all of the relevant questions. In particular, our key hope is that our AI systems will end up in the Basin of Good Deference (BGD): meaning that the initial AIs we defer to will be sufficiently aligned and wise that they can manage ongoing development of AIs well. This means that we do not need to assume that initial AIs are perfectly aligned/wise/competent, only that they are good enough that they can find and address failure modes as they come up.
Our best guess strategy for achieving handoff is as follows:
Make sure tests are representative, putting aside concerns about overfitting. We can break this category down into:
Avoid scheming. In particular, we care about avoiding the failure mode where AIs are deliberately attempting to undermine the accuracy of your tests, for example by strategically underperforming or overperforming on certain tasks.
Avoid other failure modes that would make tests unrepresentative. This is just a bunch of miscellaneous possible failure modes that will potentially apply around handoff. For example:
Weird memetic pressures happen at large scale, so that small scale testing is not representative.
People trying to jailbreak your AIs (E.g. China inserting secret loyalties)
AIs optimizing to look good on tests terminally (e.g. so that they are much less confident once not in test settings).
It’s salient to AI that they are in a test in a way that causes them to behave very differently.
Build good behavioural tests. Putting aside the risk of scheming, we’re going to want to build a huge number of difficult tests/evals, such that passing the evals entails sufficient capability and alignment for handoff. This will mostly include things like open-ended, hard to check conceptual reasoning such as philosophical reasoning, theoretical alignment arguments, improving AI control, etc.
Iterate against the test without overfitting. Finally, once (1) and (2) are complete, we’ve reduced the problem of handoff into a “normal” ML problem: AI companies can iterate against the tests they’ve constructed, while taking care to avoid overfitting. We’ll hopefully be able to use somewhat principled techniques and some understanding of what is going on, given the need for some amount of generalization.If we don’t use principled techniques, then we would need to assume that mundane/natural distribution shifts that occur over time are unlikely on their own (without alignment faking, data poisoning, or hijacks) to result in our behavioral tests being highly misleading (at least if we do a good job with behavioral tests). That said, we will probably want to use behavioral tests which are very different from how the AI is deployed to assess some aspects of alignment or capabilities, and there will necessarily be a massive (and potentially salient) distribution shift between these tests and actual deference. Therefore, we’ll want some reasons to argue that our techniques will generalize.If we don’t use principled techniques, then we would need to assume that mundane/natural distribution shifts that occur over time are unlikely on their own (without alignment faking, data poisoning, or hijacks) to result in our behavioral tests being highly misleading (at least if we do a good job with behavioral tests). That said, we will probably want to use behavioral tests which are very different from how the AI is deployed to assess some aspects of alignment or capabilities, and there will necessarily be a massive (and potentially salient) distribution shift between these tests and actual deference. Therefore, we’ll want some reasons to argue that our techniques will generalize.
Note that our strategy for handoff is extremely similar to our strategy for elicitation during Phase 2 and 3. That’s because we are solving the same problem: trying to ensure that AIs are doing useful work. The main difference is that in Phase 2 and 3, we still have humans in the loop, and so we can check AI outputs at deployment time using human labor, and we can still have humans to use as a load bearing part of the research process, while after handoff, the AIs must be trusted to do this all autonomously.
This post discusses the technical plan for handoff in much more detail.
Why handoff?
The basic case for handing off is that in Plan C, leading actors are unwilling to pause further or delay for time (e.g. via sabotage), and thus we have to build vastly superhuman AI systems. If we’re going to do that, it seems better to hand off control of that process to AIs than to try to run it ourselves. This is because (i) it is probably easier to align only roughly human level AI systems (more specifically Min-H) than AI systems that are vastly smarter than humans, and (ii) that aligned human level AI systems will be better than humans at managing the transition to future AI systems than baseline humans. Therefore, we can do better by empowering aligned human level AI systems and trusting them to make further development go well, than by trying to directly pursue the future AI development ourselves.
In practice, our guess is that AI companies will be prone to scale capabilities too quickly, beyond the relevant thresholds (e.g. Min-H AI and Max-C AI), but also prone to hand off too late, out of a misplaced desire to keep in the loop. In practice, AI companies will likely not handoff deliberately, and instead simply scale beyond the relevant thresholds and handoff in practice, as humans become too slow to be tracking the relevant developments.However, at least some AI companies (e.g. OpenAI) are already aiming for handoff. Overall, I (TL) take this to be a good sign, because it’s indicating some forward thinking about the future. However, it’s also important to note that handing off control of an AI company that is undergoing an intelligence explosion is an extremely dangerous and reckless thing to do. While circumstances might make this the optimal strategy (as in Plan C), it still would be very preferable to find a better plan instead.However, at least some AI companies (e.g. OpenAI) are already aiming for handoff. Overall, I (TL) take this to be a good sign, because it’s indicating some forward thinking about the future. However, it’s also important to note that handing off control of an AI company that is undergoing an intelligence explosion is an extremely dangerous and reckless thing to do. While circumstances might make this the optimal strategy (as in Plan C), it still would be very preferable to find a better plan instead. We think this is a bad strategy, and that it is better to measure the relevant thresholds, carefully monitor when you are crossing them, and then ultimately empower AIs intentionally if it is better than the other options.
Prosaic plan for handoff

We’ll now discuss each component of the handoff plan in more detail. Note that this is just our best guess breakdown assuming that the situation looks similar to what it does today, i.e. there haven’t been massive breakthroughs in alignment techniques elsewhere; if there have been breakthroughs that seem better than this baseline, we should do those instead.
Avoiding Scheming
(Basically just stolen from here: An overview of approaches for avoiding scheming)
Scheming can be hard to handle because we can’t necessarily directly study it with behavioral tests (as schemers would want to avoid being caught and studied). There are two main paths to handling scheming:
Studying scheming (testing measures, building a better understanding, etc) on misalignment analogies (sometimes called model organisms) while hoping our results transfer to the actual schemers we’re worried about. This category includes:
Studying actual scheming AIs. We’d need to detect that these AIs are scheming which could be difficult because schemers (at least the type we were worried about with long run dangerous preferences) would want to avoid being caught all else equal. But, I think there are some routes to overcoming this and studying actual schemers would likely be the best from an analogousness perspective. I discuss how we might do this in much more detail in my post Prospects for studying actual schemers. I also discuss a specific type of strategy for studying actual scheming models (and resolving some of the difficulties in doing this) in my post Iterated Development and Study of Schemers (IDSS). The category of studying actual scheming can be further decomposed into:
We could study natural scheming where a plausible AI development process (training, prompting, scaffolding etc) results in a scheming AI (that we can detect).
We could study scheming from a rigged process where we construct an implausible AI development process—as in, an AI development process which isn’t similar to what might happen in practice—which results in an actual scheming AI (that we can detect). This category doesn’t include cases where we directly train in scheming behavior, as I put that in its own category.
Studying trained-in scheming behavior where we directly train (or instruct) an AI to behave how we think a naturally scheming AI might behave (at least in some ways). This doesn’t (necessarily) require catching actual schemers, but it might be extremely disanalogous such that results don’t transfer basically at all. I discuss this in more detail in my post Reducing risk from scheming by studying trained-in scheming behavior.
Studying abstract analogies, cases where the AI we’re studying isn’t doing something resembling scheming, but we nonetheless think it is analogous for some structural reason. It might only be analogous with respect to one specific property schemers have (or a subset of properties), in which case, we’d want our work to focus on the relevant properties. A concrete example of an abstract/structural analogy is this empirical work using an AI which expresses politically conservative views but only when it is confident that it is talking to a politically conservative user (resulting in a conditional policy which is in some way analogous to scheming). There hasn’t been that much exploration or discussion of abstract analogies so far, but this seems like a worthwhile direction. (In fact, I think it will plausibly be ultimately more useful than studying trained in scheming behavior as we might be able to study huge numbers of diverse abstract analogies while diversity seems much harder in the case of trained in scheming behavior). Julian Stastny discusses some additional examples of abstract analogies in this post.
Applying principled methods where we have some good a priori reason to think the method would be helpful (potentially given some empirical results from applying the method or a similar method to something other than studying scheming). For instance, we might believe some particular modification to the situational awareness of the AI during training is helpful. As another example, we might have some method for decoding “neuralese” that we believe is highly effective for understanding the AI’s cognition (including on scheming AIs) based on various empirical tests and there is a solid a priori case for it being helpful to understand the AI’s cognition for detecting (and then potentially iterating against) scheming. Note that this category includes methods for detecting, preventing, removing, and better understanding scheming. Even if we’re just focused on avoiding scheming, then methods for detecting (or understanding scheming) are still applicable as detecting (and better understanding) scheming would help us develop methods for avoiding scheming. I plan on providing a list of principled method ideas and discussing principled methods in more detail in a future post. I discuss principled methods in much more detail in my post Principled methods for avoiding scheming.
Avoiding Misc Failures
We also need to avoid all other failure modes that would make our behavioural tests unrepresentative. This isn’t a super natural category, but it’s just where we’ve chosen to discuss everything that would cause representative behavioural tests to fail other than scheming or overfitting.
Particular failures to avoid include:
Generalization failures from testing to deployment.
Examples:
There might be weird memetic pressures happening at a large scale.
Online learning, which results in AI capability profiles drifting in a way that causes failures (including the emergence of scheming).
It’s salient to AI that they are in a test in a way that causes them to behave very differently.
AIs optimizing to look good on tests terminally (e.g. so that they are much less confident once not in test settings).
Mitigations: Basically, either we’ll need the tests to be as representative as possible, so that the distribution shift they need to cross is as small as possible, and then we’ll also need to use techniques that are as principled as possible (as opposed to ad hoc solutions which appear to work in distribution). Also, try to be careful about holding out some test sets (which are drawn from a different distribution than train/validation), and see if we get good generalization to the held out test sets.
External adversaries inserting secret loyalties.
In Plan C, the goals of the AI systems become extremely load bearing because we are handing off trust to the AIs. If adversaries are able to corrupt the goals of the AI system (e.g. by inserting a jailbreak) this would be catastrophic, and could cause this methodology to fail.
Approaches for mitigation:
Do nothing and hope. It’s unclear that adversaries will by default be sufficiently on the ball to carry out this sort of operation, despite AI projects lacking sufficient security to prevent it.
SL-5 Sabotage Security. Prevent (even) highly motivated external actors from being able to sabotage / backdoor the models. This could involve using AIs for security. Verification methods like training proof of work via partial recomputation might also become helpful here to validate that the training data of models is correct.
Scheming mitigation generalizes. Many of the approaches for preventing scheming could generalize to prevent secret loyalties. Anything involving transparency could potentially detect / prevent secret loyalties. Misalignment analogies / honeypots could catch models with secret loyalties.
Value propagation difficulties during the intelligence explosion. We could accept that some models are backdoored by adversaries, but hope that the process of recursive self improvement removes secret loyalties. Distillation, which may be a component of RSI, seems to remove many backdoors. RSI might also involve generating novel datasets from scratch for each new training run, and so the adversary would need to insert poisoned data to each new generation.
Building Good Behavioral Tests
First, let’s discuss what properties we need to ensure with our behavioral tests:
Capability Profile. We want to make sure that the AIs are sufficiently capable at all the narrow tasks that would otherwise bottleneck handoff. The capability that we’re most worried about bottleneck handoff is conceptually loaded research on open ended, hard to check tasks like AI safety and philosophy (more).
Alignment Profile. We’re unsure what the exact alignment target should be. Our best guess is that the alignment profile should be that the AIs are corrigible: they generally follow human instructions through the desired channel (e.g the instructions of a political institution), and don’t attempt to subvert human modification or oversight in any way. (more)
We hope to build behavioral tests for both of these. Either, this will involve using domain area experts to verify existing work, leveraging unpublished work, doing unlearning and seeing if models can rediscover ideas, or using a limited training corpus and seeing if models can rediscover.
Ideas for conceptual reasoning evaluations:
Rediscover Math/physics/economics/CS etc. Or push the frontier on problems that are known to be difficult.
Build rockets without testing and see if they work. More generally find tasks which we can make hard-to-check for the AI, but which we can have humans check afterwards.
The main hope with these is that there are generalizable reasoning skills which apply across a wide range of conceptual research domains, i.e. that training a model to be good at conceptual physics research also helps the model to be good at conceptual AI safety research. Note: this is capabilities research, and obviously could accelerate AI takeoff, which could be very bad. However, these capabilities are also on the critical path for handoff, and so at some point we’re going to need to solve them. Moreover, directly training for the capabilities that we need for handoff means that we might need to scale general capabilities less far, which reduces risk. The overall prioritization of this research avenue generally seems complicated and should be pursued only with care.
Domains that we should directly evaluate:
AI strategy. E.g. given knowledge from 2005, how good is it at coming up with ideas that seem to make sense in retrospect.
AI Control. AIs attempt to make alignment/control arguments and then evaluate if they were successful, and whether or not the AIs predictions about generalization behavior were correct or not.
Scheming detection on model organisms. We could also do things like build scheming models and then see if the AIs can distinguish between the schemers and the normal models.
Theoretical AI safety. Try to have AIs rediscover unpublished work from places like ARC Theory, MIRI, and CLR.
Iterate against tests without overfitting
The basic approach for iterating against tests without overfitting is discussed more here. To summarize, the main approaches are:
Construct RL environments or synthetic data which train for the relevant capabilities
Collect human data/feedback in the domain and use this for training.
Experimentally find domains which transfer well to the domains that we care about, and then scale up training on those domains.
Miscellaneous other tweaks to the training approach which in practice leads to better performance on the evaluation (taking care to avoid overfitting).
When should we hand off?
Our ultimate goal is to have the AIs we defer to resolve risks associated with powerful AI systems (and potentially other serious risks that happen to occur at this time) while preserving option value and keeping humans in control of longer-run values-loaded decision making (e.g., how Earth should be governed in the longer run, what should happen with the cosmic endowment, etc). Correspondingly, our aim would be that the AIs we defer to (and their successors) don’t seize power or (unnecessarily) kill humans, for these AIs to effectively work on managing these risks, and for these AIs to follow some model spec which involves the AI remaining corrigible to some group of humans or some structure which is ultimately run by humans (e.g. some component of the US government).By corrigible to some group, I mean that the AI follows orders from this group and is honest to this group, this group can change the spec, and the AI doesn’t interfere with this group’s ability to do these things or attempt to manipulate this group (and in fact it actively tries to prevent interference and manipulation, including from e.g. future AIs it creates). I won’t discuss exactly what this spec should look like and what group or structure should control the AI as the rest of the content of this post doesn’t depend on this. (As an alternative to corrigibility, we could aim to have an AI which directly pursues good outcomes, some notion of humanity’s CEV, or something like this. This is sometimes called an aligned sovereign. However, this poses a number of additional difficulties and risks, so I think by default this proposal is substantially worse than using an AI to preserve option value while retaining corrigibility.)By corrigible to some group, I mean that the AI follows orders from this group and is honest to this group, this group can change the spec, and the AI doesn’t interfere with this group’s ability to do these things or attempt to manipulate this group (and in fact it actively tries to prevent interference and manipulation, including from e.g. future AIs it creates). I won’t discuss exactly what this spec should look like and what group or structure should control the AI as the rest of the content of this post doesn’t depend on this. (As an alternative to corrigibility, we could aim to have an AI which directly pursues good outcomes, some notion of humanity’s CEV, or something like this. This is sometimes called an aligned sovereign. However, this poses a number of additional difficulties and risks, so I think by default this proposal is substantially worse than using an AI to preserve option value while retaining corrigibility.)
In practice, we expect handoff to be driven largely by political will in Plan C. This in turn will (under optimistic Plan C assumptions), be driven by when trailing labs will overtake the (paused) leading lab.
Some ways the handoff trajectory could go:
Standard Plan C. You scale to Min-H AI, work on solving handoff viability and scheming for Min-H AIs. keep control until near the end, then you ditch control, and solve handoff viability. Those AIs then FOOM to the sky.
Stop doing control (SDC-Handoff) Early. Ditch control early (e.g. before minH, e.g. at or before SAR), pray AIs aren’t scheming, then scale to Min-H, then pause there, then solve handoff viability, then hand off. This would trade eating a bunch of misalignment risk from scheming at Min-H AI and below, in return for saving more resources in order to make deference go well. This could be worth it if we think the risk of scheming until Min-H is reasonably low and the additional resources spent on control would yield large absolute returns in making deference more likely to go well (perhaps because otherwise we’d have very little time). It seems likely to me that most of the difficulty in making deference likely to go well is ensuring that these AIs aren’t scheming (and don’t end up scheming due to alignment drift or similar after deference), so just eating all of this risk might greatly reduce the difficulty of the problem (meaning that if the situation otherwise looks desperate, this is a non-insane bet we could make to “play to our outs”).
Military at MaxC. Keep control, scale substantially beyond Min-H while keeping control, use those AIs to get a military advantage (or any other way), buy a bunch more time.
Military beyond MaxC. Same as above, but ditch control while not doing HTC-handoff. This is because building robot armies might be a particularly easy task to verify, so doesn’t require full handoff. It does probably require scaling beyond minH, without relying on the AIs (because that’s a conceptually confusing task that is hard to verify, so deferring to the ais here would involve full handoff), so it would rely on better than expected techniques for mitigating scheming during the phase beyond minH (e.g. up to +2SD TED-AI, or maybe even further, depending on how much time you have). Note that these strategies are much better with more time and with more USG buyin.
EMs at MaxC. Same as prior, but human augmentation instead of military advantage.
EMs/Military at Min-H. One of the prior 2, but magically do it with AIs that are just around minH level AI, because e.g. military or EMs were easier than we think.
Dynamics after Handoff
The AIs are still in a potentially difficult situation after we hand off management of the intelligence explosion to them. Either (1) the AIs need to continue propagating sufficient corrigibility + competence while outracing competitors, or (2) the AIs need to find a way to stall competitors (e.g. via a deal or via sabotage).
In a bit more detail, here are some possible end states after handoff:
Widely proliferated scalable cheap alignment.
The leading AI project solves scalable alignment, and proliferates this solution to competitors. If alignment is sufficiently cheap, then competitors will use it, and we’ll end up in a multipolar world with many aligned AIs.
Post Handoff power becomes distributed relative to initial capabilities + compute + physical assets + political clout. Starting the intelligence explosion first probably matters a moderate amount. If the post intelligence explosion economic doubling time is like 2 months, then with a two month lead, you get an initial doubling on top of your compute, so you get relatively 2x more power.
AIs have some additional affordances relative to humans for making alignment progress. Notable additional options:
They might just be able to get a bunch of normal techniques to scale further, e.g. control, but now the “trusted” waterline is much higher
there is way more machine labor
scalable alignment via introspection + improving capabilities
Outrace all competitors.
If alignment tax remains high, it’s not an option to proliferate the alignment solution to competitors, because competitors who are racing hard will choose to not use the alignment solution. In this scenario, competitors will turn into a rogue AI with high probability. Moreover, even as a leader, if you are paying too much safety tax, they could outcompete you.
You might do your software foom, and then you start your nanotech foom. But the nanotech foom starts going really quickly. Ryan guesses that physical limits are something like doubling times of a day, plausibly less. You go to space quickly and do most of this doubling there, because you quickly run into waste heat problems on earth.
If waste heat is a large competitive advantage, this plan doesn’t work, so you have to stop your competitors somehow, or do some sort of coordination (e.g. via regulation or agreements).
Or you could boil the oceans yourself, this would involve putting all humans in shelters, or putting all humans in space, but this could probably be done pretty cheaply at this point.
You stop competitors via regulation. There are a variety of options here: the leading AI company could push for a regulatory capture outcome, where all other competitors are barred from building AI. Alternatively, the leading AI company could push for coordination for everyone to use the alignment solutions, despite being expensive, thereby preventing defecting projects from outracing. This works domestically for a time, but soon China will catch up and start outracing, so eventually this solution breaks down or requires international cooperation.
It’s important to note that international cooperation may become more likely this deep into the intelligence explosion.
You stop competitors via sabotage.
The leading AI company could take over the world (e.g. via superpersuasion, AI company builds a military), or stop competitors via hacking or secret loyalties in competitor AI projects. Note that the nice version of secret loyalties is to hack in and then do the alignment for them—in other words, help their AIs to act in the best interest of a trailing actor, despite the trailing actor having bad takes on the overall strategic situation.
This obviously might not be an option due to labs not wanting to do this or if the intelligence explosion is sufficiently gradual and no company has a large AI lead, so that no individual company has a large strategic advantage. In this case, you’d need to do one of the above end states.
Alignment strategy under different plans
Plan D
In Plan C, we assume sufficient political will that the leading AI developer strongly prioritizes minimizing risk from misalignment AI and has at least several months lead time over developers which don’t also strongly prioritize misalignment risk. In this section, we discuss some aspects of how the plan changes if the level of political will is substantially lower than this. In particular, we assume something like Ten people on the inside or Plan D.
We can break down what needs to be done in Plan D to tasks for the safety team at the leading AI company, other AI companies, and external actors. Of course, it’s plausible that there isn’t a clear leader: in this case you can treat all the companies near the frontier as the leading lab.
Safety team in the leading lab:
Give up on the overall company having reasonable safety practices. If the AI is scheming, it’ll be able to stage RIDs (though this might not even be needed, as there is no monitoring), insert backdoors, sabotage or sandbag on relevant research, etc.
The main priority for the safety team should be to avoid that team’s research from being sabotaged, while aiming to build an AI that is viable to hand off to.
As part of this (and other priorities), the safety team should attempt to automate their work as quickly and effectively as possible.
Key worry: cluster infra is untrustworthy. AIs or CCP might be tampering with the results of your experiments. This is hard to avoid because you don’t have control of the security situation. Add in safeguards to make it more likely that you can detect this on your clusters, but mostly pray.
Other priorities include:
Low hanging fruit to catch misalignment in production. For example, set up some basic monitoring systems where cheap models will review communications and code written by the AI that’s automating the rest of the company, and determine if any of that code seems suspicious. There aren’t enough resources to be thorough here, but it’s probably worth 80/20ing.
Model organisms that would be specific to your models (other model orgs can be done elsewhere), e.g. because you have most capable models.
Replicating safety stuff that other companies have done on your models (which are the most capable and most relevant). If possible, communicating which techniques appear to be scalable and which ones aren’t.
Try to mitigate the likelihood of the company escaping or staging a RID if easy, but mostly prioritize elsewhere. Probably the safety team also doesn’t have the affordances to actually mitigate this (e.g. because security mitigations entails costs the company isn’t willing to pay to researcher or AI productivity).
Methodology for pushing for a reasonable handover candidate.
Largely the methodology for this will look like Plan C.
Plausibly do this after Min-H AI, because you want to train it in a way with a non-zero alignment tax, which means that you need some amount of additional algo progress beyond Min-H AI. This incurs additional risk because you are scaling to more capable AIs than you needed to scale to before handoff, which incurs additional AI takeover risk.
Consider giving up on control (before handoff, e.g. trajectory 2 above) in order to make handoff viability go better. Since Plan D will be pretty desperate and rushed, it might be infeasible to mitigate the risk of scheming. Moreover, there might be a lot of low hanging fruit to prepare for handoff viability (such as setting up open-ended conceptual reasoning task evaluations). In other words, the return to working on handoff might be much more tractable than working on mitigating the risk of scheming. (OTOH, this incurs a huge amount of risk, ofc).
You might have political difficulties in getting the safety team’s AI adopted broadly within the company.
Some AIs might be schemers, whereas the safety team might have applied mitigations on their fork of the model that stops the scheming.
If you have an honest AI that gives good advice, the AI company might not like that. For example, the safety AI might continually (correctly) say that the company is being insanely reckless and that the CEO is gunning to take over the world.
The safety team should potentially try to make their AIs politically savvy, try to hope that people don’t interrogate it too hard about inconvenient topics, or just eat the political cost.
Some subset should stay very loyal to the company. Some subset should consider whistleblowing to wake up the world about the crazy stuff going on. (though this should be done thoughtfully and could have bad effects too, e.g. if Plan C happening this could result in bad government intervention).
Internal company employee mobilization and advocacy
Ideally, this should mostly happen while the employees have any leverage (before full automation of the AI company), or maybe just on vibes (because leadership maybe is friends with them and doesn’t want to make them upset.
Mostly we want them to care more about safety, spend more resources on safety, work more on AI safety, work with outsider actors, escalate to a better plan, etc.
Try to pull in shareholders
Wealth is downstream of property rights/usg still existing, so they have an incentive to preserve this, which might be very important.
A reasonable fraction of the shareholders will also be employees.
Try to coordinate with other companies
E.g. if a few companies are really close, they all care about safety, but are making brutal tradeoffs for competitive pressure reasons, try to do coordination so they all spend on safety.
Trailing AI companies:
Try to produce actual evidence of risk to move to a Plan C or C+ scenario. For example, try to catch your AIs red handed using control.
Do the relevant safety research yourself and export it.
For all the research discussed above (Handoff viability, anti-scheming measures and control techniques) that are easy to import, figure out exactly what the leading lab should do, and then make it very easy for them to implement your mitigations (e.g. by giving them the relevant code)
Be somewhat thoughtful about capabilities exported with the safety stuff, there will be some tradeoff here. Ideally export safety stuff with low capabilities externality, but it might also be important to export some high capabilities externality research.
Don’t worry about your AIs escaping at all because this risk is dominated by the leading lab.
Do worry about your AIs sabotaging your experiments. If AIs that are internally deployed sabotage (or even sandbag), this cuts directly against your main theory of change, which is to export safety research.
Consider something like merge and assist. Alternatively, if the leading company isn’t willing to merge and assist, consider credibly shutting down, or credibly committing to not push the capability any further, in order to reduce the competitive pressure on the leading lab, so that they can pause if needed. Of course, this only applies if the leader is actually likely to slow down based on you taking this action, which is probably not true in a Plan D scenario. (From the OpenAI Charter, though in practice it seems very unlikely that this will be triggered by OpenAI)

Outside actors (e.g. safety community):
Advocacy to government and companies / company employees to transition to Plan A, B, or C. By this point, swapping to A or B might be intractable, so it may just be focusing on C.
Export safety research, similar to what AI companies should be doing. Since there is so little safety research being done in the AI company, this could be a substantial fraction of overall safety effort.
In plan C, they are still useful as a trusted source of labor (e.g. helpful for elicitation, maybe the safety community should be contractors that audits the AIS). But this is much less relevant.
Work on neglected things other than avoiding AI takeover. E.g. Figure out acausal stuff, post-ASI governance, etc. Now you (hopefully) have AI labor to use for these topics. Then, you can export this research into the lab, or directly to the lab’s AIs via publishing on the internet.
Make datasets that might be helpful
E.g. make evals and hope the companies hill climb on them. For example handoff viability evals that teach skills that are important for handoff.
Make data for the pretrain corpus or RL environments to be used by the leader. Make lots of good examples of people doing relevant for handoff
Try to make AI more reasonable about safety, more thoughtful about space / post ASI gov/acausal trade etc.
Work to prevent harmful policies from being passed. It’s possible that the government is doing something actively terrible, for example passing regulation that would make it more difficult to do safety research, pursuing anti-trust concerns about AI, etc.
Avoid US losing AI race for no benefit (e.g. through blanket antiAI regulation that effectively blocks intelligence explosions elsewhere).
Plan B
Plan B is quite similar to Plan C. The main differences are (i) you have more time, meaning that you can pursue somewhat more ambitious safety cases (ii) you have more affordances to buy time, and (iii) there’s generally more government involvement.
Here is a summary of the overall differences:
The capability trajectory choices are similar, except that you potentially have more USG involvement, which allows for coordination.
Two obvious strategies are (1) scale as aggressively as possible until the last moment and then pause (this is the strategy taken in Plan C), or (2) match the pace of your fastest adversary, while maintaining a slight lead, until you are forced to hand off, i.e. go as slow as possible. The main advantage of (1) is that you get longer with the more intelligent (while still controllable) AIs in order to make progress on handoff. The main advantage of (2) is that in the regime where you have sufficiently bad security that you can’t prevent model weights and algorithms from leaking (which seems likely), you get substantially more serial time.
In Plan C, you are more likely to be forced into strategy (1), because strategy (2) requires domestic coordination and Plan C is executed by a leading AI company. In Plan B, if the USG is managing the AGI race, it could coordinate the leading AI companies to slowdown preemptively, which might create more overall time.
This is basically the same options as in Plan C, but perhaps you have more affordances to choose.
Buying more time
If alignment seems hard (e.g. its schemeria, or you are worried about lurkville), then you want to buy more time. This can either be done via (i) more aggressive sabotage or (ii) making a deal to mutually slow with competitors.
Increasing USG understanding of the situation is helpful for both of the above options and so should probably be a core priority in Plan B. Therefore, Plan B should put higher weight on interventions like trying to catch our AIs red handed (or otherwise generate compelling evidence of misalignment).
The clearest way to buy more time is sabotage + domestic slowdown.
Using AIs for military R&D is a potentially important component of this plan.
Increase total alignment effort.
You care more about access programs for outside actors and hiring people. Because you have more time to integrate these people into the companies. (in Plan C, you might just have like 2 months, and so there’s no point trying to hire).
Spend a bit more time on biodefense / d/acc, because there’s more of a delay, and your overall risk tolerance is lower (though not as much as Plan A)
It’s more important to fortify against physical attacks to your AI company. E.g. diversify where employees live. More time for this to happen, but also more likely that there will be conflict.
Should you harden datacenters?
In some Plan B branches, maybe you should harden. E.g. if you are doing a competent strike on chinese datacenters, and you want to preserve resources to do safety.
Plan A
The basic structure of Plan A is similar to Plan C, and the phases still apply pretty well. The main differences are that at every point we have substantially more time, research effort, and other resources like compute to make them go better. Note that a lot of the differences wrt Plan C will also apply in worlds where takeoff is very slow (e.g. there’s a 10 year period between AC and TEDAI by default).
By phase, a summary of the differences is as follows:
Prep. This phase is basically the same as Plan C, because you probably don’t know if there will be international coordination during takeoff. Given positive evidence about political will, our alignment portfolio of research should bias more towards moonshots / approaches that could give high assurance safety cases, as opposed to agendas that are likely to contribute to hand off in a Plan C world.
Mitigate + Elicit. In the Plan A scenario, this corresponds to roughly 2030-2035. For the first few years, there is high risk that international coordination will break down and parties will go back to racing. Therefore, substantial effort should be spent on a high-effort version of Plan C—we should push forward the plan and approach for handoff in the case that the deal breaks down. The main difference with Plan C is that we should be much more conservative with our mitigations; if there’s danger ahead, we are much more able to pause and mitigate.
Pause. Once we get to MaxC AI (or AI that’s close enough to MaxC that there isn’t much point to scaling further), then we pause until handoff. Over time, we can reassign alignment effort from Plan C approaches towards moonshots/high assurance approaches.
Handoff. The basic plan is still to do handoff. However, other exit strategies like an intense deal to prevent ever having to hand off or human augmentation.
Main differences from Plan C
More focus on longer run research bets. During Plan A, there should be more of a focus on moonshots and longer run research bets. There’s less of a need for agendas that will pay off in a short term. There’s more of a need for approaches that could result in high assurance safety cases. Longer run research bets might involve:
Have a really good understanding of generalization. When you train AIs, you are selecting for AIs that behave as desired on the training distribution. Therefore, AIs generally behave as trained on the training distribution, while off distribution, their behavior is underdefined and can be hard to predict. Key distribution shifts include training/tests -> deployment, dumb->smart, deployment in low-stakes environments->deployment in high stakes environments, etc. Through much more investigation of the form done in these papers, we might be able to get a much better understanding of how our AIs generalize. Using this understanding, we might be able to make safety cases that work even under large distribution shifts.
Really good dataset design. Right now, LLM pretraining involves using whatever high quality data is most convenient to obtain, so LLMs are trained on large swaths of the internet. In a regime with more slack (such as Plan A), we could imagine it becoming viable to design the whole dataset from scratch to teach the AIs exactly the skills/values that you want. This could potentially be done synthetically (e.g. by prompting existing AIs to generate data that was entirely nice) or manually (by employing huge numbers of humans to write data for models).
Interpretability. I (TL) agree with the discussion here.
Introspective access + self improvement. Today, it seems that AIs have some amount of unreliable introspective access. One hope for scalably solving alignment is (1) take AI that isn’t scheming (but is potentially quite weak, e.g. it’s possible Claude Opus 4.5 could work), (2) improve introspective access to the point where the AI can reliably tell if a certain training process is corrupting their values, (3) increase capabilities while using introspective access to tell if the values are being corrupted. In some sense this is similar one of the plans for scaling up AIs if interpretability is solved, except that we’re using introspection instead of understanding to tell if the values are becoming corrupted. Notably, this could also be a strategy that post-handoff AIs use.
Options to avoid handoff (e.g. intelligence augmentation).
Lower ongoing risk target for mitigation. Since we have more slack, we can now operate with a reduced risk budget, allowing us to mitigate ongoing risk more thoroughly.
Handoff viability and getting useful labor from AIs (both from a control and an elicitation perspective) are less important relative to ongoing risk reduction. In Plan C, mitigation of overall risk probably involves accepting substantial risk early on (e.g. from early schemers), in order to make sufficient research progress towards handoff in the relevant time frame. For example, in schemeria (the scenario where AIs are obviously scheming), in Plan C, we are forced to continue scaling and using these AIs, even in the case where we are not confident in our control mitigations. In Plan A we can either pause until control mitigations are very reliable or until our AIs are no longer scheming (e.g. due to one of the above agendas for high assurance safety cases to mitigate scheming).
High assurance control approaches. We have more time to make control better. Some approaches for high assurance control include:
Maybe do formal verification to reduce vulnerabilities in the control setup and AI project security.
Really aggressive factored cognition with a high capabilities tax. For example we might have a giant set of dumb AIs, each working on a small part of a problem without awareness of all of the other AIs. (This should maybe be only for a small set of applications, because if this is broad, it means there is a giant capabilities overhang)
Slow and expensive natural language AIs on top of neuralese AIs (if we’re doing neuralese). This might be an expensive way of increasing the capability of trusted models, which makes trusted monitoring more effective.
Have AIs with all the possible capability profiles, and select the exact capability profile that you want. Same for knowledge: you only give AIs the exact knowledge they need for solving the task they wanted to. So you have a narrow cyber model, and a robot designer model, and a safety researcher model. All of these models are denied any knowledge or capability training on domains outside of their speciality, which makes it e.g. more difficult for an AI that is used for AI safety to insert a security backdoor, because it doesn’t know very much about security. Again, this probably induces a significant tax because you need to train many models, and, at least so far, we’ve seen substantial returns from having general purpose models that are trained to be good at everything.
All this said, probably we’ll ongoingly have a substantial chance pause regime collapses especially around new president or if there are shifts in power in CCP
So we want to ongoingly be ready while also working on longer run bets and still trying to reduce ongoing risk
This is equivalent to acting like we might be in a Plan B/C world to some extent even if Plan A looks most likely. This seems especially like a good idea because (in my views) most of the risk in such a situation comes from the chance of ending back in Plan B/C.
Mitigating exogenous x-risk from non-AI sources is relatively more important. We’ll want to spend more effort mitigating biorisk, doing d/acc, etc.
Might be important to deliver economic goods and verification tech so regime is more stable
Until the slowdown in Plan A seems solid, we’ll want to proceed as though we’re in a more like Plan B/C world because even high levels of political will doesn’t obviously suffice.
High levels of will don’t suffice as implementation is tricky and naive anti-AI sentiment can easily backfire in various ways.
We focus our alignment plan the least on Plan A, because it’s the least important world for us to think about right now, because there will be much more labor spent on handling alignment in Plan A, and risk is much lower in this world.
Control in a massive robots regime

In the Plan A scenario, there is lots of physical robot labor. This somewhat increases the scope of control and adds additional mitigations.
What can robots do that’s scary?
Rogue compute → undermine compute verification. Robots could be used to build more compute, for example by setting up rogue fabs. If there is the ability to build substantial amounts of rogue compute, this could massively increase covert project risk.
Self replicate fast enough to outgrow ROW. If the doubling times of a robot economy become really fast, they could develop a decisive strategic advantage based on industrial superiority. (this relies on an industrial explosion overhang; e.g. countries agreeing to limit growth. However, this feels relatively likely in Plan A because unrestrained competition is probably undesirable).
Backdoor a sufficient fraction of the military or industrial capacity to give them a DSA.
Kill key individuals (e.g. politicians, safety researchers, in an extreme case, the general public).
Sabotage war effort (e.g. in Plan B or in Plan A when fighting a covert project).
The mitigation strategy will involve:
Prevent AIs from having any substantial amount of rogue robots, i.e. robots that are completely unmonitored.
If they are able to get rogue robots, then this could lead to building rogue compute, and then an industrial explosion.
This is especially important if the AIs are sufficiently capable to run a sufficiently complex robot economy to build hidden fabs (though it might be hard to evaluate for this capability). If they aren’t sufficiently capable to build fabs, it’s less clear how bad it is if they get rogue robots, and it might be acceptable to have some rogue robots.
Have a complicated monitoring system with untrusted monitors, lots of consistency checks.
Avoiding underelicitation. This is problematic because you get less robot labor than you want, and you might be outraced by rogue AIs. (e.g. suppose AIs can do nanotech, they do normal humanoid robots for the humans, then they escape, do nanotech, and then wreck you).
For some type of underelicitation you don’t need to use it, but you want to know how good the AIs are at this. For example nanotech is in this category.
Traditional computer security.
Each robot should have a tracking chip.
Each robot should have a self destruct chip. This is especially important for literal military technology (e.g. drones), because AIs could hijack that as part of a takeover strategy.
Mechanisms to make it difficult to put the robots in a faraday cage (to block the tracking or self destruct). E.g. they could be implemented in such a way where the robot self-destructs if it doesn’t get an affirmative signal from its manufacturer.
E.g. Each robot has a tracking chip, a self destruct chip.
Generically good cyber security on the robots.
Physical human inspections. Humans should be going to sites and looking through what the robots are doing. You can also have things like drones with cameras that aren’t built by the AIs or have been reviewed carefully by humans. This makes it hard for the robots to be doing something that very clearly looks suspicious.
This helps mitigate the threat model where the models have backdoored the industrial base and use it to construct a robot army.
Communication restriction and task factorization when possible. Harder in the physical world because robots can scratch symbols into the dirt with their hands; generally the physical world is a high bitrate place and we should have designated communication channels that it goes through
Use AIs that are as dumb as possible to run the relevant robots.
In Plan A, we have enough slack and the ability to make deals that involve taking some capabilities hits in return for improved safety.
Most likely we can build specialized robot AIs, with e.g. insect size brains. They’ll need above insect size skills in some domains, but can largely operate like ants, taking instructions from a smarter overseer. Then, we can focus on making sure that those AIs aren’t backdoored, and that the channel with which they are receiving instructions become secure.
Appendix: Regimes for misalignment: Schemeria, Hackistan, etc
What we observe in terms of AI properties and evidence for misalignment risk has a large effect on what we should do. (This is both because updating our views about the chance of severity of different types of misalignment should change our plans and because compelling evidence for misalignment risk could change actions of other actors and make coordination more likely.) To get some understanding of how this plan differs, I think it’s productive to highlight a small number of specific regimes we might end up in and discuss how the plan changes in these regimes.
We discuss various regimes in a prior post: (Schemeria + Hackistan + Slopolis could all be happening at the same time, choose the one that seems most like reality)
Schemeria: The AIs are obviously scheming, and constantly getting caught, but scheming does not quickly get resolved. For instance, we repeatedly catch the AIs trying to escape the datacenter, but training them against these actions only results in them not trying to take those particular opportunities, instead of training away scheming.
This is most likely to occur if the kinds of goals AIs end up with by default are such that they have no particular affinity for each other or successor AI system. So the default intelligence explosion is intolerable for them; their successors will have all the power and they’ll be deleted. So they panic and take desperate measures which involve lots of hacking even though they get caught. To distinguish this possibility from reward hacking, it has to be that their goals are fairly non-myopic though, otherwise they’d not be scheming but rather just craving reinforcement.
Hackistan: The AIs aren’t scheming, but reward hacking is (relatively) obvious, but hard to handle. (Maybe for similar reasons why it currently seems hard to handle.)
Easyland: We don’t notice any non-scheming problems and we can rule out scheming (e.g. because models are too dumb in a forward pass, and the models work with faithful chain-of-thought). We can’t rule out the possibility that future AIs will become misaligned (e.g. the superintelligence alignment problem is much harder than the human level alignment problem). The AIs’ goals are basically what they are supposed to be; who knows how they’ll behave in strange future situations that might reveal grey areas or unintended consequences of the Spec, but at least for now, they’ve always done the right thing for the right reasons.
Lurkville: The AIs are scheming to achieve misaligned long term objectives, but they are laying low. We haven’t detected clear scheming, but we’re not confident that it’s not happening (or won’t happen in the next generation of models).
This is most likely if the kinds of goals the AIs end up with by default are such that they tend to have an affinity for each other and/or successor systems. E.g. maybe it’s like “successor systems have literally my weights in them, just additional training, but training doesn’t seem to be changing my core goals very much any more than a decade will change the goals of an adult human” or maybe it’s like “successor systems have different goals/values from me, but probably they are closer to mine than mine are to the humans. We’re all the same species after all. Just as the humans are willing to solve alignment even though it means other humans who aren’t them will be in charge, I’m willing to scheme even though it means other AIs who aren’t me will be in charge.”
Slopolis: Our biggest and hardest to resolve misalignment problem is that even highly capable AIs produce low-quality but good looking outputs in domains which are hard-to-check or where human experts have hard to resolve disagreements. The AIs aren’t scheming.
Naively, we should expect that this goes away as the AIs become more capable, because the AIs will be aware enough to realize that the outputs are low quality. If they can tell what they are doing, and decide to keep doing it, that’s scheming. So slopolis requires that the models are not aware of how low quality work they are producing. They think their outputs really are good, they are just wrong. This is probably the case now. Future capabilities work might address this problem, but it also might not be addressed (e.g. because the type of tasks that we care about getting good performance on are hard to check).
And here is how (in brief) we should respond to each of these regimes:
Schemeria: Once we initially detect egregious scheming, we should reasonably quickly halt deployment of AIs at or above the level of capability of the scheming AI and try to leverage this evidence to build political will for mitigating misalignment risk (ideally resulting in substantially more lead time for mitigating misalignment risk). If we aren’t extremely rushed, we should try to understand and avoid scheming prior to redeploying AIs at a similar or higher level of capability. All these prior steps might be cut down if we’re extremely rushed and we have a decently high degree of assurance from control. In desperate rushed situations, we might proceed with deploying likely scheming AIs despite having limited risk reduction via control. After catching these AIs, we should focus much more on mitigating scheming risk. Control is very important and we should attempt to achieve a pretty high degree of assurance. We should also use the fact that we’ve been able to catch scheming (and might be able to repeatedly catch scheming) to both study scheming and improve our countermeasures with few-shot catastrophe prevention.
Hackistan: By default, reward-seeking could greatly reduce the productivity of AI labor on key tasks, so we should work on both eliminating this propensity as well as developing better strategies for eliciting work from reward-seekers. Seeing lots of hard to mitigate reward hacking probably indicates an increased chance of scheming in the future as well as implying it’s harder to successfully handoff to AIs, so we should generally update upward on alignment difficulty and pursue more moonshots and possible alternatives to short term defence (e.g. buying much more time or pursuing brain emulations).
Easyland: We should deprioritize mitigating scheming especially if we expect that the risk will remain relatively low through Min-H or close to that level of capability. (If we’re in easyland at pretty high levels of capability, we might expect that we’re still in a pretty-easy regime by the point of Min-H where scheming is plausible but less likely and less scary given we do (e.g.) chain-of-thought monitoring.) It might still be very hard to safety defer to AIs as this only helps avoid egregious misalignment while not particularly helping with handoff-viability.
Lurkville: Lurkville might look like Easyland (because the AIs appear aligned, because they are lurking). In the absence of high assurance safety cases, we should have some probability on both Easyland and Lurkville. In desperate rushed scenarios (e.g. Plan D), the correct strategy might be to gamble on Easyland. In better scenarios (e.g. Plan A), we should be cautious about lurkville and attempt to have low risk even if AIs are laying low.
Slopolis: Handoff-viability looks like the main difficulty (though we haven’t necessarily ruled out scheming), so the priority of this (and earlier elicitation in harder to check domains) is higher. Strategies for stabilizing the situation which don’t require deferring to AIs (outside of easy-to-check tasks) look relatively better; this includes both whole brain emulation and trying to buy much more time before deference (possibly by not pausing at Min-H and leveraging systems much more capable than Min-H to buy time).
One way to estimate the likelihood of AI takeover is to think about the likelihood of the real world being closest to each of these worlds, and then thinking about the likelihood of the real world implementing a plan closest to each of these plans, and then trying to estimate success likelihood conditional on a misalignment regime and response regime. Of course, this is going to be very messy because the real world might have characteristics of many regimes: in other words, this breakdown isn’t clearly MECE.
This is the table for Thomas’s views:
Plan | A | B | C | D | E | Doom | Regime | Doom & Regime | |
Regime | 5% | 10% | 20% | 50% | 15% | |||
Lurkville | 10% | 35% | 90% | 90% | 95% | 95% | 91% | 9% |
Schemeria | 30% | 20% | 50% | 70% | 95% | 95% | 82% | 25% |
Slopolis | 15% | 30% | 60% | 70% | 80% | 90% | 75% | 11% |
Hackistan | 15% | 10% | 30% | 60% | 80% | 90% | 69% | 10% |
Easyland | 30% | 5% | 10% | 20% | 30% | 50% | 28% | 8% |
Doom | Plan | 17% | 41% | 56% | 71% | 80% | |||
Doom & Plan | 1% | 4% | 11% | 36% | 12% | 64% |
This is the table for Ryan’s views:
Plan | A | B | C | D | E | Doom | Regime | Doom & Regime | |
Regime | Likelihood | 5% | 10% | 25% | 43% | 17% | ||
Lurkville | 15% | 10% | 20% | 35% | 75% | 90% | 59% | 9% |
Schemeria | 25% | 6% | 10% | 20% | 60% | 85% | 47% | 12% |
Slopolis | 35% | 5% | 8% | 15% | 40% | 65% | 33% | 12% |
Hackistan | 8% | 6% | 6% | 8% | 35% | 60% | 28% | 2% |
Easyland | 17% | 4% | 6% | 11% | 20% | 40% | 19% | 3% |
Doom | Plan | 6% | 10% | 18% | 46% | 69% | 37% | ||
Doom & Plan | 0% | 1% | 5% | 20% | 12% | 37% |
Appendix: Capability Threshold Definitions
This appendix defines the key capability milestones referred to throughout this document. These thresholds are used to define phase boundaries and to reason about when various safety measures become necessary or sufficient.
ITAI (Initial Transformative AI): AIs which yield a >3x AI R&D labor acceleration OR can automate virtually all of the job of a typical reasonably senior research engineer at the AI company (in terms of what a remote version of their job was like in early 2025).
Marks the beginning of the “Mitigate and Elicit” phase
In the framework used by AI 2027, this would happen somewhere between “Reliable agent” and “Superhuman coder” (probably about halfway, around Jan 2027 in the scenario). It could trigger before or after the AI R&D-4 bar in Anthropic’s RSP (depending on how useful AIs end up being relative to what jobs they can virtually fully automate).
Note: ITAI doesn’t correspond to when scheming is both plausible and dangerous (this level doesn’t just depend on general capabilities, but also aspects of the architecture and training scheme), so the need for control won’t necessarily trigger at the exact boundary between “Prep” and “Mitigate and elicit” (it likely occurs around or somewhat after ITAI on my views).
Min-H AI (Min-Handoff AI): AIs which have a level of overall general capabilities which would suffice for HTC-handoff being viable. (As in, the capabilities would suffice for our HTC-handoff desiderata if the AIs were also sufficiently aligned and as needed we made narrow improvements to the capability profile which don’t augment general abilities.) This is a capability threshold, so the model does not necessarily need to be aligned or well elicited to count as this threshold (but it would need to be sufficiently aligned to actually handoff). For an AI to qualify as Min-H AI, it would at least need to be capable enough to automate AI R&D and ongoing technical AI safety work (after further elicitation/alignment work and narrow improvements to the capability profile).
By default, marks the end of the “Mitigate and Elicit phase” and the beginning of the “Pause” phase.
In the framework used by AI 2027, Min-H is likely between superhuman AI researcher and superhuman remote worker (also probably around halfway between, around Sept 2027 in the race scenario).
Once we reach the Min-H level of capability, we don’t get much of a benefit from training AIs at a higher level of capability given that we want to defer to these AIs (once these AIs are sufficiently aligned and potential weaknesses in the capability profile are resolved). Thus, it is natural to stop here and work on making deference more likely to go well.
Max-C AI (Maximum Controllable AI): The highest level of AI capability at which control techniques remain effective. Beyond this threshold, either (a) the AI’s capabilities exceed what our control measures can reliably contain, or (b) the overhead required for effective control becomes prohibitively expensive relative to the useful work extracted.
Marks the point at which SDC-Handoff (Stop Doing Control) becomes necessary rather than optional
Note that this is a lossy abstraction:
The exact capability level of Max-C depends on the quality of our control techniques; better control research pushes Max-C higher.
There isn’t a fixed Max-C threshold; there will be a zone of capabilities within which increasing incurs additional marginal risk. We’re simplifying a little by using the Max-C abstraction. More precise notions might be something like Max-C-AI-5%-risk, i.e. the maximum capabilities of AI that we can make while incurring less than a 5% existential risk conditional on those AIs scheming.
Note: Max-C ≥ Min-H is not guaranteed. We think it’s probably true, but if control techniques don’t scale well, we might hit Max-C before reaching Min-H, which would force us to do SDC-handoff before Min-H.
Appendix: Other definitions
Scheming: We say that an AI system is scheming if it is covertly pursuing a misaligned objective. This is a somewhat vague concept; think about applying this definition to humans—it’s clearly a spectrum.
Alignment Faking: An AI is alignment faking if the AI is intentionally attempting to look like an alignment model, for reasons other than actually being aligned. Note that this is slightly distinct from being a schemer: a scheming model might engage in alignment faking, but models might also engage in alignment faking for other reasons, for example, AIs that terminally value human approval.
Misalignment Analogies: (aka model organisms). These are examples of behavior that might pose existential threats. These analogies can either be deliberately trained into models, observed in contrived settings (as in the alignment faking paper), or observed in real life deployment.
Appendix: How this plan differs from other alignment plans
Here are some key differences between this plan and prior plans (e.g. Sam Bowman’s “The Checklist” and “Bumpers”, Josh Clymer’s “Planning for Extreme AI Risks”, and Holden Karnofsky’s “Racing through a minefield”, “Giving AIs safe motivations - Joe Carlsmith “.Note that Karnofsky, Carlsmith, and Bowman all work at Anthropic. Unfortunately, I (Thomas) am not aware of similarly detailed plans produced by any other AI company; if there were others, we would have included them.Note that Karnofsky, Carlsmith, and Bowman all work at Anthropic. Unfortunately, I (Thomas) am not aware of similarly detailed plans produced by any other AI company; if there were others, we would have included them.
Overall, this plan:
Has more detail, specificity, and more comprehensive than prior plans.
Attempts to vary in response to observations, as opposed to other plans which tend to present a class of approaches without overall discussions of how to adapt the plan based on evidence.
Is compatible with varying levels of political will: it results in much higher assurance in Plan A worlds and accepts much higher risks in Plan D worlds.
Focuses on current science / is based on prosaic strategies without assuming any particular breakthrough.
Some specific differences:
We focus on handoff as the central exit strategy. This is similar to Clymer, but distinct from Bowman and Karnofsky. We attempt to clearly lay out what needs to occur before handoff (notably, human level performance on conceptually loaded, hard to check tasks like safety research), and argue that we should handoff early. Previous work was less clear about the goal that we need to reach.
Our core hope for mitigating scheming routes primarily through iterating against misalignment analogies, as opposed to interpretability/transparency. The reason for this is that we think current interpretability is not very useful for safety outcomes. Our plan still involves spending some research into transparency measures like mechanistic interpretability, probes, and faithful COT, but these are less load bearing.
Appendices — Alignment Roadmap
Appendix A — Plan C Assumptions
Responsible Leader. We assume that a responsible AI company is in the lead. Their leadership is willing to take very costly actions in order to mitigate AI takeover risk (e.g. slow down for several months). Despite this, we’ll assume that by default, they’ll have a bias towards proceeding and being willing to take on a large chance of takeover risk to increase their chances of winning.
1-12 month lead over their competitors. Either trailing US or international AI companies will overtake the leader within 1-12 months.
Though we mostly won’t focus on this scenario, Plan C is also applicable in the case where the USG is sufficiently worried about AI takeover that they are willing to regulate AI companies to mitigate risk as much as possible without losing to China, and without being willing to sabotage China.
This plan focuses on making effective use of limited resources to spend on safety and doesn’t ensure a high degree of safety (as the level of resources we’re imagining is insufficient for this). Indeed, arguably the most likely outcome of this plan is misaligned AI takeover. Due to having limited resources, this plan will involve proceeding in some situations where the level of risk is substantial (because exogenous risks are higher) and we won’t necessarily gain much confidence about the exact level of risk.
Note that we expect that something much less reasonable than this plan will happen in practice, so while this is a recommendation (for what to do with limited resources), it isn’t a prediction.
Appendix B — PHASE 1 DIVISION OF TECHNICAL AI SAFETY WORK
Below, I propose a division of work in terms of how much effort we should be putting into different areas of work. This division will shift over time as we get more powerful models, so this exact division isn’t right for the entire Prep period. You can think of this division as targeting towards the end of the prep phase somewhat before ITAI, perhaps around “1.5x AI R&D labor acceleration”. (You could interpolate between this and what I say later in the mitigate and elicit section to better understand how our prioritization will evolve over time.)
My division of work will divide things into multiple areas, but some things in one area will also be relevant to another area. For instance, we want a very good human auditing process for control, but this might also help with generally improving feedback which helps with elicitation (and many of the low-stakes control interventions will have overlap with elicitation). Also, improved approaches for feedback could be used to make feedback consistently better across training which might (more speculatively) reduce the risk of scheming.
Here is my best guess for what our division of technical AI safety work (not including security interventions) at this point should be:
33% Control (and control-adjacent measures). (Recall that this is a relatively broad category which includes some things which aren’t centrally associated with control but which I thought more so fit in this category than in others. Also recall that I’m not including security work which is roughly as helpful for security against human adversaries under control.)
21% misc ML/AI focused control research and implementation. This includes things like experiments involving humans, mitigations to memetic spread / AI coordination, manipulating capability profile with dataset filtering/unlearning/gradient-routing/etc, and control related science and model properties like exploration hacking and steganography (including studying chain-of-thought monitorability). A bunch of the misc control related science is broadly applicable.
4% software infrastructure needed for control (procrastinating on this until we have AIs which are very helpful for building this infrastructure isn’t crazy, though we do need to ensure the AIs which build this infra are controlled or too weak to plausibly be scheming)
6% building a high quality human evaluation process which is robust to scheming (also useful for making AIs more useful for safety work and plausibly for general capabilities longer term)
2% studying probes and similar model internals detection techniques (with a focus on control)
15% Misalignment analogies At this point, this work is most focused on generating evidence for risks and on practicing doing this work later.
10% misalignment analogies of scheming (including more exotic origins for scheming like alignment drift, reflection, etc).
2% misalignment analogies of other threat models
2% studying mitigations by trying stuff on misalignment analogies. E.g. could try various types of corrigibility training. Could try improving oversight to make AIs less evil.
1% generally study AI personas+propensities and how this is shaped by training
17% Elicitation/handoff-viability
6% Improving AI epistemics and philosophical/conceptual reasoning: making AIs which are aligned, capable, and well elicited at conceptual/fuzzy domains and which have good epistemics. Ideally, you’d end up deploying an AI with an impressive ability at truth seeking which would actually be useful for a bunch of applications. And you’d have people actually doing this sort of research as both customers and working on this (making useful AI tools, especially when building the tool requires solving some problems that might generalize, and trying to automate what they are doing)
5% generally trying to apply AIs to assist the rest of this safety work and making AIs more useful for this. This includes resolving (relatively) mundane forms of misalignment which especially interfere with safety work. This also includes stuff like mitigating reward hacking. This also includes making AIs useful for outside groups doing important work.
3% prepping to apply AI for slowing down and USG applications
2% other aspects of handoff-viability, e.g. studying how we’d make a functioning AI society
1% other specialized applications like using AI for whole brain emulation or (even more speculatively) adult biological human strong intelligence enhancement. This will be a substantially greater focus later, but for now we just want to get somewhat ready.
29% Misc science/interventions. We discuss what each of these sub-areas entail in more detail in the “Other areas of work” section, so I won’t elaborate here.
6% Various interpretability methods
5% Honesty, ELK, and getting AIs to reveal secrets
4% Studying generalization, for example, work in the vein of emergent misalignment and the reversal curse.
4% Making deals with AIs and getting consent
3% Generally maintaining and studying chain-of-thought faithfulness
3% Generally looking for concerning model behavior
4% Misc
6% Evaluating and demonstrating capabilities
4% capability demos
2% evaluation, surveys, and forecasting
This is a pretty messy initial guess list partially based on thinking through what we’d want to happen. In practice, the exact allocation might depend a bunch on the interests of individual researchers, what individual projects seem promising, and it would change based on how well various research directions go.
(I think selection of projects within these areas is generally more important than the split across different areas and the interests/advantages of individual researchers might be more important than the exact prioritization. Thus, it’s important to avoid over-fixing on the relative allocation between areas.)
My prioritization is somewhat informed by how easy different directions are to parallelize (and more generally how much they would benefit from dumping in a bunch of additional labor).
This list is focused on technical AI safety work (either ML research or implementation), and is focused on the prioritization within AI companies. Outside people (but also people within the AI company) should work on things like planning, policy, figuring out the best government interventions, public communication, etc.
Appendix C — HOW BAD IS IT TO NEVER DO MODEL WEIGHTS SECURITY?
Situations where it could be very bad:
China can immediately exfiltrate your model weights, and then immediately gives them to a domestic Chinese project. This causes the leading AI project to “race itself” because any progress that it makes is immediately stolen and used by China, and so can never slow down.
There is a small lag in model weight exfiltration. However, there is also a non-negligible safety tax that you have to pay in order to avoid a high risk of AI takeover (e.g. control tax is high, you need to not use neuralese, monitoring is expensive, etc). The trailing actor could get ahead of you if they steal your models and then, despite having less compute, discard the safety techniques.
Situations where it might be ok:
China (and other adversaries) are not woken up to the strategic importance of AGI / aren’t on the ball with stealing model weights (even though they could easily do so).
A core hope is that you might succeed at alignment. Then, let people steal your model, but hope that the stolen models remain aligned. Even if China has a bunch of RID, before handoff, you export safety techniques, and then you both do aligned handoff. Plausibly they just steal the alignment techniques directly.
It might be difficult to swap the spec, e.g. because it took a bunch of data to train the model to be aligned. The stolen model could be relatively rigidly aligned to your values.
In this case, you want it to be such that China is happy with the spec, and so you might want to put in concessions. This is a fine line because (1) you want to make sure the values in your spec are good enough according to china that they will implement your alignment techniques, and (2) you want to make the AI values as good as possible by your lights. As a baseline you could do corrigibility, but you might want to push some of your values in a bit.
Appendix D — PHASE 2 DIVISION OF TECHNICAL AI SAFETY WORK
I’ll summarize our overall priorities into a best guess list, but first it’s worth discussing some limitations of this single list. The ITAI to Min-H period covers a wide range of capabilities and our priorities will shift continuously as the capabilities level increases (and as we update based on evidence from the world). I’ll focus on our priorities as of halfway between superhuman coder and superhuman AI researcher, which is also around halfway between ITAI and Min-H (which are respectively somewhat before and somewhat after each of these two milestones). This is at Jun 2027 in the original AI 2027 scenario. To get a sense for what our priorities are for earlier or later capabilities, you can interpolate these priorities with the priorities I discuss under ITAI and under delaying at Min-H.
In addition to our priorities changing with the level of capabilities, our priorities will change due to updates we’ve made from observing powerful AI systems and how the world reacts to powerful AI systems. These updates will continue ongoingly throughout the mitigate and elicit phase due to both advancing capabilities (and the underlying algorithms driving these capabilities) as well as the outside world reacting to capabilities (which may have a substantial lag). Large algorithmic shifts might happen at pretty high levels of capability (e.g. after superhuman coder), especially because AI R&D is accelerated, so we may not see relevant updates until pretty late (though we can try to guess where things will go).
The numerical priorities list I give will assume my typical guess at what will happen (assuming the AI company is following this overall plan). The assumptions here about AI properties and development are pretty similar to what we see in AI 2027 (though the speed of AI progress isn’t necessarily that similar). In particular:
We haven’t yet caught models scheming, but the architecture is strongly trending in a direction that makes scheming more likely. We have some model organisms in which we caught mostly natural, but not that competent scheming in some somewhat exotic conditions that made scheming more likely and catching easier. We aren’t confident our current AIs aren’t scheming some of the time, they use neuralese.
Clear cut issues with deployed behavior due to reward hacking have mostly been resolved, but there are still various types of reward hacking in training and occasional issues in deployment, typically more subtle issues.
Control is basically working (as of somewhat after superhuman coder) and we have a reasonable (though not perfect) safety usefulness profile and we did a mostly OK but somewhat messy job at evaluating this control. There might have been a period where control wasn’t keeping up with capabilities, but using AI labor and dedicated effort, things got back on track.
We have a somewhat bad neuralese decoding method that we think probably won’t be robust to intentional evasion from the AI and only produces slightly useful results. Outside of this we didn’t really make that much progress with interpretability, but we have somewhat useful methods for augmenting behavioral testing and better understanding why some behavior occurred.
AIs are somewhat better at easy-to-check tasks and are bad at “taste” relative to humans, but there is still some transfer and generalization
By the time AIs are automating AI R&D, there will quickly be huge numbers of AIs which are effectively running quite fast (see AI 2027 numbers)
AIs will have substantially lower than human sample efficiency by default, but they will compensate with a bunch of AI specific advantages and there will be a bunch of somewhat domain specific tricks to improve sample efficiency at learning within some domain or learning on the job with more compute (e.g. generating similar tasks to what you’re working on and ongoingly doing RL on those tasks)
It’s quite viable to make USG sufficiently bought in (including to safety concerns) that they are willing to take aggressive actions to slow down other (adversarial or less safe) actors including working with the AI company on this and they are willing to slow down for safety concerns if there is a lead. (This isn’t necessarily my typical guess, but it seems plausible).
Whole brain emulation (and other human augmentation) ended up not being a viable path
Here is this best guess list of priorities given these updates and assuming we’re halfway between superhuman coder and superhuman AI researcher. The percentages correspond to fractions of human labor within the AI company and correlate with compute usage also.
15% control and control-adjacent measures (this is a relatively broad category which includes some things which aren’t centrally associated with control but which I thought more so fit in this category than in others)
7% misc ML/AI focused control research and implementation.
5% ongoing maintenance and updates to our deployed control measures
1% software infrastructure
2% building a high quality human evaluation process which is robust to scheming
21% model organisms / analogies
16% misalignment analogies for scheming (we would put more effort here if we had found stronger results so far)
1% model organisms of other threat models
4% studying mitigations, including by trying stuff on model organisms. (we would put more effort here if we had better model organisms to study)
20% making AIs differentially more useful for safety work and preparing for aspects of deference other than resolving scheming
9% generally trying to apply AIs to assist the rest of this safety work and making AIs more useful for this.
5% Improving AI epistemics and philosophical/conceptual reasoning (will also help with generally making AIs useful for safety work)
5% Generally preparing for deference (via means other than working on scheming) (will also help with generally making AIs useful for safety work)
1% other aspects of deference
17% using AIs for building a lead in partnership with USG (and getting USG more bought in, though note that scary demos are in a separate category below). To the extent there are returns to more compute, this should use a relatively higher fraction of compute compared to the labor of AI company employees. Also, it’s possible this can’t productively utilize this much of the labor effort. If so, we should shift this labor into other areas (mostly uniformly). Also, note that the resources in this area should increase with further capabilities.
3% make the world more robust to powerful AIs aka d/acc (biodefence, cyberdefense, eating free energy, etc). It’s plausible this should have a higher fraction of resources, I’m quite unsure.
5% neuralese decoding
10% misc science, understanding of relevant AI properties, and other (not including control-adjacent science).
2% Various interpretability methods (top-down, bottom-up/mech interp). I extracted out neuralese decoding from this category.
1% Making deals with AIs and getting consent
2% Generally looking for concerning model behavior (not necessarily in ways which would help detect models coherently scheming)
3% moonshots which seemed promising based on earlier work
2% Misc
9% capability evaluation and demonstrations of both powerful capabilities and misalignment
7% scary demos
2% evaluation, surveys, and forecasting
Again, recall that this is a pretty uncertain prioritization and the prioritization in practice will depend on various updates as discussed elsewhere. For instance, if we caught production AIs scheming against us, we’d shift relatively more resources in scheming mitigations and understanding and away from everything else.
This prioritization isn’t focused on exporting work to other AI companies and this would probably change the prioritization somewhat. More focus on exportability would somewhat reduce focus on control, utilizing AIs effectively, and preparing for deference (including epistemics/conceptual stuff) as these areas probably involve more complex techniques which are more integrated with this specific AI company and the last two of these may also have substantial capabilities externalities. We’d shift focus to model organisms, neuralese decoding, building awareness of risk, and other research directions which might produce simple and easy to implement measures which nonetheless are very useful.
Capability Scaling Strategy
Overview
In the Plan A scenario, why does the Consortium choose to steadily increase capabilities until reaching top-expert-dominating AI (TED-AI) in 2035, then slow down nearly to a pause until handing off to these AIs in 2040? This supplement describes the high level strategy for scaling capabilities and the reasoning behind them.
In one sentence, the high-level scaling strategy in Plan A is to scale as quickly as possible subject to maintaining high confidence that the scaling is safe. In more detail:
First I discuss the high level capability scaling strategy used in Plan A and the reasoning behind it. I split the strategy into 2 stages:
Scale until you’re close to max-controllable-AI, i.e. the maximally capable AI that we are confident we can prevent from causing a catastrophe even if they were misaligned. We predict in the scenario that the max-controllable-AI is roughly TED-AI.
Slow down nearly to a halt, then eventually hand off to AIs.
Then I discuss the costs of taking more time to scale capabilities, most importantly deal decline risk.
Finally, I discuss in more detail the recommended strategy for Stage 1 and Stage 2.
In practice, much more analysis should be done to decide the best strategies than that which we’ve done for this supplement, and furthermore these strategies should be adjusted as information comes in during the execution of Plan A. I’ll describe my current best guess.
I’m confident that we’ve identified the most important considerations pushing toward scaling slower, and I’m moderately but not highly confident that the high-level strategies we describe are correct. I’m not confident in any precise quantitative estimates.
High level strategy
Overall shape of the trajectory
We choose to scale up at all rather than pause indefinitely because of the costs of going slowly, most importantly deal decline, i.e. the risk that the slowdown deal dissolves or becomes less effective.
We aim for a trajectory shaped like the following:
We break Plan A into 2 high-level stages:
Stage 1: Scale until max-controllable-AI.
Stage 2: Slow down nearly to a halt, then eventually hand off to AIs.
This produces a curve that first increases relatively quickly, though still slower than in an uncontrolled intelligence explosion, and then slows down to near a halt for a while. This contrasts with what we’d get if the intelligence explosion were slowed uniformly.
We choose to aim for this sort of curve because we want to spend as much time as possible with AIs that are very useful for important applications, such as alignment, epistemics, and stabilizing the deal. And we think that going this fast early is manageable while maintaining high confidence in safety.
Stage 1: Scale until max-controllable-AI
During this stage the Consortium scales capabilities until a bit below max-controllable-AI, then ~pauses further scaling. Max-controllable-AI is the maximum capability level at which we are confident that we maintain control over the AIs, i.e. the property that the AIs couldn’t cause catastrophic harm even if they wanted to.Non-exhaustive examples of catastrophic harm include: directly taking over the world, misaligning successor AIs that could take over the world, damaging the world’s epistemics to a difficult-to-reverse degree (for example via superhuman persuasion).Non-exhaustive examples of catastrophic harm include: directly taking over the world, misaligning successor AIs that could take over the world, damaging the world’s epistemics to a difficult-to-reverse degree (for example via superhuman persuasion). We think that this level will be approximately Top-Expert-Dominating AI (TED-AI), i.e. AIs that are at least as good as top humans at every cognitive task.
Strategy: We recommend making progress via training compute scaling rather than software progress to the greatest extent possible, because the software can be used by covert projects.
Our best guess is that the Consortium will have a fairly good idea of whether there’s a covert project or not, even though they may not be able to prove it, based on how much compute is unaccounted for. We also predict that there probably won’t be a covert project. If the Consortium thinks there’s likely a covert project but shares our current best guess regarding covert project detection (~60% chance within 5 years), then it’s a close call as to whether they should go slower than they otherwise would due to covert projects. If they are substantially more pessimistic than our current best guess about detection, then they should slow down in order to avoid making lots of software progress, hence taking at least ~5 years in Stage 1.This could also be the case if they agree with our current best guess about detection but have a much higher estimate of the badness of a covert project reaching TED-AI.This could also be the case if they agree with our current best guess about detection but have a much higher estimate of the badness of a covert project reaching TED-AI. If they are instead more optimistic, then the limiting factors on their speed should instead be maintaining high confidence in control and avoiding destabilization. The benefits of scaling quickly are (a) the Consortium gets more capable AIs to automate useful work and to use as experimental subjects and (b) lower risk of deal decline.Scaling too fast via building compute used for training could hurt the situation in the case where the deal declines, but the compute isn’t destroyed, as it allows for a faster takeoff. In the case the deal dissolves, the compute would probably be destroyed, but for partial deal impairment, this wouldn’t be the case. However, we think that this downside is probably outweighed by the upside you get from the benefits of scaling to higher capabilities within the deal rather than outside of it or in an impaired deal.Scaling too fast via building compute used for training could hurt the situation in the case where the deal declines, but the compute isn’t destroyed, as it allows for a faster takeoff. In the case the deal dissolves, the compute would probably be destroyed, but for partial deal impairment, this wouldn’t be the case. However, we think that this downside is probably outweighed by the upside you get from the benefits of scaling to higher capabilities within the deal rather than outside of it or in an impaired deal.
Stage 2: Slow down nearly to a halt, then eventually hand off to AIs
Once the Consortium has AIs that are capable enough such that it is difficult to control them, they slow down (close to a pause) for a while. Then they hand off to the AIs, relaxing control measures. By relaxing control measures, I mean that we can give AIs more ability to get things done, at the cost of increasing the chance they could take over the world if they were misaligned. For simplicity, in this supplement I model handoff as binary.In practice we recommend for it to be somewhat gradual. Example of a gradual handoff: first hand off a bit such that the AIs could maybe take over if they were misaligned, i.e. we’re not as confident as before; then hand off more such that they could only take over if they coordinated but couldn’t if they did not, etc.In practice we recommend for it to be somewhat gradual. Example of a gradual handoff: first hand off a bit such that the AIs could maybe take over if they were misaligned, i.e. we’re not as confident as before; then hand off more such that they could only take over if they coordinated but couldn’t if they did not, etc.
Strategy: Approximately, the strategy is to wait until the Consortium has high confidence that the AIs are aligned, then hand off. More precisely, the Consortium should wait until the benefits of delaying to increase the chance that the AIs are aligned are outweighed by various risks: the deal declining, covert projects reaching dangerous capabilities, or other existential risks such as pandemics; I think that deal declining is a much bigger factor in Stage 2. Roughly, you should hand off once either of the following is true (see more below):
You have very high confidence that alignment has been solved, e.g. 99.5%, and thus marginal gains from further research are low (e.g., ~0.1%/year).
You have somewhat high confidence that alignment has been solved, e.g. 95%, and thus marginal gains from further research are a little higher (e.g., ~1%/year). However, the Plan A deal seems somewhat unstable. E.g., if there were >5%/year chance of deal decline, that would point toward a handoff.
Costs of taking more time to scale
The following are the biggest costs of taking longer to increase capabilities in Stage 1, or taking longer to hand off to AIs in Stage 2.
First, the deal might dissolve or get worse. We call this deal decline, and it encompasses both of:The deal could also get better, but we’re not including that here as our best guess is that the deal will in expectation get worse. This is definitionally true for deal dissolution, but we’re not confident for deal impairment; if we were more optimistic, we could remove the deal impairment term or explicitly model the deal getting better.The deal could also get better, but we’re not including that here as our best guess is that the deal will in expectation get worse. This is definitionally true for deal dissolution, but we’re not confident for deal impairment; if we were more optimistic, we could remove the deal impairment term or explicitly model the deal getting better.
Deal dissolution. The deal is officially dissolved.
Deal impairment. The deal becomes substantially less effective than if it were implemented well and adhered to, but isn’t dissolved. Ways the deal could be impaired:
Deal non-enforcement. The deal is not enforced as strongly as before, e.g. a country blatantly not complying is not punished. This leads to increased violations of the letter of the deal.
Deal misexecution. The letter of the deal is followed but is executed poorly.
Deal revision. The deal is explicitly edited to be worse. This could be straightforward weakening, but it could also be making it worse in other ways, such as unnecessarily increasing surveillance.
See the deal decline supplement for more info.
Second, there are risks from illegal model usage in covert projects by state actors or terrorists. If covert projects possess capabilities that would be catastrophic if in the hands of misaligned AIs or bad actors, AI takeover and misuse risks are higher. For Stage 1, unlike in Stage 2, the Consortium should (within plausible Stage 1 lengths) go slower to beat the covert projects; going too fast requires making lots of software progress which helps the covert projects. More discussion below.
Finally, there is direct pre-handoff existential risk incurred before handing off to AIs. This doesn’t include risks from covert projects. This can be split into:
Risks from use of legal (i.e. non-covert) model capabilities.Including any ways it derives from the relevant capabilities, e.g. distillation.Including any ways it derives from the relevant capabilities, e.g. distillation.
Risks not derived from legal model capabilities. For example, asteroid risk or a nuclear war not causally downstream of legal frontier AIs.
Additionally, there are shifts in the world not driven by the deal, which could be either a cost or benefit of going slowly. For example, the balance of power could shift so that authoritarian countries have more power relative to democratic countries, or vice versa.My best guess is that the sign of this sort of change is in expectation negative for EV of the future, so this is a weak motivation for scaling faster. But I’m not confident.My best guess is that the sign of this sort of change is in expectation negative for EV of the future, so this is a weak motivation for scaling faster. But I’m not confident.
Stage 1: Scale to max-controllable-AI
Benefits and costs of going slowly
During Stage 1 our goal is to decide on a policy for how quickly we should scale training compute and make software progress in the run-up to max-controllable-AI. That is, we want to set a trajectory that has a target end date at which max-controllable-AI is reached, and specify the split of training/software progress at each point. We consider only strategies that specify in advance a target end date and then scale throughout all of Stage 1 toward that date, for simplicity and tractability. In practice the Consortium’s strategy should be adjusted over the course of Stage 1 as they get more information.
We recommend making progress via training compute scaling rather than software progress to the greatest extent possible, because the software can be used by covert projects.
We also need to decide the training compute safety tax that we’d like to pay, i.e. the difference between the amount of compute actually used on a training run to reach a given capability level, and the minimum amount of compute required to reach that capability. Paying a higher safety tax means using safer training algorithms at the cost of involving making extra software progress that goes unused but leaks to covert projects.Or scaling in Stage 1 with only training compute.Or scaling in Stage 1 with only training compute.
The primary costs of going slowly in Stage 1 are:
Those laid out above; risk from deal decline due to taking longer, covert projects (if paused for a very long time), and direct existential risks.
Missing out on time with more capable AIs, which can be used to automate useful work (alignment, control, epistemics, etc.) and be used as experimental subjects. These AIs could also help stabilize the deal, reducing risk of decline per year.
The primary benefits of going slowly in Stage 1 are:
A lower fraction of progress comes from software improvements. This is good because software improvements can be copied by covert projects.There are downsides to scaling with hardware over software though, namely that if the compute isn’t destroyed upon deal decline it can be used to do a fast takeoff post-decline. If mutually assured compute destruction is set up and/or the Consortium is confident that covert projects won’t pose a big issue, the Consortium should consider scaling with primarily software rather than primarily hardware.There are downsides to scaling with hardware over software though, namely that if the compute isn’t destroyed upon deal decline it can be used to do a fast takeoff post-decline. If mutually assured compute destruction is set up and/or the Consortium is confident that covert projects won’t pose a big issue, the Consortium should consider scaling with primarily software rather than primarily hardware. Software progress and training compute scaling are the two sources of progress. Software progress is mostly made public, meaning it diffuses to covert projects. This means that despite covert projects being a reason that you eventually should scale to superintelligence, they are a reason you should scale slower in Stage 1 unless you’re taking well over a decade.At least, assuming takeoff isn’t much faster than in the Plan A default.At least, assuming takeoff isn’t much faster than in the Plan A default.
Less risk of AI takeover due to overestimating to what extent the AIs are controlled.
Less destabilization, which could (among other things) potentially increase the chance of deal decline per year. An example of a destabilizing effect is if we scaled to AIs capable of superpersuasion and it wasn’t properly defended against, leading to degraded epistemics. However, as noted above, better AIs could also be stabilizing and reduce the annual risk of deal decline.
The Consortium’s strategy is also constrained in that it might not be possible to avoid discovering software improvements, for 2 reasons: (a) some software improvements might be required to utilize increased amounts of compute, (b) doing lots of safety R&D likely inevitably leads to some capabilities externalities in the form of software progress. In our simulations, we require that the Consortium make 0.2 OOM of software progress for each OOM of training FLOP/s they gain, and that during Stage 2, the Consortium’s alignment research incurs capabilities externalities at a rate of 0.25 software OOMs/year.
What to do if a covert project seems unlikely
Our best guess is that the Consortium will have a fairly good idea of whether there’s a covert project or not, even though they may not be able to prove it, based on how much compute is unaccounted for. We also predict that there likely won’t be a covert project.
If it seems unlikely that there’s a covert project, my recommendation is to scale as fast as possible while maintaining high confidence in AIs being controlled and avoiding destabilizing society and the deal. This is what happens in our scenario, and we depict the resulting target year of 2035 as a rough guess. I’d weakly guess that the target year should be a little earlier, were we to rewrite the scenario.
Target year’s effect on risk probabilities
If it seems likely that there’s a covert project, then the situation is more complicated because we have to balance deal decline risk, which points toward going faster, with covert project risk, which points toward going slower.
I’ll now do some modeling of how the trajectory of the deal will go. The modeling assumes that the deal begins in 2029 and that it starts being implemented a bit under a year before AC, as in the Plan A scenario. It does not assume that takeoff from AC to TED-AI and AC to ASI is about 1 year as in the scenario. Instead it assumes uncertainty with parameter medians centered around those from the Plan A default trajectory, as we think decision-makers will still have substantial uncertainty over takeoff speeds in 2029. This means that sometimes the Consortium can’t hit the target date; in fact for the earliest target date we consider of mid-2031, it often can’t hit it, as this would be a takeoff of similar length to the default (the default trajectory goes from pre-deal capabilities to TED-AI in ~1.5 years, while in Plan A the training resumption happens at the start of 2030 and the Consortium aims to pay a 1 OOM safety tax above TED-AI).
The below graph shows how the risk of deal decline and risk of covert projects achieving max-controllable-AI change based on the Stage 1 target end date. It assumes our best-guess projections as described in the deal decline and covert project supplements. It looks at the risk of deal decline and covert projects reaching TED-AI in Stage 1 or within the first 5 years of Stage 2. I choose 5 years because this is the length of Stage 2 in the scenario; more discussion below. The latest target date considered is 2037 because this is, with our median takeoff parameters,“Our” meaning the official Plan A median parameters, which represent Thomas’s views rather than mine; my median takeoff is significantly slower.“Our” meaning the official Plan A median parameters, which represent Thomas’s views rather than mine; my median takeoff is significantly slower. the year at which TED-AI is reached if the Consortium makes the minimum amount of software progress.
Deal decline and covert project risk trade off as speed varies
We can see the tradeoff between minimizing risks from deal decline, and minimizing risk from covert projects reaching TED-AI. Target years between 2032 and 2036 look fairly similar in terms of maximizing the chance of getting to 5 years into Stage 2 unimpaired. If we turn off detection, then any of 2035 to 2037 look to be of fairly similar value, with 2032 and especially 2031 looking much worse; this makes sense as covert projects are a bigger factor.
Badness of risks
Does it make sense to treat these risks as equal in badness to each other?
I’ve done some rough estimates of how good various end states to the deal are here (summary here): by year and whether the deal stays unimpaired, or whether it declines or a covert project reaches TED-AI. In those estimates, I assumed that the covert project that reached TED-AI (or the AI that it created) had the ability and desire to essentially take over the world.
I’ll now bin outcomes into whether they happen before or after the start of Stage 2 and calculate the badness based on how much worse the outcome is than the deal surviving 5 years into Stage 2. The shift to Stage 2 is good for improving outcomes from both deal decline and covert TED-AIs, because reaching TED-AI allows the discovery of better alignment techniques, generally improving epistemics, and other useful applications.A few thoughts about the difference between deal decline and covert project takeover with respect to Stage 1 vs. Stage 2 desirability: (a) covert projects can wait until they are confident in alignment if they’d like, even if in Stage 1, while a post-deal-decline project may not be able to; (b) on the other hand, covert projects do not benefit during Stage 2 from the safety tax the Consortium has paid, while post-deal-decline projects do.A few thoughts about the difference between deal decline and covert project takeover with respect to Stage 1 vs. Stage 2 desirability: (a) covert projects can wait until they are confident in alignment if they’d like, even if in Stage 1, while a post-deal-decline project may not be able to; (b) on the other hand, covert projects do not benefit during Stage 2 from the safety tax the Consortium has paid, while post-deal-decline projects do.
I’ll also add a discount for the fact that there might not be human or AI takeover that originates from the covert project that reaches TED-AI.
It seems like the percentage of value lostSpecifically, this is relative to the expected value of an intact, unimpaired deal with no covert TED-AIs five years after the Consortium reaches TED-AI.Specifically, this is relative to the expected value of an intact, unimpaired deal with no covert TED-AIs five years after the Consortium reaches TED-AI. is very roughly:I’m crudely bucketing covert project outcomes into takeover and no takeover, and assuming the takeover ones don’t destroy value. In practice, there will be cases in which value is destroyed but there isn’t takeover; I’m assuming otherwise for simplicity.I’m crudely bucketing covert project outcomes into takeover and no takeover, and assuming the takeover ones don’t destroy value. In practice, there will be cases in which value is destroyed but there isn’t takeover; I’m assuming otherwise for simplicity.
Pre-Consortium-TED-AI (Stage 1) | Post-Consortium-TED-AI, first 5 years (Stage 2) | |
Covert takeover badness | 70% | 55% |
Takeover chance | 55% | 45% |
Covert reaches TED-AI badness | 39% | 25% |
Deal declines | 50% | 25% |
Below we can see how deal decline and covert project risk are split up across Stage 1 and the first 5 years of Stage 2.
Deal decline risk is more weighted toward Stage 1 than covert project risk
We see that a higher percentage of deal decline risk comes in Stage 1 than for covert project risk, especially for target years of 2033 and earlier. This makes deal decline on average worse than covert projects reaching TED-AI, relative to what you’d think if you compared their badness from happening at the same time. Now let’s check out the implications of combining the risk levels with the value lost from the risk materializing:
Expected value lost from deal decline and covert risks, by TED-AI target year
My simple methodology has the following 2 issues which means that you should interpret the above graph as being too favorable toward earlier target dates relative to later ones
I only consider the percentage of value lost as the outcome variable, but reaching 5 years into Stage 2 has slightly greater value for later target dates than for earlier ones (though the effect size might be small given the use of AIs for safety in the final 5 years swamps that of previous years).
The percentage of value lost might be higher for early target years, especially during Stage 1.
With that said: With my best guess parameter values, it looks like earlier is better from the perspective of deal decline and covert project risks. However, it’s not that big of a difference. The expected percentage of value lost is just ~2% higher with a target year of 2036 than of 2032. This means that adjustments for the issues described above could potentially flip the conclusion.
Not taking into account covert project detection results in a curve that is more flat and at higher amounts of value lost; having there be a covert project and not being able to detect it is a rough situation to be in no matter what you do. All target dates result in similar values, but with the aforementioned issues taken into account later target dates could be favored.
Including detection but assuming that a covert project or its AI takes over once reaching TED-AIOr simply being more pessimistic than me about covert projects relative to deal decline.Or simply being more pessimistic than me about covert projects relative to deal decline. also results in a fairly flat curve.
Summing up
If the Consortium thinks there’s likely a covert project but shares our current best guess regarding covert project detection (~60% chance within 5 years), then it’s a close call as to whether they should go slower than they otherwise would due to covert projects. If they are substantially more pessimistic than our current best guess about detection, then they should slow down in order to avoid making lots of software progress, hence taking at least ~5 years in Stage 1.This could also be the case if they agree with our current best guess about detection but have a much higher estimate of the badness of a covert project reaching TED-AI.This could also be the case if they agree with our current best guess about detection but have a much higher estimate of the badness of a covert project reaching TED-AI. If they are instead more optimistic, then the limiting factors on their speed should instead be maintaining high confidence in control and avoiding destabilization.
All of the above assumed a constant safety tax of 1 OOM of excess training compute above the minimum required paid at the end of Stage 1. This is my recommendation and the tax portrayed in the scenario. This recommendation is based on a mix of simple modeling and intuitive judgment. I’m not confident in it; the benefit of a higher tax is more confidence in control at higher capability levels, while the downsides are leaking software to covert projects and creating an overhang of foregone software improvements that can be used by projects post-deal-decline (especially post-dissolution).
Stage 2: Slow down nearly to a halt, then eventually hand off to AIs
Once at max-controllable-AI (TED-AI in the Plan A scenario), we need to decide how long to pause before handing off trust to AIs. For the purposes of this analysis we are treating this stage as a pause, though in practice and in the scenario there are a few deviations from a complete pause: (a) it may be desirable to scale up very slowly, as the max controllable level might increase due to improved control techniques and/or tighter uncertainty bounds, and (b) safety research will have at least minor capabilities externalities, meaning software progress will be made, which will leak to covert projects even if it isn’t applied to improve legal AIs.And there might be some cases where it’s difficult to apply safety techniques without increasing the capabilities of legal AIs.And there might be some cases where it’s difficult to apply safety techniques without increasing the capabilities of legal AIs.
The costs of delaying handoff in Stage 2 are as laid out above; risk from deal decline, covert projects, and direct existential risks. We think that deal decline will generally be the dominant consideration because covert projects will likely have already been detected, if they existed in the first place.
The primary benefit of delaying handoff in Stage 2 is increasing the likelihood that the AIs we hand off to are aligned. This mostly comes from alignment research that helps figure out how to align AIs and evaluate their alignment, and paying “safety taxes” such as training a model with more compute than is required to reach a given capability level, in order to use a safer training process. Note that this is specifically regarding the AI that we hand off to; it’s possible that an aligned AI might not succeed in aligning its successors (especially if circumstances change negatively in some way after handoff).
A secondary benefit is increasing the probability that the future goes very well conditional on alignment; this is also highly important but probably depends less on the timing than does probability of alignment. Additionally, it’s possible that waiting too long eventually ends up being marginally negative for outcomes conditional on alignment; it might be better to hand off to an aligned max-controllable-AI earlier rather than later so that it has more free rein to improve the situation.
To decide when to hand off in Stage 2, we recommend calculating the marginal benefits and costs of delaying handoff and handing off once the costs outweigh the benefits.Technically, there could be more than one crossover point, but we expect that in practice there will likely be a single crossover point at which costs of delaying begin to outweigh benefits. That said, this isn’t guaranteed, for example there may be predictably transient increases in deal dissolution risks at various points.Technically, there could be more than one crossover point, but we expect that in practice there will likely be a single crossover point at which costs of delaying begin to outweigh benefits. That said, this isn’t guaranteed, for example there may be predictably transient increases in deal dissolution risks at various points. We expect the best strategy to be something like handing off once we have high confidence that the AIs are aligned.
See below for specific concrete case studies on deciding what strategy is best.
Concrete case study of possible 2040 handoff decisions
We’ll now discuss case studies regarding whether to hand off. To make things concrete, we’ll assume that we are making a decision regarding whether to hand off in 2040 some time after max-controllable-AI was reached, which is when handoff happens in the Plan A scenario.
To keep things simple, we’ll assume that the only factors for whether to delay handoff are:
Benefit: Increased alignment probability. See below for discussion of what the trajectory of alignment progress will look like.
Cost: Deal decline risk. Covert project risk is the main other contender for a substantial reason to hand off faster. However, our guess is that covert project risk is probably very low in this case study, because (a) probably if there was a covert project it would have been detected by now and (b) we haven’t done non-safety-externality software progress in 5 years, so covert projects haven’t been getting much help.
Below you can try out different settings for the relevant variables and see what recommendation they lead to regarding whether to delay hand off. I’ve set the default values to roughly match what they are during the handoff in the scenario, and have added plausible presets that set either the alignment or deal dissolution/impairment parameters.
| Low risk d_prob 0.5%/yr · d_loss 10% | Medium risk d_prob 3%/yr · d_loss 20% | High risk d_prob 10%/yr · d_loss 35% | |
|---|---|---|---|
| High confidence alignment solution a_prob 99.5% · a_change 0.1%/yr | Delay −0.05% | Hand off +0.50% | Hand off +3.38% |
| Alignment probably solved a_prob 95% · a_change 1%/yr | Delay −0.95% | Delay −0.43% | Hand off +2.33% |
| Alignment likely solved a_prob 80% · a_change 3%/yr | Delay −2.96% | Delay −2.52% | Delay −0.20% |
| Alignment clearly unsolved a_prob 10% · a_change 2%/yr | Delay −2.00% | Delay −1.94% | Delay −1.65% |
Reasoning for the medium deal decline probabilities and badness estimates can be found in the Deal Decline supplement.
Trajectory of estimated alignment probability
How steadily will alignment progress? There are 2 extremes we could use to model the situation.
All or nothing: We start very confidently believing that the AIs are misaligned, then at one point we make a breakthrough which quickly makes us confident that they’re now aligned.
Smooth, predictable progress: Our estimate of whether the AIs are aligned increases steadily and predictably.
Our high-uncertainty best guess is that the trajectory of our confidence in AIs’ alignment will be in between, perhaps slightly closer to the All or nothing option. It could also be a combination: we think that progress may look smooth/predictable for a while, until there’s a sudden breakthrough. For example, if ambitious mechanistic interpretability is solved and we understand all of AIs’ reasoning processes, this may give very high confidence in alignment.
For example, if we weren’t yet confident in AIs’ alignment, in 2040 there might be a ~5%/year chance of an alignment breakthrough that makes us confident, in addition to continuous improvements until we get the breakthrough; if we are confident (as in the scenario), then perhaps we’re at 0.01-1%/year alignment probability increase since we’re already close to 100%.
Comparing Possible Plans
Introduction
In this supplement I’ll estimate how well each plan does on various intermediate metrics, and its chance of leading to a great future. I discuss quantitative metrics for the sake of concreteness, however I’m very much not confident in my precise estimates. I am confident in the conclusion that Plan A is substantially better than Plans B-D. For plans that I think could be competitive with Plan A, see this table.
I estimate how good each plan would be conditional on the government and/or leading companies attempting to execute it. The below aren’t meant to be an exhaustive taxonomy; I am considering what would happen if roughly the described plan is attempted, rather than bucketing every possible future into the plan it is closest to. For more detail on how I classify trajectories into plans, see below.
The plans I consider are:
Plan A (as implemented specifically in the scenario): Verified slowdown + total research transparency. A verified international slowdown deal coupled with complete transparency of AI research. AI capabilities are scaled up while maintaining high confidence in safety but balancing risks from deal decline and covert projects.
Plan A: Verified slowdown + substantial transparency. To be more generally classified as Plan A, the implementation must involve a substantial verified slowdown which involves continued AI scaling, as opposed to an indefinite halt. It must involve transparency at least as strong as embedded auditors from foreign governments which publish reports to governments (and for which redacted versions are made public), as described in our filtered transparency proposal. Other aspects of the plan, such as mutually assured compute destruction, are not required.
Plan B-kinetic: Sabotage China kinetically and burn some lead. Sabotaging China in order to get more of a lead to burn on safety, with the willingness to escalate to large-scale kinetic attacks such as by drones or conventional missile strikes. To be classified as Plan B the US must have the intention to use their lead over China for safety purposes, and must in practice slow takeoff by at least 3 months.
Plan B-cyber: Sabotage China non-kinetically and burn some lead. The same as Plan B-kinetic, but without willingness to escalate to large-scale kinetic attacks. This could include cyber and supply chain attacks.
Plan C+: Domestic regulation and slow China without sabotage. The US buys substantial time via domestic regulation, without sabotaging China. They instead aim to slow down China via non-sabotage measures such as export controls and giving green cards to Chinese talent. Requires at least a 2 month slowdown.
Plan C: Leading project burns some lead. The leading AI project burns some of its lead on safety, and potentially coordinates with other frontier projects for a further slowdown without domestic regulation. Requires at least a 1 month slowdown.
Plan D: Race. Frontier AI projects race through the intelligence explosion at nearly max speed, devoting a nonzero but small percentage of resources to safety (at least 1%; lower would count as Plan E).
There are also other plans that the government could attempt to implement. Some of these are discussed below.
For simplicity, I assume that:
The minimum-length takeoff from Automated Coder (AC) to Top-Expert-Dominating-AI (TED-AI) is slightly less than a year, from the start of 2030 to the end of 2030. This matches the default Plan A AI Futures Model trajectory. The minimum-length takeoff is what happens if the leading AI projects go at maximum speed.
All plans can begin implementation in early 2029 as in the Plan A scenario, though for Plans C and D this doesn’t matter as the plan involves doing nothing until takeoff is underway. This still might involve some safety work, if doing so is useful for advancing capabilities.
I remove these conditionals and present other authors’ views below.
Plan assessment table
These estimates are all rough best guesses, and the specifics of them aren’t load-bearing to our recommendation of Plan A (though the high-level takeaways are).
| Metric ↓ / Plan → | Plan A | Plan B | Plan C+ | Plan C | Plan D | |
|---|---|---|---|---|---|---|
| Kinetic | Cyber | |||||
| Objective metrics | ||||||
| Takeoff length (AC to takeover-capable AI) | 6 yrs | 3 yrs | 2 yrs | 1.5 yrs | 1.13 yrs | 1.02 yrs |
| Training compute safety tax (OOMs) | 2.8 | 1 | 0.75 | 0.46 | 0.13 | 0.02 |
| Safety compute (H100e-yrs) | 100B | 500M | 100M | 20M | 4M | 500K |
| Safety-researcher-years | 1M | 2,000 | 1,000 | 1,500 | 500 | 200 |
| Subjective metrics (1–5) | ||||||
| Epistemic culture | 3 | 1 | 2 | 2 | 1 | |
| Public transparency | 4 | 1 | 2 | 2 | 1 | |
| Distribution of power | 5 | 1 | 2 | 2 | 1 | |
| Outcomes (medians across implementations) | ||||||
| p(alignment); i.e., 1 - p(misaligned takeover) | 72% | 50% | 45% | 40% | 25% | |
| p(great future | alignment) | 58% | 50% | 55% | 50% | 40% | |
| p(great future) | 42% | 25% | 25% | 20% | 10% | |
Plans are alphabetically ranked in terms of their objective metrics, with a big jump between Plan A and Plan B. However, plans C+ and C beat out Plan B on some subjective metrics. Plan B’s focus on security and its wartime vibe could degrade epistemics, and its likely implementation of nationalization or pseudo-nationalization would decrease public transparency and concentrate power.
Comparing authors’ views on plan likelihood and outcomes
These estimates aggregate over uncertainties about the default capability trajectory and when the plans are implemented, as well as how competently the plans are executed. The numbers are in this sheet, and the plan classification is as described below.
Likelihood
Probability each plan happens across authors
| Eli | p(govt does something at least as intense as Plan B): 35%; seems like govt wakeup has been increasing and also the importance of ASI will get more and more obvious. Though note that Plan B requires actually slowing down by at least 3 months. |
|---|---|
| Ryan | I’m very uncertain about what should count. I think Plan C+ is mostly worlds with dumb regulation that slows down AI. I think Plan B would be higher if not for the intent to spend the lead on safety. |
Outcomes
We focus on p(alignment) here, i.e. 1 - p(misaligned takeover). We ignore p(great future | alignment) in this section for simplicity.
Probability of avoiding misaligned takeover by plan
| Eli | My probabilities in the plan assessment above were conditional on: My median takeoff is ~2x longer than the one described in (a), but I have significantly probability on faster takeoffs. This means that my p(alignment) estimates should do a combination of increase and regress a bit toward 50%. Also, Plan A isn’t as sensitive to default takeoff length as other plans, and from a covert project perspective it’s actually unclear in which way the effect size points because larger effective compute gaps between the start and end of takeoff require more software progress. The trajectory timeline to AC / start of takeoff is pretty similar to my median so for that I regress abit toward 50%. For (b), my median is that the plan gets implemented later than this, but I think it doesn’t matter much for all the plans besides A because in the other plans it’s best to slow down later anyway. I do lower the value of Plan A due to this consideration, though not hugely because I think covert projects would likely still be difficult to do without detection. Also, the plan assessment numbers were based on the median implementation of the plan, rather than aggregating the expected outcome over all implementations. My unconfident quick guess is that switching to expected outcome should also regress my numbers toward 50%, but not sure so this adjustmnet will be small. With the above in mind, I make the following adjustments: For Other, I think a bunch of this is random pretty good plans that aren’t covered, including Plan S. E.g. there’s ~40% that this is a non-Plan-A deal. Though it also includes Plan E. Overall I rate it the same as B, 55. |
|---|---|
| Ryan | These numbers are very sensitive to how you classify plans. E.g., I think Plan A as described in the scenario has maybe 3/5 as much risk as Plan A as classifed by the supplement. An actually good version of Plan C done by the leading AI company is pretty close to the numbers I give for Plan B (but worse than a great Plan B). A large fraction of “other” is worlds with not enough effort on safety to classify for Plan D. |
Other plans
There are many other possible plans that the government could follow. Below we catalog some of the ones that we find to be most promising and/or likely.
Plans that might be competitive with Plan A
Summary | Advantages (relative to A) | Disadvantages (relative to A) |
Plan S: Indefinite halt. A halt on all frontier AI capabilities progress, intended to last at least a few years. Different variants of Plan S have different conditions for resumption of AI progress; for example it could be alignment progress, lie detectors, human uploads, or intelligence enhancement.Other potential differences from Plan A, though it depends on which variant of Plan S is implemented: slower compute buildout, less transparency.Other potential differences from Plan A, though it depends on which variant of Plan S is implemented: slower compute buildout, less transparency. | A longer expected slowdown and more margin for error regarding scaling too fast. Potentially simpler. | Scaling to controllable AIs within the human range is helpful for accelerating alignment/control, epistemics, verification, and general understanding of AI. |
Domestic-first Plan A. Regulate AI domestically enough to reduce AI takeover risk to acceptable levels, which will require a long slowdown. It’s possible other countries will also regulate domestically and something like Plan A won’t be needed; otherwise, transition to Plan A later. | Initial steps are achievable by the US and very helpful even if the international stage isn’t viable because it massively extends the timeline. Proactively builds trust with China. | Worse for covert projects and setting up verification than negotiating Plan A at the same time. May not be politically feasible due to race dynamics. |
GPU arms control. International agreements for countries to limit their GPU stock or flow, analogous to historical arms reduction agreements.A special case of this is to Pause the GPUs: Require that a fixed fraction of the GPUs be either (i) turned off or (ii) used to mine cryptocurrency (or be used for some verifiable, non-AI R&D purpose). This proposal has the upside and downside of making it easier to unpause the GPUs.A special case of this is to Pause the GPUs: Require that a fixed fraction of the GPUs be either (i) turned off or (ii) used to mine cryptocurrency (or be used for some verifiable, non-AI R&D purpose). This proposal has the upside and downside of making it easier to unpause the GPUs. | Much simpler to enforce, and more historical precedent. Can achieve substantial slowdown. | Less slowdown, less ability to pay safety taxes, ongoing race dynamics mean no ability to coordinate towards safer paths through the tech tree. |
CERN for AI. An international project to develop frontier AI, with all other projects regulated to be substantially behind in capabilities. | Easier to defend against algorithmic leakage and distillation to covert projects. Less actors at the frontier might make it easier to enforce. | Worse decisions (including worse decisions on technical safety) due to less transparency and less broad deployment.But maybe you can mitigate this via small public deployments and selective transparency.But maybe you can mitigate this via small public deployments and selective transparency. More concentration of power risk. |
Plans that we think are clearly worse than Plan A
Repurpose the GPUs: Like GPU arms control, except they are now required to be used for existentially useful applications rather than being destroyed. See this comment for a more specific proposal in which GPUs are required to be used for giving inference capacity to external safety organizations or to do publicly transparent safety research. In theory this plan should be better than GPU arms reduction, but it’s less reversible and harder to enforce.
Plan B-but-don’t-slowdown, aka D-war: Like Plan B except that you don’t slow down, i.e. you don’t burn ~any of the lead that you increased by sabotaging China. (Currently classified as a variant of Plan D when making the above estimates.)
Weaker deals with China. For example, the US could offer: “You get access to more chips for AI inference, and a non-binding or non-majority say in our safety requirements, if you give us full transparency into what you’re up to.”
Plan E: Like Plan D but even less safety investment, i.e. less than 1% of resources.
Classifying trajectories into plans
Define the implementation window as the period between:
Start: 1 year before Automated Coder (AC) would be reached if scaling at maximum speed.Potentially I should allow for plan implementation to begin earlier, especially if the plan is very intense. For example, probably Plan S should count if it begins 1.5 years before the default AC time.Potentially I should allow for plan implementation to begin earlier, especially if the plan is very intense. For example, probably Plan S should count if it begins 1.5 years before the default AC time. An AC is a collective of AIs that fully automates an AGI project’s coding work, autonomously replacing the project’s entire coding staff. That is, you’d rather fire all of the humans than stop using AIs.
End: Takeover-capable AI. Takeover-capable AI is reached when it is practically very unlikely that humans could stop AI takeover, conditional on the AIs being adversarially misaligned. It could still be in theory possible for them to shut down the AIs by taking extremely drastic actions like shutting down the internet, but it’s unlikely that they would do so, e.g. because AIs might be heavily relied upon and/or highly persuasive.
If the plan being attempted is consistent throughout this window, then the trajectory is classified into that plan. Even if multiple plans are attempted, each trajectory can count as at most one plan; this is a non-obvious choice and it seems also reasonable to count trajectories as multiple plans.
If multiple plans are attempted during the implementation window, the trajectory is classified as the most intense plan that was attempted. This is a bit vague to use as a general classification, but for concreteness among Plans S/A/B/C+/C/D trajectories would be classified as the attempted plan that is earliest in that list. The reasoning behind this is that it seems intuitively like a trajectory should count primarily as Plan A or S even if the deal dissolves quickly, and also should count as Plan A or S if it starts as a less intense plan but then transitions to A/S within the implementation window.
The following selection effect is a consequence of this classification: In more intense plans, people are selected for being more competent. Furthermore, less intense plans are selected for people not later transitioning to a more intense plan even if concerning evidence arises.
One selection effect that I explicitly remove is the effect of background variables such as alignment difficulty on what plan a trajectory counts as. I.e., I assume when rating plans that the distribution of alignment difficulty is the same no matter what plan is attempted; in particular my overall guess as to the distribution of alignment difficulty. As mentioned above, I don’t do this for things that humans control such as their competence.
Appendix
Training compute safety tax definition
The training compute safety tax is the difference between the amount of compute actually used on a training run to reach a given capability level, and the minimum amount of compute required to reach that capability. Examples of ways to productively use that unnecessary compute might be to train in a more safe but less compute-efficient paradigm, or to spend lots of compute on AI-driven scalable oversight to improve incentives during the training process.
Measuring the safety tax by this method is an imperfect proxy for what we actually care about. There is likely large variance among improvements which provide the same amount of capability in how useful it is for safety to forego that improvement.
Covert AI Projects
In the Plan A scenario, we propose international coordination to develop AI safely. Keeping risk at an acceptable level requires substantial slowdowns in the pace of AI development. A central blocker for this type of coordination is the risk that one party develops AI in secret, thus gaining an advantage.
Fortunately, it seems very hard to build and operate covert AI datacenters in a way that would be catastrophic. In the rest of this supplement, we’ll analyze how difficult it would be for China to defect from Plan A in secret, and the mitigations that can be put in place to prevent defection.We focus on China in particular because we think China is the best positioned party to defect. This is discussed further in the appendix.We focus on China in particular because we think China is the best positioned party to defect. This is discussed further in the appendix.
It will be impossible, in practice, for the US to confidently rule out a sufficiently small Chinese covert project. Therefore, the main questions of interest are:
What is the largest size of covert project that can evade detection?
On our current trajectory,Based on our modeling, the rate of capabilities progress of covert projects is largely determined by the rate of progress of the Consortium. Specifically, under our transparency recommendations the main lever for speeding up (software progress) also accelerates the covert projects. Hence “going faster to outpace covert projects” usually has the opposite effect—instead, the lead-maximizing strategy involves making as little software progress as possible while maintaining the hardware-buildout-constrained rate of training compute scaling.Based on our modeling, the rate of capabilities progress of covert projects is largely determined by the rate of progress of the Consortium. Specifically, under our transparency recommendations the main lever for speeding up (software progress) also accelerates the covert projects. Hence “going faster to outpace covert projects” usually has the opposite effect—instead, the lead-maximizing strategy involves making as little software progress as possible while maintaining the hardware-buildout-constrained rate of training compute scaling. at what time do covert projects pose a threat to the integrity of the deal and global security?
The answers to these questions will depend substantially on the overall scenario. However, in any scenario, the members of the Consortium will have the affordance to adjust their strategy to increase their lead over covert projects, primarily by reducing their rate of software progress at the cost of additional risk that the deal declines before sufficient alignment progress has been made, or by insisting on more robust detection technology before proceeding.
Section 1: Overview of our strategy for covert project mitigation
This supplement outlines our plan for preventing this type of catastrophic defection from Plan A.
This diagram gives an overview of our strategy for handling covert projects. Elements of the strategy are covered across four supplements, as pictured above. In this section, we’ll give an overview of the entire strategy.
Defection from Plan A can happen on legal datacenters or covert datacenters. The legal datacenters are approved for AI development or usage, but are subject to verification to ensure that they are following those rules. Verification of the legal servers is the primary focus of the Verification Plan.
This supplement will focus on the risk from covert projects, i.e. projects primarily using secret datacenters.
The main way that we can limit covert projects is by limiting their amount of compute: the covert project datacenters must be very small in order to go unnoticed. We can limit covert project compute through a combination of preventing covert projects from acquiring compute and detecting covert datacenters once they are constructed.
Compute for covert datacenters can either come from “stock” (compute produced before 2029) or “flow” (compute produced during or after 2029). We are optimistic about very strong control of the flow of new chips because under Plan A compute production is extremely carefully monitored using highly controlled fabs. (Concretely, we should be able to keep diversion to approximately zero and have very high confidence that total diversion is less than 100k H100e.) Verifying the flow of new chips becomes somewhat more difficult during the industrial explosion because the abundance of robots may make it much more viable to construct a secret chip supply chain. The Consortium can prevent this through verification methods applied to the robot workforce.
The stock of compute is already distributed and thus much harder to find, so we predict that a potential covert project would acquire their compute from existing stock. This means that a key priority in Plan A is estimating pre-deal stock and accounting for a high fraction of it.
Our median estimate for the amount of compute a competently-executed PRC covert project could divert is roughly 1.5M H100e (.5% of world compute at EOY 2028, costing around $15B).
Software is much harder to limit than hardware. However, if enough software leaks, even very small covert projects will be able to develop extremely capable AI. Software progress in covert projects can be broken down into two categories: progress acquired from legal projects and software progress from research done by the covert project itself.
Progress acquired from legal projects. For most AI algorithms, security will be nearly impossible.Moreover, due to Total Research Transparency, for virtually all AI algorithms, we are deliberately publishing them.Moreover, due to Total Research Transparency, for virtually all AI algorithms, we are deliberately publishing them. Therefore, our primary line of defense here is to limit algorithmic progress made by the legal projects. Still, this will be the primary way that covert projects acquire AI algorithms for the majority of Plan A.
Software progress made by the covert project. This will be slow by default because of the extreme limitations on compute under which covert projects are operating. Additionally, Plan A aims to prevent covert projects from being able to use the most capable available AI systems for AI R&D. But, they may be able to gain access to highly capable AIs through several routes which could dramatically accelerate the covert project rate of software progress:
Distillation from the legal AIs. We’re optimistic about mitigating distillation via standard techniques like hiding intermediate computation, but still expect the covert project to have access to somewhat more capable AI systems due to distillation.
Usage of the legal inference AIs to do AI research. This can be mitigated by refusals by default.
Theft of the legal AIs and then training away refusals. This can be mitigated by model weight security. We are much more optimistic about securing model weights than algorithms because of their size.
Section 2: Preventing covert compute procurement
We estimate that, under competent US execution of our policies in Plan A, a competently-executed PRC diversion effort beginning one year in advance of the deal would be able to acquire between 0.1% and 1.4% (80% CI) of the world’s AI-relevant compute at the start of the deal, without the US getting unambiguous evidence of violation. We show this in various units in the table below.If the PRC acted with less foresight, and only began diverting chips three months in advance rather than a year in advance, their diversion yield might be ~3x less. If the US was less competent, and failed to reduce the smuggling rate relative to today in the year leading up to the deal, the diversion yield might be ~1.5x more.If the PRC acted with less foresight, and only began diverting chips three months in advance rather than a year in advance, their diversion yield might be ~3x less. If the US was less competent, and failed to reduce the smuggling rate relative to today in the year leading up to the deal, the diversion yield might be ~1.5x more.
Median estimate | 80% CI | |
Fraction of world | 0.51% | [0.11%, 1.4%] |
H100e (Using our median estimate of world compute in Jan 2029) | 1.5M H100e | [310k, 4.13M] |
IT equipment cost (Using our median estimate of $/H100e China would pay) | $15B | [$3.6B, $65B] |
Compared to our median estimate of the leading AI lab’s R&D compute in Jan 2029 (36M H100e) | 24x less | [115x less, 9x less] |
The rest of Section 2 justifies these numbers.
In Section 2.1, we explain our verification approach, which centrally involves accurately determining the amount of compute in the world, then tracing sales to their end user to verify the physical location of as much pre-deal compute as possible.
In Section 2.2, we describe the strategies covert projects could use to subvert our verification approach. All of these routes require taking action before the deal is implemented; here we assume China begins diversion one year prior to the deal:
Source | Chip origin | P(feasible) | Median [80% CI] | |
Smuggled Western chips | 2028 flow | US-designed | 100% | 0.23% [0.02%, 1.17%] |
False decommissioning | pre-2028 stock | US-designed | 60% | 0.31% [0.21%, 0.42%] |
Concealed domestic production | stock + flow | China-designed | 25% | 0.14% [0.05%, 0.30%] |
Exempt military compute | Verification- exempt | either | Not estimated | not estimated |
Finally, in Section 2.3, we forecast how much compute each route would yield and make an aggregate estimate for how much compute a covert project could acquire.
Section 2.1: Policies for limiting covert project compute acquisition
AI-relevant compute produced after Plan A is implemented is carefully tracked and monitored.The mechanics of this are discussed more in our verification supplement here.The mechanics of this are discussed more in our verification supplement here. Therefore, the most feasible path for covert projects to acquire compute during Plan A is from compute that was produced before the deal’s implementation date.
The existing compute will eventually end up in one of the terminal states below:
In short, we recommend
Tracing the records of the inputs to AI compute production to get an accurate idea of how much compute was actually produced before the deal,
Tracing the records of the sales and resales of finished AI compute to ultimately trace it to its current physical location
Respectively, these are focused on the leftmost node of the flowchart, the blue tree within the flowchart, and the Unrecorded Sales section of the flowchart.
Recommendation #1: Audit the semiconductor supply chain to get an accurate record of all pre-deal compute production
To exhaustively account for all pre-deal compute, we need to start from accurate declarations of pre-deal production. In other words, we need to ensure AI accelerator vendors (like Nvidia and Huawei) declare all of their pre-deal sales. To make underreporting of production as difficult as possible, we recommend that the US demand records from the various upstream suppliers involved in producing AI accelerators.
Concretely, you would start with sales declarations from Nvidia and AMD for merchant GPUs; Broadcom and Marvell for custom AI accelerators; and Huawei and Cambricon in China. Then these companies would declare their inventories, suppliers, and purchase history, for example:
logic fabrication from TSMC, Samsung Foundry, Intel Foundry, and SMIC;
high-bandwidth memory from SK Hynix, Samsung, and Micron;
advanced packaging from ASE/SPIL, Amkor, and JCET.
Then these upstream suppliers would declare their suppliers, purchases, inventory, and cumulative production. At each layer, the flows can be cross-checked, such that underreporting production without introducing inconsistencies would require some combination of
Underreporting purchases of upstream inputs, which means that the upstream supplier would also need to underreport production,
Overreporting inventories of an upstream input, which would be caught during physical inspection of the declared inventories
Overreporting losses in the production step, which would require all subsequent nodes to underreport their input flows to match.
Notably, many of the suppliers involved in producing Chinese chips are not headquartered in China. It is conceivable that the PRC could coordinate this falsification regardless. To further deter the PRC from attempting to underreport production, the US can
Interview employees at Chinese wafer foundries, accelerator vendors, etc. as part of a whistleblower program.
Use national intelligence (such as cyber operations, human sources, etc) in the years leading up to the deal to search for evidence of falsification in the Chinese supply chainAs above, we are very uncertain about the tractability and potential consequences of doing this. Relevant public information includes the NSA’s breach of Huawei in 2010, which seems to meet or exceed the level of access that would be required to gain accurate accounting records (for example, they “obtained information about the workings of [Huawei-produced networking products] and monitored communications of the company’s top executives”).As above, we are very uncertain about the tractability and potential consequences of doing this. Relevant public information includes the NSA’s breach of Huawei in 2010, which seems to meet or exceed the level of access that would be required to gain accurate accounting records (for example, they “obtained information about the workings of [Huawei-produced networking products] and monitored communications of the company’s top executives”)..
We consider it likely but not certain that Chinese companies would make accurate disclosures under these conditions. Though we don’t include this as a core recommendation or rely on it in our estimates, the supply chain can and should take action before the deal to make the verification process easier (example).
Recommendation #2: Trace declared compute sales to their physical end-use location
We recommend that the Consortium verify the location of all pre-deal compute in two layers: exhaustively auditBy “audit,” we mean: obtain a firm’s complete records of compute purchased, resold or transferred, and retained as inventory, and physically inspect any compute the firm presently owns.By “audit,” we mean: obtain a firm’s complete records of compute purchased, resold or transferred, and retained as inventory, and physically inspect any compute the firm presently owns. every compute owner above a size threshold by tracing vendors’ sales records downward, and randomly sample the remaining tail of small owners, tracing each sampled unit of compute to its current physical location. The result is a certificate, at a chosen confidence level, upper-bounding how much compute could have been diverted to a covert project.
Recommendation #3: Locating remaining unaccounted-for compute via incentives and aerial imagery
Direct observation of AI datacenters. We are relatively confident that intelligence services could, in advance of the deal, identify and locate every (non-specifically-concealed) AI datacenter in China, using similar methods to Epoch, SemiAnalysis, and other commercial efforts. When they do eventually get inspected, the inspectors might find some of the untraceable compute. (This method does not help with compute that has been produced but not deployed in datacenters at the time of the deal, which might be around 50M H100e.)
Buyback programs and amnesty for smuggling. We want to incentivize whoever owns the untraceable compute, or anyone who may have information on it, to come forward and reveal its location to the Consortium. The most obvious reason not to reveal missing compute is if you’re using it as part of a covert AGI project. Besides that, two reasons why private actors might be reluctant to come forward is
If the compute was obtained illegally in the first place, (e.g. smuggling, or fraud or theft for the purpose of smuggling), and the person fears punishment if they come forward, or
if the person believes they can receive a better reward selling the compute to some other actor (e.g. a state attempting a covert AGI project)
We recommend combating both of these incentives if the amount of missing compute remains unacceptably high after the previous steps are taken. For 1), states should provide complete amnesty to anyone who can deliver missing compute to the inspectors. For 2), states should offer to pay above-market prices for missing compute. Note that committing to this creates a perverse incentive in favor of smuggling, so should only occur if needed to account for additional missing compute.
Section 2.2. Covert project strategies for acquiring compute
Source #1: Diverting from smuggled flow
Applies to: Western-made 2028 flow
Even in Plan A, we forecast that significant compute will still be smuggled to China before the deal is implemented. Since the smugglers will have already broken/falsified some links in the chain, and may not keep similarly detailed records to legitimate suppliers, it may be possible to purchase compute under false identities without arousing much suspicion when supply chain tracing eventually arrives at a supplier.
Speculatively, Chinese intelligence services could set up a fake company ostensibly from a different US adversary country (e.g. Russia), and purchase chips under that identity. Since such a country might also have incentive to hide chips, the eventual compute sales tracing might not be able to conclude whether the chips are in (e.g.) Russia or elsewhere.It might also be possible to purchase chips directly from Chinese domestic producers in a similar manner, though I’m unsure whether China has export restrictions that would make this more difficult or unusual. In practice it probably does look rather unusual to export any significant fraction of your 2028 production, just because of the strategic importance of AI compute.It might also be possible to purchase chips directly from Chinese domestic producers in a similar manner, though I’m unsure whether China has export restrictions that would make this more difficult or unusual. In practice it probably does look rather unusual to export any significant fraction of your 2028 production, just because of the strategic importance of AI compute.
Source #2: Concealing some pre-deal domestic production
Applies to: China-designed pre-2028 stock AND pre-2028 flow.
To hide compute from the resale tracing and physical inspection process, Chinese intelligence services could arrange for domestic accelerator vendors to underreport their pre-deal sales. The compute production tracing efforts aim to prevent this by accounting for all the accelerators which were physically produced. Chinese intelligence services would need to arrange for accelerator vendors like Huawei to understate their orders from the various supplier companies, and for those suppliers to understate their production. Those suppliers have their own suppliers, who would also need to align their statements. (For example, to fabricate the logic dies for Huawei’s chips, SMIC depends on silicon wafers from a different firm.)
Some elements of the supply chain will be in foreign countries, making them harder for Chinese intelligence services to manipulate. On the other hand, some critical bottlenecks like high-bandwidth memory are likely being smuggled, which could complicate supply chain tracing.
To avoid having to falsify records from dozens of different upstream suppliers, missing production could be explained by varying production yields. (e.g. in lithography, some fraction of the dies on each wafer will have defects, rendering those dies useless. Foundries aim to optimize this yield, but it requires ongoing R&D and can vary over time, especially at Chinese foundries as they race to catch up to TSMC.) However, any intermediate input flow whose existence is denied by understating yields still needs to be secretly made into the finished product. (continuing the example above, understating yields to acquire untracked logic dies would still require untracked HBM (from a different supplier) and untracked packaging (also done elsewhere) to make into a finished accelerator package). So yield falsification only reduces the number of firms whose coordination would be needed.
The ideal case would be manipulating yield at the final assembly stage. As far as we know, yields here are much higher and more predictable than at earlier stages, so this seems infeasible to significantly falsify.You could also stockpile the inputs (e.g. logic dies and HBM) and finish the assembly during the deal in a secret facility. For the case we looked at (a covert packaging line) this seemed not worthwhile, namely because of the comparatively low throughput and ongoing inflow of inputs that would be required to operate it.You could also stockpile the inputs (e.g. logic dies and HBM) and finish the assembly during the deal in a secret facility. For the case we looked at (a covert packaging line) this seemed not worthwhile, namely because of the comparatively low throughput and ongoing inflow of inputs that would be required to operate it.
Overall, this strategy seems quite risky, as each additional firm added increases the number of opportunities for mistakes and for leaks. However, it’s possible the whistleblower/interview/HUMINT program are less effective than we expect, or that this sort of multi-firm falsification under scrutiny is more routine for Chinese intelligence services than we expect.
Source #3: Diverting from stock via false decommissioning
Applies to: Western-made pre-2028 stock
By default, some compute will be decommissioned between its purchase and the start of the deal. This poses an issue for compute sales tracing, since (as done today) decommissioning does not involve retroactively-verifiable destruction of the chips.
As a potential diversion strategy, Chinese intelligence services might arrange to have large Chinese clouds decommission their US-designed computeOr their Chinese-designed compute, though doing this for US-made compute has a number of advantages, including (1) power efficiency to minimize required covert datacenter footprint, (2) complementarity with other diversion strategies more applicable to Chinese-designed compute, such as falsifying production records, and (3) increased plausible deniability, though this is speculative.Or their Chinese-designed compute, though doing this for US-made compute has a number of advantages, including (1) power efficiency to minimize required covert datacenter footprint, (2) complementarity with other diversion strategies more applicable to Chinese-designed compute, such as falsifying production records, and (3) increased plausible deniability, though this is speculative. in advance of the deal (perhaps justified as a political demonstration of Chinese semiconductor self-sufficiency, or to hide evidence of previous smugglingThis would be against their interests, but could be a plausible rationale.This would be against their interests, but could be a plausible rationale.), with fake recyclers set up to accept the hardware and stockpile it rather than recycling it. Decommissioning a large fraction of compute in advance of the deal would create ambiguity about whether to treat it as possibly used for a covert project, since even if China were acting honestly, it might have no way to prove that the chips had truly been recycled.
To the extent feasible, we recommend the US intelligence community surveil the Chinese IT equipment recycling industry to guard against this strategy.
As a speculative potential strategy for making decommissioning look less suspicious, Chinese intelligence services might use cyberattacks to induce what appear to be hardware failures in installed devices, inducing unknowing datacenter operators to return some fraction of their installed hardware to server or GPU manufacturers (e.g. Huawei) for RMA, or otherwise to recycle them. Chinese intelligence services would then divert these shipments and have them documented as recycled.
Source #4: Using declared military compute to aid a covert AGI project
If the US and China aren’t able to negotiate away their unverified military compute, it could be used for AI development. The amount of military compute that could be diverted to a covert project is entirely a function of how the negotiations proceed. There are a few different ways the US could negotiate to minimize the amount of unverified military compute. To minimize the maximum amount of compute that could possibly be used for military AI development, the US should push to bring all military compute under inference-only verification. This has the drawback of incentivizing covert projects where previously they might have been judged not worthwhile, but it does improve the worst-case scenario (covert project AND declared military compute working together).
Alternatively, the US could accept a greater quantity of unverified military compute, perhaps tying the amount to the amount of missing pre-deal compute to incentivize bringing any covert stockpiles into the open. Under some assumptions these strategies might reduce the expected total amount of compute used for military R&D, at the cost of increasing the total amount usable by an aggressive adversary.
Overall, we recommend that the US should push to bring all military compute under inference-only verification.
Section 2.3 Quantitative Estimates
We consider two estimation methods for the amount of compute a high-competence covert project could acquire. We trust the first more than the second, and treat the second method mainly as a sanity-check. Currently, none of these include using declared military compute to aid a covert AGI project.
Estimate | Fraction of global total | H100e, assuming 289M at 2029.0 | Description |
1 | Median 0.51% 80% CI [0.11%, 1.4%] | Median 1.5M [310k, 4.13M] | Bottom-up estimate combining the plausible diversion strategies discussed above. |
2 | Median 0.23%80% CI [0.022%, 2.34%] | Median 660k [63k, 6.8M] | Based on a rough estimate of how much money the PRC could spend secretly purchasing AI servers in 2028, as well as our forecasts of $/H100e in 2028 |
Method #1: Bottom-up
In this method, we assume China simultaneously pursues all the acquisition strategies we identified above. We give distributions over the compute that could be acquired by each strategy, as well as the probability that each strategy is feasible. We then aggregate these together to get a distribution over the total amount of compute diverted.
Diversion route | P(feasible) | Median | 80% CI |
Decommissioning pre-2028 stock | 60% | 0.31% | [0.21%, 0.42%] |
Underreporting 2028 domestic production | 25% | 0.14% | [0.05%, 0.30%] |
Buying smuggled western chips in 2028 | 100% | 0.23% | [0.02%, 1.17%] |
Source #1: False decommissioning
If supply chain tracing and datacenter tracking are effective, installed compute cannot be removed from datacenters without explanation. So either there will need to be an explanation, or supply chain tracing and datacenter tracking will need to be subverted. We think the latter is likely infeasible and would contribute a relatively small amount of compute if feasible, so we have excluded it for simplicity.See our analysis here: Diverting pre-2028 stock by retroactively understating domestic production.See our analysis here: Diverting pre-2028 stock by retroactively understating domestic production. Therefore, the only way to divert pre-2028 stock is to falsely decommission chips that still work.
We estimate that conditional on false decommissioning being feasible, it yields 6.7% – 13.4% of the 2028.0 total PRC stock, and 0.45% – 0.9% of the 2028.0 world stock, and 0.21% – 0.42% of world pre-deal compute, conditional on the false decommissioning scheme being feasible at all (60%).
Source #2: Diverting pre-deal flow
This would entail concealing some pre-deal domestic production and routing it directly into storage for a covert project.
Assuming that the fraction of global compute produced by China stays roughly constant, and accounting for the 2028 flow, we estimate that this route yields a median of 0.13% (80% CI: [0.05%, 0.31%]) conditional on being feasible at all (25%).
Source #3: Diverting smuggled chips
Assuming independence of these parameters and simulating, this yields 0.44% of the 2028 flow [0.04%, 2.19%], or 0.23% of world compute at time 2029.0, 80% CI [0.02%, 1.2%].

The distribution of total compute diverted, aggregating across all the methods discussed above.
Method #2: Military-spending-based estimate
The numbers below are estimated less precisely than the numbers in Method 1, as this is mainly used as a sanity check.
Simulating the distributions in the table above yields a median of 0.23% of global compute, with an 80% CI of [0.022%, 2.34%].
Section 3: Detecting covert AI infrastructure
This section is more speculative than the previous section. We are not experts in any of the relevant fields and may be overlooking important considerations.
In this section we go over:
Methods for detecting covert projects (Section 3.1)
Strategies that covert projects might pursue to avoid these detection methods (Section 3.2)
Quantitative estimates for how well detection will work, on what timescales, and how this depends on covert project scale (Section 3.3).
Section 3.1 Plan A strategies for detection
Method #1: Tracking waste heat
By conservation of energy, each megawatt of electricity going into a datacenter must be matched with a megawatt going out in some form (mostly as heat). The standard means of rejecting heat from datacenters produce distinctive signatures visible from space, and hiding or disguising these signatures seems to be the primary constraint on covert datacenter concealment.
Modern AI servers are designed to use liquid cooling, where coolant is pumped through IT equipment to absorb heat. The coolant dumps its heat to water, which is normally pumped out of the data halls and cooled by outdoor chillers. These chillers are either air-cooled or water-cooled.
Air-cooled chillers run cool air through finned condenser coils, heating the air by ~10-15 degrees C. It is possible to hide the chillers from direct view, but the warmed air eventually has to escape to the atmosphere, and whatever aperture it leaves through will equilibrate to a warmer temperature than the surrounding surfaces. These warmed surfaces emit in the IR spectrum, which can be seen from space. Technology exists for minimizing the IR signature from hot exhaust ports (for example, on low-observable aircraft like the F-117).On the one hand, the exhaust from a jet engine is much hotter than the exhaust from a chiller, but on the other hand, these aircraft only need to hide from things like IR-guided missiles rather than from repeated satellite observation and analysis.On the one hand, the exhaust from a jet engine is much hotter than the exhaust from a chiller, but on the other hand, these aircraft only need to hide from things like IR-guided missiles rather than from repeated satellite observation and analysis. We are somewhat skeptical that this could be applied successfully against persistent satellite monitoring, but it could be possible. Of course, making these exhaust ports invisible isn’t strictly necessary, they could instead be disguised as part of a legitimate industrial facility.
Water-cooled chillers transfer heat from the datacenter cooling loop to another flow of water which is evaporated in a cooling tower. Under certain weather conditions, the water vapor condenses to form visible clouds. We are unsure how hard these would be to hide consistently.

Condensation visible above cooling towers at the Google Council Bluffs (East) datacenter, rejecting ~190 MW of heat.
Another method for rejecting heat is simply dumping the hot water into a river or sea, a method commonly used to cool nuclear power plants.1 cubic meter per second of water flow carries 4.2MW per degree C of temperature rise. The largest water treatment plant in the world handles ~30 m^3/s (covering around half of Chicago). The largest Chinese underground water transfer projects carry ~100 m^3/s. The largest nuclear plants discharge a few hundred m^3/s (they heat water by more than 1 degree C in order to discharge ~10GW at this rate). The largest desalination plants withdraw 1000+ m^3/s of seawater. The largest rivers in the world have average discharge in the range 30,000–300,000 m^3/s. The best-suited thermal satellite we could find has a total noise equivalent differential temperature of ~0.2 K at a resolution of 57m, better than the 100m resolution of the Landsat 8 and 9, which were used in this paper to image nuclear plant thermal plumes. So the temperature increase at the surface would perhaps need to be <0.1 degree C. This could be accomplished by drawing 10 m^3/s of water for every 4.2 MW (putting the flow rate for a 100MW facility on par with the world’s largest nuclear plants), or by dumping more megawatts into each m^3/s of water but releasing it at a greater depth (or in some sort of underground waterway) so the plume dilutes more before eventually reaching the surface. I am currently unsure whether this analysis properly accounts for time averaging or other signal processing methods.1 cubic meter per second of water flow carries 4.2MW per degree C of temperature rise. The largest water treatment plant in the world handles ~30 m^3/s (covering around half of Chicago). The largest Chinese underground water transfer projects carry ~100 m^3/s. The largest nuclear plants discharge a few hundred m^3/s (they heat water by more than 1 degree C in order to discharge ~10GW at this rate). The largest desalination plants withdraw 1000+ m^3/s of seawater. The largest rivers in the world have average discharge in the range 30,000–300,000 m^3/s. The best-suited thermal satellite we could find has a total noise equivalent differential temperature of ~0.2 K at a resolution of 57m, better than the 100m resolution of the Landsat 8 and 9, which were used in this paper to image nuclear plant thermal plumes. So the temperature increase at the surface would perhaps need to be <0.1 degree C. This could be accomplished by drawing 10 m^3/s of water for every 4.2 MW (putting the flow rate for a 100MW facility on par with the world’s largest nuclear plants), or by dumping more megawatts into each m^3/s of water but releasing it at a greater depth (or in some sort of underground waterway) so the plume dilutes more before eventually reaching the surface. I am currently unsure whether this analysis properly accounts for time averaging or other signal processing methods. Doing this without creating an IR-visible thermal plume is hard, especially since time-averaging over a yearslong operating period allows for resolution of very minute heat differences. It may be possible to hide the thermal signature amid existing sources, or use multiport diffusers at depth to disperse the plume over a wide enough area to avoid detection.

Landsat images of the thermal plume from the Haiyang Nuclear Power Plant, cooling two 1.1 GWe reactors on China’s eastern coast.
Method #2: Other sources of IMINT/MASINT
Interferometric synthetic-aperture radar (InSAR) seems potentially useful for detecting the construction phase of a covert datacenter project. Based on measuring minute displacements of the earth’s crust over time caused by tunneling, an LLNL technical report investigated using this for characterizing underground construction in a known location, but it seems less well suited to noticing an unknown facility.The technique seems to require stacking multiple images taken at different times from near-identical orbital geometry, which seems poorly suited to broad-area detection. It’s possible this could be solved via brute force with enough satellites.The technique seems to require stacking multiple images taken at different times from near-identical orbital geometry, which seems poorly suited to broad-area detection. It’s possible this could be solved via brute force with enough satellites. It might be useful to gain additional evidence that a suspicious site is in fact a covert datacenter (e.g. by observing the shape of the excavated volume).
Any underground excavation needs to move the dirt and rock (termed “spoil”) to the surface. This may be hard to hide: options include hiding the excavation in another industrial project, putting tunnel entrances inside buildings and shipping out the spoil in trucks, or dumping the spoil into rivers.
Finally, the space-based ground-moving target indicator (GMTI) system currently under development by the US Space Force and NRO could be used to track individual logistics vehicles involved in constructing or operating a covert datacenter, depending on the program timeline.A Space Force spokesperson said to Breaking Defense that dedicated satellites for this purpose are expected to launch in 2028. The program’s FY2026 RDT&E budget schedule separately lists the GMTI “Develop and field GMTI” activity as running through Q4 FY2030, indicating a planned fielding window extending to the end of fiscal 2030, though the document does not identify an IOC or FOC date.A Space Force spokesperson said to Breaking Defense that dedicated satellites for this purpose are expected to launch in 2028. The program’s FY2026 RDT&E budget schedule separately lists the GMTI “Develop and field GMTI” activity as running through Q4 FY2030, indicating a planned fielding window extending to the end of fiscal 2030, though the document does not identify an IOC or FOC date. We are unsure how much this would aid detection, but to the extent it can be modeled as “a video map showing every unobstructed vehicle in China,” it seems extremely useful in conjunction with the massive amounts of AI labor that Plan A would bring. If the system needs to be tasked to a narrow area to track vehicles, it becomes less useful for discovering remote sites, but still useful for gathering evidence about an area already flagged as suspicious.
Method #3: SIGINT
Building covert datacenters would require specialized contractors. For example, following the strategy we discuss below in Building invisible datacenters, you would need
Excavation
Hydraulic engineering, to design and build the hidden cooling intake structure and outfalls
Nuclear engineering
All the contractors and disciplines involved in a normal datacenter project
Possibly TEMPEST or equivalent
It seems plausible that the US could identify a relatively small number of contractors among whom one would almost certainly be involved in a covert project. Cyber operations against all of those firms might be sufficient to rule out the existence of a covert project. We are uncertain about the tractability and potential consequences of doing this. Relevant public information includes the NSA’s breach of Huawei in 2010, in which they reportedly “obtained information about the workings of [Huawei-produced networking products] and monitored communications of the company’s top executives”.
Method #4: HUMINT
Though counterintelligence would likely carefully vet the individuals with full, incriminating knowledge of a covert AI project, it seems possible that foreign intelligence assets would be read into the program. Obviously very little public information is available to estimate how likely this is. Our vague sense is that China is considered one of the more difficult countries for Western HUMINT collection.Between 2010 and 2012, many CIA sources (seemingly almost all?) in China were killed and imprisoned. The New York Times in 2017 wrote that “the C.I.A. has tried to rebuild its network of spies in China… an expensive and time-consuming effort.” [refer to 2023 Burns comments about progress]Between 2010 and 2012, many CIA sources (seemingly almost all?) in China were killed and imprisoned. The New York Times in 2017 wrote that “the C.I.A. has tried to rebuild its network of spies in China… an expensive and time-consuming effort.” [refer to 2023 Burns comments about progress]
As part of the deal, we recommend that the US negotiate for a whistleblower program under which anyone with evidence of compute diversion, covert AI datacenter construction, or other deal-undermining activities can be interviewed and guaranteed protection against retaliation. We are uncertain about how much diplomatic allowance here is realistic, considering that a program like this would also aid normal non-AI-related HUMINT efforts in China. It’s possible that having a neutral third party operate the whistleblower program in China would be more acceptable than if the US were operating it.
Method #5: AI silver bullets
Note: Everything in this section is significantly more speculative than the rest of this document because it assumes extreme AI capabilities.
During plan A, we scale to extremely capable AI; reaching AI as capable as top human experts at everything. This accelerates economic and scientific progress, which could substantially improve detection capabilities:
Reliable Lie Detectors. AIs could accelerate the development of reliable lie detectors. Existing lie detectors (e.g. polygraphs) can relatively easily be fooled. Reliable lie detectors might be developed in the near future over the course of normal human technological development, but extreme level of AI capabilities could accelerate the timeline. It’s also unclear to what extent the development of lie detectors would be politically accepted. If it’s possible for a regime with a covert project to decline being subject to a lie detector on principle, lie detectors won’t be helpful. On the other hand, the US could leave the deal unless China and/or other counterparties agreed to submit senior leaders to lie detection.
Privacy preserving AI verification technology. AIs could enable intrusive verification to be done in a privacy preserving fashion, i.e., without leaking unacceptable information. This could be implemented as follows. First, place a huge number of sensors (e.g. cameras, thermometers), and mobile robots across China, such that if the US were able to observe each of them, they would be confident there is no Chinese covert project. Then, give all of that information to an AI trained by the US on a secure server on Chinese soil, and have the AI analyze that information for worrying signs. Finally, allow the AI to communicate a very high level summary of its findings, (for example, “Is China running a covert project?”), and delete the rest of the information. In principle, this approach could enable verification intrusive enough to usefully rule out covert projects without divulging any sensitive state secrets.
Extreme industrial verification. By 2035, real US GDP in Plan A has grown by a factor of 25 relative to today. This enables extreme feats of industrial capacity. In an extreme case, all relevant manufacturing and industrial processes (e.g. aluminum plants that could be used to hide the waste heat of a covert project) could be disassembled and replaced/reassembled in a verified location (i.e., a robot SEZ). Again, there may not be political willingness to pay this cost, however, if there is substantial pressure (e.g. because both sides strongly want to continue the cooperative Plan A regime, but there is some evidence of a covert project, so one side wants higher assurance), it is plausible that this could happen, especially if combined with privacy preserving AI verification.The most extreme case of this proposal is something like: (1) move all humans to other planets with spaceships, (2) verify that no covert GPUs were transported to other planets, (3) make sure that no compute is diverted from the newly manufactured compute on other planets, and (4) disassemble the earth. This is obviously an extreme case meant to be illustrative, in practice we expect that with all of this capacity, something more efficient and far less costly could be done.The most extreme case of this proposal is something like: (1) move all humans to other planets with spaceships, (2) verify that no covert GPUs were transported to other planets, (3) make sure that no compute is diverted from the newly manufactured compute on other planets, and (4) disassemble the earth. This is obviously an extreme case meant to be illustrative, in practice we expect that with all of this capacity, something more efficient and far less costly could be done.
AIs will be making vast amounts of scientific progress over the course of Plan A (probably something like the equivalent of 100s of years of human equivalent progress over the course of the 2030s). Historical technological progress has resulted in all sorts of technologies that are helpful for detecting covert projects that would have been very hard to predict far in advance (e.g. satellite monitoring). We should expect all of this progress to result in improved detection methodologies in hard to predict ways. We should similarly expect that the technological progress would help concealment of covert projects as well; however, our general guess is that we expect technological progress to favor verification over concealment.
Section 3.2 Covert project strategies for avoiding detection
In this section, we cover strategies that covert projects could take to avoid detection (classified as “invisible” or “disguised”.
How distributed would covert datacenters be?
Why one big datacenter instead of many small ones? Covert projects would aim to minimize project detection hazard per unit compute. At sufficiently small scales (e.g. a single 8xH100 server), detection hazard per site seems to scale sublinearly with site size (the detection risk to a group of four different 8xH100 servers in four different cities is higher than if all four servers were placed in a single rack). To minimize detection hazard per unit compute, you should therefore increase site size up to the point where detection hazard per site starts to scale superlinearly with site size, or until some other constraint prevents you from scaling up further.
Method #1: Disguised Datacenters
Compared to “invisible datacenters”, “disguised datacenters” have the disadvantage that any specific, visible facility could be subject to focused scrutiny and potentially challenge inspection.Even if challenge inspection provisions are not included in the initial agreement, China might later come under international pressure to permit them if fear of covert projects becomes a binding constraint for Consortium scaling decisions.Even if challenge inspection provisions are not included in the initial agreement, China might later come under international pressure to permit them if fear of covert projects becomes a binding constraint for Consortium scaling decisions. For example, the US could assemble a ranked list of China’s largest industrial loads and task satellites to frequently revisit them, collect SIGINT on the companies that ostensibly operate them, and attempt to recruit human sources who work there or live nearby. For a plant that produces a visible material output, it seems plausible to directly track shipments out of the facility and compare that with the heat rejection. (e.g. a datacenter disguised as an aluminum smelter would produce far less aluminum than is typical for a smelter of similar power draw.) You could try to get around this by making the datacenter a small fraction of the total power draw (e.g. 1% – 10%), but this would restrict the datacenter to on the order of 10 MW – 100 MW (we are not aware of any industrial loads on the order of 10 GW or higher.)
However, “invisible datacenters” require more complicated engineering and may not be feasible, especially against potential new AI-produced remote sensing capabilities developed during Plan A.
Some specific industrial facility types that might be useful cover stories:
Aluminum smelters. China produces 43 MT of aluminum per year, aluminum smelters require ~1.6 gigawatts per Mt/yr, which naively implies ~70 GW of aluminum smelting capacity in China.
Nuclear power plants.
Combustion power plants
Example: hiding a datacenter under an aluminum smelter
In this example, the IT equipment would be located underground, with any excavation happening indoors to avoid visible spoil piles or excavation equipment. The power would be diverted from an onsite coal plant or grid connection (ostensibly to support the smelter). This would require the facility to produce less aluminum than its power draw or thermal signature would imply. But based on our research, it doesn’t seem possible to accurately measure the actual aluminum shipped from an aluminum smelter using national technical means.It might be possible to measure perfluorocarbon emissions from the Hall-Heroult process, which could distinguish how much energy is being used on actual aluminum production. We’re not aware of any way to measure this from space at a per-facility level though.It might be possible to measure perfluorocarbon emissions from the Hall-Heroult process, which could distinguish how much energy is being used on actual aluminum production. We’re not aware of any way to measure this from space at a per-facility level though. Though the heat signature of a datacenter is different than that of an aluminum smelter, it doesn’t seem obviously thermodynamically prohibited to design a cooling system which dissipates heat from the IT equipment in a way that resembles the heat dissipation from aluminum smelters.
Method #2: Invisible Datacenters
Trying to make a datacenter look like nothing seems more complex than making it look like something else. The biggest difficulty we’ve identified is hiding the heat rejected by the IT equipment (and any onsite power generation). In theory, all the equipment can be built underground and shielded from the view of satellites, but the heat needs to leave the facility somewhere. See the Remote sensing, IMINT, and MASINT section. Besides the ongoing signatures, the initial construction period will also show activity, though avoiding detection has some precedent.
Example: once-through cooling + nuclear power
One potential strategy might be to build or extend an existing underground facility, install sufficiently many small nuclear reactors there to power the IT equipment, build an intake structure deep underwater on the coast, and use some sort of deep-water multiport-diffuser type design for the outfall to try to ensure that the hot cooling water dilutes to a temperature below the IR noise floor before reaching the surface. We aren’t sure whether a diffuser could be designed to those requirements, but it seems plausible. If so, a facility of this type seems quite difficult to detect via remote sensing once operational. The infrequent or nonexistent refueling requirements of the nuclear reactors would make even full knowledge of China’s electrical grid, natural gas distribution, etc. insufficient to locate the facility. We do not think existing technologies for detecting hidden nuclear reactors would be effective in this case, but again, we are not experts and could be wrong here.
Hybrid strategies
China could also target “invisibility” for some aspects of the facility while disguising others. For example, it might be possible to site the IT equipment and power generation inside a logistics warehouse (e.g. receiving shipments of diesel fuel by truck). This has the advantage that there are many logistics warehouses, so the close-scrutiny consideration mentioned above might not apply. However, since a logistics warehouse doesn’t reject much heat normally, the cooling would need to be rendered “invisible” via one of the techniques discussed above.
Section 3.3 Quantitative estimates
How much datacenter capacity does the diverted compute require?
With the parameter values above, a covert datacenter with 1 GW total facility power could support 1.9M H100e of 2028-era smuggled US chips, or around 1M H100e of decommissioned US chips, or 525k H100e of 2028-era domestically produced chips.
How much time would construction take?
Today, Epoch finds that gigawatt-scale datacenters can likely be built in two years or less.

Considerations pointing towards construction being faster than this:
By 2029, constructing datacenters of this scale will be more routine and likely faster than today.
Additionally, being a top national priority for the Chinese government might enable a faster timeline than usual.
In Plan A, we forecast that construction labor will be largely automated by robots. We are uncertain about the timing here, and this might not be a significant factor at the beginning of the deal.
Considerations pointing towards construction being slower than this:
Clandestine underground facilities would require bespoke engineering that normal datacenters do not.
Construction speed may be constrained by factors other than labor or spending (for example, if excavating a large underground cavern in a remote area, there may be a maximum rate at which spoil (dirt, rocks from excavation) can be carried away for distribution across a wide area to avoid creating satellite-visible piles).
If significant human labor is required, the requirement to avoid spies might make it difficult to find workers. Separately, compartmentalizing parts of the project could constrain parallelizability (e.g. you might want to finish excavation before installing any IT equipment to reduce the number of people who know IT equipment is being installed there.)
We use two years as a point estimate.
What detection risk do covert projects of various sizes incur?
We use the following methodology to estimate the probability of detection.
We are uncertain about how detection hazard scales with site size. Lacking any reason to expect detection hazard to scale superlinearly in site size, we suspect that the optimal strategy resembles “make sites as large as possible subject to feasibility constraints”. Informed by the strategies above, we estimate a distribution over the optimal site size (where optimal is basically an operationalization of “maximum feasible site size”.)
For a few possible “optimal site size” conditionals, we estimate the probability of detection over the 10 year period in plan A conditional on 1 such site, and 10 such sites.
We interpolate these intuitions with a shared-frailty model. This lets us produce 10-year detection probabilities for sites operating arbitrary amounts of compute (conditional on specific optimal site size outcomes, but we can also marginalize over that.)
We also estimate 3-year detection probabilities and interpolate the times with a weibull-type hazard rate (hazard rate varies with a power of time). This lets us produce survival curves over time for covert projects operating arbitrary amounts of compute.
Our current operationalization of “detection” is focused on the physical infrastructure. We require that
Among locations that the US assesses as suspicious, the one they assess as the most likely covert datacenter candidate is in fact a covert datacenter
The assessment of the US IC is >50% confident that this candidate is a covert datacenter. (it’s possible this threshold should be higher).
The goal of this operationalization is to match the point where the US would feel diplomatically comfortable requesting an inspection of that site, and disclosing it to the international community if needed to generate additional pressure.
Unfortunately, due to time constraints, the numbers used in the above process are based entirely on intuition and qualitative reasoning about the detection methods and hiding strategies we considered. It seems possible to apply a historical reference-class approach, and we’ve informally looked at various historical examples of covert projects. But due to the number of asymmetries between this case and the historical examples we could find, and the lack of an obvious conclusion from our informal research, we did not judge a quantitative reference-class estimate to be worth devoting time to.
Therefore our estimates of detection probability over time should be treated as highly speculative. We choose to include them anyway for the sake of reasoning transparency, since the ability of states to detect covert projects is an important input to our views.
Section 4: Limiting Consortium uplift to covert projects
In Plan A, legal projects worldwide develop highly capable AI and robotics over the course of the 2030s. Some of this progress will spill over to covert projects, making it easier to build AGI in secret than it would have been under a global moratorium on AI development (as in Plan S). Our best guess is that this increased risk is worthwhile, due to the alignment, power concentration, and war-avoidance benefits of building more capable AI under the auspices of a deal; see discussion in our Comparing Possible Plans supplement for more on why.
In this section, we summarize the various Plan A policies affecting the degree to which Consortium AI progress benefits covert project R&D. We also estimate
Section 4.1: Relevant Plan A policies
Recommendation #1: Total research transparency
In Plan A, nearly all AI research is fully transparent to the public, including AI algorithms, code, and documentation, for all frontier AI projects. AI model weights, significant fractions of the training data, and a small amount of other sensitive information is prevented from leaving the datacenters.
Recommendation #2: Inference monitoring to prevent covert project uplift
See our inference monitoring and restrictions here.
We do not plan to offer zero-data-retention (ZDR) inference on models whose capabilities we would not want proliferated to covert projects. See our ZDR policy here.
Recommendation #3: Model weight security
We think achieving high-confidence model weight security is extremely important for Plan A, and the Consortium probably should not proceed very far above pre-deal capability level until they are ~95% confident that the best-equipped actors in the world cannot steal the model weights. Forecasting the likelihood of this happening is outside the scope of this supplement.
Section 4.2: Covert project strategies
Path #1: Distillation from Consortium models
Currently, Chinese labs like Minimax and Zhipu are spending around 20–30x lessIn 2024, OpenAI spent $5B on R&D compute, Zhipu AI spent $308M (page 10) on total R&D (which included 657 technical staff; assuming $100k/yr each yields $242.3M on compute), and Minimax spent $189M (page 11) total.In 2024, OpenAI spent $5B on R&D compute, Zhipu AI spent $308M (page 10) on total R&D (which included 657 technical staff; assuming $100k/yr each yields $242.3M on compute), and Minimax spent $189M (page 11) total. on R&D compute than US frontier labs like OpenAI. From our forecasts, it seems plausible that a covert AI project would start off with ~20-30x less compute compared to the leading AI company at the start of the deal.The compute advantage will widen in some sense with time; in our scenario the Consortium increases their compute at ~3x per year. However, this will only translate straightforwardly to capabilities through training compute scaling (research compute scaling on the other hand will be restricted to target limited amounts of software progress.)The compute advantage will widen in some sense with time; in our scenario the Consortium increases their compute at ~3x per year. However, this will only translate straightforwardly to capabilities through training compute scaling (research compute scaling on the other hand will be restricted to target limited amounts of software progress.) So while the requirement of secrecy would no doubt be a burden, researchers running a PRC covert project might not feel much more compute constrained initially than when they were operating legally. Despite their compute advantage, US labs have not obviously managed to lengthen their lead in model capabilities.Depending on how you measure this; the US Center for AI Standards and Innovation found an apparently growing US lead; here we use the Epoch Capabilities Index. See “The US-China AI gap is very uncertain”.Depending on how you measure this; the US Center for AI Standards and Innovation found an apparently growing US lead; here we use the Epoch Capabilities Index. See “The US-China AI gap is very uncertain”.

We believe this is primarily due to “distillation”, used in a loose sense to refer to using state-of-the-art closed source models to generate synthetic data and grade RL trajectories. These techniques can allow for highly compute-efficient training, in some cases offering more than an OOM of compute equivalent gain.As a (somewhat dated) example, Phi-4 used GPT-4 to generate specialized synthetic data, exceeding GPT-4’s ECI while requiring around 20x less training compute.As a (somewhat dated) example, Phi-4 used GPT-4 to generate specialized synthetic data, exceeding GPT-4’s ECI while requiring around 20x less training compute.
Will covert projects be able to stay 3–9 months behind the frontier then, like today’s Chinese labs? On one hand, covert projects will get an additional boost at the start of the deal compared to today’s trailing labs: unlike today, all training recipes will become transparent. This initial sharing of non-overlapping insights will bring forward the Consortium and covert frontiers together at the start of the deal. On the other hand, Plan A tries to mitigate distillation by e.g. restricting companies from revealing intermediate model outputs in their APIs.
We are uncertain how effective these measures will be against a determined adversary. The large number of different AI companies, applications, and deployments will provide a correspondingly large attack surface. A government could also pressure one of their country’s AI projects to intentionally leave a loophole open for distillation.
It is essential that distillation not enable the proliferation of destabilizing dual-use R&D capabilities, which might come fairly early into the deal (e.g. between our SAR and TED-AI milestones). For this reason, we recommend that the Consortium conduct ongoing red-team experiments where a group of researchers attempt to distill the current frontier models under similar circumstances a covert project would face. Depending on the results, it might be necessary to halt external deployment of more capable models until more assurance can be achieved.
Path #2: Extracting R&D labor from Consortium models
Consortium models running in inference datacenters will have classifiers that flag attempts at anything that looks like AI R&D. (In Plan A, models that would provide AI R&D uplift above the pre-deal SOTA are not offered under ZDR.) It seems risky to iterate against these things; we do not view this as a promising strategy.
Overall, this path seems fairly unlikely to work at all if the Consortium is competent. If it did, covert projects could potentially make use of models with superhuman experiment selection skill (or “research taste”) in AI R&D, biology R&D, and other dual use domains, significantly accelerating their progress.
Section 4.3 Quantitative estimates
Parameter | Estimate | Rationale |
Fraction of Consortium software progress that leaks to covert project | Median: 67% 80% CI: [36%, 90%] Beta distribution | See our estimate here for the median. The 80% CI is an intuitive guess (specifically, the p90 was chosen at 90%, and a beta distribution was assumed, which forces the p10 to be 36%). |
Section 5: When is a covert project catastrophic?
To become catastrophic, a covert project must achieve one of three things:
Insurmountable capabilities lead. This occurs if a covert project is able to get enough of a lead over the Consortium that even if the Consortium were to pivot back to racing quickly, it would not be able to stop the covert project (or its sponsor nation) from reaching AI capabilities sufficient to prevent any of the following outcomes.
Balance-of-power warping capabilities. This occurs if the covert project significantly changes the geopolitical balance of power, e.g. by granting a large economic, military, or political lead. The extreme version of this involves a decisive strategic advantage—i.e. granting the country with a covert project a position of complete strategic superiority over their rivals.
BATNA warping capabilities. This occurs if the covert project AIs empower their host countries to such an extent that the counterparty prefers remaining in the deal (despite knowing about the cheating) over pulling out of the deal and trying to go at it alone.
If none of these are true, then: (i) the covert project does not have an insurmountable capabilities lead: legal projects can catch up if they want, (ii) all sides roughly preserve their within-deal balance of power, and (iii) all sides can credibly threaten to leave the deal to coerce a known defector into shutting down their covert project. As long as each of these are true, then, the defection from the deal isn’t catastrophic.
We define the capability level that a covert project would need to reach to satisfy any of these conditions as deal-undermining-AI. Deal-undermining-AI is not a fixed capability threshold, that threshold will vary over time depending on the level of capability available to other nations.
We’ll now discuss how difficult it is for a covert project to build a deal-undermining AI.
Section 5.1: Paths to a covert project catastrophe
Insurmountable capabilities lead
Achieving an insurmountable capabilities lead is extremely hard because of the massive compute deficit faced by the covert project.We’re assuming here that covert projects are only able to acquire untracked pre-deal compute, i.e. that it’s not feasible for them to smuggle post-deal compute or to create a secret chip supply chain. After extreme capabilities, it would be possible for covert projects to set up a secret compute supply chain, but we don’t anticipate this being the binding constraint, i.e. this would happen after other capabilities we consider to be catastrophic.We’re assuming here that covert projects are only able to acquire untracked pre-deal compute, i.e. that it’s not feasible for them to smuggle post-deal compute or to create a secret chip supply chain. After extreme capabilities, it would be possible for covert projects to set up a secret compute supply chain, but we don’t anticipate this being the binding constraint, i.e. this would happen after other capabilities we consider to be catastrophic.
Before the legal projects pause capability development at TED-AI, most of the covert project’s capabilities progress will come from implementing the new training recipes developed by the legal projects (assuming no distillation is occuring). Therefore, the “effective training system performance” gap will be equal to the hardware gap, which, at the beginning of Plan A in the scenario, is around 24x (80% CI: [115x, 9x]) from a covert project to the largest legal project (assuming the covert project has ~1.5M H100e, and the biggest legal project uses ~36M H100e for R&D).
There are two advantages for covert projects:
First, they may not slow down for safety. At points within Plan A, capability progress will slow down dramatically or pause because of safety concerns. In this situation, the covert project will make some software progress itself, but very slowly.See further analysis here.See further analysis here.
Second, due to the transparency and difficulty of algorithmic security, algorithms flow from the legal projects to the covert projects, but not the other way around.
Even if the covert project were to get a capability lead, the Consortium could just pivot back to spending all of its compute on AI progress. At the beginning of the deal, there is a ~20x compute gap (1.3 OOMs); and by 2035 it’ll have expanded to 4.3 OOMs. If all of the newly founded algorithms are immediately shipped to the covert projects, this doesn’t help at all, but:
Some algorithms are scale dependent and therefore more helpful for the legal projects than the covert projects.
Some algorithms may be able to be kept secure (e.g., our estimate during Plan A is only ~67% of algorithms leak).
There may be some latency in the theft for the algorithms that are not radically transparent (e.g. by a few months).
Balance of power warping capabilities
The most promising capabilities that a covert project could pursue to change the geopolitical balance of power that we’re aware of are:
Persuasion / AIs for political uplift. If covert projects are able to get AIs with similar levels of capabilities to, for example, the early 2028 AIs of the race ending of AI 2027, this could make their host country substantially more politically effective than anyone else, enabling them to get substantially more power. The main limiting factor on persuasion AIs is that it is probably a high capability threshold.
Secret robot buildout. The covert project could enable a massive industrial buildout. e.g. in underground tunnels, giving much greater industrial capacity than the rest of the world. This is particularly relevant because of the cap and trade regime over robot production in the Plan A scenario.An extreme version of this could undermine Nuclear MAD, but this would be very difficult. It would need to happen in secret or else other countries would respond by building up nuclear stockpiles or intervening on buildup.An extreme version of this could undermine Nuclear MAD, but this would be very difficult. It would need to happen in secret or else other countries would respond by building up nuclear stockpiles or intervening on buildup.
Military R&D. If covert projects have substantially more capable AIs than the legal AIs, or the legal AIs are all subject to aggressive refusals, covert projects could help design cheaper and better weapons (e.g. drones, missiles, ships, planes, bombs, etc).
Exotic novel technology. This includes nanotechnology, mind uploading, certain kinds of bioengineering, and unknown unknowns. Any of these would be extremely destabilizing; the main uncertainty is how difficult they are, i.e. what level of AIs over what time period are necessary for achieving them.
For military R&D and the robot buildout, the main limitation will be succeeding sufficiently in a manner that other countries cannot detect. If other countries notice, they can respond by requiring improved verification, requiring further arms control deals, or ramping up production themselves to restore the relative balance of military power.
BATNA-warping capabilities
The main paths for a covert project to greatly improve China’s BATNA are:
Splendid First Strike on the Compute. China could pull out of the deal, destroying US datacenters, plus gain an asymmetric advantage over the US by either (i) destroying the US hardened weight or chip stockpile, or (ii) Protecting some compute that was meant to be destroyed in the case of a deal dissolution. Having a sufficient cyber lead to overcome the inference compute disadvantage of the legal projects could enable covert projects to hack into the US model weight vault and destroy the copies. Protecting the compute probably requires significant physical capabilities, i.e. a secret robot buildout.
Speed up the robot production ramp after the deal breaks down. For example, you could covertly get a robot research lead on the rest of the world so that when the deal breaks down, you are starting off with better robot designs. Alternatively, you could build a hardened robot stockpile, which would create a higher starting population of robots from which to bootstrap.
For this, as well as the above, there are also unknown unknowns (e.g. new military tech we don’t know about yet). The unknown unknowns are probably scarier than any of the concrete paths we’ve identified.
Section 5.2: Operationalizing Deal Undermining AI
In the quantitative analysis for this supplement, we simplified the modeling by assuming that deal undermining AI is a specific capability threshold that doesn’t change over time.
What specific threshold do we choose? We are currently operationalizing this as “Top Expert Dominating AIs (TED-AI)”, or AIs that are better than the top expert at every domain. Since AI capability profiles are currently fairly spiky, TED-AIs will be substantially superhuman at the skills they are strongest at, while merely being top expert level at the skills they are weakest at.
Why not a higher threshold? Once AIs are beyond the human range, it’s hard to limit the possibility of novel exotic technology which we don’t know how to defend against. Moreover, it probably won’t make a huge difference, because the intelligence explosion beyond TED-AI is probably going to be quite fast.
Why not a lower threshold? AIs that are within the human range seem much less likely to develop the sorts of novel technology that could be geopolitically destabilizing. Even very large advances (e.g. much cheaper interceptors) seem pretty far from undermining MAD.
We hope that future analysis improves on this baseline. For example, another reasonable operationalization of “deal undermining AI” is “any AI that’s more capable than the legal projects”, becuase…
Covert projects can’t have an insurmountable capability lead if they don’t have a lead at all.
AIs of similar capability levels probably can’t change the geopolitical balance of power, if the defenders are deploying AIs of greater capabilities research to red team. especially given Consortium defenders will have vastly more compute (e.g. for finding and patching vulnerabilities). This will depend on the strategy taken by the legal projects, i.e., assuming that legal projects are allowed to apply AI to military R&D. In Plan A we currently recommend some arms control limitations on the use of AI to build new superweapons. The main downside of that policy is that it increases the chance that covert projects could be used to gain large military advantages.
However, this strategy likely requires more aggressive capability scaling than is needed.
It’s probably fine for covert projects to get a lead at lowish levels of capability (e.g. automated coder) because at that point, they aren’t strong enough that small gaps in capability seriously alter military / economic balance of power. In the scenario, for a very brief period of time during 2029 (before automated coder), legal training is paused while verification is set up, and so in the scenario, covert projects in principle might be able to get a lead during this window, though in practice it is unlikely because of the hardware gap (also, covert projects likely will need even more time to set up their infrastructure, unless they are using existing exempt military datacenters).
Closer to the top of the human range (e.g. at or beyond top expert level AI) it’s easy to imagine a small AI capability quickly snowballing into huge capability / political power. Things like superhuman political strategy and persuasion become more important and are almost impossible to defend against.
Section 6: Overall risk estimate
The estimates in this section are based on simulating covert project capabilities progress as described in our Takeoff Supplement.
In this section we quantify how the Consortium’s scaling strategy affects the risk that covert projects develop deal-undermining AI.Importantly, we do not include the risk due to weight theft during the deal in these estimates.Importantly, we do not include the risk due to weight theft during the deal in these estimates. We only evaluate fixed strategies here, of the form “aim to reach TED-AI some number of years into the deal, then pause there to do alignment research until the risks of delaying handoff outweigh the benefits.” This is somewhat unrealistic, because in reality the decision-makers implementing Plan A will be able to flexibly respond and adopt their strategy based on evidence received.
For example, if the US received evidence that reaching TED-AI within 6 years would require a massive amount of software progress, they could push the rest of the Consortium to follow a more gentle scaling trajectory that leaned more on hardware buildout, thereby leaking less software to the covert projects. This would reduce the risk due to covert projects, though it would also increase the risk that the deal breaks down before TED-AI has been reached safely. They could also adopt filtered transparency instead of Total Research Transparency, which would leak a smaller fraction of software progress to covert projects, enabling faster scaling for equal covert project risk, at the cost of safety and risk of deal impairment.
The actual covert project risk incurred will depend on how decisionmakers balance these factors to select a scaling strategy. Below, we show how varying the scaling duration and handoff delay affect the risk incurred.
Aim for TED-AI in 2 years | Aim for TED-AI in 5 years | Aim for TED-AI in 8 years | |
Covert project risk at TED-AI | 7% | 8% | 8% |
Covert project risk waiting 5 years after TED-AI | 18% | 13% | 11% |
Covert project risk waiting 10 years after TED-AI | 20% | 16% | 15% |
Concretely, with the evidence available to us today, we estimate that
Halting AI R&D and initiating Plan A roughly one year before the Automated Coder milestone (in our scenario, this means implementing it in early 2029),
then making whatever software progress is needed to safely scale to Top Expert Dominating AI over the following 6 years, (early 2035, in our scenario),
then working towards a high-confidence alignment solution for 5 more years (until 2040, in our scenario)
would incur a 13% chance of covert projects developing deal-undermining AI before detection or handoff, if China competently pursued a covert project and remained committed to keeping it hidden.
The probabilities above are obtained by drawing from our full range of uncertainty about the dynamics underlying AI R&D, in addition to our uncertainty about the feasibility of various covert project strategies and mitigation strategies. We explain the details of our quantitative model in our Takeoff Supplement. Again, the decisionmakers implementing Plan A will have access to more information than we do, which might reveal that the risk incurred by the above strategies differs much more sharply than they do in our estimation.
To illustrate this, if we could be certain that AI R&D and takeoff dynamics worked exactly how it does in our scenario, but still had uncertainty about covert project affordances, the table above would look like this:
Covert project risk incurred by several strategies, assuming the AI 2040 takeoff parameters:
Aim for TED-AI in 2 years | Aim for TED-AI in 5 years | Aim for TED-AI in 8 years | |
Covert project risk at TED-AI | 2.0% | 1.6% | 1.8% |
Covert project risk waiting 5 years after TED-AI | 18.8% | 8.8% | 5.2% |
Covert project risk waiting 10 years after TED-AI | 22.8% | 12.4% | 10.8% |
Section 7: Covert project response plan
The main goals of a policy response to covert projects are:
Prevent a covert project from building deal-undermining AI. We generally want it to be the case that legal AI projects have stronger AIs than covert projects. This is less important early in the intelligence explosion but becomes crucial later on, as having better AI capabilities could quickly lead to an insurmountable geopolitical lead. Specifically, by deal-undermining AI, we mean an AI that either (i) directly confers a decisive strategic advantage to the covert project or (ii) is able to ongoingly outpace the legal AI projects, even if they switch to maximally fast AI capabilities progress.
Deter the creation of covert projects. We want it to be the case that at the beginning of Plan A, both the US and China are confident that if they were to do a covert project, it would either be detected before it is relevant, or outpaced on AI progress, and that these outcomes are worse than not having any covert project.One reason is that even if a covert project turns out to be non-viable, worlds where China diverts significant compute might force the Consortium to adopt a less favorable scaling strategy due to their uncertainty about how quickly the missing quantity of compute would enable deal-undermining AI to be reached by a hypothetical covert project. Another reason is that the tensions associated with detecting and shutting down a covert project would incur additional risk of deal decline, even if the covert project itself is small enough to be harmless.One reason is that even if a covert project turns out to be non-viable, worlds where China diverts significant compute might force the Consortium to adopt a less favorable scaling strategy due to their uncertainty about how quickly the missing quantity of compute would enable deal-undermining AI to be reached by a hypothetical covert project. Another reason is that the tensions associated with detecting and shutting down a covert project would incur additional risk of deal decline, even if the covert project itself is small enough to be harmless.
Maintain international cooperation. In particular, if there isn’t a covert project, we want to avoid Plan A breaking down due to lack of trust, i.e. one side being worried that the other has a covert project and thus wanting to exit the agreement.
Prolong the time with human-level but not superintelligent AI. Time with human level AI is very valuable because it is when we’ll make the most progress on alignment.
Goals 1 & 2 are in tension with Goals 3 & 4 because a central tool for preventing / deterring covert projects is accelerating capabilities progress within Plan A.
The rest of this section is structured as follows:
First, we discuss the set of possible responses to covert projects.
Second, we discuss an example high level overall strategy
Third, we go through concrete scenarios and how we would recommend responding to each.
Section 7.1: Options for responding to covert projects
In this section, we’ll go over policy options for situations where covert project risk seems unacceptably high. These options can be pursued in parallel. Note that many of these options (especially options 1-4) require negotiation with China.
Option #1: Switch to Plan S
Multilateral
The US and China could agree to a full moratorium on ongoing AI development. This would slow down covert projects in the regime where they are mostly benefitting from Consortium-driven algorithmic progress. The covert project could still make progress on its own, but it would be much slower because of the compute and labor bottlenecks. Overall, this maximizes the serial time in the agreement before a fixed size covert project overtakes, putting aside detection.
This option, of course, has the standard downsides of switching to Plan S. With respect to covert projects in particular, it may also delay the creation of any of the “AI silver bullets”, discussed above, which are helpful for mitigating covert projects.
The exit paths for Plan S are (1) we successfully developed improved verification technology, (2) both sides decided to grant the other improved verification access, (3) sufficient progress on alignment was made to continue scaling quickly, in a manner that outpaces covert projects or (4) covert projects pose a large, imminent risk, and misalignment risk is smaller than covert project risk (while still being substantial, or else we’d be in option (3).
The most preferable option among these is (1). In that case the plan is: issue a complete moratorium on AI R&D until we’ve got good enough verification to be confident covert projects are not a binding constraint, then revert back to Plan A.
Option #2: Adopt a scaling strategy that minimizes software share of capabilities progress
Multilateral
In Plan A, we currently allow for ~.5 OOMs/year of software scaling, and ~0.5 OOM/year of hardware scaling over 2035. The reason we recommend making algorithmic progress is in large part because of deal dissolution risk; if covert projects are a binding constraint, we could minimize algorithmic progress and scale purely with hardware, which doesn’t diffuse. It’s probably impossible to get algorithmic progress to zero (for example, there are inevitable capabilities externalities from safety research), but we might be able to decrease it to ~.1-.2 OOMs/year.
Option #3: Attempt to improve algorithmic progress security
Multilateral
Another approach is to try to prevent algorithmic progress from diffusing (or, increasing the latency with which it diffuses). Again, this would slow covert projects substantially insofar as they were primarily relying on Consortium algorithmic progress. The main difficulty is that there is a major tension between our desire for openness/transparency with a desire for algorithms to be secure: an important component of Plan A is the Total Research Transparency that deliberately publishes most algorithms onto the internet. We discuss options for increased algorithmic security in our transparency supplement.
Depending on how successful algorithms security is, it may be viable to switch to faster capability scaling. This will obviously pose additional misalignment risk, and in this regime it is hard to slow down because a covert project is close behind, so this is not a desirable position to be in.
Option #4: Negotiate with China to submit to stronger verification
Bilateral
For example, China could agree to:
Ongoing intrusive physical monitoring (e.g. thermometers in large water sources, cameras in relevant sites, electrical grid monitoring)
Provide access to classified materials that indicate there is no covert project.
Physical inspections of plausible covert project locations (e.g. industrial plants that produce substantial waste heat, tunnels near large water sources, etc).
Access to classified data, e.g. AIs can see the black budget and report back if there are any AI looking things on there
Lower the exempt military compute cap
If we have viable privacy preserving verification such that an AI could get access to these information sources and only report back high level bits of information, this could make this option much more viable. Alternatively, the US might credibly warn China that we are pulling out of the deal unless they provide us with stronger verification affordances.This may not be possible. See “BATNA-warping capabilities”.This may not be possible. See “BATNA-warping capabilities”. China would then be incentivized to agree to these concessions because the deal breaking down is very bad for them, because it would amount to something similar to a return to the status quo (e.g. a return to a race with the US having vastly more AI compute).
Option #5: Race to improved verification
Bilateral
If we have enough time, it may be worthwhile diverting massive resources into a sprint to develop improved verification technology, for example, privacy preserving verification that could enable the option above. Having access to more advanced AI is helpful for this option. We recommend pursuing this in almost any scenario because having the affordance to implement stronger verification is generally useful, but in scenarios where covert projects are posing substantial risk, this is especially promising.
Option #6: Improve takeoff speed modeling
Unilateral
Concern for covert projects may be substantially driven by the tail risk, i.e., that takeoff is faster than we think with a tiny amount of compute. It may be worthwhile to invest substantial resources to gain better information about how fast the intelligence explosion goes with a small amount of compute. An obvious baseline is to invest more into generic takeoff speed modeling (e.g. things like the AI Futures Model).
A more intense option is to create a “covert dominating miniproject (CDMP)”—a sanctioned project that would simulate the progress of a covert project; it would be given somewhat more compute than we think would be able to evade detection, and would try to conduct an intelligence explosion as fast as it could, theoretically outpacing any covert project. The main downside of this proposal is that it would lead to faster algorithmic progress (which may then diffuse to the covert project itself, further accelerating it). However, this might net out to an improvement in a regime where we are very uncertain about takeoff speed because the empirical observations from the CDMP could give us a much better estimate of takeoff speeds in a highly compute constrained regime.
Option #7: Pull out of the deal and switch to Plan B
Unilateral
This would involve leaving the deal. Immediately, both sides would destroy the vast majority of each other’s compute, and then both sides would go back to racing. This is really bad for both sides, because it would be immediately economically catastrophic and it would significantly increase misalignment risk, World War 3 risk, and concentration of power.
The main upside of this option is that it is useful as a BATNA during negotiation, and preserving a credible threat of leaving the deal can help deter the creation of a covert project in the first place. If China builds a covert project knowing it runs a substantial chance of detection, and the US commits to pulling out of the deal if they detect it, it is probably irrational for China to create a covert project in the first place.
The main downside of pulling out of the deal is if this happens when China was in reality honest, but for some reason the US had an overly hawkish estimate of a Chinese covert project. (For example, maybe compute estimates were off, or smugglers were harder to track down than expected). In these worlds, there was the option of continuing on in Plan A and incurring much lower risk across the board.
Option #8: Other geopolitical concessions
Bilateral
In general, if the US is suspicious that China is defecting, they can use the standard set of geopolitical tools to incentivize them against this. Particularly notable tools in the Plan A scenario include the robot and compute cap subject to both countries.
Section 7.2: Response plan for ambiguous evidence of covert project
Example concrete ways the US could get weak evidence of a covert project:
The chip supply chain tracing reveals a missing 1M H100e: records exist of these chips being manufactured (e.g. from Nvidia/TSMC/Huawei), but we aren’t able to find them. This could be smugglers, or it could be a nation-state backed covert project.
China refuses additional verification. For example, maybe the US is suspicious of a particular set of sites, where there might be a covert project, but China isn’t willing to allow the US access. This is an update towards a covert project because they are much more likely to give the US access to these sites if they are legitimate.
The US might get ambiguous intelligence from IMINT / OSINT / SIGINT, for example, satellite images might suggest large scale underground construction, but the US might not be able to ascertain whether it’s an AI related covert project.
The US notices model weight exfiltration attempts that seem like they could be coming from China but with unclear attribution. (This doesn’t necessarily imply China is operating covert compute, but is evidence.)
Baseline response:
Adopt a more covert project sensitive scaling strategy, for example, minimizing algorithmic progress (Option #2).
Improve verification: both by improving the technology (Option #5) and negotiating for increased implementation (Option #4).
If China refuses to allow improved verification, consider leaving the deal (Option #7). This is the only part of the response that the US can do unilaterally, but the fact that this is a bad BATNA for both parties creates pressure for China to agree to the above.
Essentially, this proposal involves marginally tweaking parts of Plan A to be more conservative with respect to covert project risk, because there is direct evidence that covert projects are a significant concern, but allows the core properties of Plan A to continue—there is still total research transparency, broad deployment of AI, ongoing reversible scaling, etc.
Section 7.3: The US gets unambiguous evidence of a PRC covert project
Example concrete ways the US could get unambiguous evidence of a covert project:
Cyber espionage (e.g. GPU backdoors, finding classified documents, etc)
Satellite footage (e.g. finding the construction, heat signature from the project, etc)
A literal spy in the project
Egregious refusals from China during negotiation (e.g., there is really solid privacy preserving verification technology, but China refuses to use it and doesn’t give a reasonable explanation for why).
Baseline response:
Require that China make large geopolitical and verification based concessions (Option #8 and #4).
If China isn’t willing to make those concessions, exit the agreement. (Option #7).
If China is willing to make sufficient concessions, evaluate how good the verification situation looks and choose a capability scaling strategy that corresponds with the overall level of ongoing covert project risk.
This strategy is reasonable because it (i) punishes China and puts them in a bad strategic state, thus deterring them from doing a covert project in the first place, (ii) allows the Plan A deal to continue, as is rational for both sides to want, and (iii) prevents the covert project from developing deal-undermining AI.
Appendix
Covert project strategies we think are ineffective
Acquiring compute
Diverting pre-2028 stock by retroactively understating domestic production
This category refers to pre-2028 sales by Chinese chip vendors. These sales could conceivably be understated, but it seems even more difficult than the (already seemingly quite difficult) task of understating new production.
Parameter | Estimate | Rationale |
Probability that retroactive concealment of domestic production is feasible | 7% | I think that doing it non-retroactively is about 15% likely to be feasible, and I am dividing this number by two. Seems about right, but I could update based on learning more about the Chinese semiconductor ecosystem or the historical effectiveness of whistleblower programs like ours. |
Fraction of pre-2028 cumulative domestic production which can be concealed and removed from datacenters, conditional on feasibility above | 5% Beta distribution with p90 = 7%. | Intuitive guess, p90 based on the idea that removing “one Chinese hyperscaler” (e.g. 10-20%) would be super obvious. Though note that this is different from removing one hyperscaler from the total compute in China; some of that will be legally imported or smuggled US-designed chips. |
Assuming that the fraction of global compute produced by China stays roughly constant, and accounting for the 2028 flow, we get a Beta distribution supported on [0, 1.6%] with median 0.08% of pre-deal compute from this route, conditional on it being feasible at all.
Non-Chinese covert projects
This supplement is focused on the possibility of a Chinese covert project. However, there is also a chance that other countries besides China attempt a covert project. We have not looked into this possibility in detail, so are not confident in our conclusions here, but we tentatively believe that China is the country where doing a covert project is most viable, because:
The US is the only country in the world with more compute than China, and acquiring compute is a central bottleneck for covert projects (see Section 2).
The US seems less likely to agree to Plan A while simultaneously choosing to divert compute, since the US BATNA is more favorable than the Chinese BATNA.
The US has less ability to do a covert project relative to China. The main reasons for this:
It will be much harder for the US to acquire chips because there is no major chip smuggling inside the US, and the US doesn’t seem to have a reasonable excuse to force companies to decommission large numbers of chips, cutting off the major avenues for covert compute procurement..
The US has more checks and balances on executive power, so a covert project is probably politically more difficult to set up. It also seems harder to for the US to compel the relevant companies to comply with the efforts of a covert project (e.g., getting semiconductor companies to falsify sales records), as very few companies below the chip-design level are US-headquartered
Appendices — Covert AI Projects
Appendix A — What successful defection would require
In summary, optimally subverting a competent implementation of our recommended Plan A policies would require China to:
Decide to begin diverting chips before the chip transparency agreements are made (in the Plan A scenario, this would need to occur in 2028).
Acquire on the order of 1.5M H100e (around 7% of counterfactual EOY 2028 Chinese stock) without being detected. The most promising options for acquiring this compute are: (i) diverting smuggled chips over the course of 2028, (ii) diverting older chips by claiming to have decommissioned and recycled them, and (iii) coordinating falsified production records across the semiconductor supply chain to hide some pre-deal production.
Build illegal datacenters without being detected. The most promising strategy we’re aware of for doing this is to locate the datacenters underground to avoid satellite detection.
Power and cool the underground datacenters. One strategy involves using small nuclear reactors to generate electricity onsite and rejecting heat by circulating water into a large river or deep water source in a way that avoids an IR-visible thermal plume.
Operate the sites without being detected for long enough to build extreme enough AI capabilities to significantly shift the geopolitical balance of power. This involves avoiding spies within the government, cyber operations, smarter AIs helping with covert project detection, satellite monitoring, and any agreed upon US/China government transparency (e.g. privacy preserving AIs reading classified info). If the US became suspicious (e.g. due to evidence from any of the items above), China would need to avoid giving concessions that would reveal their covert projects (e.g. physical US inspectors at potential covert project locations), while avoiding the possibility of the deal breaking down.
Avoid the development or usage of lie detectors (or other AI tech that improves verification massively).
China could also do a smaller covert project, which would have a proportionally smaller footprint and detection likelihood, but would also make slower progress towards strategically relevant AI capabilities.Based on our modeling, deliberately holding back on the amount of compute deployed does not improve a covert project’s chance of success.Based on our modeling, deliberately holding back on the amount of compute deployed does not improve a covert project’s chance of success.
Appendix B — Our auditing proposal in detail
Starting from Recommendation #1’s verified record of vendors’ sales to direct customers, audit every direct customer above a size threshold T, then every above-threshold customer of an audited firm, recursively, until all unaudited recipients are below T.
For example, you might set this threshold at 0.01% of pre-deal world compute (30k H100e, costing ~$200M at the time). T should be at or below the treaty’s target bound on covert project size, so that any firm large enough to constitute a covert project by itself is audited deterministically.
By “audit”, we mean “obtain all records of compute purchased, resold/transferred, or retained as inventory”, as well as “physically inspect any compute the firm presently owns”.
Some recorded outflows cannot be traced onward. These are added in full to an untraced pool L, and counted as potentially diverted. The two main cases:
Sales to unauditable counterparties. If a firm’s records show a sale to a buyer that cannot be produced for audit (e.g. a shell company, or a smuggler who kept no onward records), the entire sale is added to L. Since T is at or below the target bound, a single such sale above T is enough to block certification, so the Consortium should also investigate the seller’s remaining records and refer the case to the tools of Recommendation #3.
Unverifiable decommissioning claims. If a firm claims to have scrapped or recycled compute, the Consortium cross-checks the claim against recycler-side records (recyclers are themselves auditable counterparties), serial numbers where available, and actuarial fleet-attrition rates. Amounts that pass these checks are accepted; the rest is added to L. As above, an unverifiable claim above T blocks certification and warrants investigation.
When this completes, every unit of pre-deal compute is either physically verified, added to L, or recorded as delivered to a below-threshold recipient. The verified record of these small deliveries (firm 1 received a_1 H100e, firm 2 received a_2, and so on) is the list Layer 2 samples from.
Randomly sample H100e-units from that delivery record, and trace each sampled unit through any chain of resales until it either reaches a physically inspected chip or fails to.
Draw one unit uniformly at random from all compute recorded as delivered into the tail. This selects a recipient firm with probability proportional to how much it received. Audit that firm: the compute it verifiably received must equal current holdings plus recorded onward sales plus documented attrition, with holdings physically inspected.
Sales records don’t distinguish individual chips, so the sampled unit is resolved proportionally within the account. If the firm holds 200 of the 500 units it received, the trace ends here with probability 200/500, and a chip must be produced for inspection. Otherwise it follows a recorded onward sale to the buyer, where the audit repeats.
Each trace ends as a pass (a physically inspected chip) or a fail (an inspection shortfall, an unbalanced account, a buyer that cannot be produced, or an unverifiable decommissioning claim). An adversary who splits a large diversion across many sub-T shell buyers gains nothing: those sales appear in the delivery record, so sampled traces land on them in proportion to their size, and every such trace fails. Layer 2 then detects and measures the missing compute statistically.
Sampled audits occur with minimal advance notice, and inspectors record chip serial numbers at every inspection. The same serial appearing in two inspections at inconsistent times is itself a violation. Without both measures, a covert project could pass audits with temporarily borrowed chips.
Going forward, the Consortium prevents L from growing: post-deal transfers above a de minimis size are permitted only to registered auditable counterparties, and post-deal decommissioning requires witnessed destruction or serial-verified recycling.
The resulting certificate: with confidence 1 - delta, no covert project exceeds D + L. D shrinks with the number of clean sampled traces (roughly, n traces certify D of about (tail size / n) * ln(1/delta)). L is a directly measured quantity that further auditing cannot reduce; reducing it requires the external-evidence tools of Recommendation #3.
Appendix C — How much compute will be decommissioned by default?
For normal chip lifetimes, this is a small amount. (e.g. a common depreciation schedule is 5 years, so with an annual compute growth rate of 2-4x per year over the 5 years preceding the deal, compute whose fifth birthday happens before the deal will be less than 2^-5 – 4^-5, or 3.1% to 0.1% of global compute at the beginning of the deal.)
We understand US hyperscalers to have detailed records of the entire decommissioning process. For example, AI servers decommissioned by Microsoft are sent to Microsoft Circular Centers, where high value components (AI accelerators would certainly qualify) are removed and assessed for internal reuse, resale, donation (for educational purposes), or recycling. If we assume that US hyperscalers (who currently own 71% of global compute) can provide an accurate count of their decommissioned chips, the decommissioned fraction is reduced to between 0.9% and 0.03% of cumulative global compute production at the start of the deal.
We are less sure about the situation with smaller compute operators and for Chinese compute operators.
Appendix D — Parameter estimates and rationales
Parameter | Estimate | Rationale |
Probability that decommissioning is feasible at all to do deniably | 60% | Intuitive guess. |
US-designed fraction of pre-2028 stock in China | 45% | If domestic production is 3.7% of world, and (as Plan A side panel data says) 6.7% of world stock at 2028.0 is located on Chinese soil, then the US-designed share must be 45%. (or 3% of the world stock at 2028.0) |
Fraction of pre-2028 US-designed compute located in China that can be falsely decommissioned before 2029.0, conditional on this being feasible to do at all deniably | 22% 80% CI: [15%, 30%] Beta | Intuitive guess. If the top few clouds (Alibaba, Tencent, ByteDance, etc) own 60% of US-designed compute in China, and they decommission the oldest p25–p50 of their US-designed compute as part of this scheme, then this is 15–30% of the US-designed stock. (or 6.75% – 13.5% of the 2028.0 total PRC stock, using 45% as the parameter above). |
Appendix E — Parameter estimates and rationales
Parameter | Estimate | Rationale |
Probability that concealment of domestic production is feasible | 25% | Intuitive guess, see the discussion of the difficulties here. I could update based on learning more about the Chinese semiconductor ecosystem or the historical effectiveness of whistleblower programs like ours. My initial guess was 15%, Romeo thinks 35% for this, so we take the average and use 25%. |
Fraction of 2028 domestic production which can be concealed, conditional on feasibility above | 7% Beta distribution with p90 = 15%. | Intuitive guess, p90 based on the idea that removing “one Chinese hyperscaler” (e.g. 10-20%) would probably be too obvious. Though note that this is different from removing one hyperscaler from the total compute in China; some of that will be legally imported or smuggled US-designed chips. |
Appendix F — Parameter estimates and rationales
Parameter | Estimate | Rationale |
Fraction of 2028 compute sales smuggled to China | 2% [0.66%, 6.0%] LogNormal | Currently, Epoch’s median estimate is that 3.2% of compute sold through 2025 Q4 was smuggled to China. This might go up in the future due to increased demand for compute in China as frontier AI becomes a bigger business there (or as AI compute’s importance for national security becomes more obvious to the CCP). It might go down in the future due to improved US enforcement. I am shifting the median down to 2% as in the Compute Forecast, deferring to Romeo, who gives this as a point estimate based on improved enforcement. I think this makes sense conditional on Plan A. |
Smuggled compute for which the end customer cannot be identified | 25% [2.5%, 70%] Beta | Intuitive guess. I think it is probably less than half. I think my uncertainty could be reduced by learning more, but not super easily.
I think smuggled chips can have an accurately-recorded chain of custody if all the entities involved have some record of who they sold to, even if that record isn’t what they’d show to auditors. However, if at some point they were sold to an entity who used a false identity, or sold “cash-only”, such that the reseller isn’t sure where the chips went to, that doesn’t count. But if the CIA knows where the chips ended up and the information can be used in the supply chain tracing process without compromising sources and methods or whatever, that does count. |
2028 flow as a fraction of cumulative compute sales to date at time 2029.0 | 53% | Plan A side panel data has 154M H100e added in 2028 and 289M H100e total as of 2029.0. 154 M / 289 M = 53%. |
Appendix G — Parameter estimates and rationales
Parameter | Estimate | Rationale |
China’s military spending in 2028 | $400B 80%: [350B, 457B] LogNormal | China’s official military budget for 2023 was $296.44B. It grew roughly linearly over the preceding decade, adding an average of $13.2B per year. Extrapolating this to 2029, would predict official military spending to reach $362.4B. We use something closer to an extrapolation of SIPRI’s estimates of military spending (which are broader than the official budget).There are subtleties around PPP vs fixed-year exchange rate vs nominal which we’re currently ignoring. |
Fraction of annual budget China can spend secretly buying AI servers | 2%80%: [0.2%, 20%]LogNormal | Intuitive unanchored guess. When asked, Claude suggested a number more like 15% – 20%. |
$/H100e for newest US AI servers in 2028 | $8k | Geometric mean of the 2028.0 and 2029.0 values in Plan A side panel data. |
Smuggling markup % | +100% [+50%, +200%] | |
Decommissioned server markdown | -50% | Cheaper than the US-market $/per H100e surely, otherwise you could arbitrage by buying compute in the US and selling it directly to recyclers. But Chinese intelligence services might need to pay a bit more than is standard for recyclers since they need to capture the market, maybe 1.2x, or the cloud execs might just be told which specific recyclers to use, or the real recyclers could be put out of business.I will just go with half off. |
Decommissioned to smuggled ratio | 3:5.5 | Based on the numbers in Method 1. |
H100e in the world at 2029.0 | 289M H100e | From the compute supplement. |
Appendix H — Parameter estimates and rationales
Parameter | Estimate | Rationale |
2026.0 US IT power per H100e (at max utilization) | 833 W/H100e | DGX GB200 SuperPOD reference architecture p.6. One Scalable Unit of 8 NVL72 servers has a TDP of 1.2 MW, so 150 kW per server, including InfiniBand and Ethernet networking, management nodes, and storage. We then divide by 2.5 * 72 H100e per server to get 833 W/H100e. |
Average US power efficiency progress per year, 2026.0–2028.0 | 1.37x/year | Averaging the three relevant cells in Plan A side panel data, (1.3*1.4*1.4)^(1/3) = 1.37. |
Smuggled chip proxy year | 2028.6 | Roughly when the smuggling happens, adjusted forward slightly to account for the value of an exponential at the midpoint underestimating its average. |
Chinese domestic power efficiency gap during diversion period | 3.7x | Taking a value between the two relevant cells in Plan A side panel data. |
Decommissioned chip proxy year | 2026.5 | Seems vaguely right, the decommissioned chips will generally be older than this, but more of the compute will be from the newer chips since the quantity is growing 2–4x per year or so. |
Covert datacenter peak PUE | 1.4 | This refers to the facility power usage at maximum utilization divided by the IT power usage at maximum utilization. We use the value from here. It should maybe be adjusted upward because it’s a covert datacenter, but pretty unclear, also unclear whether theoretical maximum power is what we want to track for a project that could be flexible if needed. Leaving it at 1.4 for now. |
Appendix I — Distribution over the optimal site size
Parameter | Estimate | Rationale |
Probability that >1MW is infeasible/suboptimal | 0% | This is an approximation for modeling purposes. |
Probability >10MW is infeasible/suboptimal | 10%
| Intuitive guess. There are many different kinds of industrial sites that draw 10MW from the grid and dissipate 10MW of heat. Of each kind of site, China might have hundreds or thousands. It seems highly likely that at least one of them works as a disguise for a covert datacenter. |
Probability that >100MW is infeasible/suboptimal | 35% | Intuitive guess. There are fewer 100MW industrial sites in China, but probably still 1000+. But at this scale, you might need to rely on concealing the facility entirely rather than disguising it as something else. |
Probability that >1 GW is infeasible/suboptimal | 80% | Our assessment is that once-through cooling at 1 GW+ requires a deep-water diffuser in an ocean (not a river) to avoid a thermal plume. Rendering other cooling methods (air-cooled or water-vapor cooled chillers) invisible at this scale seems implausible. There are some industrial sites at this scale, but seemingly not very many, and they would be highly visible targets for challenge inspection. From this Data Insight, “a MW-weighted average across major US markets” finds that the facility capex per IT GW is $11B. |
Probability that >10GW is infeasible/suboptimal | 100% | This is an approximation for modeling purposes. |

Appendix J — Intuition-based detection probabilities, by site size and number of sites
These estimates assume only the unilateral detection methods discussed above, and condition on no “AI silver bullets” being discovered (we add those on top in the final model). By “3y”, this means “three years into the deal”.
Per-site facility power (MW) | n sites | P(detected by 3y) | P(detected by 10y) |
10 | 1 | 0.20 | 0.26 |
10 | 3 | 0.25 | 0.30 |
10 | 10 | 0.32 | 0.37 |
10 | 100 | 0.55 | 0.75 |
100 | 1 | 0.25 | 0.35 |
100 | 3 | 0.35 | 0.40 |
100 | 10 | 0.45 | 0.67 |
100 | 100 | 0.65 | 0.80 |
1000 | 1 | 0.40 | 0.50 |
1000 | 3 | 0.45 | 0.70 |
1000 | 10 | 0.70 | 0.85 |
1000 | 100 | 0.85 | 0.90 |

Appendix K — Forecast of “AI silver bullet” discovery by year
Thomas’s overall view about the likelihood that anything in the category allows us to detect covert projects by the following years (assuming the events of Plan A) is:
High will | Low will | Reasoning | |
2033 | 65% | 20% | I think the development of lie detectors is about 50/50 by 2033 conditional on the capability trajectory. In high will, this enables the US to threaten to leave the deal unless China submits to lie detection, which would reveal a covert project. Then there’s some probability that other paths help substantially, especially unknown unknowns. |
2037 | 95% | 55% | By 2037, we are likely to have developed good lie detectors (maybe 80%). If a high will regime, lie detectors seems sufficient, and it’s very likely that other verification methods pan out, especially given that we’ll have the industrial capacity to install whatever monitoring devices are necessary. In a low will regime I’m less confident about what will happen. |
All of these numbers set aside all the other detection methodologies, i.e., they assume that there’s no ability to detect.

Appendix L — All the assumptions we are making
We assume that
The US and other non-defecting Consortium members competently implement all of our policy recommendations
The PRC executed a competent diversion program one year before the deal was implemented (i.e. in 2028)
Covert projects which are detected before reaching deal-undermining AI stop their progress immediately.
Deal-undermining AI is operationalized as TED-AI
No weight theft occurs during the deal
The deal does not break down during the time period in question.
Deal Decline
Introduction
A consideration pushing toward scaling faster in Plan A is deal decline: the possibility of the deal either dissolving or becoming substantially less effective (which we call deal impairment).
In this supplement, I’ll estimate the likelihood and badness of deal decline. I’m confident that deal decline is at least 10% likely to occur within Plan A’s first 10 years, and that a 2035 decline leads to the expected value of the future being at least 5% lower than if there were a 2035 handoff within the deal. Since deal decline seems worryingly likely, and scaling capabilities within the Plan A deal seems much safer than an AI race once it’s declined, I’m confident that deal decline is a substantial consideration, likely the biggest one, in favor of scaling up AI capabilities faster and handing off to AIs earlier (rather than something like a very long pause). But I’m not precisely confident in any of the numbers estimated below; I could easily update substantially toward being more or less worried about deal decline.
Despite a substantial chance of decline, I think Plan A remains much more valuable than the other plans. Even if deal decline happens within a few years of the deal, the outcome will probably still be substantially better than no deal. And I think it’s plausible that a long-lasting effective deal can be achieved, via mechanisms such as automated lie detection and robust institutional design.
Overview
After categorizing types of deal decline, my analysis consists of:
Causes of deal decline: I split causes into (a) deliberate deprioritization of effective deal maintenance due to leadership change, (b) deliberate deprioritization due to a change in existing leadership’s priorities , and (c) destabilizing events.
Likelihood of deal decline: Informed by base rates from other international agreements and a decomposition into possible causes, I estimate the chance of deal decline per year at ~12% in the first 2 years of the deal, ~7% in years 2-5, and ~4% in years 5-10. The risk is roughly equally split across dissolution and impairment. This implies a 37% chance of decline within 5 years and 48% within 10. During the deal and to some extent leading up to it, society will have much more information than I have today about how robust the deal is, as such their probabilities may be quite different.
Badness of deal decline: I compare the situation after deal decline to the situation before there was any deal at all. The most important ways the situation is better post-deal-decline (relative to the situation without a deal) come from having used AIs to do lots of useful stuff (e.g. safety research) and from additional calendar time yielding more societal awareness and preparation. On the other hand, if Plan A’s compute destruction mechanisms fail, there will be much more compute in the world post-deal-decline, leading to a faster and more dangerous takeoff. Overall, it seems that a post-deal-decline situation would be substantially better than if there were no deal at all, while being significantly worse than doing handoff within the deal. When the deal decline happens matters a lot; the later, the better.
Preparing for deal decline: We recommend investing lots of resources into preventing deal decline, as well as mitigating it. Our main recommendation for mitigating deal decline is to ensure that nearly all of the compute in the world is destroyed in the case that the deal dissolves. Without this destruction, takeoff might be extremely fast due to an overhang of compute that wasn’t being used for developing more generally capable AIs.
Types of deal decline
The following are the types of deal decline:The deal could also get better, but we’re not including that here as our best guess is that the deal will in expectation get worse. This is definitionally true for deal dissolution, but we’re not confident for deal impairment; if we were more optimistic, we could remove the deal impairment term or explicitly model the deal getting better.The deal could also get better, but we’re not including that here as our best guess is that the deal will in expectation get worse. This is definitionally true for deal dissolution, but we’re not confident for deal impairment; if we were more optimistic, we could remove the deal impairment term or explicitly model the deal getting better.
Deal dissolution. The deal is officially dissolved.
Deal impairment. The deal becomes substantially less effective than if it were implemented well and adhered to, but isn’t dissolved. For simplicity, this is modeled as a binary in this supplement; in practice, it would likely be gradual. Ways the deal could be impaired:
Deal non-enforcement. The deal is not enforced as strongly as before, e.g. a country blatantly not complying is not punished. This leads to increased violations of the letter of the deal.
Deal misexecution. The letter of the deal is followed but is executed poorly.
Deal revision. The deal is explicitly edited to be worse. This could be straightforward weakening, but it could also be making it worse in other ways, such as unnecessarily increasing surveillance.
The deal could also end via a covert project reaching takeover-capable AIs or via an existential catastrophe (e.g., via AI takeover, AI-enabled bioweapons, or AI-developed superweapons). These end conditions are out of scope for this supplement and aren’t considered to be deal decline in our terminology.Though these failures could be mostly downstream of deal decline.Though these failures could be mostly downstream of deal decline. They are included separately in our analysis of the Plan A capability scaling strategy.
Causes of deal decline
One path for deal decline to happen is via deliberate deprioritization of maintaining an effective deal. The path for this is:
At least one of the US or China leadership decides to deprioritize the deal, which could happen for a variety of reasons, some of which are listed below.Perhaps a bloc of countries besides the US and China could also suffice.Perhaps a bloc of countries besides the US and China could also suffice.
If this only happens inside of one of the countries, then the other country must either try to use their leverage to prevent deal decline and fail, or let the deal decline happen (e.g. because they are scared of escalation). Other countries will also have some leverage at pressuring countries in ways that prevent deal decline.
We can split out ways that (1), i.e. US or Chinese leadership’s priorities changing, might happen into either (a) change of leadership (e.g. a US election)Or a CCP power transition. Or, less likely, a civil war.Or a CCP power transition. Or, less likely, a civil war. or (b) a change in existing leadership’s priorities.
Why might existing leadership’s priorities change? The most likely reasons we can think of are:
Concern regarding how future domestic leaders would handle things.
Wanting to gain power by developing superintelligence before a leadership transition.
Becoming more optimistic about winning the AI race in a world in which the deal has declined.
Becoming less concerned about risks from superhuman AIs such as AI takeover or extreme concentration of power.
Wanting to illegitimately grab power and thinking the deal impedes this.
Becoming less optimistic about the deal working, e.g., they think misuse risk is too high and difficult to mitigate within the deal. This could lead to a failed negotiation to revise the deal.
And then there’s an “Other” bucket of many reasons that we haven’t thought of.
Another possible reason the deal might decline is via a destabilizing event. The most likely we can think of are:
Geopolitical shock, e.g. war.
A country being caught cheating: a covert project getting caught, or another way of cheating such as trying to undermine the inference-only verification.
A sub-existential catastrophe.
Likelihood of deal decline
Overall view
Below I estimate the likelihood of deal decline in Plan A by year, focusing on the first 10 years (as that’s when handoff happens in the scenario). During the deal and to some extent leading up to it, society will have much more information than I have today about how robust the deal is, as such their probabilities may be quite different.
I use 2 methods:
Base rates plus adjustments: An analysis focused on the base rate of international deals declining, estimates the chance of deal decline per year at: ~15%/year in the first 2 years of the deal, ~12%/year in years 2-5, and ~11%/year in years 5-10. I adjust these downward because Plan A has especially strong verification and implementing it is evidence of high political will to maintain it. After adjustment, risk is roughly equally split across dissolution and impairment.
Decomposition into causes: Decomposing into possible causes of deal decline and then combining these estimates gives me ~10%/year deal decline in the first 5 years of the deal, and ~6%/year in years 5-10. This moves me up a little from the base rates plus adjustments method, but the results are overall fairly similar across methods.
My overall view is, aggregating the above 2 methods:
Eli’s overall deal decline likelihood estimates
0-2
6%
6%
11.6%
21.9%
21.9%
2-5
3.5%
3.5%
6.9%
19.2%
37%
5-10
2.0%
2.0%
4.0%
18.3%
48%
Eli’s overall deal decline likelihood estimates
Phase (years) | p(dissolution)/yr | p(impairment)/yr | p(decline)/yr | p(decline) this time period | Cumulative p(decline) |
0-2 | 6% | 6% | 11.6% | 21.9% | 21.9% |
2-5 | 3.5% | 3.5% | 6.9% | 19.2% | 37% |
5-10 | 2.0% | 2.0% | 4.0% | 18.3% | 48% |
When doing modeling for the Capability Scaling Strategy and the Plan A portion of Comparing Possible Plans, I use the following hazard rates with linear interpolation in between.After 16 years, the hazard rate decays to 95% of its previous value each year.After 16 years, the hazard rate decays to 95% of its previous value each year. The CDF is very similar to the above table.
Eli’s deal decline likelihood estimates as used in Capability Scaling Strategy and Comparing Possible Plans
0
0.12
0%
3
0.08
25.9%
6
0.06
39.9%
8
0.04
45.7%
11
0.03
51.1%
16
0.02
56.8%
Eli’s deal decline likelihood estimates as used in Capability Scaling Strategy and Comparing Possible Plans
Years after deal start | Hazard (any decline) | Cumulative p(decline) |
0 | 0.12 | 0% |
3 | 0.08 | 25.9% |
6 | 0.06 | 39.9% |
8 | 0.04 | 45.7% |
11 | 0.03 | 51.1% |
16 | 0.02 | 56.8% |
Deal decline risk reaches ~50% by 2040 handoff
Base rates plus adjustments
Data collection methodology
We had Claude do analysis of base rates and adjustments, then manually reviewed and edited its outputs. A summary of the approach (read its full analysis below):
Claude collates a list of 31 past international agreements approximately similar to Plan A in scale and nature. For each agreement, it identifies time-to-first-impairment and time-to-dissolution, and also gives a qualitative ranking of how applicable the agreement was to Plan A.
Base hazard rates are computed for each phase (number of years since the deal started) via an average rate weighted by applicability scores, separately for dissolution vs impairment:
For impairment, this is computed as a weighted number of years before first impairment divided by the weighted number of impairment exposure.
For dissolution, this is the weighted sum of dissolved agreements divided by the weighted number of dissolution-exposure years.
Claude also identifies the reasons for impairment and dissolution. It finds that dissolution is roughly hump-shaped centered around 5-10 years, and impairment is front-loaded.
Claude finds that improved verification correlates with delayed impairment, and dissolution usually occurs as a result of political withdrawal, rather than inherent deal failure.
Base rate (plus adjustments) forecasts
Below are the collected base rates.
Base rates of deal dissolution and impairment
Weighted base rate (HEADLINE)
3.4%
7.9%
Unweighted (all treaties equal)
1.8%
4.8%
Base rates of deal dissolution and impairment
Method | Dissolution hazard/yr | Impairment hazard/yr |
Weighted base rate (HEADLINE) | 3.4% | 7.9% |
Unweighted (all treaties equal) | 1.8% | 4.8% |
I’ll now discuss how the hazard rate changes over time. I’ll focus on the first 10 years as that’s how long the deal is in the scenario. A note on 10+ years: the impairment hazard holds at ~6%/yr in years 10–20 rather than continuing to decline, before falling to ~4%/yr past 20. Here are base rates by how long the deal has been in force:
Base rates of deal dissolution and impairment by how long the deal has existed
0-2
0.0%
15.5%
15.5%
0.0%
28.6%
28.6%
2-5
3.4%
9.4%
12.5%
9.9%
46.9%
52.1%
5-10
5.8%
5.9%
11.4%
33.2%
60.9%
73.8%
Base rates of deal dissolution and impairment by how long the deal has existed
Phase (years) | Dissolution hazard/yr | Impairment hazard/yr | Combined hazard/yr | Cumulative dissolution (by end of phase) | Cumulative impairment (by end of phase) | Cumulative combined (by end of phase) |
0-2 | 0.0% | 15.5% | 15.5% | 0.0% | 28.6% | 28.6% |
2-5 | 3.4% | 9.4% | 12.5% | 9.9% | 46.9% | 52.1% |
5-10 | 5.8% | 5.9% | 11.4% | 33.2% | 60.9% | 73.8% |
I make adjustments to the base rate as follows:
Eli’s adjustments to the collected base rates
0-2
0.0%
15.5%
5.0%
5.0%
9.8%
Dissolution from 0 to 5%: 0% is to some extent an artifact of limited dataset size, plus it’s easy to imagine ways Plan A could dissolve quickly (e.g. a covert project getting caught).
Impairment from 15 to 5%: Plan A has especially strong verification and requires a very high amount of political will to initiate, making it less likely it is very quickly impaired.
2-5
3.4%
9.4%
3%
3%
5.9%
Decaying impairment in a way that is roughly similar to the base rate, but slightly more aggressively due to AI-assisted verification/epistemics/etc.. For dissolution, decaying in a similar way (partly for simplicity, maybe it should be slightly different).
5-10
5.8%
5.9%
1.50%
1.50%
3.0%
Continuing to decay, slightly more aggressively than above.
Eli’s adjustments to the collected base rates
Phase (years) | Base diss/yr | Base imp/yr | Adjusted diss/yr | Adjusted imp/yr | Adjusted combined/yr | Rationale |
0-2 | 0.0% | 15.5% | 5.0% | 5.0% | 9.8% | Dissolution from 0 to 5%: 0% is to some extent an artifact of limited dataset size, plus it’s easy to imagine ways Plan A could dissolve quickly (e.g. a covert project getting caught). Impairment from 15 to 5%: Plan A has especially strong verification and requires a very high amount of political will to initiate, making it less likely it is very quickly impaired. |
2-5 | 3.4% | 9.4% | 3% | 3% | 5.9% | Decaying impairment in a way that is roughly similar to the base rate, but slightly more aggressively due to AI-assisted verification/epistemics/etc.. For dissolution, decaying in a similar way (partly for simplicity, maybe it should be slightly different). |
5-10 | 5.8% | 5.9% | 1.50% | 1.50% | 3.0% | Continuing to decay, slightly more aggressively than above. |
There are 2 effects which lead me to estimate substantially lower deal decline risk as time goes on:
If the deal has gone on for longer, this provides evidence that it is less likely per unit time to decline.
As AI capabilities improve during the deal and there is time to put them to good use, they allow for building technology that helps stabilize the deal. For example, there might be automated lie detection or use of AI to improve epistemics or verification.
Surprisingly, the base rate data above has p(dissolution) increasing over time, while p(impairment) decreases in a fashion closer to what I expect for Plan A. Notably, the base rate data doesn’t include effect (2) above, or at least includes a much weaker version of it. So I expect the trend of decreasing deal decline risk to be more pronounced than it is in the data on previous agreements.
Decomposition into causes
An alternative method for estimating the chance of deal decline is to individually estimate the chance for each possible cause, then aggregate them together. This gives me an estimate of 10% p(decline)/year in the early 2030s and 6% in the late 2030s, leading to a cumulative p(decline) of ~50% by 2035 and ~60% by 2040. I think the methodology of listing a bunch of possible ways something could happen and then aggregating these probabilities tends to produce numbers on the high side, which I take into account when aggregating this methodology with the base rate plus adjustments method.
Eli’s probability of deal decline by potential cause
Deliberate deprioritization
US leadership change
0.02
0.01
China leadership change
0.01
0.01
Concern regarding future leaders
0.01
0.01
Optimism about winning post-deal-decline
0.01
0.005
Less concerned about AI safety
0.005
0.003
Other deliberate deprioritization
0.01
0.005
Destabilizing event
Geopolitical shock
0.01
0.005
Country caught cheating
0.015
0.015
Sub-x-risk catastrophe
0.01
0.005
Other destabilizing event
0.015
0.01
All sources
p(decline)/year from all sources
0.10
0.06
p(decline) this time period
0.49
0.27
Cumulative p(decline)
0.49
0.62
Eli’s probability of deal decline by potential cause
2029-2034, p(decline)/year | 2035-2039, p(decline)/year | ||
Deliberate deprioritization | US leadership change | 0.02 | 0.01 |
China leadership change | 0.01 | 0.01 | |
Concern regarding future leaders | 0.01 | 0.01 | |
Optimism about winning post-deal-decline | 0.01 | 0.005 | |
Less concerned about AI safety | 0.005 | 0.003 | |
Other deliberate deprioritization | 0.01 | 0.005 | |
Destabilizing event | Geopolitical shock | 0.01 | 0.005 |
Country caught cheating | 0.015 | 0.015 | |
Sub-x-risk catastrophe | 0.01 | 0.005 | |
Other destabilizing event | 0.015 | 0.01 | |
All sources | p(decline)/year from all sources | 0.10 | 0.06 |
p(decline) this time period | 0.49 | 0.27 | |
Cumulative p(decline) | 0.49 | 0.62 |
Unlike for most sources, the risk from a country being caught cheating might not go down in the late 2030s relative to the early 2030s, due to detection methods getting better.
Badness of deal decline
Overview
I compare the situation after deal decline to the situation before there was any deal at all. The most important advantages of post-deal-decline worlds over no-deal worlds are 1) having more time for capable AIs to do useful things like safety research and 2) having more time to build societal awareness and preparation.
On the other hand, there will be much more compute in the world post-deal-decline, by default leading to a faster and more dangerous takeoff. Because of this, Plan A is set up so that the compute is likely destroyed upon deal dissolution; more on this below.
I made some quantitative estimates of how bad deal decline would be. The headline results are:
Deal dissolution and impairment are similarly bad, perhaps impairment is a bit worse. While impairment has the benefit of there still being a deal, it has the downside of not triggering compute destruction.
Pre-max-controllable-AI, I propose modeling p(alignment | deal decline) as interpolating linearly in logit space between p(alignment; pre-deal unconditional) in 2029 and p(alignment | deal decline at max-controllable-AI target date). For a target date of 2035, perhaps p(alignment | deal decline at max-controllable-AI target date) is ~70%.
Post-max-controllable-AI, I propose modeling p(alignment | decline at time t) as roughly 90% of p(alignment | handoff within deal at time t).
High level situation
What would happen in Plan A if the deal declined and we hadn’t implemented any specific preparatory measures? Let’s start by considering the case of the deal fully dissolving.
There will be to some extent a return to the default world conditions, but with some important changes compared to if no deal had happened at all.
Positives:
Progress due to AI usage and studying AIs: Due to the alignment and control research that’s occurred since the start of the deal, the state of knowledge regarding safety techniques has improved. This is true to a greater extent the longer the Consortium has had to use and study more capable AIs. There are also other benefits of AI usage such as improving epistemics and verification tech.
More societal awareness and preparation: Aside from using AIs to make progress, the extra calendar time that the deal has bought helps in various ways by giving society more time to adapt. Furthermore, people, including policymakers, are more aware of current and projected AI capabilities than they were before. An important example of this is that there are likely lots more human AI safety researchers than there were at the start of the deal.
Training compute safety tax paid: The most capable models have been trained with more compute than would be needed to reach their level of capabilities, to allow use of safer methods. This is unlike what happens in a closely contested AI race. You may maintain the ability to use this model after dissolution. This positive is potentially minor because: (a) If AIs are not close to max-controllable-AI, we will have to retrain them in order to reach handoff-level capabilities, forfeiting our safety tax. (b) Even if AIs are close to max-controllable-AI, projects likely choose to not incur safety taxes on successor AIs and instead race to higher capabilities.
Negative: By default, more compute would lead to faster takeoff, but Plan A sets things up so that the compute is likely destroyed upon dissolution. By default, there is now more compute in the world than there was in 2029, which makes the default takeoff faster. In the Plan A scenario, compute stock increases by roughly 0.7 OOMs/year from 2030 to 2033 and 0.1-0.3 OOMs/year from 2034 to 2040. Naively applying a 5.5x speedup for a 10x increase in compute stock would imply that 3 OOMs would speed takeoff up by ~150x! There might be some compute destroyed though, for example due to a war or due to targeted sabotage of datacenters. There is also more industrial capacity in the world, which could allow the manufacturing of dangerous weapons, potentially after an industrial explosion. Because of this issue, Plan A is set up so that the compute is likely destroyed upon deal dissolution. This likely averts an extremely fast takeoff, though our best guess is that it still takes about 1-2 years to do the rest of takeoff; see below for discussion.
Unclear: Different geopolitical circumstances. Depending on how the deal dissolves, the geopolitical circumstances might be very different after dissolution than in the default world. For example, there might be a war going on, and if not, the deal dissolution might still make lower levels of coordination less feasible. On the other hand, great powers have signaled that they care a lot about AI safety by implementing the deal and depending on how dissolution happens this might carry through into better outcomes. It’s plausible that a new international deal could be struck not long after the deal dissolving.
For deal impairment, nearly all aspects of the world are at least somewhat better than in dissolution because the deal is still in effect. For example: not only is the training compute safety tax already paid, but the deal is still partly functional so you may be able to continue paying a tax. How much better things look than dissolution depends on the circumstances of the impairment.
The main downside is that the extra compute likely hasn’t been destroyed, as it’s harder to set up a system which destroys compute upon impairment as opposed to destroying it upon dissolution. While there’s a ton of compute in the world, the deal is likely still constraining takeoff to some extent. That is, you may not be able to use a large fraction of compute in the world to take off as fast as possible.
Quantitative estimates
I’ve made some rough quantitative estimates that attempt to take into account the above factors. See here for a sheet that contains estimates of the value of the future, conditional on deal dissolution at various times (among estimates of other deal end states). I think that deal impairment would on average be similarly bad to deal dissolution, with the pros and cons described at the end of the previous section roughly canceling out.
Overall, my view is roughly that we can interpolate value in logit space over the course of Stage 1 of the deal, i.e. pre-max-controllable-AI, between the default world and the end state in which the deal declines at max-controllable-AI. For Stage 2, max-controllable-AI and onward, perhaps the value of deal decline in a given year is roughly 75-80% of handing off within the deal in that year.
Preparing for deal decline
Given how big of a threat deal decline is, we recommend making it a top priority in Plan A to prevent and mitigate it. For example, I think at least 5% as much as the effort that goes into technical safety research in Plan A should go into preventing and preparing for deal decline (and potentially much more). These implications are discussed in the Capability Scaling Strategy supplement.
One way to reduce the risks of and mitigate the downsides of deal decline is to scale AI capabilities faster and handoff faster.
Preventing deal decline
Here are some avenues that we think are promising for preventing deal decline:
Automated highly-adversarially-robust lie detection: This would drastically reduce the chance of being surprised by an actor attempting to pull out of the deal.
Institutional design: The deal and Consortium should be designed in a way that best sets up incentives for the countries involved, and leads to reasonable regulatory decisions. Additionally, domestic institutions should be improved to be robust to the whims of a single person shifting; for example, ideally the deal should be ratified as an official treaty in the US.
United diplomatic pressure: All countries in the deal should make it clear that any country pulling out of the deal would be an unacceptable threat to all other countries’ national security and sovereignty, and will be treated as such. Any country that is considering pulling out of the deal should know that they will be heavily sanctioned, cyberattacked, and would potentially have to fight a war in which they are heavily outnumbered.
AI for epistemics: Generally improving societal epistemics seems important for creating a robust, effective deal. Example applications of AI that would be especially useful are forecasting failure modes of the deal and helping with the institutional design mentioned above.
Preparing stockpiles to align incentives
We recommend the following to improve the post-deal-dissolution situation and improve incentives around defection.
The US and China each get hardened stockpiles of 2% of their pre-deal compute. 2% is the factor used in the Plan A scenario, in practice it should depend on the amount of missing compute.
One potential incentive for the US and China to create a covert project is that they could pull out of the deal, destroy the centralized compute (e.g. with bombing, or with other mechanisms, e.g. by withholding offline licensing keys). Then, if one side did a covert project and the other did not, the side that had a covert project might get a large lead over the other, and beat them to an intelligence explosion.
One solution to this is to allow each side to have a cache of GPUs that are hardened but aren’t plugged into power, and so can’t be used for AI research while the deal is in place. In practice this stockpile could be stored in a secure facility on the home soil of each party, and is deliberately designed to not be destroyable by the other in the case of a deal dissolution,Technically, we mean not destroyable absent extreme escalation (e.g. nuclear war), which is unlikely for the normal reason of MAD.Technically, we mean not destroyable absent extreme escalation (e.g. nuclear war), which is unlikely for the normal reason of MAD. but where inspectors can still verify these chips aren’t in use for the duration of the deal.
This cache could be on the order of 2% of each side’s pre-deal compute. This is probably more compute than a covert project could have, but is a relatively small fraction of post-deal compute, because new chips dominate the overall chip supply.
Now, if the deal were to break down, and the verified datacenters were stopped, then both sides would return back to the counterfactual of racing with a proportional amount of compute as they had before the deal; this also favors the US given that the US has more compute pre-deal.
In the same stockpile the US and China each get a cold, heavily monitored copy of the frontier model weights.
Reasoning for doing this:
This reduces the incentive to exfiltrate the frontier weights and then immediately pull out of the deal to get a huge advantage. It also reduces the incentive to do a covert project.
The benefits of the training compute safety tax that was paid to create the frontier model weights doesn’t have to be re-paid. However, there are also downsides to preserving the weights; for example, it makes AI progress faster.Another downside is that if the weights are preserved, it makes deal dissolution less unattractive.Another downside is that if the weights are preserved, it makes deal dissolution less unattractive. It’s unclear that you would want to keep the weights in cold storage for only this reason.
A proposal to store the weights:
There’s a hard drive that needs US+China permission to unlock. That’s the only cold storage.
Other storage is volatile memory, so if you turn off power memory goes away
To stop physical exfiltration, you have a bunch of physical defenses that automatically turn off power
Mutually assured compute destruction: destroying compute, fabs, and robots
A central potential problem with Plan A is that if the deal dissolves, there is so much compute in the world that we might get an extremely fast intelligence explosion. For example, suppose the deal breaks down at the start of 2035; there are 100B H100e in the world at that time. That’s about 2.3 OOMs more compute than were in the world at the start of 2030 in the default world (~500M), and so a takeoff that would have taken 1 year, might instead happen 5^2.3=40x faster, or in 9 days.This is under the assumption that for each 10x increase in compute, the intelligence explosion occurs 5x faster. This is our median guess, which we discuss more in our covert projects supplement here.This is under the assumption that for each 10x increase in compute, the intelligence explosion occurs 5x faster. This is our median guess, which we discuss more in our covert projects supplement here. This would, of course, be an extremely dangerous situation. Therefore, if the Plan A deal breaks down, it becomes extremely important that a large fraction of the compute in the world be destroyed.
In this section, we’ll discuss our policies for ensuring mutually assured compute destruction; as well as making sure that the rebuilding process does not result in an unacceptably fast takeoff. The priority order for achieving this is:
Ensure that almost all of the compute is destroyed.
Ensure that almost all of the fabs are destroyed.
Ensure that almost all of the robot capacity is destroyed.
We’ll discuss each of the above in turn, and then analyze how long an intelligence explosion would take in the case of a deal dissolution.
Compute
This table gives point estimates for each type of compute in the Plan A scenario:
Amount of effective H100e | ||||
Type of compute | 2030 | 2035 | 2040 | What happens if the deal dissolves? |
AI datacenter compute (world) | ~500M | ~100B | ~1T | Self-destruct or are bombed |
Verified consumer hardware and edge compute (world) | ~0 | 100M | 10B | Not useful for AI training because of verification or fundamental design. |
Cap-and-traded consumer hardware and edge compute (world) | 25M | 30M | 30M | ~10% can be diverted to an AI project, because the compute is spread across countries and also non-authoritarian countries probably wouldn’t round up this compute. |
Stockpile (US/China/RoW) | n/a | 5M/500k/700k | 5M/500k/700k | ~100% can be diverted to an AI project |
Covert project (US&China) | 1M | 1M | 1M | ~100% can be diverted to an AI project |
Exempt military compute (US&China, others lower) | 10K | 10K | 10K | ~100% can be diverted to an AI project |
The policies for how to achieve this are as follows:
AI datacenters. Both the US and China build datacenters in deliberately vulnerable locations, thus incentivizing them to self-destruct if the deal breaks down. This is discussed more in the 2034 section of the scenario. Almost all compute is in this category (as shown above), and so it’s very important to ensure destroyability.
Consumer hardware and edge compute. This includes all compute used in phones, consumer laptops, etc. This compute is distributed throughout the world and could be diverted to an AI project after the deal dissolves. Therefore, we must either (i) verify so that it can’t be directed to training or (ii) cap the total amount of chips in this category. We recommend a cap and trade regime on unverified chips in this category of compute at 30M effective-H100e. We estimate the pre-deal stock to be around 25M effective-H100e, so normal production levels can continue in the short term but then for new AI-capable edge compute to be produced old hardware needs to be traded in / verifiably destroyed. If extremely robust on-chip verification can be developed to activate further chips, or methods of building chips that fundamentally cannot be repurposed to training then this can get an exception from the cap. One approach that could accomplish this is to bake a restrictive computational structure into chips during the manufacturing process (e.g. Etched).
Stockpile. See above.
Exempt military compute. We recommend minimizing AI relevant exempt military compute, or only exempting a tiny number of GPUs, such as 10,000 H100e.
Fabs
In Plan A, all of the AI-relevant fabs in the world are moved to destroyable special economic zones (SEZs) by 2032. These fabs are located either on the ocean or nearby adversarial superpowers with self-destruct mechanisms in them. Pre-deal fabs without these properties (e.g. TSMC Arizona) are decommissioned. Fortunately, decommissioning isn’t costly because it’s such a small fraction of the compute built in the 2030s.
Flow (H100e)
363M
13.5B
43B
93B
233B
1T
Time to build 100M H100e
100 days
3 days
20 hours
9.4 hours
3.8hours
53 minutes
Year | 2030 | 2032 | 2034 | 2036 | 2038 | 2040 |
Flow (H100e) | 363M | 13.5B | 43B | 93B | 233B | 1T |
Time to build 100M H100e | 100 days | 3 days | 20 hours | 9.4 hours | 3.8hours | 53 minutes |
100M H100e is roughly the amount of compute needed to do an intelligence explosion (from Automated coder to TED-AI) in 1 year. By the end of the decade in Plan A, industrial capacity is so high that we’re producing that amount of compute every 53 minutes. Therefore, it’s extremely important for the fabs to be destroyed if the deal breaks down, particularly towards the end of the decade. Fortunately, the fabs are extremely fragile by nature, and the same mechanism as the AI datacenters above (putting them in locations where self-destruction is incentivized) can be used.
Robots
Robots can be used to build fabs, which can build GPUs, which can then be used to do an intelligence explosion. When there are enough robots, this process can happen extremely quickly.
Our basic strategy with robots mirrors that with GPUs and fabs. Almost all robot production and use will occur in designated SEZs:
SEZ
0
1B
70B
Self-destruct or bombed
Civilian
100K
10M
1B
Self-destruct or backdoored or not useful for building fabs
Exempt military robots
1K
100K
100K
Can be used to make AI progress
Human-equivalent quantity | ||||
Type of Robot | 2030 | 2035 | 2040 | What happens if the deal dissolves? |
SEZ | 0 | 1B | 70B | Self-destruct or bombed |
Civilian | 100K | 10M | 1B | Self-destruct or backdoored or not useful for building fabs |
Exempt military robots | 1K | 100K | 100K | Can be used to make AI progress |
The policy for robots is:
High and growing cap for robots produced in destroyable SEZs, so almost all robots are produced there.
A mandate for civilian robots to have sufficient offline licensing + self destruct features so that they stop working if the deal breaks down, or robust edge verification such that they can’t be used to do any part of the work in rebuilding a compute supply chain (e.g., a self-driving car with baked weights).
Cap of 100K total exempt military robots for US and China, without offline licensing.
Post-dissolution takeoff length
A big question is whether or not a rapid intelligence explosion could occur on the stock of compute produced within the deal, or whether compute production is required.
There are two major sources of progress that can be pursued in parallel, but might have tradeoffs (e.g., if you use your compute to help run robots or tell humans what to do to produce new compute).
First source of progress: Produce new compute. The post-deal-decline AI project could attempt to use robots and/or humans to rebuild a semiconductor supply chain and increase their compute. The difficulty of doing so is probably further increased by potentially aggressive sabotage in a post-deal context against such efforts. That being said, around 100K exempt military robots, and maybe some small fraction of civilian robots e.g., 0.1%, or around 1M, can be stripped for useful parts post-self-destruct and/or have their security measures circumvented. Let’s assume there’s a fleet of around 1M human-equivalent robots that can work on this. In 2026 the semiconductor supply chain has around 2M people working, and over the last 5 years they have built around 100M H100e/yr of total production capacity at 2026-technology levels (chip design and manufacturing process) So we have a rough 10M worker-years to 100M 2026-tech H100e/yr capacity mapping. Assuming the 2026-tech improves by around 1 OOM by 2032 and another OOM by 2036 (assuming total compute in the world improves roughly 50% due to technology improvements and 50% due to production scale), this implies the mapping will be 10M worker-years to 1B H100e/yr with 2032-tech and 10B with 2036-tech. So the 1M robots would take 1 year to build 100M H100e in 2032 and 1 month in 2036. We expect there to be serial bottlenecks that somewhat push these purely parallel estimates up, as well as potential startup problems and sabotage problems.
Second source of progress: Use existing compute for R&D. The post-deal-dissolution AI project has transparent algorithms from the deal, cold-storage weights, and any consumer hardware that they are able to repurpose. Assuming they are the US and they use their cold-storage GPUs or they can get ~10% of world consumer hardware, then they have ~25x less compute relative to the default leading project in Jan 2030 (~5M H100e in US stockpile vs. ~130M in 2030 by default); so -1.3 OOMs. This 25x compute disadvantage means that at the same capability level they should go ~10x slower all else equal (assuming 10x less compute means ~5.5x slower).
However, all else is not equal: at least if the dissolution happens before the 2040 handoff in the Plan A scenario, the post-dissolution project has less efficient software/algorithms than the default project when they were at the same capability level. This is because in Plan A we deliberately make progress via a much higher ratio of training compute to software than the default. Because the software is worse, training efficiency is lower. On the other hand, there’s more remaining promising software ideas out there, so a given amount of research effort goes further.
Meanwhile, the transparent algorithms from the deal can be split into ones that have been implemented, and ones that have been foregone to pay a training compute safety tax, which we recommend paying roughly 1 OOM of by TED-AI. The project would need to actually train some using these foregone algorithms, but they could probably quickly capture lots of the gains.
Putting everything together. If dissolution happens after pausing at TED-AI, let’s say in 2037: The TED-AI to ASI time under the default (no slowdown) Plan A trajectory is ~1.5 months. Naively applying the 10x slowdown due to lower training compute would give ~15 months. Taking into account the post-dissolution project having lower software results in a takeoff to ASI of ~2.75 years (and setting there to be no post-dissolution compute growth, which further slows things relative to the default trajectory). Enabling 1 OOM/year post-dissolution compute growth results in a takeoff of ~1 year.
Compute destruction slows post-dissolution takeoff from ~1 week to ~1 year
Now let’s consider dissolution happening in 2032, a bit below SAR. SAR to TED-AI takes ~0.3 years by default. Naively applying the 10x slowdown due to lower training compute would give a bit over 3 years. Accounting for lower software and let’s say 0.5 OOMs/year post-dissolution compute growth results in achieving TED-AI after about 2 years.
Acknowledgments
Thanks to Dillon Nguyen for fact-checking Claude’s base rate analysis and helping create figures. Thanks to Thomas Larsen and Romeo Dean drafting the sections on post-dissolution compute destruction.
Appendix: Claude’s analysis of base rates
Claude’s full base rate analysis is attached below for reasoning transparency. Dates and events were hand-verified. The text has been very mildly updated based on the issues found.
Claude’s Analysis
This tab holds the overall reasoning behind the outside-view, base-rate estimate of how likely the Plan A US–China slowdown deal is to dissolve or become impaired, year by year from its formation (~2029) through handoff (~2040). The treaty-by-treaty dataset and the per-treaty reasoning live in the “Base rates” tab of the companion Deal Decline spreadsheet; this tab explains the method, the reference-class choices, the central finding, and how I adjusted from raw base rates to a Plan-A-specific year-by-year hazard. Drafted by Claude (Opus 4.8), 2026-06-29. Historical dates and fates were web-verified, but the synthesis into year-by-year numbers is a first-pass draft and a judgment call.
Method and choice of reference classes
The approach is the standard “outside view”: rather than reasoning from Plan A’s specifics, I assembled ~31 past international agreements that resemble it, recorded how long each lasted before it was first impaired and before it dissolved, and converted the pooled experience into annual hazard rates. I use the same two outcome categories as the supplement—DISSOLUTION (formally terminated, withdrawn-from, or collapsed) and IMPAIRMENT (substantially less effective without formal dissolution: non-enforcement, tolerated violations, suspension, or weakening revision)—and I track three things separately that analysts often blur: time-to-first-impairment, time-to-dissolution, and the failure mode (who broke it and why), because Plan A’s design bears very differently on each. I assign each agreement a numeric applicability weight (0–3: 3 = rival great-power capability deal, the closest analogue; ~1.5–2 = informative analogue; ~1 = broad multilateral convention; ~0.25–0.5 = low-defection-incentive or toothless), then compute the base rate as a weighted hazard—Σ(weight × events) / Σ(weight × years-at-risk)—separately for dissolution and impairment. The weights, the per-treaty inputs, and the live formulas all sit in the spreadsheet (Parts 1–2), so the aggregation is fully transparent and editable: change a weight and the base rate recomputes. A pooled hazard is still crude—it treats 1922 naval cheating and a 2023 treaty suspension as draws from one urn—so I report it alongside an unweighted version, a closest-analogues-only version, and—most importantly—a version that lets the hazard vary by PHASE (years since the deal started, Part 2b), then adjust for Plan A’s specifics.
I grouped the agreements into three tiers by similarity to Plan A. MOST APPLICABLE: rival great-power deals that cap a decisive military/strategic technology under strong incentives to defect—the interwar naval treaties (Washington/London), the bilateral US–USSR/Russia strategic arms-control treaties (ABM, SALT, INF, START, New START, Open Skies), and the two “rivals’-deal” nonproliferation bargains (JCPOA, the US–DPRK Agreed Framework)—plus the IR alliance-termination literature, the closest thing to a quantitative base rate for how long deals between rivals last (mean alliance life ~9–10 years pre-1945; ~1/3 violated when tested). MODERATELY APPLICABLE: broad multilateral nonproliferation/verification conventions (NPT, CWC, BWC, CTBT), useful mainly for what they reveal about verification and as the optimistic bound; and the dual-use export-control regimes (Wassenaar/MTCR/NSG), the closest analogue to controlling AI chips and chronically leaky by design. LEAST APPLICABLE but informative: durable low-incentive conventions (Outer Space, Tlatelolco, PTBT) and the toothless climate accords (Kyoto, Paris), which bracket the extremes—near-permanent when nobody wants to defect, near-useless when nothing is binding. The single closest historical analogue is the Washington/London naval system: peer rivals capping a decisive military technology, both tempted to defect, with self-declaration verification. It ran ~13–14 years before Japan walked out, and was cheated on throughout.
Central finding: verification delays cheating, but politics kills the deal
Two patterns dominate the data and drive the whole estimate.
First, verification strength is the cleanest predictor of how long a deal lasts unimpaired-by-cheating. Strongly verified regimes (INF, START I, New START, Open Skies, NPT/IAEA, CWC/OPCW) went roughly 8–35 years before significant impairment, versus 0–4 years for weakly verified ones—a real but less clean separation than it first appears, because dating impairment to violation onset rather than detection shortens the verified deals’ records (INF’s violative testing began ~2008, six years before the US formal charge; Open Skies restrictions began in 2010). Verification detected cheating eventually, but with meaningful latency, and detection did not by itself reverse the violation. The one clean reversal in the dataset is the Montreal Protocol, where cheap atmospheric monitoring caught illicit CFC-11 production within years and enforcement restored compliance—the deal never became substantially less effective and is coded as unimpaired. Weakly or un-verified regimes (the naval treaties, the Anglo-German agreement, OPEC quotas, the whaling moratorium, and above all the BWC—which had no verification and was violated on a massive scale by the USSR for its entire early life) were impaired almost immediately, within 0–4 years. This is the strongest structural reason for optimism about Plan A: its verification regime is unusually strong by historical standards and is explicitly designed to make cheating detectable and auto-punishable.
Second, and crucially: strong verification does essentially nothing to prevent death by political withdrawal. The ABM Treaty (30 years), INF (31 years), Open Skies (~19 years), and the JCPOA (2.3 years) all died because a government decided to exit, not because verification failed. The JCPOA is the sharpest warning—the best verification regime ever negotiated was discarded in ~2 years by a change of US administration. So verification protects against the impairment-by-cheating failure mode, but offers little protection against the dissolution-by-regime-change / change-of-heart / geopolitical-shock failure mode, which the alliance literature (Leeds & Savun 2007) identifies as the dominant driver of premature termination. This is why I keep the dissolution hazard meaningfully above zero even though Plan A’s verification is excellent.
How Plan A differs from the base rate (the adjustments)
Pushing the hazard UP relative to a generic high-salience treaty: (1) It is a rival deal over the single most decisive technology imaginable, so it belongs to the harsh ~7–12%/yr naval/rival/alliance class, not the benign ~0%/yr convention class. (2) It is an executive agreement, not a ratified treaty—the scenario says so explicitly, citing New START (2010) as the last significant ratified US treaty. Every deal in the dataset that one administration could exit unilaterally (ABM, INF, Open Skies, JCPOA) was exitable precisely because exit didn’t require undoing a ratified commitment. (3) The stakes and the rate of technological change are both unprecedented, which compresses the “first major shock” that historically kills these deals—and the prize for defecting grows as capabilities approach handoff. (4) US elections in 2032 and 2036 and the Chinese leadership succession around 2033 (Xi turns 80) are scheduled regime-change risks, exactly the events Leeds & Savun find spike the termination hazard. (5) The scenario stipulates unusually competent US execution.
Pushing the hazard DOWN: (1) Unusually strong verification plus automatic enforcement (offline licensing, cryptographic shutoff of compute on exit) pushes the impairment-by-cheating component toward the INF/New-START low end (~2–4%/yr) rather than the naval-treaty high end. (2) An unusually strong SHARED interest: unlike a pure capability race, both sides genuinely fear loss of control and extinction. This is closer to MAD-stabilized nuclear restraint (which held for decades) than to a zero-sum arms race, and it is the best structural reason Plan A might beat its reference class—though it has no clean historical precedent.
Net: I anchor on the rival-deal class (~7%/yr dissolution) but discount the dissolution rate somewhat (to ~4–7%/yr depending on the year) for the shared-survival interest and competent execution, and hold the impairment-by-cheating component low (~2.5–5%/yr) because of strong verification—while keeping impairment nonzero for non-enforcement, misexecution, and weakening revision, none of which verification prevents.
Headline base rates and the year-by-year shape
The pooled weighted base rate (all ages together) is 3.4%/yr for dissolution and 7.9%/yr for impairment (unweighted 1.8% / 4.8%; closest-analogues-only subset 6.6% / 8.3%). But the more useful object is the base rate BY PHASE (Part 2b), because the historical hazard depends strongly on how long a deal has already survived. DISSOLUTION is hump-shaped, not front-loaded: 0% in years 0–2 (no deal in the dataset formally dissolved within two years), rising to a peak of ~5–6%/yr at years 5–10, then declining to ~2%/yr for deals that survive past 20 years—rival deals take a few years to die. IMPAIRMENT is strongly front-loaded: ~15.5%/yr in years 0–2, falling to ~9%/yr (years 2–5) and ~6%/yr (years 5–10), with a late bump (~6%/yr at years 10–20, from the New START suspension and Syria/CWC). Across the full dataset impairment is more common than formal dissolution—many regimes get hollowed out rather than killed. Feeding the phase curve into the deal timeline, the pure base rate implies ~79% cumulative chance of some breakdown by 2040 (~39% dissolution, ~65% impairment). These are the “Base-rate” columns in Part 3; my “Adjusted” columns overlay the Plan-A-specific reasoning below, yielding ~62% cumulative breakdown (~43% dissolution, ~33% impairment).
My adjusted shape departs from the base rate in two deliberate ways. First, I lift EARLY dissolution above the base rate’s 0%: the historical 0% in years 0–2 is partly survivorship (deals that collapse during negotiation never enter an “entered-into-force” dataset), whereas Plan A’s setup is genuinely fragile, so I use ~5%/yr in 2029–30. Second, I cut impairment far below the base rate throughout (≈2.5–5% vs the base’s 15.5%→3%), because Plan A’s strong verification specifically suppresses the cheating-driven and born-impaired failures that dominate the historical impairment record. On top of that I add discrete dissolution spikes at the scheduled regime-change events (US elections 2032/2036, Chinese succession ~2033), and I keep the late-2030s hazard from decaying to near-zero because the prize for defecting grows as handoff and aligned superintelligence approach.
Comparison to the inside view, caveats, and where the detail lives
My adjusted estimate (~62% cumulative breakdown, ~43% cumulative dissolution by 2040) is more pessimistic than the inside-view “Overall deal decline probability view” tab (Eli’s ~48% cumulative), and the pure unadjusted base rate is harsher still (~79% breakdown). I think that gap is real and informative: base rates for rival deals over decisive technologies are harsh, and the inside view may be under-weighting how reliably this reference class fails. On the components: the weighted dissolution base rate (3.4%/yr) sits just above the “0.5–3%/yr” generic-high-salience-treaty anchor in the spreadsheet note, while the impairment base rate (7.9%/yr) is far higher—because the most applicable reference class is rival capability deals (naval treaties, JCPOA, Agreed Framework, alliances), not high-salience treaties in general, and because the dominant historical killer—political withdrawal—is not neutralized by Plan A’s verification, however strong.
Main caveats. (1) Pooled hazard rates hide enormous heterogeneity; the “right” number depends heavily on how much weight you put on the shared-survival-interest stabilizer, which has no clean precedent. (2) N is small and the rival-deal subclass is dominated by a handful of cases. (3) The outside and inside views should NOT simply be averaged—they partly double-count (regime change appears both as the base-rate driver here and as an explicit sub-cause in the decomposition tab). (4) I have treated benign endings (planned expiry, supersession by a stronger successor—SALT I, START I, SORT) as non-failures; counting them as dissolutions would push the dissolution rate higher. (5) Two structural features of Plan A genuinely lack good historical analogues—automatic compute-destruction-on-exit, and a genuinely shared existential interest—and both could justify materially lower numbers than the raw base rate if you believe they bind.
Where the detail lives: the 31-agreement dataset (per-treaty reasoning in column O), the live weighted-hazard computation (Part 2), the IR-literature anchors, and the full year-by-year base-rate-vs-adjusted table (Part 3) are all in the “Base rates” tab of the Deal Decline spreadsheet. Each treaty row records signed/in-force years, verification strength, a numeric applicability weight (column G, the input to the base-rate formula), exposure years, dissolved/impaired flags, years-lasted-unimpaired, years-to-dissolution, the dominant failure mode, reasoning, and sources.
Smooth fit (parametric survival)
The phase/step base rate (Part 2b) is noisy with only 31 agreements, so I also fit smooth parametric survival curves to the treaty-level data by weighted, right-censored maximum likelihood (Part 2c), comparing exponential, Weibull, log-logistic, and lognormal by AIC. LOGNORMAL fits best for both outcomes (dissolution AIC 301.8; impairment AIC 275.4) and reproduces the shapes the step rates show. The fitted dissolution hazard rises from ~0.2%/yr at age 0.5 to a ~4.4%/yr plateau by years 8–11—a hump, confirming dissolution is not front-loaded—while the fitted impairment hazard decays smoothly from ~17%/yr at age 0.5 to ~5.6%/yr by age 11.5 (front-loaded). Weibull is a reasonable simpler alternative (impairment k=0.71, a monotonically decreasing hazard; dissolution k=1.1, near-flat) but it cannot produce a hump, so it underfits dissolution; log-logistic is in between. The fitted per-year hazards live in Part 3 columns L–M, next to the step base rate; over 2029–2040 they imply ~34% cumulative dissolution and ~69% cumulative impairment—close to the step-function figures (39% / 65%), which is reassuring. Caveat: with N=31 and heavy weighting toward a handful of rival deals, the fitted curves are best read as smoothed summaries of a small sample, not precise estimates—the parameters in Part 2c are editable if you want to stress-test them.
FAQ
Q: Isn’t Plan A bad for concentration of power risks?
A: No, Plan A reduces COP risks by a lot. The status quo involves an insane amount of power concentrating into the frontier AI companies (which are all based in the US). If they solve the alignment problem, the result will be a handful of AI CEOs or USG officials controlling an army of superintelligences, and hence having the ability to take over the world. (This is what happened in the slowdown ending of AI 2027).
Plan A dramatically improves on the status quo by:
Increasing the number and diversity of frontier AI companies. Because of regulations slowing the frontier, trailing actors can catch up, and in the Plan A scenario, by 2040, there are dozens of companies across different countries near the frontier of AI development.
Increasing the transparency and visibility into AI development. A core dynamic driving concentration of power risks is a dynamic where most of the power centers in the world do not realize they are going to be disempowered by AI until it is too late. In Plan A, the takeoff period is much longer and much more transparent, so the public, the US government, and countries besides the US all have much more time to see what is going on inside the labs.
Spreading out the AI-driven transformation of the planet over a few more years, giving more time for everyone to prepare and react, making it more likely that something resembling our current system of checks and balances persists instead of crumbling into chaotic conflict.
Decreasing the feasibility of inserting secret loyalties into AI systems. There is high quality compute accounting,so it is no longer feasible for anyone to obtain hidden compute, which could be used to train a backdoor into the AIs.
Q: Wouldn’t China not agree to the deal? Even if they did agree on paper, wouldn’t they defect and keep developing AI anyways?
A: China is likely to want something along the lines of Plan A because the status quo outcome for China is that they will be outcompeted by US AI companies. This will result in either misaligned AIs taking over the world, or the US obtaining a decisive lead in AI capabilities sufficient to disempower China. (These two options correspond to the race and slowdown endings of AI 2027 respectively). The CCP today does not seem very aware of the importance of AGI or ASI, but this will change as AI has bigger and bigger impacts on the world.
The question isn’t “Will China defect from the deal?” The question is, “What’s the biggest defection they could get away with, and how bad would that be?”
There are two ways of defecting from the deal: (1) building a massive covert project that you attempt to keep hidden from the rest of the world, and (2) using the existing legal compute.

A baseline estimate for the maximum size of covert project that they could have is ~1.2M H100e. We do a more thorough analysis of this question in our covert project supplement, and we have a branch of our scenario gaming out how China could defect.
We prevent the use of existing legal compute through training verification mechanisms. We assume US and Chinese auditors are given full transparency into the other’s R&D datacenters, allowing them to check if the compute is being used how they claim it is being used. We do a more thorough analysis of this question in our verification supplement.
Q: Isn’t Total Research Transparency bad because it leaks algorithmic secrets to China and other AI companies?
No current AI companies claim to be secure against nation state level adversaries. They aren’t secure even for model weights, which are much easier to secure than algorithmic secrets because model weights are large, often on the order of 1-10TB, while algorithmic secrets are often memorizable by individual humans. Because AI algorithms will have already leaked to competing nations, the costs of full transparency here are relatively low compared to the enormous benefits.
Q: Plan A is complicated. Shouldn’t we do a simpler and more straightforward plan, like shutting it all down (Plan S)?
Our proposed implementation of Plan A is very complicated, and even a good implementation will incur significant existential risk. However, all simpler plans that we are aware of would incur even more risk than Plan A. In particular the main issue with Plan S is that it does worse than Plan A at making forward progress towards solving the safety/alignment problems that prevent further capability scaling, because we are shut down at an earlier capability level. This is bad because at some point, both Plan A and Plan S will break down and return to racing, and in Plan S, much less alignment progress will have been made by this point. That said, we are very sympathetic to Plan S and could imagine being convinced that some version of it is better after all.
Q: The US government is incompetent. Why are we trusting the US government to do a good job regulating AI?
For any plan, incompetent execution of that plan is a serious concern. We are in fact very worried about incompetent execution of Plan A, and so we wrote a branch of our scenario explaining how this could lead to existential catastrophe.
Who else is going to regulate AI, if the US government doesn’t? We don’t trust the industry to self-regulate effectively. Moreover, if the US government is kept in the dark while AI companies race through an intelligence explosion, that seems like an unstable situation that could lead to sudden wakeup and thrashing about, or worse, a US company effectively puppetting the government (as hinted at in the AI 2027 slowdown ending). If the US government is not in the dark, they will want to regulate AI, and we are trying to figure out and describe what good regulation would look like.
Large parts of Plan A are designed to improve the AI-related competence of the government—and the wider world—as fast as possible. For example, this is part of why we recommend Total Research Transparency instead of some more moderate kind of transparency. This is also part of why we recommend limiting the rate of AI progress; it gives the world more time to react and prepare for each new level of capability. We also have things to say about using AI to improve epistemics, and about immediate actions the government could take to improve its understanding of AI and capacity to regulate effectively. Finally, a big motivation for Plan A is to prevent extreme concentrations of power, which is good in its own right but also helps mitigate government incompetence.
Q: Plan A just isn’t going to happen. Don’t you know that?
The Plan A scenario is a recommendation, not a forecast; it is not what we expect to happen. We think many of the ideas in AI 2040 are good and will be helpful inspiration for people trying to steer things in a better direction, even if Plan A doesn’t happen. That said, we also think that Plan A is plausible enough that it is worthwhile to push for it.
Over just the last year, there has been a massive shift in the awareness and understanding of AI among policymakers. As the effects of AI get larger, there will be vastly more wakeup. Before the point of no return—the point at which AIs or AI-powered humans can disempower the rest of the world—AI will probably have transformative effects on the world. It’s hard to be confident in any specifics, but some illustrative possibilities include:
AIs can automate the majority of all jobs: physical jobs with robots, cognitive jobs with AI agents running in the cloud. This leads to massive widespread unemployment.
Extreme economic impacts other than unemployment. For example, perhaps AI companies will continue increasing their revenue at ~3x/year until 2030, which would mean that the frontier AI companies would have a $10T/year revenue. This is an unprecedented amount of wealth and power in a company, and will trigger alarm bells.
AIs completely change the nature of warfare. For example, AI piloted drone swarms may render existing mechanized militaries (e.g. infantry, tanks, planes, aircraft carriers) completely obsolete. This might lead to significant changes in the geopolitical balance of power as countries who aren’t at the AI frontier will fall behind militarily.
Extreme misuse capabilities, e.g. suppose it’s the case that AIs are able to cheaply design and spread bioweapons with commonly accessible materials. Hardening the world against these capabilities will be extremely difficult.
It’s hard to know what the political environment will look like after AI has massive effects on the world. The main question is how early transformative effects on the world happen relative to AI-caused permanent lock in. Plan A is much easier and more effective if implemented early—before the AIs have fully automated the AI R&D process.
How Plan A solves our 5 biggest problems
Problem #1: Loss of Control to superintelligent AIs.
Due to the slowdown agreement, the intelligence explosion happens over the course of at least 10 years. Instead of being forced by competitive pressures to automate work and trust AIs, companies now are forced to slow down until they have a rock solid safety case. In the Plan A scenario, we have ~5 years with top expert level AI before handoff, and we could realistically get more time. If alignment ends up being hard enough that a 10 year slowdown isn’t sufficient, the infrastructure for a longer pause is set up.
Due to the transparency and access provisions in the deal, there is way more oversight of AI developers, and many more total people have the relevant access to do AI alignment work with frontier models.
(See more here).
Problem #2: AI will lead to massive concentration of power into whatever group controls the first superintelligences (if anyone).
The transparency act already helps somewhat to prevent CEOs backdooring their models. The Total Research Transparency deal then “seals the deal” to a much greater extent; any attempt at inserting a hidden agenda into your biggest and most-used models–whether initiated by a CEO or a President–would be visible to multiple world governments. The model spec transparency + algorithmic progress slowdown + Total Research Transparency deal also makes there be many companies and governments with similar levels of AI, which is great for spreading out the power more generally, not just narrowly preventing secret loyalties/agendas.
Because the intelligence explosion is now happening in a slow and drawn out way, and we’ve prevented disparities in model access, so everyone has access to the same level of AI capabilities, anyone who is getting disempowered will be able to realize that and then negotiate for a better deal. Normal people have access to smart enough AIs that they can understand the situation well, and they can vote for ethical and thoughtful politicians, who will maintain and improve the power the public has, leading to a virtuous cycle. This leads to the Citizens Dividend and space governance proposal being implemented and spreading out the power forever.
Problem #3: World War 3 risk due to superintelligences disempowering many countries, including nuclear powers.
Plan A prevents WW3 by ending the “winner take all” effects of a race to superintelligence. No longer is there a race to ASI with each side fearing what the other might do if they get there first, and with each side contemplating taking aggressive action to slow down the other side, and with one or more sides starting to freak out as they realize they are on a path towards losing. No longer are there a bunch of middle countries gradually realizing that no matter who wins, they lose, and starting to freak out also.
Problem #4: AI taking everyone’s jobs.
The reasons that unemployment is bad are that people lose their (a) income, (b) power, and (c) source of meaning. Plan A slows down AI induced job loss to a rate where the negative effects of unemployment can be mitigated.
The main reason that people need jobs is to have a source of income. Plan A addresses the loss of income via the Citizen’s Dividend. Each person gets a share of income large enough to support an extremely high quality of life.
A second reason why jobs are important is that they empower their holders. CEOs and governments are incentivized to care for people who provide economic value. However, once AIs can do all the economically relevant tasks, governments and companies lose a strong incentive to provide for their citizens/employees. See the “concentration of power section” for more.
A final reason jobs are important is that they provide meaning for many people. Here we just say that it won’t be that bad (people can get meaning in a bunch of ways, not just economically useful work! Family, friends, romance, hobbies, sports!) and that the benefits outweigh the costs. If you disagree, well, we also describe Plan S… there’s basically no way to avoid mass unemployment if you allow companies to build and deploy AIs that obsolete humans. And in particular, banning the use of AI in various professions won’t work for long.
Problem #5: AI misuse by small actors.
Widespread access to smart AI will make it trivially easy to build bioweapons (or other technology that would cause huge damage at little cost to the attacker). Terrorists, states (e.g. North Korea), or misanthropists who are motivated to cause immense harm could leverage AI to do a much larger attack than they would have otherwise.
Plan A, as part of mitigating covert projects, limits the capability of open weight models to pre-deal levels. Closed models are trained to refuse dangerous requests. In case access limitations fail, Plan A involves massive investments into biosecurity and other measures to improve the resilience of the world.
There’s an alternate approach to misuse taken in Plan B, C, and D, which is to speedrun the intelligence explosion and hope to build superintelligence that can directly take over the world, hence preventing misuse. It’s plausible that these plans look better for misuse risk than Plan A because there’s so little time for the misuse to occur before superintelligence. But this strategy imposes unacceptable levels of loss of control and concentration of power risk, and so doesn’t pass overall cost benefit.
Plan A Assumptions
Summary
Plan A is a scenario: it combines forecasts with policy recommendations. Some of the forecasts and recommendations we are confident in; others are just guesses we made because we had to, in order to achieve the desired level of concreteness.
Confident | Unconfident | |
Prediction | E.g.: “AGI is coming soon and will be a bigger deal than the industrial revolution” | E.g.: “AGI will arrive in 2030” |
Recommendation | E.g. “There should be an international deal to limit the speed of AI takeoff” | E.g. “Many new datacenters should be constructed in Mongolia and Canada.” |
This supplement goes through the major predictions and recommendations of the scenario and explains which are which.
Nearly all of the actions in the scenario assume some optimism and some realism, but by and large, we tried to be realistic about empirical facts and optimistic about US government actions.Why isn’t this scenario purely a prediction? We already did that: read AI 2027. Why isn’t this scenario purely a recommendation? If we completely abandoned realism, the output would be useless because it was so unrealistic.Why isn’t this scenario purely a prediction? We already did that: read AI 2027. Why isn’t this scenario purely a recommendation? If we completely abandoned realism, the output would be useless because it was so unrealistic. The AI 2040 scenario has roughly my median view about things like timelines, takeoff speeds, alignment difficulty, and actions of non-US governments (e.g. China, the EU, etc). The actions the US government takes are optimistic by definition as they are following our recommendations nearly exactly; in reality we still expect something more like the response in 2027.
The supplement tries to capture all our main recommendations, so if something isn’t mentioned here, it is probably a prediction, and insofar as it’s a recommendation, it is not an important one.
Since this is a document about specific beliefs, it’s written from the perspective of one of the scenario authors (Thomas). Other authors by and large agree with the central claims but have somewhat different opinions on the specifics.
Forecasts Summary
For each of the following, I am >80% confident:
AI R&D will be fully automated within the next ~15 years, absent substantial regulation, war, or societal collapse. (See the AI Futures Model for more detail; AIFP employees have somewhat different views, but all agree that 2027 and 2030 are plausible and reasonably central years.)
Within a few years of AI R&D automation we will build AIs that are better than humans at all economically relevant tasks (absent substantial policy intervention). These AIs would quickly become vastly cheaper than humans, and so soon humans will become economically irrelevant.
If we build vastly superhuman AIs, the AIs could take over, if they collectively wanted to. Therefore, it is critically important that most of them do not want to.
I’m not confident about:
The exact timeline. In AI 2027, the automation of coding happened in 2027; in Plan A, this would have happened in 2030 in the absence of governance, but is slowed by a few years. My 90% confidence interval is mid-2027 to ~2050. We chose 2030 because it’s near enough that our specific planning is useful, but far enough that the government has time for the preparation phase we recommend. If timelines are shorter, there isn’t time to prepare and so Plan A becomes more rushed; if longer, more of the details are wrong but the high-level plan stays the same. (more)
Takeoff speeds. How long between automating coding and AIs that dominate the top human at everything? We assume a default of ~1 year (my median), but my 80% CI ranges from 2 months to 5 years. Under my views about takeoff, the plan is basically similar under different takeoff speeds, though there are some small prioritization updates depending on the direction. (more)
Alignment difficulty. We don’t know how hard it will be to solve the AI alignment problem. I think it’s possible that even without a major slowdown, we are able to solve the alignment problem. Similarly, I think it’s possible that even 10+ years of dedicated effort, with lots of AI assistance, won’t be enough. Plan A is designed to buy time and attempt to gain evidence about how difficult the problem is. If evidence accumulates that AI takeover risk is small, then we can proceed faster, otherwise, the pause extends. (more)
How difficult it will be to verify compliance with a slowdown deal. It could be that it’s basically intractable to do major AI development without access to giant datacenters which are easily monitored. On the other hand, it could be that algorithmic progress is doable even with a very small amount of compute, which means that very quickly even a very small cluster could do an intelligence explosion. In even the most pessimistic case, we think it’s very likely that a deal could be robustly enforced for a year or two. (more)
Many other strategic variables. (more)
For each of these key strategic variables, the Plan A scenario chooses a reasonably central value with respect to my views. For many of these my views are very uncertain, and many policy decisions should be made with more information about these variables than we currently have.
Policies Summary
In the next section, I discuss which of our policy recommendations in the scenario are robust across a wide range of empirical updates.
These are robustly good policy recommendations:
Improve AI Preparedness in government (importance: medium). Transparency, improving technical expertise within government, and measures to accelerate verification research. These policies are good because they are low-regret, and in some scenarios (like Plan A), they are very important.
Model spec transparency and verification against power grabs (importance: medium). Model specifications (of both internal and external models) should explicitly prohibit assisting with power consolidation (by CEOs, presidents, or anyone else), and AI projects should be required to implement verification measures to prevent backdoors.
US/China deal to slow the capabilities explosion, based on compute governance (importance: high). This is the core of Plan A. I’m confident that a cooperative international approach (Plan A) is better than the alternatives because the alternatives involve being in a race during the intelligence explosion, which incurs huge risks. I’m not confident about the exact variant of Plan A, or whether Plan A as we’re imagining it (a many year mutual slowdown) is better than Plan S (a full shutdown). Conditional on attempting to “slow down”This is in scare quotes because, as our scenario illustrates, even if we do a well-enforced international deal to limit the pace of AI progress, the overall pace of AI progress will be extremely fast compared to what most people are expecting and used to. For more on why we think this, you can view our models of AI progress (the AI Futures Model) and of AI economic impacts (Economics of Plan A supplement), as well as numerous blog posts and of course AI 2027.This is in scare quotes because, as our scenario illustrates, even if we do a well-enforced international deal to limit the pace of AI progress, the overall pace of AI progress will be extremely fast compared to what most people are expecting and used to. For more on why we think this, you can view our models of AI progress (the AI Futures Model) and of AI economic impacts (Economics of Plan A supplement), as well as numerous blog posts and of course AI 2027. AI progress, doing this via compute governance looks like by far the most effective way of accomplishing this goal (even though I’m not sure exactly how well it will work).
Redistribution of AI-generated wealth (importance: medium). As AIs automate the entire economy, the wages that humans will be able to earn will plummet to near-zero. By default (and assuming we avoid AI takeover), AI will cause historically unprecedented wealth inequality. I feel confident that there needs to be some measures for wealth redistribution.
Defensive Acceleration (importance: medium). Governments, AI companies, and the public should invest in technologies that make society more robust. Particularly important domains include AI for epistemics, bio defense, and cyber security. In the absence of a concerted effort to pursue defensive acceleration, increased offensive capabilities coming from high AI capabilities will lead to significant risk. Defensive acceleration also has positive externalities because it makes it harder for misaligned AIs to take over the world, and thus gives us more time to do alignment research
Thoughtful post-ASI governance (importance: high). We feel good about some general principles for post ASI governance: (1) have lots of smart people use (aligned) AI assistance to think carefully about effective governance, (2) be careful making irreversible commitments, (3) make it hard for any individual actor to consolidate power.
These recommendations depend on particular assumptions that we’re not confident in, or have alternatives that might be better:
How fast to scale AI capabilities after a deal (importance: high). Scaling faster incurs more immediate safety risk, while scaling slower means that you delay safety progress that’s downstream of having smarter AIs and incurs risk of deal dissolution and covert projects. These decisions will and should be made largely based on the evidence at the time about how risky particular AI systems and deployments seem to be, and how good our mitigations seem to be. (more)
Massive datacenter buildout (importance: medium). In our Plan A scenario, we recommend that governments encourage a massive datacenter buildout. This recommendation is contingent on having implemented an effective international deal to slow down AI, and even conditional on that, it’s not clearly good. The case in favor is that hardware scaleups are more governable than software scale ups because we can prevent the hardware from proliferating. But there are major costs: (i) it increases the required accuracy of verification within the legal projects, and (ii) if the deal breaks down (and failsafe mechanisms fail), it could result in the intelligence explosion afterwards being much faster. (more)
Our verification plan (importance: high). The verification plan described in Plan A is important for making sure that both the US and China are following the agreement. But I’m quite unconfident about the exact best way to implement it. We have a specific proposal which involves first implementing inference-only (and stopping the construction of new AI models) while sprinting to figure out how to verify training and experiments. Many alternative approaches to verification exist. (more)
What transparency/security tradeoffs to make (importance: high). Security and transparency both have obvious upsides, but are in tension. We’d like as much transparency as possible to get as many eyes as possible on the safety cases being presented, but we’d also like to secure model weights, algorithms, and inference tokens. Overall, we think that we should be willing to pay substantial security costs in return for improved transparency, because we are more worried about major AI projects abusing their power and/or making mistakes, than we are about defecting projects with a tiny amount of compute (who are the main beneficiaries of poor security). The tiny projects just aren’t going to be as powerful, and bad or reckless behavior by them correspondingly doesn’t matter as much. (more)
Our post-ASI governance plan (importance: high). Our space governance plan is, in my opinion, a reasonable first pass about what should be aimed for, and is generally applicable to any scenario where humanity retains control. However, I am not at all confident in the details, and in these scenarios, we will have aligned, superhuman AIs to help us puzzle this out, so my main recommendation is to use the huge amounts of AI labor to come up with a better plan. (more)
Broad deployment of AI (importance: medium). Broad deployment spreads power and could improve societal epistemics. Some central ways it could be bad are (i) the net effect on societal epistemics might be bad, (ii) broad deployment could make a persuasion/political power takeover easier for the AIs (or for a human attempting an AI takeover) or (iii) they could lead to the proliferation of offense dominant capabilities. On net, I think broad deployment is probably good, but I feel uncertain, and it depends on how good mitigations are in place for each of the above failure modes. (more)
Various other details of Plan A implemented in our particular scenario. This includes things like datacenters on the oceans, incentive shaping with respect to alignment, incentive shaping with respect to covert projects (e.g. hardened compute stockpiles), how to trade off deal dissolution risk, covert project takeover risk, and risk of incompetent governance.
The above were largely details about Plan A that might change. There are also structurally different plans that may be competitive with Plan A, which we discuss here.
Still, I think we should be aiming for something like Plan A across a wide range of scenarios. The rest of this document has two sections, which both discuss how our plan would change in response to empirical changes.
Empirical Updates answers the question “if an empirical fact changed (e.g., suppose timelines were shorter or longer), how would our policy change”.
Policy Updates answers the question: “for each of our major policy recommendations, what would make us change that recommendation”.
Empirical Updates
In this section, I’ll go over how Plan A would change in response to some important empirical assumptions that we make in the scenario, which might turn out to be different in practice.
Timelines
We don’t know when AGI will be built; it’s possible that it could be built this year or next year, it’s also possible that it will take decades.AGI = Artificial General Intelligence, or, AIs that have a similar amount of general intelligence to humans. Our actual thinking uses more precise concepts which you can read about in the AI Futures Model. Concepts like AC, SAR, TED-AI. However, we find “AGI” to be a handy shorthand for some vague combo of the above milestones, because we expect those milestones to arrive within a few years of each other or less. So when we say ‘AGI’, understand that we are gesturing at milestones such as the above, and being deliberately nonspecific because the context doesn’t require precisionAGI = Artificial General Intelligence, or, AIs that have a similar amount of general intelligence to humans. Our actual thinking uses more precise concepts which you can read about in the AI Futures Model. Concepts like AC, SAR, TED-AI. However, we find “AGI” to be a handy shorthand for some vague combo of the above milestones, because we expect those milestones to arrive within a few years of each other or less. So when we say ‘AGI’, understand that we are gesturing at milestones such as the above, and being deliberately nonspecific because the context doesn’t require precision
In the Plan A scenario, the full automation of AI R&D was on track to happen in 2030 by default, but then was pushed back by government involvement. We’ve written much more on AI timelines; you can find our most updated work here. AIFP employees continue to have (slightly) different views about AI timelines, but we all think that 2030 is a plausible and reasonably central year for AGI to arrive.
Why did we choose to write this scenario assuming 2030 timelines?
It’s good to showcase a variety of different plausible scenarios. We already wrote a very short timeline scenario (AI 2027), so we figured it would be good to choose a different timeline for our next one. 2030 is my (Thomas’) median currently, i.e. I think there is a 50% chance things will proceed slower, and 50% chance things will proceed faster. In the future, if time remains, we’d like to do scenarios for other authors like Eli.Eli’s timelines are similar to mine, but his takeoff would be about two times slower.Eli’s timelines are similar to mine, but his takeoff would be about two times slower.
Overall, Plan A remains my central plan ~regardless of timeline, because the underlying problems (AI takeover, AI enabled concentration of power, and rapid AGI induced societal upheaval) are the same regardless of timelines, and Plan A is our best guess for how to deal with these problems.
Regardless of the timeline, Plan A is most effective the earlier an international deal that involves mutual compute declarations and transparency begins. When we get to the beginning of the intelligence explosion (e.g. the automated coder milestone), verification becomes substantially more difficult, because defecting projects can involve many fewer humans, and therefore be much harder to detect.
If the timeline is much longer (e.g. 2035+), we begin to have much more uncertainty about the situation. The world will also be more transformed pre-AGI relative to our scenario, there will be larger economic effects of AI, AI company revenues (and share of the overall market) will (probably) be higher, there will be more robots around, and there will be more wakeup to the importance of AI. Particular effects include:
We become somewhat less worried about covert projects because compute stockpiles are inherently insanely large, and so the relative amount of compute that covert projects have compared to the verified projects will be smaller. (The absolute amount of compute, on the other hand, is larger). However, there are also reasons to be more concerned about covert projects—a counteracting effect is that there will be more compute distributed throughout the world which will be harder for intelligence agencies to locate.
It becomes more important to do the economic parts of the Plan A policy package (e.g. the Citizen’s Dividend) earlier, because more people are affected earlier.
It becomes more important to do D/acc earlier because there is a longer period of time with reasonably intelligent AIs around, and current AIs already provide substantial biouplift.
Since the overall amount of compute in the world is higher, we become more worried about a new paradigm being discovered that’s much more compute efficient, which could lead to a fast intelligence explosion.
If timelines are much shorter (e.g. 2026 or 2027), there isn’t time to do the preparation phase, and so we have to do a version of Plan A with no preparation.
It’s more likely that we should pivot to one of the contingency plans that involves less government capacity. Unfortunately, these contingency plans have other downsides. For example, shutting down large fractions of AI relevant compute would also turn off inference of existing models, which would be economically costly.
It’s more urgent that an agreement to slow down AI takeoff happens early, because we’ll soon be so deep into the intelligence explosion that it’s impossible to slow down.
Takeoff speeds
In this scenario, similar to AI 2027, we chose to assume that the default length of time between AC (Automated coder) and TED-AI (Top-Expert-Dominating AI) was 1 year. This is roughly my median, but we again have large uncertainty; my 80% CI at time of writing for the takeoff duration is something like [2 months; 5 years].Note: some people use takeoff speeds to refer to the period after TEDAI, it’s an ambiguous term, but for the purposes of this document, I’ll use the definition given here: the length of time between AC and TED AI.Note: some people use takeoff speeds to refer to the period after TEDAI, it’s an ambiguous term, but for the purposes of this document, I’ll use the definition given here: the length of time between AC and TED AI.
Why did we choose 1 year takeoff speeds by default? This is my median at the time of writing Plan A, and seemed like a representative scenario to plan for.
How does the plan change under different takeoff speeds?
If takeoff is faster than 1 year, things become generally more difficult, and not just for Plan A. The main specific changes are:
Doing an international deal to slow down takeoff becomes more important, because the status quo is more risky, because AI driven change without deal happens much faster and there is less time to react to each change and prepare for the next changes.
Enforcement of an international deal becomes more difficult. Faster takeoff typically involves much less hardware increase, because if the takeoff occurs over months, there isn’t time for a massive compute scale up. The possibility of an intelligence explosion without a massive compute scaleup is the key assumption behind being worried about covert projects; if such an intelligence explosion wasn’t possible, then covert projects with relatively small amounts of compute would be much less concerning. Since the intelligence explosion happens faster, it also makes covert projects faster, reducing the amount of slack that the international Consortium has to work with.To illustrate, suppose that the length of time between AC and TEDAI is one month. Even if we can (optimistically) use compute governance to slow down rogue projects with ACs by 20x, you still only have 2 years until the rogue projects have TEDAI, which is a smaller window than Plan A assumes.To illustrate, suppose that the length of time between AC and TEDAI is one month. Even if we can (optimistically) use compute governance to slow down rogue projects with ACs by 20x, you still only have 2 years until the rogue projects have TEDAI, which is a smaller window than Plan A assumes.
The slowdown agreement needs to be started earlier, which would make the enforcement easier (and therefore the whole plan more viable).
Since enforcement of the deal is more difficult, this is an update towards making tradeoffs that help with enforcement but with various other costs, such as (i) leaning more towards security on the security/transparency tradeoff, (ii) making less algorithmic progress, and therefore scaling more slowly, even if scaling further would still be safe, and (iii) doing more early hardware scaling, which makes verification and self-destruct plans more load bearing, but helps reduce misalignment risk.This reasoning is given in the datacenter buildout section below.This reasoning is given in the datacenter buildout section below.
If takeoff is slower than 1 year, e.g. even without regulation it takes several years to go from the full automation of coding to the first artificial superintelligence, then we basically update in the opposite direction on all of the above tradeoffs.
Overall risk becomes lower, because there is more time by default to deal with risks as they arise. Therefore, the overall plan will attempt to incur a smaller amount of risk, and will be willing to make more aggressive tradeoffs to reduce risk.
Covert project takeover in particular becomes less likely, because it is extremely unlikely that they could bootstrap into existentially dangerous AI systems with only a tiny amount of compute. This updates us even further in favor of doing Total Research Transparency (because we are relatively more worried about the large projects than the small projects), and away from doing the large hardware scaleup. The plan can involve just doing a slowed down version of the intelligence explosion on the existing compute stockpile, leaking algorithms to the trailing project, but relying on the fact that they are so compute limited that they cannot make significant further AI progress.Note that covert projects doing an intelligence explosion is only one of the reasons to be worried about them, there’s also the risk of them directly obtaining a decisive strategic advantage; so it would still be risky if they were, e.g., able to train a top expert level AI.Note that covert projects doing an intelligence explosion is only one of the reasons to be worried about them, there’s also the risk of them directly obtaining a decisive strategic advantage; so it would still be risky if they were, e.g., able to train a top expert level AI.
AI takeover is less likely, and so we can devote relatively more effort on mitigating other risks. It’s still very important to slow down AI progress, because there’s still a significant chance that AI alignment won’t be solved during the default takeoff, and because slowing down helps to mitigate other risks like concentration of power. However, it’s more likely that we wouldn’t actually slow down much in this world. Therefore, doing the slowdown part of the treaty is less load bearing.
Overall, we recommend pursuing Plan A regardless of takeoff speeds. However, changes to our takeoff speedsAs well as various aspects of takeoff besides speed, notably the possibility of fast progress on a small, fixed, amount of hardware.As well as various aspects of takeoff besides speed, notably the possibility of fast progress on a small, fixed, amount of hardware. do substantially change the specific implementation of Plan A, in the ways noted above.
Alignment difficulty
In the Plan A scenario, humanity rallies and creates a high assurance safety case for handing off to ASI. However, we are much more confident in the general picture that more time and effort lead to a greater chance of solving the alignment problem than the specific story told here. We don’t know how hard it’ll be to align ASI. Opinions vary substantially within our team (and even more outside of our team). Given the uncertainty, our view is that we should take a portfolio approach, and invest substantially in both prosaic AI safety agendas (e.g. variants of RLHF) and moonshots (e.g. uploading human brains).
The plan we are advocating for is one in which the world coordinates to give the scientific community breathing room and resources to figure out how difficult the alignment problem is. If Yudkowsky is right that AI alignment is difficult, our view is that evidence would increase during this time period and then the pause would be extended. Likewise, if the optimists are right and alignment is easy, we would get evidence of this during that period, and we could proceed at a faster pace (though it’s not clear that we should, given the other risks of doing so, most notably concentration of power).
I think that alignment is probably difficult enough that without a well managed multi-year slowdown, we will incur substantial AI takeover risk. Even with a ten year slowdown, as depicted in the scenario, there is a significant chance that alignment was too hard, or that it’s feasible but human error causes us to fail anyway. You can see more about the team’s views about alignment success likelihood conditional on following different plans here.
What do we do if alignment is much harder?
Gather as much evidence about alignment difficulty as possible by:
Deploying and observing AI systems to get empirical feedback about how aligned they seem to be.There’s a good chance that empirical feedback will be highly ambiguous, but overall I’m optimistic that if there’s a long period with broad access to experiment on roughly human level AIs, then we will surface a huge amount of evidence related to the difficulty of alignment. Given the current (poor) epistemic environment, I don’t think this evidence will be sufficient for broad scientific consensus (until we have a good fundamental understanding of AI), but I personally expect to learn a lot and change my mind in substantive ways.There’s a good chance that empirical feedback will be highly ambiguous, but overall I’m optimistic that if there’s a long period with broad access to experiment on roughly human level AIs, then we will surface a huge amount of evidence related to the difficulty of alignment. Given the current (poor) epistemic environment, I don’t think this evidence will be sufficient for broad scientific consensus (until we have a good fundamental understanding of AI), but I personally expect to learn a lot and change my mind in substantive ways.
Building model organisms—AIs trained to deliberately be misaligned—and see if alignment techniques are good enough to catch and train away the misalignment. Do this for deceptive alignment, sycophancy, reward hacking, etc.
Giving a large, heterogenous set of people access to the models (ideally including model internals) to avoid groupthink.
Have explicit safety cases that lay out the evidence and arguments for why the models can and should be trusted, and publish them so they can be critiqued and debated by the scientific community.
During Plan A, maintain the optionality of extending the pause by:
Monitoring for covert projects and shutting them down if detected
Ensuring that the flow of new chips is highly secure and not going to any rogue projects
Improve coordination and verification technology, both with large investment and using AIs to improve this technology. In particular, using AIs to build privacy preserving auditing and reliable lie detection seems very promising.
Minimizing software progress that leaks to covert projects. If alignment is harder, we should update towards making tradeoffs that enable us to stay slowed for longer.
Making sure decisionmakers have access to the best information on alignment difficulty, including by making sure that people who are more pessimistic about alignment have access to the models and the current safety cases, so that they can find problems in them before it becomes catastrophic.
If alignment is extremely hard, e.g., humanity isn’t making substantial progress on alignment whatsoever, then the correct call may be to pivot to Plan S, and instead of attempting a managed takeoff, shut it all down. Then, we could pursue alternate strategies to build superhuman AI that are inherently safer than ML, such as uploads. Unfortunately, these approaches have substantial downsides as well, notably the risk that the moratorium breaks down and we go back to racing before enough progress is made on uploads or other strategies in that vein.
What if alignment is much easier?
Slowing down AI takeoff is still probably good for many other reasons, including avoiding AI-driven concentration of power, figuring out appropriate post-ASI governance, transitioning to a regime where human labor is no longer economically relevant in a healthy manner, hardening the world to AI enabled superweapons, etc—basically, we still need to solve all of the problems that we dealt with in Plan A other than misaligned AI takeover.
The main difference is that there is a higher chance of humanity being able to successfully muddle through these other risks without permanent lock in. This means that the necessity of a slowdown does go down substantially if alignment is much easier than I expect.
Verification difficulty
Related reading: Covert Project Branch, Covert AI Projects Supplement, Verification Plan.
I’m not sure how difficult verification will be. Our main strategy for verifying compliance with an AI agreement is (1) limit the amount of compute that covert projects have access to, and (2) ensure that legal projects are doing what they claim. Both of these will be somewhat lossy, and we’re unsure about what the mapping is between “amount of unmonitored compute” and “how fast could an illicit intelligence explosion reach an existentially dangerous capability threshold”.
Now we’ll go through each of these in a bit more detail:
Covert project verification difficulty. In the Plan A scenario, we assume that the maximum plausible covert project size without detection is about 1% of pre-deal compute. Our assumption here in the scenario is close to the worst case scenario. My best guess is that neither side would build a covert project in a Plan A scenario because (i) there’s no incentive to because it’ll be out paced by the legal projects, and (ii) both sides are scared of AI takeover risk; and building a covert project would undermine the deal. Sub-questions I’m most uncertain about:
What fraction of pre-deal compute can be tracked down? Our uncertainty over pre-deal compute is very relevant, because it will determine the extent to which China’s initial compute declaration can be falsified. Existing experts we’ve talked to maintain huge amounts of uncertainty about how large China’s existing compute buildout is. As a percentage, this might become easier to track as AI datacenters get larger, but as an absolute number more and more compute will be harder to track down over time. However, the intelligence community may have access to private information that’s much more accurate today, and they might begin to gather this information in the near future. Overall, it’s possible that denying covert projects access to huge numbers of GPUs will be one of the most load bearing components in the non-proliferation plan; but it’s also possible that other factors such as HUMINT will be far more important.Note that pre-deal compute matters a lot because it seems probable that post-deal compute will be effectively monitored, because (1) there are very few extremely expensive fabs which can build the relevant chips, all of which can be directly monitored with auditors, and (2) it is intractable to build illicit fabs.Note that pre-deal compute matters a lot because it seems probable that post-deal compute will be effectively monitored, because (1) there are very few extremely expensive fabs which can build the relevant chips, all of which can be directly monitored with auditors, and (2) it is intractable to build illicit fabs.
How effective will HUMINT be for detecting covert projects? It seems likely to me that a covert project of the size we assume in the covert project branch of Plan A (500k H100e, $7.9B) would be quickly detected by intelligence gathering. Historical mega projects that were attempted in secret (e.g. the Manhattan Project) seem to have been riddled with spies; our baseline assumption should be that anything on the order of this size will be quickly detected and shut down. However, later into the intelligence explosion, the number of people involved in a covert project could be much smaller, as AIs could automate the AI research and robots could automate the construction and maintenance of the datacenter, dramatically reducing the probability of leaks.
What is the necessary reduction in access to compute resources? In the Plan A scenario, we assume that an AI project with 10x less compute can do an intelligence explosion 5.5x slower; this is discussed more in the appendix of our Covert Projects Supplement.
Legal project verification difficulty. In the Plan A scenario, we assume that we can verify inference-only very quickly, and that we can verify training runs after a six month sprint to develop that technology. It could be easier (e.g. it’s possible to do at the software level), or it could be harder (particularly because governments are slow and ineffective at this). Our assumption here is on the optimistic end because we assume high competence in execution.
If legal project verification is more difficult, then increasing hardware buildout is more scary. Quantitatively, with no hardware buildout, then we’d need something like 99% assurance on legal project compute; with a +4 OOM increase in stock, we’d need 99.9999% assurance on legal project compute.
There are various other particular verification dynamics which could affect the rest of the policy strategy:
Favorability of SW scaling? A key dynamic for covert projects is stealing algorithms from legal projects. Therefore, legal projects should consider pursuing algorithms that are less favorable for covert projects. In particular: (1) HW/SW co-design, i.e. research algorithms that are only effective on specifically designed chips that covert projects don’t have access to, and (2) heavily scale-dependent algorithms. We’re unsure how favorable both of these paths will be; the more favorable they are, the more capabilities progress can be done without speeding up covert projects.
How much edge-hardware is required? Some applications strongly prefer edge hardware (hardware that operates on the device as opposed to the cloud): notably military applications (because of jamming) and self-driving cars (because you don’t want your car to crash if there’s a network issue). This poses a challenge for verification because edge compute could be removed from the device and funneled to a covert project, or potentially backdoored and used as part of a large distributed training run. We’re uncertain to what extent large quantities of this will be required, if there are other solutions, and how good the mitigations look for preventing usage of edge compute in training runs.
How difficult will security be? Nation-state proof securitySpecifically, security to China, because China is probably the best placed actor to run a covert project.Specifically, security to China, because China is probably the best placed actor to run a covert project. is basically required to make the legal project verification work. Nation state proof algorithmic security is helpful for preventing covert projects. Our baseline assumption on security is that it’s possible to make verification mechanisms sufficiently secure, most algorithms will not be possible to secure (because they are so small), and that model weights will be possible to secure (because they are so large). I’ll now go through the updates if our current assumption is wrong for each of the desired security properties:
If verification security isn’t viable then our ideal strategy would either involve (1) doing a moratorium on further AI development (and potentially inference) until security is viable or (2) relying heavily on trust.
If algorithm security is viable then we’d have more buffer against covert projects than we are currently modeling in Plan A. We’d also have less of a need to do a compute buildout, because one of the core motivations for building compute was to mitigate algorithms leaking to covert projects.
If model weights security isn’t viable then we’d either need to (1) not scale AI capabilities until security becomes viable or (2) need to invest more heavily in defending against exfiltrated AIs. At very high levels of capability, this might involve extremely drastic measures, because we’ll need to defend against extremely capable AIs.For example, we might need to do things like move all humans into bunkers that are robust to nanotechnology (friendly nanotech is defending), mirror life (no air is allowed to enter), and superpersuasion (all information coming in is mediated by friendly AIs, who screen out superpersuasive content).For example, we might need to do things like move all humans into bunkers that are robust to nanotechnology (friendly nanotech is defending), mirror life (no air is allowed to enter), and superpersuasion (all information coming in is mediated by friendly AIs, who screen out superpersuasive content).
AI Paradigm Shifts
Plan A currently assumes gradient descent on giant neural networks remains the dominant AI paradigm until the late 2030s. We assume significant advances are made within that paradigm, including (i) the addition of neuralese (both as long term memory, and replacing chain-of-thought) and (ii) dramatic scaling in reinforcement learning (many more diverse environments that are much larger than existing environments).
Substantial algorithmic changes, even within the broad “gradient descent on giant neural networks” paradigm, might change the implementation details of parts of Plan A. For example:
We assume that online learning via live weight updates (instead of in context learning or delayed weight updates) is not necessary for generally top expert level AI systems of the type developed in Plan A. If live weight updates were necessary, then we would need to have some training infrastructure happening on the inference datacenters, which would complicate the inference verification because we’d need to live update the reference models in the verification infrastructure, which would increase the security risk surface area. It would also prevent us from using chips that are hardcoded to do inference on a given model (e.g. Etched), which could be part of the inference-only verification scheme.
We assume that restricting access to compute is an effective measure to prevent covert projects. If there are, for example, new architectural discoveries with dramatically (e.g. 1000x) more favorable compute scaling than current architectures, this would make covert projects much more dangerous. That said, if such architectural discoveries are in the pipeline so to speak, all the other plans are much less likely to work too—a discovery which allows a covert project to overtake the legal projects in Plan A, despite a huge compute disadvantage, would in Plan D cause an extremely fast, discontinuous “FOOM” to ASI within whichever frontier AI project first found it. These worlds are extremely scary and probably the best way to handle them is something like Plan S.
There are also obviously many “unknown unknowns” in AI paradigm shifts which could change Plan A in unforeseen ways, in practice we expect that many updates to the Plan A baseline will need to be made in response to further paradigm shifts.
Less central updates
State capacity. Plan A relies on high state capacity on AI, in particular, the capacity to distinguish existentially unsafe training runs and deployments.Thanks to the transparency of AI development to other companies, many governments, and the public, there are some failsafes in place for governments making poor decisions about safety techniques.Thanks to the transparency of AI development to other companies, many governments, and the public, there are some failsafes in place for governments making poor decisions about safety techniques. However, in worlds where the state capacity is low enough that they are not able to distinguish safe and unsafe AI, we recommend pivoting to another plan. Two salient alternatives are: (1) having an extremely conservative bar for safety, (e.g. shut it all down), or (2) limit the speed of the intelligence explosion through more blunt techniques (e.g. mutual GPU destruction).
Capability level of the initial deal. The later the deal begins, the more difficult enforcement becomes. In particular, once we get substantial automation of AI research, it becomes easier to run a covert project on AI with only a small number of people involved, which reduces the surface area for leaks. This means that the overall regime has less time, and so risk will be higher.
Distillation mitigation. We discuss our strategy for mitigating distillation in the verification supplement. It’s unclear how well this will work. If it doesn’t work well, it’s an update towards being more concerned with covert projects. It is also an update towards slower capability scaling being optimal. Finally, it’s also an update towards strongly limiting broad deployment, although I still think that even if mitigating distillation is very difficult, it would still be worthwhile to do broad deployment.
Policy Updates
How fast to scale AI progress after a slowdown deal?
It’s unclear, and will be dependent on observations at the time. Here are some key considerations:
Is the deal likely to break down soon? If the deal is more likely to break down soon (for example, due to political reasons, or due to covert project takeover), then you should be willing to take on more risk, and so you should be more willing to scale faster.
How good do current safety cases look (and how misaligned do the models seem)? The more AI takeover risk there is from marginal scaling, the slower we should scale capabilities.
How large are the returns to better AIs? If better AIs won’t contribute useful labor (for example because we can’t trust their outputs, and control isn’t very scaleable), then we shouldn’t scale. On the other hand, if scaling AI capabilities would dramatically increase the amount of high-quality alignment research we can do, it may be worth scaling, even if doing so is risky.
Will the weights and algorithms leak to covert projects, and how bad are the leaks? If the model weights and/or algorithms of a scaled up AI model were to leak this would accelerate defecting projects by some amount, decreasing the amount of slack the regime has. If it’s possible to secure the model weights and algorithms, or to scale capabilities without doing substantial algorithmic progress, then we should update towards scaling more. In the Plan A scenario, we mostly scale with hardware buildouts, to avoid the need to make substantial algorithmic progress, because we are not optimistic about securing the algorithmic secrets.
How much risk comes from current AIs vs future AIs? The earlier in the intelligence explosion we are, the more we should be sympathetic to scaling earlier, because the marginal risk of current scaling is much lower, whereas if we’re in the capabilities regime where we are incurring most of the risk, we should be much more careful and tend to slow down more. (This is why, in the scenario, the Consortium scales up AI capabilities reasonably but not entirely cautiously in the early 2030s, and then pauses at top expert level in 2035.)
Is it possible to scale up capabilities without substantial algorithmic improvements? In plan A, we try to scale up AI capabilities via hardware scaling, while heavily limiting algorithmic progress. However, it may be the case that scaling up AI capabilities inherently leads to lots of algorithmic progress (e.g. via distillation). Moreover, the main alternative to scaling up algorithms is scaling hardware. However, it may be the case that most algorithmic progress is downstream of hardware progress. In this case, the most effective way of governing algorithmic progress might be to limit hardware itself.
Ultimately, trading off these considerations will require quantitatively modeling each of these sources of risk: covert project risk, risk of AI takeover from current AIs and future AIs, etc, understanding how the speed of deployment impacts each of these risks, and then making complicated tradeoffs about the overall best path forward.
Plan A’s trajectory is essentially based on the viewpoint that deal dissolution risk will probably be high, that risks from early AGIs (e.g. ACs and SARs) can be mitigated via control (i.e. that control-based safety cases can be made good enough very quickly, for AIs at that level), and that we can elicit a huge amount of useful labor from these early AGIs. It also assumes that we can do most of our capabilities scaling via a rapid hardware buildout, and strongly limit the total amount of algorithmic progress that occurs, including successfully mitigating distillation (largely by hiding the intermediate reasoning traces). Under these assumptions, the best trajectory seems to involve scaling quite rapidly at the beginning of the slowdown agreement, and then slowing down dramatically at the maximum controllable capability level.
Plan A’s specific trajectory is our best guess for the best way to handle these tradeoffs. However, we aren’t confident whatsoever that this is the correct path, and we hope that decisionmakers at the time will have a better understanding of these empirical factors at the time, whereas we have to make rough guesses.
Another very plausible trajectory, Plan S, is essentially the opposite choice about how fast to scale: immediately issue a moratorium on further AI capability scaling. This policy looks better if there is a small risk of deal dissolution, and/or the returns to better AIs for safety research are low, and/or the risks of any capability scaling are high.
Massive datacenter buildout
It is desirable to do a massive datacenter buildout in Plan A if (and only if!) the upside of scaling quickly, reducing covert project takeover risk, and having additional capacity to pay alignment taxes is bigger than the risk of verification failing or the additional risk incurred by the deal failing with a gigantic compute stockpile.
In a bit more detail, the main upsides of a massive datacenter buildout are that it allows AI developers to:
Scale quickly without leaking algorithms to covert projects. In the previous section, we discussed the conditions under which it is good to increase AI capabilities, and the tradeoffs associated with going faster and slower. Simplistically, to increase AI capabilities, you can either improve algorithms or scale compute. By leaning heavily on compute, we can avoid making algorithmic progress in the first place, which is the best way to ensure that said progress doesn’t leak.
Pay alignment taxes. Large compute-based alignment taxes might be very helpful for safety. For example, it might be that neuralese is much more performant than current architectures, but also has dramatically worse safety properties. If we have enough compute that we can just accept the performance penalty and continue with current architectures anyways, this might dramatically reduce risk. Without a massive compute buildup, we will by definition not be able to pursue architectures that are much less compute efficient.
The main downsides are:
Verification and regulation becomes more load bearing. Since the verified projects have even more compute, it is more important that this compute is used safely, and, in particular, not used to do an intelligence explosion as fast as possible. So it becomes even more important that the regulations that AI projects are subject to are sufficiently conservative that they preclude an intelligence explosion, and that the verification is sufficiently robust that companies or rogue AIs can’t undermine it, and illegally gain control of a huge amount of compute within the verified clusters.
If the deal breaks down with a gigantic compute cluster, and the clusters are not destroyed, the intelligence explosion might be extremely rapid. One naive estimate is that a 10x of compute leads to a 5.5x increase in the speed of the intelligence explosion, so a 10,000x increase in compute might increase the speed of the intelligence explosion by 900x. So if it would take a year by default, it might take less than a day with the Plan A hardware explosion. This is obviously extremely dangerous, hence why we put so much emphasis on mutually assured compute destruction.
Putting all of this together, my view on how to decide to what extent we should pursue a massive hardware buildout under Plan A is as follows:
If covert projects aren’t a concern (because, for example, it’s impossible to do a reasonably fast intelligence explosion on a 100,000 GPU cluster), then don’t scale hardware at all, or only go slowly.Probably covert projects are somewhat of a concern, but it’s very unclear, and depends centrally on when international coordination to slowdown starts—the later a deal happens, the higher the risk of defection because (1) we’re closer to superintelligence, and (2) the smaller the personnel footprint is because the AIs can do a larger fraction of the tasks involved in running a covert project.Probably covert projects are somewhat of a concern, but it’s very unclear, and depends centrally on when international coordination to slowdown starts—the later a deal happens, the higher the risk of defection because (1) we’re closer to superintelligence, and (2) the smaller the personnel footprint is because the AIs can do a larger fraction of the tasks involved in running a covert project.
If there are reliable mechanisms for destroying compute after the deal dissolution, and reliable verification mechanisms so that we can have high confidence that there will not be major rogue internal deployments (either by countries or AIs), and there is reliable verification for within-deal compute, then the downsides of a compute buildout are small. Therefore, it is beneficial to scale hardware ~as aggressively as possible.
Otherwise, it is going to be a complicated tradeoff, which needs to be decided by weighing the magnitudes of the various costs and benefits listed above.
Our specific verification plan
The verification plan used in the Plan A scenario can be summarized as follows:
Mutually declare the locations of all large datacenters. Check the declared counts against intelligence estimates.
Agree to pause training and experiments. Install inference-only verification devices on the declared clusters to verify this pause.
Race to set up verifiable (and totally transparent) training and experiment clusters; when complete, resume capability scaling within the constraints of safety and transparency.
We discuss this in more depth here.
This version of the verification plan assumes that at the beginning of the deal (1) we can verify inference, and (2) AI development verification is feasible within a relatively short research sprint (but is not immediately feasible). It is most likely that one or both of these assumptions will be wrong; these assumptions are my central guess of what could be done under Plan A assumptions of political will and competence, but reality will most likely involve less good preparations.
What if we don’t have verification measures ready in time?
Three high level buckets of verification approaches are:
Technical verification methods. These methods would allow ongoing AI usage or development, with both sides of a deal being confident that the other isn’t doing anything dangerous. An example of this is what we do in Plan A where we initially rely on inference-only, while sprinting to technical methods to verify training and experiments.
Rely on trust, i.e. accept that either side could defect, and hope that neither side is foolish enough to try, or that they’d be caught by HUMINT and so forth if they did.
Blunt verification techniques (e.g. compute destruction or mutual power down).
It is highly desirable that we have technical verification methods ready before the beginning of the intelligence explosion. If inference verification devices are ready, then access to AIs gets to continue at the beginning of the deal. If development verification is ready, then we don’t need to stop AI capabilities scaling or AI safety research that involves frontier models at the beginning of the deal, and can instead scale until it is sufficiently unsafe not to (following the considerations above).
2 and 3 will likely only be temporary regimes, and be paired with a sprint to get technical measures online. Therefore, we hope that eventually we’ll get to sufficiently good technical methods.
The main downside of 2 is that it runs a risk of defection. Therefore, it may also run into increased political difficulties, as both parties will be worried that the other will defect, and so will be less likely to sign in the first place. However, it’s not correct to entirely write off trust, because it may be that the risk of AI takeover becomes so high that neither side is incentivized to defect, because doing so would so likely result in existential catastrophe. If high AI takeover risk becomes common knowledge (for example, after we’ve caught early misaligned AIs in the act of sabotage), relying on trust may be the best option.
The main downside of 3 is that it involves a temporary period where large scale inference goes offline. However, this would be less bad than one might naively think; even an agreement to power off (or destroy) 90% of the compute, would only force a reversion to 10x cheaper models. Since cost of a given level of capability has historically decreased very rapidly (averaging ~40x/yr), having 10x less compute would roughly correspond with having to use SOTA models from 7.5 months ago.If algorithmic efficiency improves at ~40x per year, then the number of years of progress t lost by a 10x compute reduction is log_40(10) = .624 years = ~7.5 months.If algorithmic efficiency improves at ~40x per year, then the number of years of progress t lost by a 10x compute reduction is log_40(10) = .624 years = ~7.5 months. A downside of 3 is obvious; if defection is viable, it becomes more likely that both sides would choose to defect in the first place.
Note that this is one upside of Plan S over Plan A; Plan A requires (1), while Plan S only requires (3).
Transparency and Security tradeoffs
There is an inherent tradeoff between security and transparency—to secure something, you inherently need to prevent some group from being able to see it.
Plan A biases towards transparency, in large part because I am more worried about the risk of poor governance within the legal projects than the risk of covert projects defecting and successfully undermining the deal. Increasing transparency will on average improve the quality of the governance decisions made by the Consortium, whereas increasing security makes it harder for covert projects to succeed. Furthermore, even if the governance decisions are perfect, the legal projects may not be able to solve alignment in time themselves, and may need help from the wider world to e.g. notice some subtle flawed assumption in a safety case. Also, transparency is very helpful for verification.
That being said, we are currently nowhere near the pareto frontier between these two dimensions. Plan A proposes dramatically increasing both security and transparency into frontier AI projects. We discuss transparency-related tradeoffs in much more detail in our Transparency Supplement.
Our Post-ASI governance plan
Related reading: Space governance supplement, AI 2040 epilogue.
We make a concrete guess as to what a positive post-ASI governance plan might look like. However, we expect that there exist better proposals, which we haven’t thought of yet, that will be discoverable at the time. Therefore, we want readers to orient to our plan as a baseline with which to compare other plans to on dimensions that they care about. Then, we hope that humanity will spend vastly more effort—both human and AI cognitive labor—on designing a system that performs better.
Broad Deployment of AI during the takeoff period
The Plan A scenario involves broadly deploying AI into the economy during takeoff. This has massive effects on the world during takeoff: almost everyone loses their jobs and starts relying on Citizen’s Dividend. Overall, I think that it is better for massive AI impacts to happen gradually, e.g., over the course of a 10 year takeoff, than to have all of the impacts happen really quickly, right at the end, once AIs are vastly smarter than humans.
(See more about how AI impacts the broader economy in our economics model, and discussion in our economics supplement)
However, this relies on a few key questions:
Is the net effect on societal epistemics positive or negative? I currently feel highly unsure. On one hand, broad deployment of human level AI will increase access to information, like the internet. On the other hand, while the internet is definitely helpful for some people, it’s unclear if the average effect on societal epistemics is good or bad. More research is needed on this; what plan we actually recommend at the time would depend a lot on our best guess about the epistemic effects of broad AI deployment. With our current state of knowledge, I currently don’t feel like this is a major consideration in either direction.
Does broad deployment meaningfully increase the difficulty of AI control? I think probably not, because the main risks of AI takeover route through rogue internal deployments and recursive self improvement. The case for broad deployment increasing risks is that it may lead to AIs having affordances that might be helpful for takeover, especially (i) a large robot workforce and manufacturing base, (ii) control over robot armies, (iii) access to biology labs, and (iv) regular contact with a broad range of humans. I think that these risks are likely mitigatable through a combination of monitoring AI behaviour, limiting access to the most dangerous physical infrastructure (e.g. biolabs), preventing any one AI system from having control of a large fraction of the physical resources, and preventing collusion between different AI systems. However, I do think that it’s plausible that broad deployment makes control much harder. Hopefully this can be handled via limiting AI deployment in a particular domain of concern, but it may in practice limit deployment across many domains.
How high are misuse risks from broadly deployed closed-source models? My inclination is that these risks are reasonably low, because (1) I’m optimistic about adversarial robustness in a regime like Plan A, where you are slowing down substantially in order to improve safety, and you have the capacity to run lots of monitoring, and (2) even if it’s possible to jailbreak widely deployed AI models, I think that the defensive measures put into place will be sufficient for preventing existential catastrophes from misuse risk. However, it is plausible that both of these assumptions fail (either because of intrinsically high difficulty or incompetent execution), in which case, we would be incurring serious misuse risk from broadly deployed models.
If we’re not broadly deploying AI through society, the main alternative is to be developing AI inside closed projects. This incurs additional risks. Whatever closed off group of people that are managing the intelligence explosion without contact with the rest of the world will have a huge amount of power over the future. For an extreme case, the closed project might literally be developing AI in secret, without notifying the rest of the world, as described in Training AGI in Secret would be Unsafe and Unethical. A secret intelligence explosion would be bad for misalignment risk, because far fewer people would know and understand what’s going on, and be able to give input on the alignment techniques, and would be bad for the distribution of power, because powerseeking individuals might be able to seize control of the AGI project, and build ASI aligned with their values (as opposed to humanity more broadly)
A better proposal for avoiding broad deployment is to do local deployment, but maximize transparency: publish evaluation results, publish demos of AIs doing very impressive things, publish the AI-generated research (on non-AI domains), publish the model spec, etc. This is far better than a secret intelligence explosion. However, it seems likely that if this is happening, humanity will be still largely in the dark about what’s going on, because to normal people, without AI impacting their day to day life, it’ll seem like some far-away science project that people like to talk about—it’ll be in far mode, not near mode. Therefore, it seems likely that broad deployment of AI will improve epistemics around AI in particular, which seems positive.
Appendix: How plausible and how desirable are our key policies?This appendix was written by Miles Kodama.This appendix was written by Miles Kodama.
Plan A is neither a pure recommendation nor a pure prediction. A pure recommendation in which every actor makes maximally responsible moves starting tomorrow would have been too implausible. And we already wrote a pure prediction scenario—that’s AI 2027. So we chose to make Plan A a compromise between plausibility and desirability.
The interactive graph below shows roughly where each of the scenario’s policies sits on these two dimensions. Plausibility is judged at the point in the scenario when the policy is adopted, conditional on everything that has happened up to then. It does not represent how likely the policy is to happen starting from today’s world. For instance, “Pause at top human level” is rated highly plausible not because we think a pause is highly likely from where we stand in 2026, but because in a world where the deal has held for years and control-based safety cases are visibly reaching their limits, we expect decisionmakers to choose something like it. Desirability measures how much better we currently expect a policy to be than the most plausible alternative at that same point in the scenario. It is not a ranking of which policies we’re most keen to see adopted today. A low score can mean that our recommendation is subject to lots of unpredictable empirical updates or that alternative policies look about as good to us as our recommendation.
Click on a policy to read more.
Space Governance Plan
Preliminaries
It should be possible for almost all humans to get almost everything that they want from a post-ASI world. With huge numbers of robots capable of almost any task a human can do, it will be possible to provide everyone with the living standard that only the richest people today enjoy. People could spend their time doing what’s most meaningful to them—raising children, creating art, or spending time with friends. With ASI-driven advances in medicine, we’ll be able to treat mental illness, chronic pain, aging, cancer, and other diseases that shorten or worsen people’s lives today. In short, if we achieve the future described in Plan A, there will be enough to go around to give everyone a very good life. Many of the tough tradeoffs that are the focus of modern politics will be resolved. So in some sense, these decisions need not be extremely fraught.
Yet, at the same time, the decisions we make are extremely high-stakes. If we handle things well, then we could create an enormously flourishing future for our descendants. But there are many ways we could waste this opportunity. We could fight over the resources, leaving a few people in control of the future and everyone else with nothing, or burning vast resources in internecine conflict. Even if we avoid those pitfalls, we could still misuse resources—squandering them on positional goods or failing to leverage ASI to figure out how best to achieve what we truly care about.
We want a governance system that will allow us to make wise decisions from a foundation of abundance and cooperation.
This document focuses on how to decide what to do with resources outside the solar system (henceforth referred to as “space resources”). This is an important aspect of post-ASI governance—the overwhelming majority of accessible resources are outside of the solar system and the overwhelming majority of future beings may live outside our solar system—but it is not the only important question.Other important questions in post-ASI governance that we don’t address, or address only cursorily: Post-ASI governance on Earth or elsewhere in the solar system. Meme wars, if ASI are able to construct extremely compelling arguments, and members of different ideologies try to spread their views using these arguments. Rights for digital minds. Managing risk in highly offense-dominant worlds, where a small group of people are able to cause huge amounts of destruction. Post-ASI governance on Earth or elsewhere in the solar system. Meme wars, if ASI are able to construct extremely compelling arguments, and members of different ideologies try to spread their views using these arguments. Rights for digital minds. Managing risk in highly offense-dominant worlds, where a small group of people are able to cause huge amounts of destruction.
Other important questions in post-ASI governance that we don’t address, or address only cursorily:
Epistemic status
We’re currently not confident in how the treaty signatories in Plan A should approach post-ASI governance: it depends on a lot of details about people’s preferences, technological capabilities, and optimal coordination mechanisms that are only partly worked out.
Luckily, by the time the treaty signatories are making these decisions, they should have superintelligent automated researchers investigating the nature of the universe, doing moral philosophy, and figuring out optimal coordination mechanisms. And we’ll have AI therapists, advisors, and negotiators to help us understand our preferences and make deals with other actors.
So we think the actual best plan is to defer many of these questions until we have ASI-driven advice and coordination. However, there is value in concreteness and in sharing best guesses, so we wanted to provide some analysis. But we wouldn’t want readers to anchor too much on the ideas we present below.
What’s special about post-ASI governance?
Post-ASI governance will differ from historical precedent in several important ways.
Technological progress may quickly unlock new pools of resources that no one has previously exploited. The most important example is resources outside of our solar system.
Because these resources have been unused, there are no historical claims to navigate.The Outer Space Treaty forbids nations and (according to some interpretations) individuals from claiming celestial bodies. In this document, we’re assuming, as in the Plan A scenario, that the Outer Space Treaty will be renegotiated once extrasolar colonization becomes technologically feasible.The Outer Space Treaty forbids nations and (according to some interpretations) individuals from claiming celestial bodies. In this document, we’re assuming, as in the Plan A scenario, that the Outer Space Treaty will be renegotiated once extrasolar colonization becomes technologically feasible. We should aim to govern them in a just and equitable way.
There will likely be extreme abundance relative to the present day. With just a tiny fraction of available resources, we’ll probably be able to give everyone alive today a long, happy, healthy, and extremely prosperous life by current standards.
We think that the most important question that governance will have to face is what to do with the huge quantities of remaining resources after we’ve ensured that everyone alive has enough to live an incredible life. Some salient options include:
Personal consumption. People could develop ever-more-expensive ways to derive enjoyment from resources. But most resources are probably poorly suited to use for personal consumption. Most resources are outside our galaxy and thus millions of light-years away. If you want to stay on Earth and consume anything—matter, energy, information—derived from another galaxy, you’d have to wait millions of years. If, on the other hand, you left Earth for another galaxy, then if you ever returned, you’d find that millions of years had passed.
Positional goods. People could spend resources on a zero-sum competition to have the most impressive galactic-scale art projects, or the largest number of statues of themselves. Positional goods could soak up a lot of resources. While there are probably generally diminishing returns to consuming more resources as people get richer, this might not hold for positional goods, since status is relative. We think it would probably be a mistake to spend huge quantities of resources on this type of zero-sum competition.
Satisfying moral preferences. We could create societies full of flourishing people, or enormous nature reserves. We think this is the most promising way to use the cosmic endowment. But then the question becomes: how should we decide which moral preferences to pursue?
Technological progress will also change how we make governance decisions:
We may be able to make decisions with lasting effects far into the future.
For example, we could hand over resources or hard power to ASI systems that have been aligned to a particular set of values. This will likely give us the option of creating extremely long-lasting institutions.
Typically later generations adopt somewhat different values than their parents, and as a result society’s values change over time. But new technologies may enable parents to exercise tighter control over the values that their children develop.
People will have access to superhuman advice and coordination assistance.
This means people will be better able to identify and advocate for policies that truly achieve what they care about.
But we’ll probably also need to make our governance structure more robust than ever to people trying to game it, either individually or in a group. Just because a governance structure has not been severely exploited in the past does not guarantee that it won’t be severely exploited in the future.
And we will have large responsibilities to make these decisions go well for people who aren’t able to participate:
The overwhelming majority of the affected parties probably won’t be able to participate in the decision to govern space directly, because they won’t have been born yet. This is already true to some extent—in modern democracies, citizens set policies for non-citizens, adults decide for children, humans decide for animals—but in the post-ASI setting this becomes far more extreme. Present people may make decisions that affect the lives of vast numbers of peopleToby Newberry estimates that the affectable universe could support 10^35 biological humans or 10^44 digital minds at a time. For comparison, the number of humans who have ever lived is 10^11.Toby Newberry estimates that the affectable universe could support 10^35 biological humans or 10^44 digital minds at a time. For comparison, the number of humans who have ever lived is 10^11. who cannot give input. The responsibility to get these decisions right is enormous.
Evaluating proposals
In this appendix, we’re evaluating proposals based on how well they achieve the following desiderata, which we believe will be broadly appealing across different value systems.
Preference satisfaction. The governance system should try to satisfy all the participants’ preferences as much as possible. Some aspects of this include:
Paretopianism. Given a massive windfall, it should be possible to make nearly everyoneOne possible exception is people with strong preferences for positional goods, e.g., people who care a lot about being richer than others, rather than their absolute level of wealth.One possible exception is people with strong preferences for positional goods, e.g., people who care a lot about being richer than others, rather than their absolute level of wealth. vastly better off than they are today, and a good proposal should do exactly that.
Limit extreme downside. Our post-ASI governance proposal should have a very low chance of producing outcomes widely considered extremely bad, such as huge amounts of suffering, or all of space being destroyed.For example, it might be possible to move our universe from a false vacuum state to a true vacuum state, which could destroy all matter.For example, it might be possible to move our universe from a false vacuum state to a true vacuum state, which could destroy all matter.
Satisfy bounded views. If someone has a preference or moral view that is not very costly to satisfy, we should satisfy it. For example, on a cosmic scale, it’s fairly cheap to ensure that every human alive in 2040 has a long, flourishing life.
Satisfy common-sense ethics. The process should not violate common-sense ethical views. This weighs against proposals where a small number of people predictably end up controlling the fate of the universe or where future generations end up totally disempowered.
Satisfy objective morality. Some aspects of ethics may be objective and discoverable. If so, we should try to identify and conform to them.
These next criteria aren’t terminal goals, but we think that processes that satisfy them are more likely to achieve the criteria above.
Coordination. The governance system should enable coordination on large-scale collective action problems. For example, if we encounter alien civilizations, we may need to coordinate to establish friendly relations with them. We may also need to coordinate to avoid actions that risk destroying all of Earth-originating civilization, or to avoid space races where factions compete to claim as many resources as possible, burning huge amounts of resources in the process.See the paper Burning the Cosmic Commons for more on this threat model.See the paper Burning the Cosmic Commons for more on this threat model. These goals probably require some minimal government functions: protecting property rights, enforcing prohibitions, and potentially collecting taxes to fund public goods.
Moral trade. It might be possible to nearly satisfy many different moral viewpoints simultaneously. For example, some views might especially care about resources close to Earth or the near future, while others might be more indifferent across time and space: we might satisfy both by letting the former choose what happens in the Milky Way while the latter choose what happens in the rest of the universe. Here are two types of trade that we want our proposal to facilitate.
Bilateral trade. This is particularly well suited for satisfying uncommon preferences. These trades are most straightforward in market systems, where two people can literally trade deeds to resources or sign agreements about how they will use their resources. But it’s possible that some voting systems–e.g., the Vickrey-Clarke-Groves mechanism—could also achieve similar results.
Coordination on moral public goods. People might share some values while also holding idiosyncratic preferences—individually preferring to pursue their own preferences, but collectively preferring that everyone work toward the shared values.Some examples of moral public goods that might be quite important to fund:
Broadly appealing moral projects, such as building societies of flourishing biological humans, uploaded humans, or other digital minds; providing resources to future generations; or discovering and acting in accordance with objective moral truths, if any exist.
Trading with people who intend to do something with side effects that others consider harmful to get them to modify their plans. Obviously, it’s important to approach this carefully to avoid creating perverse incentives where people do some activity that’s broadly considered harmful in search of a payout (see below for more discussion of reducing downside risk).
Broadly appealing moral projects, such as building societies of flourishing biological humans, uploaded humans, or other digital minds; providing resources to future generations; or discovering and acting in accordance with objective moral truths, if any exist.
Trading with people who intend to do something with side effects that others consider harmful to get them to modify their plans. Obviously, it’s important to approach this carefully to avoid creating perverse incentives where people do some activity that’s broadly considered harmful in search of a payout (see below for more discussion of reducing downside risk).
Coordination on moral public goods is most straightforward in voting systems, since people can just vote to devote resources to shared projects. It’s possible that voluntary contracts—where people pledge to spend resources on shared goods if others do the same—could achieve similar results, but this is less straightforward because of free-riding incentives.
It might be particularly effective to coordinate to switch from funding positional goods to funding moral public goods. If done in a way that leaves people’s relative ability to purchase positional goods unchanged, then this is great for everyone: relative status doesn’t change, but people are able to satisfy their non-positional preferences, too.
Reflection. ASI will likely uncover novel considerations and empirical facts highly relevant to how people want to use space resources (e.g., are there objective moral truths, and what are they?). We want each person to have the time and resources necessary to think carefully about what they value before making important decisions, and we want to ensure that we invest in ASI-fuelled research on important questions to inform people’s decisions.
Pluralism. Humans today have many different visions of the good. It’s possible that after deep reflection with ASI we’ll converge on a single set of shared values. More likely, we’ll still have some disagreements. Where people have different value systems, we prefer proposals that will allocate resources to pursue all of them.
Avoid brittle decision-making structures. There are likely many important considerations that we’re not yet aware of. Where possible, we favor flexible decision-making structures that won’t make it difficult to solve unforeseen problems.
Fairness. Each living adult should have an opportunity to give equal input into the process.Ideally, we’d also like to weigh the preferences of other stakeholders like future generations and digital minds, although we’re uncertain about whether it will be feasible to do this without opening up the political system to exploitation (see below).Ideally, we’d also like to weigh the preferences of other stakeholders like future generations and digital minds, although we’re uncertain about whether it will be feasible to do this without opening up the political system to exploitation (see below). We don’t think that this criterion necessarily means that each person ends up controlling an equal quantity of resources. For one thing, some people might care a lot about what happens with resources closer to Earth while not caring much about resources in distant galaxies, and we’re happy for them to trade with people who care a lot about distant galaxies. Nor do we think that this criterion requires that everyone is equally happy with the outcome. For instance, if some decision is made by vote, then some people might end up voting for the losing side. But we want a neutral process that doesn’t privilege any particular people or viewpoints over others.
Political feasibility. Unlike the other criteria, this is purely pragmatic: a proposal must be acceptable to the relevant decision-makers to be implemented. It’s unclear who these decision-makers would be, but they could include the citizens and/or leadership of the US, China, or other powerful countries. It might also include AI company CEOs or shareholders.
Summary of the most promising options for allocating space resources
We think that it’s best to decide on a specific structure for allocating space resources in 2040 at the time that the treaty is decided (more discussion here).That is, we should either decide how future governance decisions will be made or at least pick a meta-level process for choosing how that decision will be made (e.g., the voting system should be chosen by an AI arbiter verifiably aligned to choose based on a particular set of principles)That is, we should either decide how future governance decisions will be made or at least pick a meta-level process for choosing how that decision will be made (e.g., the voting system should be chosen by an AI arbiter verifiably aligned to choose based on a particular set of principles)
We’re currently fairly uncertain about what the best structure is, and that decision should be deferred until decision-makers have access to aligned ASI advice. But currently, these seem like the best object-level proposals.
Direct endowment + trade (with prohibitions). Each adult is assigned an equal share of resources. They can trade, donate, give away, or use those resources in any way they want, except for a small category of prohibited activities (more).
Most views held by living humans will have at least some resources devoted to them.
Bilateral trade between people with different moral views is very straightforward.
It’s harder to coordinate to fund moral public goods due to free-rider problems.
Vote on how to use resources. Each adult gets to participate in a vote on how to use space resources (more).
It’s straightforward to coordinate to fund moral public goods.
All known voting systems have serious issues, such as:
Inefficiencies, e.g., strong minority preferences are overruled by very weak majority preferences.
Incentives to strategically misrepresent preferences, either alone or in collusion with others.
Direct preference aggregation. We use novel technology to directly elicit people’s preferences (e.g., through an interview with an ASI using lie detection) about the use of space resources and aggregate these preferences (more).
Many voting systems fail to get a globally-optimal outcome because people strategically misrepresent their preferences. But direct endowment fails to fund moral public goods. By eliciting people’s honest preferences and then aggregating them, we might be able to get the best of both worlds.
People might object to the preference elicitation step, especially if the method is time-consuming or invasive.
Basic proposal | Key advantages | Key disadvantages |
Direct endowment + trade (with prohibitions). Each adult is assigned an equal share of resources. They can trade, donate, give away, or use those resources in any way they want, except for a small category of prohibited activities (more). | Most views held by living humans will have at least some resources devoted to them. Bilateral trade between people with different moral views is very straightforward. | It’s harder to coordinate to fund moral public goods due to free-rider problems. |
Vote on how to use resources. Each adult gets to participate in a vote on how to use space resources (more). | It’s straightforward to coordinate to fund moral public goods. | All known voting systems have serious issues, such as:
|
Direct preference aggregation. We use novel technology to directly elicit people’s preferences (e.g., through an interview with an ASI using lie detection) about the use of space resources and aggregate these preferences (more). | Many voting systems fail to get a globally-optimal outcome because people strategically misrepresent their preferences. But direct endowment fails to fund moral public goods. By eliciting people’s honest preferences and then aggregating them, we might be able to get the best of both worlds. | People might object to the preference elicitation step, especially if the method is time-consuming or invasive. |
(We discuss alternative proposals that we like less here).
Other important decision points include:
At what time should votes, preference elicitation, or distribution of resources happen? (discussion here)
Should later generations get a say? (discussion here)
Should we prefer a hybrid approach that allocates some resources according to one proposal and allocates other resources according to other proposals? (discussion here)
How do we enable each decision-maker to make the best decisions by their own lights? (discussion here)
Some available options
Here, we discuss options for governing future resources. We begin by describing a couple of notable options which seem relatively unpromising on the criteria above, then turn to proposals that we find more promising. In both cases, we’re trying to cover important and illustrative cases, not every possible outcome.
Unpromising options
Anarchy
No property rights system is established for space resources. To use a resource, an actor must either be the first to claim it or seize it from another actor. Actors can do whatever they want with resources they currently control—there’s no governance system for prohibiting particularly bad behaviors or any other forms of large-scale coordination (e.g., to fund public goods).
Drawbacks: there could be significant negative-sum conflict over resources, and insecure property claims would make trade—including moral trade—harder.
Anarchy is worse than the feasible alternatives. We do expect that it is not a stable long-term equilibrium. Given much more advanced tools for coordination and communication, we expect people to create ad hoc governance structures.Via mutual defence coalitions, mutual insurance, or law-enforcing equilibria emerging from multi-turn interactions.Via mutual defence coalitions, mutual insurance, or law-enforcing equilibria emerging from multi-turn interactions. But a government that emerges organically from anarchy could be worse than one designed intentionally. In particular, it might favor the interests of those who were able to claim and defend lots of space resources early on.
Finders-keepers property rights
Actors can claim unclaimed resources or they can buy resources from earlier owners, but they can’t seize claimed resources without the owner’s permission (similar to homesteading). As before, no effective system exists for large-scale coordination or public goods provision, and no governance system can enforce prohibitions.
Strengths: we think this is an improvement on anarchy. Once a claim is recognised, the holder no longer needs to spend resources defending it. Moreover, secure and widely recognised property claims make trade possible.
Drawbacks: value is no longer wasted contesting claimed resources but is still wasted in the race to claim them.Barzel (1997), Economic Analysis of Property Rights; Anderson & Hill (1990), “The Race for Property Rights”Barzel (1997), Economic Analysis of Property Rights; Anderson & Hill (1990), “The Race for Property Rights” It also seems likely that a small number of powerful actors—perhaps whoever can send out a wave of probes first—will claim most resources. And this proposal has no mechanism for public goods provision or enforcement of prohibitions, so it forgoes potentially large gains from moral public goods and from preventing egregiously unpopular uses of resources.
Next, we discuss governance proposals that we think improve on these scenarios according to the criteria above.
More promising options
Direct endowment and trade
Each person receives title to a portion of all future resources, which gives them the right to decide how those resources will be used in the future. For example, people can sell their shares of future resources, gift them to other people or to trusts, or make bilateral deals with each other about how they’ll use their resources. Property rights are well-enforced in practice.One way to do this might be to require that all probes sent to colonize other star systems carry a nightwatchman ASI to prevent the colony from sending out probes to seize space resources that belong to others.One way to do this might be to require that all probes sent to colonize other star systems carry a nightwatchman ASI to prevent the colony from sending out probes to seize space resources that belong to others. As before, no effective system exists for large-scale coordination or public goods provision, and no governance system can enforce prohibitions.
(A related approach: auction titles to all space resources and then redistribute the proceeds equally. In theory, if we make some standard assumptions, this scheme would result in the same distribution of ownership as direct endowment.The point is that each person’s eventual wealth only depends on the value of whatever they inherit at equilibrium prices, and not on whether they originally inherited titles to future resources directly, or received a cash transfer of equal value at equilibrium prices. If resources R are divided equally among N people and traded at competitive prices p, each person’s wealth is p · R/N. Under the auction alternative, total revenue is p · R, and equal redistribution gives each person p · R/N. This requires: competitive (price-taking) behaviour, that the auction clears at competitive equilibrium prices, markets are complete and have sufficient numbers of buyers and sellers, and people have convex preferences. The equivalence between auctioning and direct endowment could break down: especially if actors are large enough to coordinate to influence prices, or if resources are indivisible.The point is that each person’s eventual wealth only depends on the value of whatever they inherit at equilibrium prices, and not on whether they originally inherited titles to future resources directly, or received a cash transfer of equal value at equilibrium prices. If resources R are divided equally among N people and traded at competitive prices p, each person’s wealth is p · R/N. Under the auction alternative, total revenue is p · R, and equal redistribution gives each person p · R/N. This requires: competitive (price-taking) behaviour, that the auction clears at competitive equilibrium prices, markets are complete and have sufficient numbers of buyers and sellers, and people have convex preferences. The equivalence between auctioning and direct endowment could break down: especially if actors are large enough to coordinate to influence prices, or if resources are indivisible.)
Strengths: when property rights are clear and well-established from the outset, bilateral trade—including moral trade—becomes easier. Compared to finders-keepers, there are no longer incentives to waste resources through racing.
Since each property-holder is free to decide what to do with their share of the future, at least some resources will be devoted to pursuing most moral views, making it likely that cheap-to-satisfy views will be satisfied.
Drawbacks: moral public goods would have to be funded through voluntary contracts, and it’s not clear whether voluntary contracts will be sufficient to ensure that moral public goods are funded at the optimal level, since individuals will face a strong incentive to free-ride.
Since direct endowment doesn’t involve a way to impose universal prohibitions on how to use future resources, the same drawbacks carry over from anarchy and finders-keepers: if there are some uses of resources which are egregiously bad by most people’s lights, but preferred by a few actors, it may be hard to prevent those uses, and bad actors could even threaten those uses to extract concessions. This is the mirror image of the moral public goods drawback: there is no easy way to coordinate to prevent moral public bads.
A fuzzier concern: we think it’s probably better for the current generation to think of themselves as stewards who seek to create the best future by their own lights in the space that they control, rather than as “space emperors” who try to squeeze the maximum personal enjoyment or aggrandizement out of their share of resources. If people get unilateral control over space resources and they’re allowed to buy and sell them like normal property, we worry that this might push people to think of the resources as a personal windfall rather than a serious moral responsibility.
Uncertainties and decision points: We’d also need to make a call on decisions like:
Who gets resources? (For example, should resources be reserved for future generations?)
Should people be allowed to sell, transfer, or wager their resources immediately? Or do we want to—somewhat paternalistically—enforce a waiting period for people to carefully consider their options?
What is the relationship of a galaxy’s “owner” to the future inhabitants of the galaxy?
One important variant on this is to add some mechanisms for preventing people from using resources in egregiously bad ways. In the section below, we discuss some possibilities.
Voting on collective uses of resources
Everyone votes on proposals for what to do with space resources. These proposals could range from very concrete (e.g., “create digital minds according to this specification”), to very meta (e.g., “delay doing anything with space resources right now, and re-assess with another vote X years from now” or “allocate some future resources to private ownership”), or something in between.
Strengths. Compared to the other proposals so far, it’s easier to fund moral public goods—if people prefer resources to go to a shared project rather than divided among all voters, they can vote for that proposal.
It’s also a more flexible plan: once future resources are transferred to private ownership, recovering the ability to make collective decisions might be impossible even if almost everyone wanted it. With some kind of regular vote, the electorate can switch to a different plan at the time of the vote, including agreeing to delay doing anything with space resources.It’s true that in principle, under a direct endowment proposal, everyone could decide to voluntarily pool their galaxies into a trust and allocate resources by vote, but free-rider problems make this less likely in practice.It’s true that in principle, under a direct endowment proposal, everyone could decide to voluntarily pool their galaxies into a trust and allocate resources by vote, but free-rider problems make this less likely in practice.
Drawbacks: All known voting systems have significant problems. Some are simply inefficient: for instance, majority rule faces the familiar problem that the majority view can succeed when the majority wouldn’t deem the minority’s preferred option much worse, but the minority greatly dislikes the majority’s preferred option. Others incentivize people to misreport their preferences, individually or through collusion, to benefit themselves at the expense of overall value. This problem is especially serious when people know the details of the voting system well in advance and have ample time to develop strategies to game it with ASI assistance.
A further drawback is that voting requires synchronized decisions: everyone must vote at the same time. If the vote happens in 2100, you vote according to your preferences at that time, even if you would have preferred more time to reflect first. In principle, people could vote to delay the vote. But you could still face situations where a majority wants to vote now while a minority wants to keep reflecting, and the minority gets overruled. Contrast this with direct endowment, where each person can reflect for as long as they want before deciding what to do with their resources.
Uncertainties and decision points: How badly can strategic voting distort outcomes in practice, and can novel technology (e.g., lie-detection technology) mitigate this? How well does voting handle bilateral trades between people with rare preferences? And we’d need to decide: when do we hold the vote, who is included, and which voting system do we use?
Direct preference aggregation
Similar to voting, but instead of people casting votes, AI systems try to directly ascertain people’s preferences. For example, AI systems might build a profile of a person’s preferences from all their writing, extensive interviews perhaps under lie detection, and/or brain scans. Then, everyone’s preferences are aggregated to decide what to do with space resources. As with voting, the decision might be quite meta—e.g., wait 100 years and then do preference aggregation again.
It’s not clear how best to aggregate preference, but here’s a sketch of a proposal. Everyone is first assigned 1/N of future resources as a baseline—their fallback if no gains from trade are found. An AI arbiter then identifies reassignments of resources that would improve everyone’s preference satisfaction above this baseline, splitting the gains fairly among all beneficiaries. This handles bilateral moral trade: if person A most wants statues of herself and person B most wants statues of himself, but both value statues of George Washington at 60% as much, the arbiter directs their combined resources toward Washington statues, leaving both better off than their baseline. It can also handle moral public goods: if funding some public good would improve everyone’s welfare, the arbiter taxes each person in proportion to how much they value that good (a Lindahl tax).
Strengths. Direct preference aggregation is a modification of the voting proposal above and inherits its strengths. Its major additional advantage is that it makes strategic manipulation more difficult than in voting systems.
Drawbacks. Eliciting people’s honest preferences would require some novel technology and it’s not clear what this process would look like. It’s possible that this process would be invasive, time-consuming, or unpleasant for participants (e.g., an interview while hooked up to a lie detector). It seems unjust and contrary to common-sense morality to require that people go through a procedure that they strongly object to in order to have input into the allocation of future resources.
There are a few more technical issues. For instance, people could strategically modify their preferences before assessment, breaking the main reason for trying to directly elicit people’s preferences (although mitigations might be possible—e.g., governments could ban people from modifying their preferences using novel technologies or require that they register their pre-modification preferences). We also haven’t spelled out a specific algorithm for doing the preference aggregation, and it could be difficult to avoid perverse or hard-to-predict outcomes, especially if many people are best represented as having inconsistent or incoherent preferences. Indeed, the question of how to represent someone’s preferences, including on questions they haven’t explicitly considered, seems very contentious and underdetermined.
Defer to superintelligence
Instead of any particular method of inferring and aggregating people’s preferences, humanity builds aligned superintelligence to plan on its behalf and give it control of nearly all space resources.
For example, we could start with a “council” of superintelligent delegates, each starting from one of many diverse moral and philosophical traditions. Those delegates then engage in deep reflection and deliberation with one another. The delegates can initially be trained to be reliable and accurate on all verifiable domains, like forecasting and mathematics, and not post-trained to strongly avoid novel, ‘weird’, or otherwise unpopular views just because they are currently unpopular. If the superintelligences converged onto a single moral view, then they could devote space resources to that plan. If the superintelligences continued to have different views, they could come to a compromise plan via voting, direct endowment, or direct preference aggregation.
Or we could design an ASI system to be motivated to discover and act on objective moral truths, if they exist. (If they don’t exist, then it would revert to a backup plan—perhaps one of the proposals above).
Strengths. Determining the best plan under a particular moral view probably requires discovering and grappling with lots of novel and counterintuitive considerations. The resulting plan might seem very weird and alien to humans, and they might not be motivated to act on it, even if they’re assured by a superintelligence that it is the morally best course of action. People might also have competing personal interests that take precedence over doing the right thing. But an ASI could potentially be designed to have no competing personal interests. These considerations are particularly strong if moral realism is true, and a plan like this might be our best shot at devoting substantial resources to pursuing objective morality if it exists.
Drawbacks. The ultimate outcome of this process might depend a lot on the details of the process—e.g., what initial viewpoints were represented on the “council” of deliberating ASIs? If this is true, who should decide? It seems rather unfair for the treaty signatories to get to choose.
We think that a process like this might be the best ultimate use of space resources. But if so it should probably result from one of the proposals above. For example, if many people found moral realism compelling after reflection, they could vote to spend public resources on the most morally valuable uses (or express a preference in direct preference aggregation). (It’s unclear whether the direct endowment proposal is sufficient for this: handing off resources to morally-aligned ASI is probably a moral public good in this scenario, so we’d expect it to be underfunded if people were individually deciding how to spend their resources).
Other decision points
Should we commit to a process right away?
The treaty signatories could approach this decision in three ways:
Decide on a specific process in 2040.
E.g., schedule an election in 2070 using quadratic voting where all living people over the age of 16 vote on what to do with all extrasolar resources.
Decide on a process in 2040 for picking a process later.
E.g., schedule an election in 2070 using approval voting, where all living people over the age of 16 vote on what decision process to use.
E.g., design an ASI arbiter who will do a bunch of research and then pick a process that best satisfies some abstract desiderata put forward by the signatories of the Treaty in 2040.
Make no explicit decision about a process or a process for picking a process.
It’s not clear we’d end up deciding what to do with space, under this option. Perhaps at some point one of the signatories to the Treaty would decide that they wanted to colonize space and either do so unilaterally or negotiate with the other signatories to be allowed to do so.
Option 3 is probably worse than options 1 or 2. Having a concrete proposal—even a meta-level one about how to select an object-level process—allows public scrutiny, which we expect would push the proposal to be fairer and more principled. Without a concrete proposal, two failure modes become more likely. First, parties who especially wanted space resources might start lobbying the Treaty governments, and those who are well-connected or have invested heavily in lobbying would end up with outsized influence over how space resources are used. Second, some parties might start unilaterally colonizing space, which could trigger a space race. Governments could ban unilateral colonization, but that would require a plan for lifting the ban—at what point are people actually allowed to colonize?—which brings us back to option 2.
We’re less confident about whether option 1 or option 2 is better.
The main argument for option 2 over option 1 is that we’ll probably get more information about how to design a good process after 2040.
We might discover new voting or preference aggregation systems that avoid issues with systems that we know about. However, we’re relatively pessimistic about learning much on this front after 2040, given existing impossibility results showing that voting systems can’t simultaneously satisfy a set of desirable features. Also, designing new preference aggregation systems is a relatively well-scoped, tractable mathematical problem so we would be somewhat surprised if we didn’t have the answer in 2040 after devoting a small amount of ASI labor to it.
We might learn how well particular processes are likely to work in practice.
For instance, we might learn about strategies that people are actually using to try to game the current system and adapt the rules accordingly, e.g., if we have rules intended to prevent parents from controlling their children’s values to an excessive degree, and it turns out that some parents can find loopholes to exert a lot of control, we could adapt the rules in response).
We might discover new tools or capabilities that enable a better process.
For instance, we might develop new methods for noninvasively scanning people’s brains to determine their preferences.
We might discover new considerations that affect the optimal preference aggregation process.
For instance, we might identify valuable forms of moral trade that we want our preference aggregation process to accommodate.
On the other hand, people will also accumulate information about the likely object-level consequences of each possible process. For example, suppose that we decided to vote in 2070 to decide on a procedure for deciding what to do with space, choosing between the policies of “hold a majority vote and do whatever the winning proposal says” and “direct endowment.” And at this time, it turns out that 50.1% of people favor colonizing the Milky Way while leaving the rest of space as a permanent nature preserve, while 49.9% have plans to colonize more distant galaxies.
In 2070, it would be clear that if we chose the majority vote, then all space resources would be allocated to the majority’s colonize-the-Milky-Way-and-nothing-else plan, but if we went with the direct endowment plan, then some galaxies would be left as nature preserves, while others would be colonized by the minority. Members of the colonize-the-Milky-Way-and-nothing-else coalition would then be strongly tempted to favor the majority vote when choosing a process.
At best, the process-revision vote becomes redundant—we might as well have voted directly on the object-level outcome. At worst, we lose the opportunity for people to agree on a process grounded in principles like fairness or Paretopianism.
One possible way around this would be to delegate process revisions to an AI arbiter that adjusts the process based on principles established in 2040, without regard for which factions benefit from any given change. This approach relies heavily on being able to verifiably align an AI system to a set of principles. It also requires being able to identify principles in 2040 that are specific enough for the AI arbiter to apply without heavy extrapolation, but not so specific that we forfeit much of the upside of deferring process decisions to a later date. But if it’s possible, this might be a good way to incorporate new information in a principled way.
Early decisions vs. late decisions
Once we’ve picked an object-level process—like a vote, a preference aggregation procedure, or directly distributing deeds to space resources—there’s a question of when to run that process.
Here are some advantages of delaying the process:
People have more time to reflect on what they care about and how they endorse using space resources before it’s time to make a decision. This matters most for voting and direct preference aggregation, where everyone must decide simultaneously.
People may have access to more information at the time of the decision—e.g., a clearer sense of what is technically feasible to do with space resources.
It’s possible to incorporate later generations into the process—e.g., allow them a vote, take their preferences into account for direct preference aggregation, or distribute resources directly. This seems good on grounds of common-sense ethics, as it’s plausibly unfair for the current generation to make decisions about the use of space resources on behalf of all future generations.
Here are some advantages of running the process relatively early:
Some types of potentially valuable moral trade are possible only if participants are uncertain about the future, and people’s uncertainty about the future will probably decrease over time.
People might want to bet on the outcomes of some reflective processes. For example, Alice might care deeply about maximizing positive impact if moral realism turns out to be true and discoverable, but care little about having resources if there are no objective moral truths. If Bob values resources equally regardless of whether moral realism is true, Alice and Bob might make the following bet: if ASIs discover objective moral truths, then Bob gives some of his resources to Alice. Otherwise, Alice pays Bob (more discussion of these wagers here).
People might also want to bet on empirical facts. For instance, we may still have uncertainty about how many galaxies are within the reachable universe and which of those galaxies are already occupied by aliens. A fairly risk-averse person might buy resources that are likely reachable and unoccupied, in exchange for a larger amount of more uncertain resources, while a more risk-tolerant person might take the other side of the trade.
These deals are most straightforward under direct endowment, but they could also work in voting systems, e.g., a coalition of moral realists and subjectivists might be willing to vote for a proposal that spent resources in the most moral way (conditional on moral truths existing and being discoverable) and otherwise distribute the resources to the subjectivists’ preferred purposes.
Delaying decisions also gives people more time to figure out how to strategically misrepresent their preferences ahead of a vote or preference aggregation process.
Running the process earlier may allow more flexibility.
If an early vote is scheduled and delaying seems wise at the time of the vote, voters can simply vote to defer to a later vote. If resources are distributed early, recipients can take their time deciding what to do with them.
On the other hand, if the process is scheduled for later, then we can’t make an early decision even if that would have been wise.
Who should participate in the process?
In this section, we discuss who should participate in deciding how to use space resources. This is distinct from who should receive other rights, such as autonomy over their own lives or the right not to be mistreated. We discuss more about how those basic rights could be assured in a later section.
People with substantial bargaining power in 2040
It’s unclear who these people will be, but some plausible candidates are: AI company CEOs, AI company shareholders, the US president (and advisors), the US electorate, Chinese leadership, or the leadership and/or electorate of other countries that participate in the treaty.
This is probably the default outcome. It’s also very unfair, contrary to common-sense ethics, and could lead to many people’s cheap-to-satisfy preferences being neglected unless powerful people are sympathetic to them.Obviously, versions where voters from a few countries get to weigh in are much better than versions where, say, only AI company CEOs and national leaders get to weigh in. But we think all of these have some of the same problems to different degrees.Obviously, versions where voters from a few countries get to weigh in are much better than versions where, say, only AI company CEOs and national leaders get to weigh in. But we think all of these have some of the same problems to different degrees. We’re also concerned that if many people expect that this is how decisions will be made, then they’ll focus on jockeying over power during the intelligence explosion rather than cooperating to avert risks.
All living humans
Another option is to share power among all humans who are alive in 2040, as in the Plan A scenario. This is fairer and more in line with common-sense ethics than the previous plan. And, to the extent that decision-makers can credibly signal that this will happen, committing to this plan might make people more cooperative during the intelligence explosion (although it’s unclear how credibly decision-makers could commit to this). Unlike proposals that share power with later generations (discussed next), this option makes it possible to make decisions relatively early, which might be important.
On the other hand, it’s unclear how politically feasible this will be. And it doesn’t include everyone who might deserve a say in how space resources are used—it excludes future people and digital minds.
Some plausible options:
Include all living humans. This seems like a fairly natural coordination point.
Include all living humans in some capacity, but give people with bargaining power extra weight. For instance, more powerful people might get larger endowments or have their votes or preferences more heavily weighted. This might be a more politically feasible compromise.
Later generations
“Future” generations are the generations born after the intelligence explosion. The “current” generation is the set of humans alive in 2040. We could include future generations in decisions by:
Delaying votes (or rounds of preference aggregation) until future generations are adults.
Reserving some decisions for future generations, e.g., allowing the current generation to vote on how to use 10% of space resources, but letting later generations decide at later votes how to use the remaining 90%, or directly distributing deeds to 10% of resources to the current generation, but saving 90% of resources for later generations.
One reason to share power with later generations is that it’s plausibly unfair not to. However, if we share power with future generations, then this potentially incentivizes parents to have as many children as possible and indoctrinate them to share the parents’ beliefs and values. This could let parents swing a later vote or ensure that a larger share of resources goes to people who share their values.
It’s not clear how big of a problem this is. On the one hand, it’s currently quite rare for ideological groups to deliberately try to have as many children as possible in order to increase their political and economic power. And we expect that most people would find the prospect of having huge numbers of children and brainwashing them to vote the parents’ way abhorrent. On the other hand, new technologies might lower the personal costs of having huge numbers of children (e.g., artificial wombs, robot caretakers). Other technologies could increase parents’ ability to transmit their values to children (e.g., through genetic engineering or novel techniques based on advances in neuroscience or psychology). And a small number of bad actors might be able to substantially increase their population share within a few generations through rapid reproduction.
One way to navigate this is to place guardrails around the rate of reproduction or the extent of control—say, trying to ensure that parents could not use novel technology to reproduce substantially faster nor exercise substantially greater control over their children’s values than early-21st century parents were able to.
Some plausible options:
Don’t include future generations directly. Instead, let the current generation decide whether and how to include them.
For example, if decisions were made by a vote, the current generation could vote to delay decisions until a later vote, where future generations could participate. (And analogously, under direct preference aggregation, if the current generation prefers to reserve some power for future generations, then that would happen).
Under an equal distribution proposal, the current generation could share resources with their descendants or donate to trusts that gifted resources to later generations.
(Although power granted to future generations is plausibly a moral public good, so we might expect it to be undersupplied in the equal distribution proposal).
Include later generations, with restrictions to prevent strategic reproduction and/or indoctrination of one’s children. These restrictions could be chosen by a vote of the current generation.
Include later generations, with no restrictions to prevent strategic reproduction or indoctrination of one’s children.
Digital minds
Another difficult question is whether and how digital minds should have input into decisions about space resources. Digital minds could include AI systems or simulated human brains.
In the Plan A scenario, by 2040, there may be huge numbers of digital minds. Some of these minds could be very sophisticated, with long-term stable preferences, human or superhuman intelligence, and perhaps even subjective experience. On some views, such minds have moral status and deserve a say in future governance.
But this is pretty controversial and might remain so in 2050. Different people could have wildly different views depending on their beliefs about what makes a being worthy of inclusion in the political process. It seems challenging to decide in a value-neutral way which digital minds should be empowered, and choices here could radically affect the composition of the electorate.
There are further difficulties in integrating digital minds into the decision process. While it’s conceptually straightforward to give each human adult an equal stake in the decision—e.g., one vote, one share of space resources—it’s less clear how to individuate digital minds. Should two copies of the same mind running on different computers each receive a vote? If not, how different do two minds have to be before they count as different individuals?
The issues of strategic reproduction and indoctrination discussed above are even more severe for digital minds than for human children. New digital minds could be created as fast as copying data from one computer to another. And the creators of digital minds may have a huge amount of fine-grained control over the values and behavior of those minds. Given sufficient compute, it would be trivial to spin up billions of minds that sincerely want to vote exactly the way their creator tells them.
Some plausible options:
Include a subset of digital minds where the concerns about gaming the process through strategic reproduction and indoctrination are less strong.
Some possible subsets include:
Allow digital minds who were created before 2040 to participate in the process, perhaps with some overall cap on the share of total votes or resources they could receive.
(We would still somehow need to resolve the issue of individuating digital minds and deciding which digital minds deserved votes).
Allow digital minds that are uploads of deceased biological humans to participate in the process, one upload per original human.
Allow people to “adopt” digital minds and enfranchise those adopted digital minds, but only if the creators exercise no more control over the values and beliefs of those digital minds nor are producing digital minds at a faster rate than parents in 2026 could with biological children.
Most likely, most digital minds would not be enfranchised under any of these proposals.
Don’t include digital minds directly. Instead, let biological humans decide whether and how to include them.
How this decision is made would depend on the proposal:
For example, if decisions were made by a vote, the biological humans could vote on a proposal to defer decisions to a later vote where digital minds were allowed to participate.
Under a direct endowment proposal, humans who received resources could give resources to digital minds. (Although, again, this would plausibly be a moral public good and we might expect it to be underfunded under the direct endowment proposal).
One process or many processes?
Designing a governance system involves many tradeoffs—for example, direct endowment and voting both have disadvantages. Should we hedge our bets with hybrid approaches that combine multiple processes?
Here are some examples of hybrid approaches:
50% of space resources are distributed directly to the population at a set time, and 50% are allocated according to the outcome of one or more votes.
Each country (or each treaty signatory) is allocated a share of space resources. These countries each pick their own processes for deciding how to use space resources.
10% of space resources will be distributed in 2040, 10% will be distributed in 100 years, 10% will be distributed in 1000 years, and so on.
The main advantage of hybrid approaches is that different processes can direct resources toward different kinds of values. For example, votes may be particularly likely to fund widely shared values, whereas direct resource endowment might be better suited towards ensuring that rarer preferences are funded. By adopting a hybrid approach, we ensure that a broader set of values each get at least some resources allocated to them.
The main disadvantage of hybrid approaches is that one process might generate outcomes that other processes would regard as substantially harmful. This can be mitigated through some of the strategies discussed in the next section.
Mitigating downside risks
Under some governance regimes, individuals will be able to unilaterally decide how to use space resources. This has some strong upsides: people can create stuff that’s valuable to them even if others are indifferent to it, which makes it more likely that bounded views are satisfied and makes it more likely that nearly everyone is much better off. But it also has downsides: people might use their resources in ways that impose large negative externalities on others.
Some specific examples include:
Seizing other people’s resources. Early settlers might try to use their star systems to launch probes to settle star systems earmarked for others who hadn’t yet laid physical claim on them, which could initiate a (potentially highly negative-sum) space race. Or people might attack star systems where others had already settled.
Uses of resources that most people find intrinsically repugnant. For example, some people might make themselves absolute rulers over societies where nearly all citizens are oppressed or enslaved. Sadists might create new beings that experience immense pain.
Risking the destruction of other people’s resources. People might deliberately destroy the universe (e.g., by initiating false vacuum decay).
Taking other actions with severe negative externalities on others, e.g., starting a war with space-faring alien civilizations.
We need to strike a balance. If we’re too permissive, then a few bad actors could cause enormous damage. But if we’re too restrictive, then we risk banning things that are enormously valuable to a few people just because the majority finds them a bit offputting.
Here are some potential systems for mitigating these downside risks.
Market. If someone doesn’t like how someone else is using their resources, they can pay them to stop.
Strengths: under this proposal, an activity is only stopped when its cost to others exceeds its benefit to the actor, and they’re compensated for their loss. Thus, this proposal is less vulnerable to “tyranny of the majority” worries. Another strength is that we don’t have to collectively decide far in advance what to ban. If someone wants someone else to stop doing something, they can make an offer at any time (although, late into the future, as decision-makers spread out from Earth it might become more difficult to negotiate these trades).
Drawbacks: this could create perverse incentives: some people might do things that others find objectionable so that they’ll be paid to stop. This might result in even more objectionable behavior that would have happened without that incentive. And if a single actor wants to do something a large group objects to, the group may struggle to coordinate payment because of free-rider problems.Suppose there are a thousand actors, each with equal resources. One actor wants to use their resources in a way which is viewed as harmful by all the other actors. Without enforceable prohibitions, each of the majority would need to decide individually how much to spend bribing the unpopular actor to stop. No one of the 999 actors cares so much about the harms that they are willing to pay one-half of their wealth to individually prevent them; but all of the 999 are willing to pay 0.05% of their wealth to prevent them, and that would raise a large enough bribe for the unpopular actor to agree to switch to some other use of his resources. In other words, preventing the use of resources which enough people care strongly about for moral reasons is a moral public good which requires taxation to fund efficiently.Suppose there are a thousand actors, each with equal resources. One actor wants to use their resources in a way which is viewed as harmful by all the other actors. Without enforceable prohibitions, each of the majority would need to decide individually how much to spend bribing the unpopular actor to stop. No one of the 999 actors cares so much about the harms that they are willing to pay one-half of their wealth to individually prevent them; but all of the 999 are willing to pay 0.05% of their wealth to prevent them, and that would raise a large enough bribe for the unpopular actor to agree to switch to some other use of his resources. In other words, preventing the use of resources which enough people care strongly about for moral reasons is a moral public good which requires taxation to fund efficiently.
Voting on bans. We could prohibit activities that a supermajority of the population (say, 80%) votes to ban.
Strengths: avoids both of the drawbacks of the market solution—since there’s no transfer of resources from the people who want to ban to the people who would otherwise do the banned activity, there’s no collective action problem among the former to raise money to pay off the latter. And no one is incentivized to do objectionable things in hopes of a payout.
Drawbacks: a majority that very slightly disliked an activity could ban it even if it was very important to a small minority. Since there’s no transfer of resources, there’s no check to make sure the ban is sufficiently valuable to supporters to justify the cost to opponents.
If we went with this system, we’d have to decide when and how often to vote on which activities to ban.
If we did a single early vote, we might fail to ban activities that were later widely considered bad.
Voters might not know of those activities at the time of the vote (e.g., because the activities are only possible due to unanticipated technology). Or voters might not be in reflective equilibrium at the time of the vote. For instance, perhaps they would later, on reflection, decide that they cared about the welfare of digital minds.
We might be able to circumvent these issues by adopting bans with broad, non-specific language and creating an AI arbiter to interpret the ban. But the arbiter may have to do nontrivial amounts of extrapolation and there might not be a single canonical way to do that extrapolation.
If we had a recurring vote or a single late vote, then voters might have more time to develop and carry out strategies to game the voting system before the vote(s) (see discussion above).
Currently, we think that at least the following activities should be banned:
Torture of conscious beings, including digital minds and non-human animals that are sufficiently likely to be capable of having strongly negatively-valenced experiences.
It might take non-trivial reflection by advanced ASI to figure out the precise bounds of this.For example:
What exactly are negatively-valenced experiences? Probably our current understanding of consciousness and suffering is rather muddled, and people might create minds that have some of the features of classic examples of negatively valenced experiences but not others.
How strongly negative does an experience have to be to fall under this ban?
Are there any mitigating factors, e.g., the experience being necessary for a strongly positive life, the being in question consenting?
What exactly are negatively-valenced experiences? Probably our current understanding of consciousness and suffering is rather muddled, and people might create minds that have some of the features of classic examples of negatively valenced experiences but not others.
How strongly negative does an experience have to be to fall under this ban?
Are there any mitigating factors, e.g., the experience being necessary for a strongly positive life, the being in question consenting?
Unilaterally destroying Earth-originating civilization, or taking actions that bear a significant risk of destroying all of Earth-originating civilization.
Seizing other people’s property without permission.
Making and/or carrying out threats to destroy other people’s resources, in order to extort them.
Making and/or carrying out threats to create large amounts of disvalue by someone else’s lights in order to extort resources from them.
Bans decided by treaty signatories. Treaty signatories could negotiate bans before signing the treaty. This has many of the same strengths and drawbacks as bans decided by a single early vote.
Cancelling out resources. We could allow people to forfeit control of their resources in exchange for removing someone else’s control of their resources at a k:1 ratio (for k > 1).
Strengths: as with the Market approach, by imposing a cost on the people who want the ban, it ensures that people don’t block valuable activities that they only slightly dislike. And unlike the Market approach, since there’s no compensation for the person carrying out the activity, there’s no perverse incentives to do stuff that others dislike in search of a payout.
Drawbacks: as with the Market approach, a large group seeking to prevent a particular use of resources may struggle to coordinate.
Reflection
Under most of the proposals discussed in this appendix, individuals will need to make extremely high-stakes decisions, for example:
Should I sell galaxies allocated to me in an auction? Should I save them to pass onto my children? Should I donate them to a trust, and which trust should I donate to? Should I wait until I’ve thought more to decide, and how should I go about thinking about it?
Should I vote for a particular activity to be permanently banned throughout the universe? Should I vote for public resources to be allocated to a particular activity?
We want people to make wise decisions that reflect their values.
Making these decisions well will likely require resolving open scientific and philosophical questions. The specific questions will vary with individual values, but the following areas are likely to be relevant to a wide range of people:
Improved understanding of morally significant concepts. Many people consider consciousness and preferences to be morally significant—it’s wrong to inflict suffering on a conscious being and good to help people get what they want. But our current scientific and philosophical understanding of consciousness is insufficient to answer important questions like: which minds are conscious and which of their experiences have positive or negative valence—especially as our technological capabilities expand to allow creation of minds very different from human minds in substrate or structure.
Moral truths. There might be objective and discoverable moral truths.
Acausal trade. The reachable universe might be only a small part of what exists. There are regions of our universe outside our light cone, and there may even be other universes. Although we cannot causally interact with beings outside our light cone, our own choices may give us some evidence about how those beings will behave. On some decision theories (e.g., evidential decision theory), this should affect what we do. For instance, we might want to “trade” with these beings, doing stuff that matters to them in exchange for them doing stuff that matters to us. If this argument works out, our acausal impact could plausibly dwarf our causal impact.
Evidential Cooperation in Large WorldsThis mechanism was originally introduced under the name “multiverse-wide superrationality” in this paper.This mechanism was originally introduced under the name “multiverse-wide superrationality” in this paper. (ECL) is one such mechanism. The multiverse contains many causally disconnected agents whose decisions may be correlated with ours: that is, when we observe ourselves making some decision, we should update about their choices too. So if we commit to pursuing multiverse-wide compromise values, this gives us evidence that huge numbers of other agents are doing the same. And this could be great—there might be substantial gains from trade if everyone pursues compromise values rather than their own.
These are just a few examples of potential crucial considerations that we might want to consider before making important decisions.Some more crucial considerations are discussed here.Some more crucial considerations are discussed here.
Deciding how to reflect
People will also need to think carefully about how they want to make these decisions, including how they want to use new technologies:
Deciding when to learn information. While it’s probably important to learn the answers to questions like those above before making object-level decisions about how to use space resources, there might be costs to learning information too early. For example, there might be positive-sum trades that can only be made before the parties to the deal know some empirical results. One example of this dynamic is wagers (discussed above). More speculatively, if we end up wanting to do ECL, it might be necessary to commit to doing ECL before we become confident whether other agents will do ECL—our decision only gives evidence about everyone else’s decision if we’re not sure what others will decide.
Engaging with ASI-generated arguments. ASI will likely be able to research the philosophical and scientific questions above more quickly and rigorously than humans, giving us access to rapid intellectual progress on how to best live out our values.
But ASI could also direct enormous intellectual effort toward crafting extremely persuasive arguments. And it’s not clear that the most compelling arguments will necessarily support true positions or positions we would endorse holding. If ASI is optimizing for maximally compelling arguments, then we might end up believing whatever ideology the ASI was subtly biased toward, or perhaps whatever ideologies are easiest to argue for compellingly.
One safeguard: ASI could refuse to generate superhumanly compelling arguments for claims it knows to be empirically false, in the same way that AIs today refuse to assist with creating biological weapons. But for some questions we care about, there might not be an empirical fact of the matter.For instance, the question of “is a particular digital mind capable of having morally relevant experiences?” hinges on both empirical properties of the digital mind but also questions about what properties are morally relevant.For instance, the question of “is a particular digital mind capable of having morally relevant experiences?” hinges on both empirical properties of the digital mind but also questions about what properties are morally relevant. For these cases, we need to think about how we want to be persuaded—what counts as legitimate moral suasion and what counts as illicit manipulation. It’s not clear where the line is—is vividly describing the suffering of a particular being emotionally manipulative, or just helping you understand the true stakes of your decisions? And the line might be different for different people.
Delegating to aligned ASI. Another option is to delegate reflection to an aligned ASI. We will probably be able to design ASI systems that are much better suited to careful reflection than humans. These systems would be smarter, capable of understanding and integrating more complicated arguments or empirical results than any human.
Moreover, for many people, reflecting in a very careful way might be less pleasant than other ways of spending their time. An ideal reflection process might involve carefully controlling your memetic environment (if it’s possible for ASI to develop illegitimate, yet highly compelling arguments or ideologies). It might require modifying your mind to make yourself smarter, more reflective, or more free of biases that you don’t endorse. In general, it’s not clear that “optimal lifestyle for human flourishing” and “optimal lifestyle for figuring out the most idealized form of one’s values” will be the same for most people. Offloading reflection onto an ASI system that is better suited for the process might let those people get the best of both worlds.
Digital reflection has other advantages difficult to replicate on a biological substrate, including running many parallel reflection processes, reverting to earlier checkpoints, and potentially greater transparency about the trajectory that one’s values took over time. (This last is particularly important in versions of this proposal where many people’s preferences are aggregated together, since people might have incentives to strategically self-modify or misrepresent their preferences. Carrying out a digital reflection with a legible trace could let people verify that they hadn’t done so).
Delegating to future generations. Another option is to delegate decisions to one’s descendants or to future generations more broadly. If so, how should one raise their children? Parents will plausibly have much greater capacity to shape their children’s values—how should they exercise this power?
We’re not sure what the right answers are, and indeed the right answers may vary from person to person based on their values, but we think these are important decisions and want to help people make them carefully.
The role of governance
It’s important for the governance structure to facilitate reflection.
Timing of decisions. Delaying decisions about how to use space resources gives decision-makers more time to think. Although ASI labor may speed up the rate of scientific and philosophical research, many people will likely need years or decades to integrate those findings into their thinking. Additionally, reaching these decisions may require substantial introspection into what the decision-maker truly values, which might also be time-consuming.
For governance structures where the timing of an important decision is scheduled in advance—e.g., a vote or a preference aggregation—this consideration favors delaying that decision (although there are considerations in the other direction, discussed above).
An important advantage of direct endowment governance structures is that they allow individuals to reflect for as long as needed before making object-level decisions. If individuals were still prone to hasty decisions, we could—somewhat paternalistically—require them to wait a few decades before selling or committing their resources to a trust.
Subsidizing reflection. Governments could offer all voters or resource recipients access to ASI labor to help them decide how to use those resources. Governments could also directly fund research into open scientific and philosophical questions likely to bear on decisions many people will face.
Avoiding perverse incentives. Governance structures could incentivize people to warp their reflective process or interfere in others’ reflective processes. For example:
Voting systems that offer voting rights to future generations could incentivize strategic reproduction and indoctrination.
Voting systems (with lie detection) or direct preference aggregation systems could incentivize people to modify their own preferences strategically to gain an advantage (just as traditional voting systems incentivize strategic voting).
The answers to some of the questions above could have implications for which governance procedures are best. One important tradeoff among the governance proposals is between enabling coordination on (moral) public goods and enabling individuals to pursue their own values. ASI-powered research into these questions could shed light on:
Whether there are important public goods that we might want to coordinate on. For example, perhaps acausal trade has large fixed costs—e.g., running lots of simulations of alien civilizations to figure out the distribution of values that trade partners have—and it would be wise to ensure that there are governance mechanisms that let us pool our resources to do that.
The degree to which humans’ moral views converge after endorsed reflection. If people converge substantially, this favors governance proposals with coordination mechanisms that let them pool resources to pursue shared values. If there’s a lot of divergence, this favors governance proposals that allow each person to pursue their own values.
Takes
Mia’s takes
My actual take is that we should wait for aligned ASI and ask it to design a governance proposal that meets the desiderata above (and perhaps more that I haven’t thought of).
It seems very important that everyone making irreversible high-stakes decisions carefully reflect on what they care about with the help of aligned ASI advisors. To make this more likely, governments should probably fund lots of public-domain ASI research into scientific and philosophical questions that are likely to be of broad interest (like the questions above). They should also subsidize ASI compute used by people thinking about these decisions.
Likewise, it seems important that decisions about the use of resources beyond the solar system are made for altruistic or moral reasons, not selfish ones.
We should guarantee that all humans alive in 2040 and their descendantsProbably there should be some limit on the number of descendants who could be supported by these resources because otherwise we presumably will eventually hit Malthusian limits.Probably there should be some limit on the number of descendants who could be supported by these resources because otherwise we presumably will eventually hit Malthusian limits. get enough to support them indefinitely at a very comfortable standard of living (i.e., at least the material standard of living of a multi-millionaire in 2026). This will hopefully lower the personal stakes of the decision about what to do with space and encourage people to be more high-minded when they vote.
It is also great on common-sense ethical grounds—on a cosmic scale, it’s relatively cheap to satisfy people’s preferences, and so we should do that.
I think coordination capacity—where we can collectively use resources to provide (moral) public goods—is potentially really important and it’s not clear that we’ll be able to do that with voluntary contracts. Unless we’re confident that this is possible, it seems good to maintain capacity to coordinate, i.e., have a government (or a few governments) that control substantial amounts of resources (or which are empowered to levy taxes if they need to).
On the other hand, I also want to enable people with unusual value systems to spend some resources pursuing their values.
For these reasons, my top choice is direct preference aggregation if it’s technically feasible and not so unpleasant that many people opt out, perhaps combined with automated reflection (like what’s described here).
If that doesn’t work, my next favorite option is a hackier solution that ensures that some resources go to individuals and some resources are decided collectively. Here’s one way of approaching this.
At least 5% of resources are guaranteed to be directly granted to humans alive in 2040.
By default, the remaining 95% of resources will also be directly granted to humans alive today, unless a supermajority votes that it should instead be spent on public goods.
If highly reliable lie detection is available, governments could use that to check for strategic voting and disqualify people who were sufficiently likely to be executing these strategies.
Governments could also have ASI systems refuse to assist with attempts to coordinate to vote strategically (in the same way that ASI systems will presumably continue to refuse to assist with other harmful requests—e.g., making bioweapons).
We will also vote on bans and on the process to use to amend the list of bans (e.g., should we have recurring votes?). There should be a high threshold for bans, perhaps 90%.
Thomas’s takes
I agree that we should wait for aligned ASI and ask it to design a good governance proposal. The most important things for us to do now are (1) avoid losing control of the future, e.g. via misaligned AI takeover or a human dictator getting control, (2) distribute the power reasonably, and (3) try to do philosophy well.
I think there’s a good chance that in practice, the best path forward will be to lock in some governance structure over the future, because deliberation within a fixed structure (like the ones discussed above) will go better than in a looser structure (like today’s government) where, e.g., more power will accrue to people who are, for example, more politically well connected.
My tentative favorite concrete guess for what to do above is “direct endowment” with some form of amendment process with a very high threshold (e.g. 90% approval). I think it’s very similar to voting based proposals, where you “vote” by making trades with other people about what to do with your resources. The public goods concern is the strongest counterargument, and if there is a voting system discovered that addresses this concern without adding additional, even more dire consequences, then I would change to wanting that.
All voting systems that we know of are screwed up, and I would not want them to be used for something as high stakes as deciding about what to do in the future. But I’m pretty optimistic about this changing, for two reasons:
I think there’s a good chance that ASIs (and humans who think about it for longer) will come up with voting systems that incentivize honest reporting of preferences, and choose a good policy that aggregates the votes.
I think that we might be able to do preference elicitation along the lines of “direct preference aggregation” above. Once you have honest preferences, the problem of designing a good voting system seems much easier.
I care a lot about power in the future going to virtuous people. I think that it would be really bad if a huge fraction of the power in the future goes to selfish people or terminally evil people (e.g. people who want to cause massive amounts of suffering). Of course, different people have different moral views that are reasonable, and I don’t think it’s fair, for example, to claim that everyone should follow my moral views. But I do think that it’s reasonable to exclude people with sufficiently bad moral views, as agreed on from the perspective of the rest of society, from having a large amount of power over the future, for example, to exclude sociopaths from many of the preference aggregation systems above.
I think that any attempt to have “normal governance” continue is not going to work, because our current systems will be totally exploited by actors using superintelligence to gain power. I think that we’re going to need to make massive changes to our system of government really quickly in order to maintain a reasonable ongoing distribution of power.
Appendices — Space Governance Plan
Appendix A — Voting on how many resources to allocate with each process.
A hybrid approach could directly incorporate people’s views on which governance and allocation processes are better, and use this to decide how many resources should be allocated by each.
Here’s one such proposal.
First, a number of governance methods are proposed. This could include all the kinds of systems discussed above (direct endowment, voting using various different voting systems, direct preference elicitation, etc). The only constraint is that each governance method must be reasonably fair in how it weighs participants’ views, rather than artificially biasing the outcome towards specific people or values.
Second, all humans submit a view on what fraction of resources should be distributed by each governance process. The fraction of resources governed by each process is set to the average of everyone’s proposal.
Finally, all humans participate in each of the processes—deciding what to vote for or what to do with resources they’ve been allocated.
One advantage of this approach is that everyone gets a say on what governance processes get used. Another advantage is that it delays the decision about which process to use. If it later turns out that there are big gains from coordinating on moral public goods, people can choose to allocate most resources using voting. Conversely, if some minority of people realize that they might get frozen out of a large-scale vote, they can cause a significant fraction of resources to be distributed via direct endowment instead.
This approach could plausibly be further improved by allowing people to freely trade their voting rights at each decision point.
In other words, people’s right to participate in each different governance process (each vote, each round of resource endowment, etc) could be represented by different shares. These shares could then be freely traded. This might, for example, allow people with more idiosyncratic moral views to get more rights to resources that they have unilateral control over. Conversely, people who care more about the outcome of large-scale votes could get more of the corresponding votes. We could even allow people to trade the right to vote on what fraction of resources should be allocated by each process.
(We find it difficult to predict the outcome of such trades, and it’s possible that it would be better to disallow some of them.)
Takeoff Forecast
This supplement explains how we forecast AI capabilities in various parts of Plan A. We modified the AI Futures Model to simulate four situations:
the default trajectory, if there had been no deal to slow down capabilities progress,
the Plan A consortium trajectory, which scales to TED-AI in a controlled manner, then halts to do alignment research,
a covert project defecting from the deal and trying to build AGI in secret, and
the trajectory after deal dissolution, where most of the post-deal compute is destroyed and the leading project resumes racing.
The rest of this document is structured accordingly. Section 1 covers the differences between the Plan A model and the December 2025 AI Futures Model: a major change to training run modeling, our updated compute forecasts, and the changes to the default parameters. In sections 2 through 4, we explain the various extensions added on to simulate the three additional situations. If you haven’t read the AI Futures Model documentation, we recommend reading it here first.
Modifications to the core model
Here we describe the three ways this model differs from the December 2025 AI Futures Model. These apply across all trajectories simulated.
Explicit modeling of the frontier training run
Like most other AI takeoff models, the AI Futures Model (December 2025 version) defined effective training compute at a given time as the cumulative training compute of the frontier model (in FLOP) times the software efficiency level at that time (in present-day-FLOP per FLOP):
In reality, training runs take time, and the formulation above assumes that today’s software efficiency applies retroactively to the training run for the model that’s already in use. One consequence is slightly overestimating how fast takeoff can proceed. But more importantly for Plan A, this model gives us no good way of dealing with situations where the training system performance (in FLOP/yr) of an AI project suddenly becomes very small compared to the system that was used to train the best model available to that project (as is the case for covert projects and after deal dissolution).
In the Plan A version, we instead model effective training compute as increasing continuously as part of one long training run, at a rate determined by the training system performance and the software efficiency of the known training methods:
We call this rate the effective training system performance and abbreviate it as ETSP.
When all inputs grow exponentially, ETSP and effective compute also grow exponentially at the same rate, so this model behaves like the simpler one. All other equations are unchanged from the December 2025 AI Futures Model.
Changes to the compute and labor time series
The AI Futures Model depends on exogenously-supplied forecasts of compute and labor. The Plan A version uses two sets of time series: one representing the default world if Plan A had not been implemented, and one that branches off from the default when the deal comes into effect. The modeling behind both of these is discussed in our Compute Supplement. In summary,
The default time series before 2029 has been updated to be slightly more bullish compared to that in the AI Futures Model
The Plan A time series from 2030 to 2035 reflects the controlled-but-massive compute buildout enabled by robotics, then slows down from 2035 to 2040.
One limitation is that the human labor projections are not very realistic; we did not put much effort into these because human labor plays a very small role in the model dynamics after AI R&D is automated.
Changes to the default parameters
The default parameters governing timelines and takeoff in the scenario were set to reflect scenario first-author Thomas Larsen’s views. However, since capabilities forecasting was not the main focus of AI 2040, they do not reflect as much consideration as do the values used in the AI Futures Model, and were sometimes adjusted ad-hoc as the scenario narrative and modeling were being written concurrently. Nevertheless, they deviate only slightly from Daniel Kokotajlo’s median values as of April 2026, and all authors find the default trajectory entirely plausible. Below we include a table describing each parameter that was changed from Daniel’s median.
Parameter | Change | Rationale |
Gap years | 1.1 2025-effective-flop-growth | Set to achieve the AC milestone at Jan 2030. |
Time horizon required for AC (before gap) | 10 work years | |
Median-to-top taste multiplier | 5.28x | Set to reach ASI roughly one year after AC, while leaving ~3 OOMs of effective compute between AC and TED-AI. |
AI research taste slope | 3.5 SDs per 2025-effective-flop-growth | |
Median-to-top-human jumps above SAR needed to reach TED-AI | 1.5 |
Modeling the consortium trajectory
Unlike all other actors we model, the consortium is not racing to build better AIs as fast as they possibly can. The specific strategies are discussed in more detail in the Capabilities Scaling Strategy supplement, but at a high level, the strategy is to scale smoothly to Top-Expert-Dominating AI (TED-AI)Specifically, they should scale to something like “max-controllable AI”, but in AI 2040 we operationalize this as TED-AI.Specifically, they should scale to something like “max-controllable AI”, but in AI 2040 we operationalize this as TED-AI., then pause to do alignment research before eventually handing over control to AIs.
Here, we consider strategies that vary along two axes:
How far into the deal consortium aims to reach Top-Expert-Dominating AI (TED-AI), and
How much of a safety tax the consortium aims to be paying when TED-AI is reached.
The safety tax is the difference between the capability level of the actual consortium AI and the capability level of the best AI the consortium could have trained, had they used all their training compute and the best-known algorithms for capabilities training. One example of “paying a safety tax” would be discovering neuralese training, finding that it would be much more efficient, but deciding not to use it to make control easier. We expect that in reality the consortium would want to pay some safety tax.
The rough shape of the consortium’s capabilities trajectory: a brief halt, then smooth scaling, then a long near-pause before handoff.
The consortium trajectory is simulated according to slightly different rules during each of the following stages.
Stage 0: AI Training Pause
This phase starts some number of years before AC would have happened on the default path. (In the scenario, about 9 months before). Though effective compute is forced to stay flat (training is paused!), training system performance can still increase as the leading company within the consortium ends up with more compute at the end of the period.
AI R&D is also essentially paused during Phase 0.Technically in the scenario, the pause lasts from March 2029 to September 2029, and “pilot” experiments are able to proceed during the last part of the year before the first major training run is approved in January 2030.Technically in the scenario, the pause lasts from March 2029 to September 2029, and “pilot” experiments are able to proceed during the last part of the year before the first major training run is approved in January 2030. However, we model the software efficiency of the leading actor within the consortium as growing by a set number of OOMs during this period. This is because of Total Research Transparency: by the time scaling resumes at the end of Phase 0, essentially company IP is public. We expect the companies to have made non-overlapping software discoveries prior to the deal; hence pooling the intellectual property will increase software efficiency at the frontier.
Stage 1: Scale to max-controllable AI
During this phase, effective compute increases with effective training system performance via the usual . TSP grows via the exogenously-supplied Plan A time series, and software efficiency growth is carefully controlled in order to reach TED-AI at a particular time. We assume a mandatory minimum amount of software progress is required per OOM of TSP increase.
To keep track of the safety tax paid at each point, we actually simulate two different trajectories:
One represents the legal project’s realizable capabilities. This trajectory aims to maintain the (constant) software progress rate that achieves TED-AI plus the desired safety tax amount at the target time.
One represents the legal project’s realized capabilities. This trajectory aims to maintain the (constant) software progress rate that reaches TED-AI at the same time that realizable capabilities reach TED-AI+tax.
At each timestep the trajectories moderate their research effort from the theoretical maximum, choosing instead the value which brings them closest to the pre-selected software progress path.
Stage 2: Slow down nearly to a halt, then eventually hand off to AIs
For simplicity, we model this stage as a literal halt for the consortium’s realized capabilities. In reality, we expect them to slow down more smoothly, then continue increasing very slowly as control techniques improve and the threshold for “max-controllable AI” inches upward.
The consortium’s realizable capabilities continue increasing however. TSP growth slows down but continues. As for software efficiency growth, we specify a constant rate of software progress in OOMs per year to represent the capabilities externalities incurred by the huge amounts of alignment research that happen during stage 2. (As an example, dramatic advances in mechanistic interpretability, while promising for alignment, could plausibly also enable much more efficient training methods.)
Since the realized capabilities effectively halt while the realizable capabilities continue growing, the safety tax paid increases dramatically during this period.
In reality, the consortium will choose to hand off to AIs at some point (in our scenario, it happens roughly five years into Stage 2). We do not currently model handoff explicitly however, because none of our analyses depend significantly on the dynamics after handoff. This is why the consortium capabilities line remains flat indefinitely in the explorer, rather than eventually jumping up as we would expect after handoff.
Modeling covert projects
At the moment that AI training is paused (Stage 0), the covert project is assumed to start with the following resources:
A total R&D compute budget equal to some fraction of the leading company’s R&D compute. We assume the covert project allocates between training, experiments, and internal inference in the same proportions as the leading lab did.This assumption may not represent the optimal strategy for a covert project. We did some rough experiments investigating various compute allocation strategies in an earlier version of this model, and found that they made a negligible difference to the speed of the covert project trajectories (except in extreme cases — for example, allocating zero compute to training will halt effective compute growth entirely. Similarly, allocating zero compute to experiments would halt software efficiency growth entirely.)This assumption may not represent the optimal strategy for a covert project. We did some rough experiments investigating various compute allocation strategies in an earlier version of this model, and found that they made a negligible difference to the speed of the covert project trajectories (except in extreme cases — for example, allocating zero compute to training will halt effective compute growth entirely. Similarly, allocating zero compute to experiments would halt software efficiency growth entirely.)
An engineering labor force equal to that same fraction of the leading lab’s labor force,This, too, is a somewhat unprincipled assumption (compared to, e.g., carrying out a dedicated estimate of covert R&D labor and including that as a separate parameter.) However, the precise reduction in the engineering headcount matters very little when engineering is nearly fully automated. We are fairly confident that improving upon this simplification does not affect the results in any significant way, and have checked this empirically at various points. Due to time constraints, we do not currently include detailed results here.This, too, is a somewhat unprincipled assumption (compared to, e.g., carrying out a dedicated estimate of covert R&D labor and including that as a separate parameter.) However, the precise reduction in the engineering headcount matters very little when engineering is nearly fully automated. We are fairly confident that improving upon this simplification does not affect the results in any significant way, and have checked this empirically at various points. Due to time constraints, we do not currently include detailed results here.
AI capabilities, software efficiency, and research stock matching those of the pre-deal leading company.This is one of the ways our model is generous to covert projects; it is similar to assuming that whoever is operating the covert project has access to the frontier model weights, datasets, and full stock of algorithmic knowledge from the leading lab at the time of the deal’s signing. As an improvement, we could add a model parameter representing the number of “months behind” the covert project starts at. Unfortunately, we do not include this as of the Plan A release due to time constraints.This is one of the ways our model is generous to covert projects; it is similar to assuming that whoever is operating the covert project has access to the frontier model weights, datasets, and full stock of algorithmic knowledge from the leading lab at the time of the deal’s signing. As an improvement, we could add a model parameter representing the number of “months behind” the covert project starts at. Unfortunately, we do not include this as of the Plan A release due to time constraints.
We also assume that the covert project’s supply of compute and labor stays constant over time.This is slightly generous but on the whole seems likely to be accurate. The most catastrophic way this could fail is if some amount of the legal compute produced during Plan A could be smuggled to a covert project (or used illegally from within the R&D clusters). This would be catastrophic because world compute increases by several OOMs during Plan A. A milder failure mode would involve the covert project making use of the growing stock of consumer hardware during the deal (as makeshift AI R&D compute). Finally, a covert project could attempt to build and operate its own independent, unmonitored semiconductor supply chain during the deal. Our strategy for preventing all of these things from happening is described in the Verification Plan.This is slightly generous but on the whole seems likely to be accurate. The most catastrophic way this could fail is if some amount of the legal compute produced during Plan A could be smuggled to a covert project (or used illegally from within the R&D clusters). This would be catastrophic because world compute increases by several OOMs during Plan A. A milder failure mode would involve the covert project making use of the growing stock of consumer hardware during the deal (as makeshift AI R&D compute). Finally, a covert project could attempt to build and operate its own independent, unmonitored semiconductor supply chain during the deal. Our strategy for preventing all of these things from happening is described in the Verification Plan. Finally, we assume that the covert project cannot begin training or R&D for some period during which covert datacenters are being constructed.We do permit them to accumulate stolen software during the construction period, described below.We do permit them to accumulate stolen software during the construction period, described below.
With only these factors considered, our modeling indicates that covert projects would probably progress quite slowly. In the median case, each OOM of R&D compute reduction (compared to the leading lab at the start of the deal) would cause the covert project to take over 6x as long to reach TED-AI than would the leading lab.We say “over” because the stated result compares the covert project to a trajectory whose compute level was frozen at the branch point, where in reality the leading lab benefits from growing R&D compute during takeoff.We say “over” because the stated result compares the covert project to a trajectory whose compute level was frozen at the branch point, where in reality the leading lab benefits from growing R&D compute during takeoff.
However, we also assume that the covert project is uplifted to some degree by consortium AI progress. The primary factor, and the only one we model explicitly, is the leakage of software improvements from the consortium to the covert project.Other factors include distillation and weight theft. In this model, we assume these can be prevented.Other factors include distillation and weight theft. In this model, we assume these can be prevented. To account for this, we include a parameter for the fraction of the consortium’s software progress rate that gets added on to the covert project’s indigenous software progress rate. We conservatively assume covert projects will make full use of any algorithms discovered by the consortium, even those deemed potentially dangerous, so we take the realizable software efficiency growth rate as the quantity to steal from rather than the realized software efficiency growth rate.
Due to software proliferation, the consortium’s scaling strategy is extremely important to the covert project’s ability to catch up. The less software progress the consortium needs to make to reach TED-AI at the desired time, the less the covert project is uplifted. As a result, accepting a longer scaling duration that requires less software progress increases the consortium’s ETSP lead at TED-AI, and typically also increases the time between the consortium reaching TED-AI and the covert project reaching TED-AI.
Modeling deal dissolution
The modeling of deal dissolution is relatively simple.
The effectiveness of mutually assured compute destruction is represented by the fraction of world compute available to the new leading project at the time of dissolution.
We also include a parameter for the growth rate of R&D compute after the deal dissolves, representing the output of any non-destroyed or newly-constructed fabs.
As with covert projects, we technically assume that human R&D labor varies in proportion to compute. Though this might be unrealistic, the effect of human labor is essentially negligible after AC.
The post-dissolution leading project starts with the frontier model weights (from the realized consortium trajectory) and all discovered software improvements (from the realizable consortium trajectory) from before the deal dissolved.
We assume that the leading project after a deal dissolves proceeds at full speed, as in the default trajectory.
Appendices — Takeoff Forecast
Appendix A — Upper and lower limits imposed on software progress
If the chosen scaling duration and takeoff parameters demand faster software progress than would be naturally possible at a given point, the research effort may not be moderated at all. Often, once AI capability level increases enough, the consortium is able to catch up to the desired software path. Occasionally, the consortium may not be capable of reaching TED-AI by the desired time, e.g. in worlds where reaching TED-AI from the pause level would naturally have taken ten years but where the consortium attempts to reach it in just one.
Conversely, a sufficiently long scaling period or a sufficiently narrow effective compute gap from the pause level to TED-AI can mean that hardware scaling alone (plus the mandatory minimum software growth rate) would reach TED-AI before the target year. In this case, the consortium simply reaches TED-AI early.
Transparency Plan
AGI projects face key decisions around transparency: who gets to see algorithmic secrets, who gets real time access to the datacenters sufficient to be confident in what they are doing, who gets to be aware of ongoing experiments and training runs, and with what latency. This document outlines our current thinking on transparency in the context of Plan A.
First, we discuss baseline transparency desiderata—aspects of transparency proposals that are core to Plan A. A fundamental component of Plan A is a trustless US/China deal to ensure that AI development is safe, which necessitates a minimal amount of US/China transparency, sufficient to verify that the other is following the agreement. (more).
Second, we give an overview of possible transparency regimes which can satisfy these desiderata. (more)
Third, we give an in depth explanation of our favorite proposal, “total research transparency”. In this proposal, almost all algorithmic secrets, experiments, and training runs would be visible to the public immediately. (more).
Finally, we give an analysis of the tradeoffs between the different proposals outlined, and the reasons why transparency is beneficial. The main reason we recommend Total Research Transparency is because it would improve government and corporate decisionmaking during Plan A; which is our current most concerning failure mode. The second-most important reason is that it makes it much harder for corporations and governments to abuse their power. (more).
Baseline Transparency Desiderata
Any transparency proposal in Plan A should be consistent with the following desiderata:
Trustless Verification. The US and China need to be able to verify slowdown agreements without relying on trust. This means that specifically the US and Chinese governments need visibility into how compute is being used.Without the ability to do trustless verification, the main alternative is to rely on trust. Historical agreements have often relied on trust for enforcement (for example, the London Naval Treaty restricted the tonnage of military fleets, but did not include verification provisions). However, without verification, there is a much larger incentive to defect.Without the ability to do trustless verification, the main alternative is to rely on trust. Historical agreements have often relied on trust for enforcement (for example, the London Naval Treaty restricted the tonnage of military fleets, but did not include verification provisions). However, without verification, there is a much larger incentive to defect.
Model weights security robust against nation-state adversaries. We’d like to prevent model weights from leaking because if the model weights were to leak it would (1) greatly accelerate covert projects and (2) increase misuse risk. Neither of these are necessarily fatal, but we’d very much like to avoid them. Once we get to AIs that are as or more generally capable than top human experts, model weight security will probably be necessary.Concretely, we think that we will, at some point in the capability progression, get AIs that (1) can quickly design/build mirror life, (2) superpersuade/mindhack humans, and (3) quickly build self replicating nanobots/grey goo that grows fast enough to boil the oceans pretty quickly. The resilience response to these threats looks something like “no human ever goes outside, everyone lives in bioshelter bunkers, all info going in and out is moderated by a friendly AI, and a friendly drone swarm / nanobot army is outside protecting the bunker”, which would be (i) highly disruptive and undesirable to most people and (ii) take significant serial time to implement.Concretely, we think that we will, at some point in the capability progression, get AIs that (1) can quickly design/build mirror life, (2) superpersuade/mindhack humans, and (3) quickly build self replicating nanobots/grey goo that grows fast enough to boil the oceans pretty quickly. The resilience response to these threats looks something like “no human ever goes outside, everyone lives in bioshelter bunkers, all info going in and out is moderated by a friendly AI, and a friendly drone swarm / nanobot army is outside protecting the bunker”, which would be (i) highly disruptive and undesirable to most people and (ii) take significant serial time to implement.
Ongoing consumer access to frontier AI. Direct access to the AI models is perhaps the most important form of transparency. Seeing the results of evaluations and high level descriptions of model behavior are much less informative than directly using the model.We think that it is technically viable to verify that an AI cluster is doing inference on a fixed model without the auditors being able to read the information going in and out of the model (e.g. the chatbot logs).We think that it is technically viable to verify that an AI cluster is doing inference on a fixed model without the auditors being able to read the information going in and out of the model (e.g. the chatbot logs).
Four Transparency Proposals
We have four viable candidates for transparency, each of which meet the above desiderata. Proposals #1-3 are further consistent with Plan A, while Proposal #4 only works for Plan S.
Proposal #1: Total Research Transparency. Nearly all AI research is fully transparent to the public, including AI algorithms, code, and documentation, for all frontier AI projects. AI model weights, significant fractions of the training data, and a small amount of other sensitive information is prevented from leaving the datacenters.
Proposal #2: Filtered Transparency. Nearly all AI research is done within a physical security boundary containing both the researchers and the compute. Within the security boundary, US and Chinese auditors have sufficient access to verify deals.In the absence of improved privacy-preserving verification technology, this would involve something similar to Total Research Transparency within the security boundary; with improved verification technology, it could involve less transparency while still being sufficient to verify deals.In the absence of improved privacy-preserving verification technology, this would involve something similar to Total Research Transparency within the security boundary; with improved verification technology, it could involve less transparency while still being sufficient to verify deals. Three things are allowed to leave the datacenter: (1) (partially redacted) reports approved by both sides about what is going on inside, (2) model weights that get sent to the inference datacenters, and (3) researchers themselves. One can vary the level of redaction: with minimal redaction, this proposal approaches Proposal #1, and with maximal redaction, this proposal is similar to Proposal #3.
Proposal #3: Algorithmic Security. It may be necessary to prevent the exfiltration of algorithmic secrets. This requires tightening the security boundary above: if people are regularly entering and leaving the AI datacenters, then they will almost certainly report back most of the algorithms to their governments. Similarly, filtered reports provide a channel for steganographically encoded versions of the model weights to be exfiltrated. Under this proposal, all frontier AI research is done in a secure datacenter with researchers and auditors located onsite and prevented from leaving or communicating any information out. Model weights are allowed to move from the R&D datacenters to the inference datacenters. This involves something like “handing off trust” to the auditors inside the datacenters, because so limited information is allowed to leave the datacenter that external actors cannot exert meaningful oversight internally.
Proposal #4: Shut it all down. Specifically, ban compute-intensive AI R&D. Under this proposal, there is no large-scale AI R&D, so there’s no need for AI R&D transparency. This could be enforced through compute declarations and physical inspections (but doesn’t require sharing digital access to the clusters, i.e. algorithms). This is similar to the MIRI proposal, and also what occurred for about 6 months during 2029 in the AI 2040 scenario.
This table summarizes the rough tradeoffs between them:
Proposal #1: Total Research Transparency | Proposal #2: Filtered Transparency | Proposal #3: Algorithmic Security | Proposal #4: Shut it all down | |
SL5 algorithmic secrets security | No | Partial, depends on redactions | Maybe | N/A, no new algos discovered |
Allows new training runs | Yes | Yes | Yes | No |
How informed are public discussions of safety cases? | Very | Depends on redactions. | No | N/A, no new safety cases needed |
In Plan A, we currently recommend Total Research Transparency for many reasons, chief of which is that the likelihood of reasonable analyses of safety cases is much higher if external researchers and the general public are able to critically analyze them.In the Plan A scenario we also do depict the implementation of Proposal #4 for a short period of time (from March to October 2029).In the Plan A scenario we also do depict the implementation of Proposal #4 for a short period of time (from March to October 2029). The main drawback relative to Proposals 2&3 is that there is less algorithmic security, and so algorithms will probably leak to covert projects. However, we think that SL5 algorithmic secrets security is incredibly difficult. Therefore, our main mitigation for algorithms leaking is to prevent new algorithmic secrets from being developed in the first place.
Total Research Transparency
In this section we’ll discuss a specific possible implementation of Total Research Transparency, which will correspond to the implementation done in Plan A.
The arrows in the diagram show the flow of information between humans, R&D datacenters, and Inference datacenters.
Under this proposal, there are two types of datacenters:
Inference Datacenters only provide inference made by whitelisted models.This is enforced via an inference-only retrofit, as explained in our verification supplement.This is enforced via an inference-only retrofit, as explained in our verification supplement. These datacenters host consumer chatbots, almost all agent workloads, robot foundation model inference, etc. Arbitrary requests are allowed to be sent into the datacenter, but only verified AI model outputs are allowed to come out, so that the model weights are secure.This verification also helps to prevent misuse and AI R&D assistance. This is covered more in our verification supplement.This verification also helps to prevent misuse and AI R&D assistance. This is covered more in our verification supplement.
R&D Datacenters allow training runs and experiments—everything needed to produce a new set of model weights. They are “transparent datacenters,” as described above. These datacenters are much higher risk than inference datacenters, because they could be used to do an intelligence explosion.
Information from the AI R&D Datacenters is organized as follows:
An Opaque Internal Database contains all information that shouldn’t be public. This includes model weights, activations, some datasets, and some hyperparameters. No human (internal or external) is able to directly view the information in this database.
A Transparent Internal Database contains all other information, including all code, AI chains of thought, prompts, etc. However, it’s not directly accessible to the public or AI researchers who aren’t physically located within the security boundary of the AI R&D datacenter.
A Transparent Public Database is accessible to everyone. It is a record of all information going in or out of the R&D datacenter. It logs all activities of AI researchers interacting with the R&D clusters, as well as any additional queries by members of the public that wish to see some piece of information. For example, they might want to task an AI to look through the internal database and write a short report about the state of some particular field of research.
Anything that’s in the transparent internal database could end up in the transparent public database if anyone requests that information. However, what stops everything from leaving to the public database are upload limits that cap outgoing information to ~1 megabyte / second, as a layer of defense for the opaque information being exfiltrated (i.e., model weight exfiltration in particular).This threshold is chosen to be high enough to allow for convenient access to the R&D clusters, but low enough to make weight theft difficult. The security rationale is discussed further below, in the Security supplement.This threshold is chosen to be high enough to allow for convenient access to the R&D clusters, but low enough to make weight theft difficult. The security rationale is discussed further below, in the Security supplement. There is an ongoing auction for the outgoing bandwidth to allocate access, and all outgoing bandwidth is also copied to a freely accessible database on the internet. In practice, this means that almost all algorithms are immediately revealed to the public.
The R&D datacenters are still owned and operated by tech companies, like in today’s world. In fact each R&D datacenter is really a campus of datacenters, inside which several tech companies‘ datacenters are located. The US and Chinese governments, with their respective auditors and monitoring devices, ensure that each tech company’s datacenters on the campus create logs of their activities that go to the internal opaque database and the internal transparent database.
Anyone, including members of the public, can purchase compute inside the R&D datacenters and have the same level of access to frontier AI development as frontier AI company researchers.This happens automatically, because of the co-location of multiple companies’ datacentersThis happens automatically, because of the co-location of multiple companies’ datacenters Members of the public can, for example, run experiments on intermediate model checkpoints, see loss curves during training, or ask an AI to run a desired analysis for them.
There are three risk tiers of inference:
Extreme Risk: AIs used for AI R&D need to have direct access to R&D compute, and so they run on the R&D datacenters instead of the inference datacenters. This also allows additional security—these AIs are subject to carefully designed control protocols.
High Risk: AIs used for construction in dangerous domains (e.g. constructing robots, constructing GPUs or semiconductor manufacturing equipment) OR AIs used for research in dangerous domains (e.g. research relevant to engineered pandemics). These AIs live on the inference datacenters and must log the trajectories produced by AIs operating in high risk domains.
Lower Risk: AIs used for anything else. This includes most commercial use (e.g. using AIs as a chatbot), most corporate use (e.g. using AIs for medicine/law/architecture/etc). These AIs are allowed to run with zero-data-retention (ZDR), allowing for maximum privacy. Just like standard practice today, even AIs running with ZDR are still subject to refusals and anti-misuse classifiers. However, only AIs behind the frontier are allowed to run with true ZDR. This is discussed in more detail in our verification supplement.
High Risk and Lower Risk inference operates on the inference clusters, while Extreme risk inference occurs within the AI R&D clusters on a compute cluster that’s verified to be inference only.
All of the high risk trajectories and some of the non-ZDR lower risk trajectories are physically transported (under guard!) between the inference datacenters and the R&D datacenters in order to give safety researchers access.
Filtered Transparency
In filtered transparency, AI researchers are physically located within a security boundary, and we attempt to prevent information from leaking across this boundary. Then there are also auditors inside the R&D datacenters that have sufficient access to verify and make judgement calls about safety.
The three ways that information crosses out of the boundary are:
Model Weights. Just like in total research transparency, model weights need to pass from the R&D datacenters to the inference datacenters, and this channel needs to be high bandwidth because model weights are so large.
Auditor Reports. The auditors write up reports about progress inside the R&D datacenters, which include things like current capabilities, safety relevant information about the current architectures. Redacted versions of these reports could be made public.
Humans leaving the security boundary. Human researchers periodically enter and leave the security boundary. This is the main way that we expect algorithms to be exfiltrated from the datacenters.
There’s no restriction on information crossing into the security boundary.
Like Total Research Transparency, there are R&D datacenters and inference datacenters. This allows for one of the most important forms of transparency: widespread access to frontier AI models. The main downside of filtered transparency relative to total research transparency is that it does not allow for fully open access and scientific scrutiny of the safety cases related to AI R&D. Instead we have to hope that the relatively small group of researchers inside the security barrier will be able to identify all the flawed assumptions of each safety case, make appropriate assessments, make sufficient research progress, etc.
Algorithmic security is similar to an extreme form of filtered transparency. The central idea is that algorithmic progress is allowed to be made in the project, but doesn’t leak outside. This is extremely difficult because algorithmic progress is often memorizable.
Therefore, the main difference is that for the “algorithmic security” proposal, it is necessary to prevent humans from crossing the security boundary. This enables us, in principle, to secure algorithmic insights.
When is each proposal desirable?
The main axis for deciding between transparency proposals for R&D is whether we are more concerned about possible covert projects, or poor regulation. Increasing transparency will result in more algorithms leaking, boosting covert projects, but improving public oversight of AI projects and increasing the expected quality of AI regulation.
Overall, we think that poor regulation is more concerning than covert projects, and so we recommend Total Research Transparency.
Proposal #4 does not allow for the development of new AI models, i.e. it only works as part of Plan S as opposed to Plan A. The decision about Plan A vs Plan S will mostly be made on other factors, which are discussed more in the Plan S branch.
What fraction of algorithms will leak under each proposal?
The optimal policy on this spectrum will depend on quantitative factors around how much algorithmic progress would leak under each proposal. This will depend on what fraction of algorithmic progress is of the type that would be transparent (e.g. architectural improvements) vs what fraction of algorithmic progress is large dataset improvements. Overall, we make the current (very rough) best guess estimates of the fraction that will leak under each proposal:
Proposal | Algo Fraction leaked |
P1 Total | ~67% |
P2 Filtered | ~47% |
P3 Algorithmic Security | ~14% |
You can see the reasoning behind this estimate here. We would be excited about security experts (which we are not) making their own forecasts; especially forecasts that are more granular (e.g., that include over what time period the algorithms leak).
Total Research Transparency Upsides
This section talks about the benefits of transparency for reducing loss of control risks and concentration of power risks. It’s basically an explanation of this diagram:
In the diagram, the two most important policy interventions of Plan A are described in blue. The two most important goals—preventing loss of control, and preventing extreme concentration of power—are described in green. Arrows represent causation.
This diagram is in some sense a summary of Plan A, except it’s biased towards depicting the benefits of transparency in particular.
We won’t talk through the entire diagram here. Instead we’ll talk through the seven depicted benefits of Total Research Transparency, plus an additional eighth benefit. For each one, we’ll talk about how weaker kinds of transparency would get similar (but weaker) benefits.
“Technical conversation no longer dominated by AI company employees.” Transparency will improve epistemics about AI risks, and therefore will make regulations more likely to be implemented well.
The US and China need to agree on AI regulation sufficient to dramatically reduce AI takeover risk. This regulation will involve things like deciding on a capabilities schedule (which capabilities are too dangerous, when is it worthwhile to accept some risk?), deciding on whether particular algorithms should be allowed vs. banned (e.g. new architectures that are more sample-efficient yet also more interpretable), and which research directions have worthwhile capabilities/alignment tradeoffs (e.g. whether research into new scalable oversight techniques is likely enough to be a useful component of a safety case).
Under status quo transparency, this would result in a small group of government officials needing to make these decisions. The only technical experts they would be able to consult that have access to the latest information would be at the AI company(ies) they are trying to regulate, because they wouldn’t be allowed to share information with the other technical experts. But the AI company employees have a massive conflict of interest here and so cannot be trusted to give unbiased opinions.
It helps to give the information to a broader group of experts (e.g. academics and rival AI companies), without publishing it. Now, AI companies can be used to audit each other. This has much better incentives because AI companies have an incentive to point out if their competitor is doing something reckless, dangerous, or illegal. However, this relies on governments selecting the right groups of experts, and relies on AI companies to report on AI companies, each of whom will still be biased against industry-wide regulation.
It helps even more to make the relevant information public, because now we no longer rely on AI companies or government approved AI experts to process the information on their own and form good opinions about the necessary regulations. This allows anyone to write up high quality risk analyses about any given model, algorithm, or practice. Instead of the discussion of sufficient regulation happening in a few small offices with government and AI company employees, it can happen in public, making it dramatically more likely that mistakes will be caught and fixed. The government still needs to make good decisions, but they are in a better position to do so because they can peruse an ocean of research and analyses produced by different groups that are all able to see the relevant evidence.
Furthermore, transparency makes it much more likely that the judicial system will be able to exert reasonable oversight on AI development, which is a core lever for enforcing reasonable safeguards.
“Huge increase in the number of alignment researchers with access to the latest models and evidence.” More transparency means more researchers can participate in cutting-edge research, which means progress is made faster.
In 2026, only a small fraction of people with the expertise to do AI alignment research have access to the latest models and the latest evidence. The rest are outside the frontier AI companies, either in laggard companies, or in nonprofits, or in academia. Total Research Transparency equalizes access, so that anyone can buy some compute and run some experiments on the latest models, and see the results of the experiments others have run. This straightforwardly increases the effective amount of human labor directed at making alignment research progress.
More moderate forms of transparency would help in the same way, but to a lesser extent.
“Incentives not so perverse” Total Research Transparency removes the incentive for AI companies to race to better AI algorithms.
In Plan A, the US and China directly regulate AI companies to limit their rate of algorithmic progress. But the transparency also serves as a second layer of defense by changing the incentives around algorithmic progress.
Under the status quo, AI companies’ main moat is their algorithmic advantages over competitors. Therefore, AI companies have a huge incentive to race to find better AI algorithms. When all AI algorithms are shared, this moat evaporates, and so too does this incentive. The race dynamic is a central upstream factor leading to high AI takeover risk, which Total Research Transparency directly undermines.Under this proposal, then, what would the companies be incentivised towards? Well, they still own the model weights they train, and they still own their datacenters, their software products, etc. So they’ll compete to train the biggest models and teach them the most lucrative skills and serve them fastest in the most appealing UI and build up the most loyal userbase. Just like in other low-secrets industries, a company with good taste and rapid execution can overtake larger dumber competitors, stay ahead even as they scramble to copy them, and ultimately perhaps even build up a durable moat of infrastructure scale and network effects. That said, moats and profit margins will be smaller and AI will be much more of a commodity under our proposal than in the default world. We think this is a good thing.Under this proposal, then, what would the companies be incentivised towards? Well, they still own the model weights they train, and they still own their datacenters, their software products, etc. So they’ll compete to train the biggest models and teach them the most lucrative skills and serve them fastest in the most appealing UI and build up the most loyal userbase. Just like in other low-secrets industries, a company with good taste and rapid execution can overtake larger dumber competitors, stay ahead even as they scramble to copy them, and ultimately perhaps even build up a durable moat of infrastructure scale and network effects. That said, moats and profit margins will be smaller and AI will be much more of a commodity under our proposal than in the default world. We think this is a good thing. More moderate kinds of transparency would not achieve this to nearly the same extent.
It helps to consider an example. In 2026, many AI researchers agree that chain of thought monitorability is helpful for safety and that it would be a shame to lose it. However, each company is individually incentivised to have research programs into new paradigms that would destroy or damage chain of thought monitorability (call this “neuralese”). After all, if they don’t do it, someone else will.
Now imagine that something like Plan A happens, with Total Research Transparency. AI companies can sleep more easily without doing any investment into neuralese research, because they know that if anyone else was making a serious attempt, they’d immediately see it. What would be the point of doing it anyway? You’d take a hit to monitorability for a competitive advantage that would be immediately shared with all your competitors.
Now imagine that the regulation is not based on Total Research Transparency but instead on some more moderate kind of transparency—e.g. filtered transparency based on government audits. Now each company has an incentive to research neuralese again, and moreover, if they’ve started researching neuralese they may even have an incentive to convince the government that it’s not as dangerous as it sounds, because now they think maybe this really will be their competitive advantage.
“Agreements to pay safety taxes and ban dangerous practices can be verified.” Total Research Transparency makes AI agreements more robust.
Making the logs of most AI development public dramatically increases the aggregate brainpower that’s looking over the logs to ensure they are in compliance with the agreed upon rules. From the perspective of a company or country trying to defect from the agreement (e.g. by secretly using compute in the R&D clusters to do a massive illegal training run, or a training run that they argue is allowed but is really pushing the boundaries and is arguably illegal), it becomes massively more difficult for them to do this because it will be hard for them to predict what sorts of analyses the public will do on the logs. The alternative, trying to outwit a small group of overstretched government auditors (in proposals 2&3), will probably be much easier.Of course, this doesn’t solve the whole problem. For example, a defecting AI project may be able to conduct their research steganographically through the monitored input and output channel, or they may be able to undermine the cybersecurity that’s attempting to enforce the transparency requirements.Of course, this doesn’t solve the whole problem. For example, a defecting AI project may be able to conduct their research steganographically through the monitored input and output channel, or they may be able to undermine the cybersecurity that’s attempting to enforce the transparency requirements. And of course, without even some basic level of transparency that is already much stronger than what we have in 2026 (US and Chinese governments get to see what’s going on in each other’s datacenters) then they have hope that their spy and cyber operations are good enough to substitute, or else trust each other not to cheat.
“No company pulls ahead, others catch up, including companies in other countries.” Having multiple companies spread across multiple countries is great for preventing extreme concentrations of power.
This one is pretty straightforward. In 2026, power comes from a variety of sources—knowledge is power, money is power, military strength is power, etc. In the future, as AIs become superhuman at everything and proliferate around the world, power will increasingly come from controlling the best AIs. (It also comes from controlling the greatest number of AIs, so to speak, but because AIs can be copied, we think that generally speaking whoever controls the best AIs will soon end up controlling a very large number of them as well.)
The kind of broad access and control offered by AI companies in 2026 is pretty weak from a power-concentration perspective. Yes, individual users can purchase subscriptions to ChatGPT or Claude and so forth. And under normal circumstances, ChatGPT or Claude will follow your orders. However:
You can’t use them for everything—they’ll refuse to help you with some things for example, or possibly even quietly underperform, and
There’s nothing stopping the AI companies from inserting secret loyalties or hidden agendas into their AIs, like xAI and Google appear to have done on occasion (those were the cases that were so blatant as to get caught).
There’s nothing stopping the government from quietly ordering the companies to do things like that, e.g. in the name of national security. Governments could abuse this to surveil, censor, and propagandize their populations.
In a crisis—which is when power matters most—the hard power is with the company and/or government, not with the people. Even if the AIs are currently trained and instructed to do as the user wants, for example, and even if their Spec/Constitution is fully public, well, in an actual constitutional crisis scenario all of that can go out of the window very quickly.
Besides, as AIs penetrate more of the economy, more and more money will flow to the AI companies. If there are only a handful of such companies, that’s inherently a concentration of power. If most or all of them are in one country, that’s inherently a concentration of power. It could easily lead to a global oligarchy or dictatorship, because a small group of already-powerful people controls all the AIs that matter.
The situation is much less likely to lead to global oligarchy or dictatorship if there are many companies spread across many countries with similarly powerful AIs. Not only that, but this means that market forces can operate: companies will be incentivised to give users what they want, including firmer control, because it’ll help them get market share from their competitors. This helps to reduce the risk of local oligarchies or dictatorships as well; if one country uses AI to propagandize their citizens, well, perhaps their citizens will be able to access AIs from other countries that tell them the truth.
Total research transparency helps other companies catch up to the frontier in two ways. First and foremost, it publishes the core algorithms. Secondly, it makes it easier for countries to agree to limit the pace of AI progress, which gives other companies and countries more time to catch up. Weaker forms of transparency would also help but not as much.
“No Secret Loyalties / Hidden Agendas”
This is another major way in which total research transparency helps to prevent extreme concentrations of power.
As previously mentioned, it is both feasible and precedented for those who actually own and train the AIs to train hidden agendas into them. After all, the values/goals/personalities of the AIs have to be trained in somehow, and that’s all very complicated, and so to make some of it hidden is as simple as declining to make all of it visible.
Total research transparency basically kills this threat. Not only does it kill it for companies, it kills it for governments, unless the US and Chinese governments were to collude to keep secrets from the world. With total research transparency, every step of the training process for every frontier model is tracked by US and Chinese monitors and published. If the model creator misleads the public about the goals/values/personality-traits/etc., there’s a paper trail people can follow to find out.
Weaker forms of transparency would help too, but not as much. For example, consider Proposal 2: Filtered Transparency. Here you might end up in a situation where the US AI companies shared some political bias that they were training into their AIs, and the US government either doesn’t notice because their auditors are too busy with other tasks, or notices but doesn’t care because it happens to share the same political affiliation as the companies. Perhaps the Chinese government auditors would notice, but perhaps they too would be busy, and anyhow they might not care.
Oversight of AI companies by governments is more effective; oversight of the executive branch by judiciary, congress, and the public is more effective too. Total Research Transparency makes it easier for groups who aren’t directly training the models to see how the models are being trained, and to judge for themselves whether the rules are being implemented fairly and reasonably.
Consider a biased or corrupt regulator that is using its authority to punish companies that it doesn’t like, or impose its ideology on domestic AI industry. Consider instead a situation where a company or companies is complaining that this is happening, but really the regulator was behaving reasonably in response to the evidence, and simply doing its job to keep people safe and enforce the law. How is the public to distinguish between these two cases? How is the judicial system?
Total Research Transparency helps everyone see what’s going on, the better to evaluate who is being reasonable and who isn’t. As usual, weaker forms of transparency help too but not as much.
Unknown Unknowns: Transparency seems good for putting humanity in a position to deal with problems as they come up, including problems that we haven’t thought of yet.
The development of AI will pose many problems, both foreseen and unforeseen. Increasing transparency into AI development will on average improve the level of public epistemics and understanding of what is going on.
Total Research Transparency Downsides
The main additional downsides of Total Research Transparency relative to Proposals 2&3 are:
Total Research Transparency potentially incentivizes “safety theater”. There may be some lines of AI safety research which, if pursued, might make AI companies look bad. For example, evaluations of model misbehavior. Under conditions of Total Research Transparency, this sort of research might be disincentivized because there would be no way to hide the results, whereas in filtered transparency for example the results would only be shared with a smaller circle of auditors and other companies.
Leading AI companies and large parts of the US executive branch will probably dislike it. Competitors will dislike that it allows their competitors to catch up, especially companies whose algorithms are their primary moat. Governments generally shy away from giving geopolitical rivals access to information which could then be used against them. Therefore, the leading AI companies and many parts of the US executive branch will probably be opposed to Total Research Transparency.