You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI-enabled crisis instability currently lacks detailed threat models which rigorously map out concrete pathways that lead from today's systems to crises that run beyond human control, what capabilities and deployment conditions are required for each of these pathways, and how these different pathways and threats interact with each other. A lot of threat modelling is underspecified and narrowly scoped, in terms of the concrete pressures, incentives, and mechanisms, and also in the interplay between the behaviour of the delegated agents themselves and the broader strategic and security landscape.
This project is an attempt to fill this gap and create detailed models of how exactly the delegation of strategic decisions to AI agents could destabilise crises and agreement-making between states. I intend to spend 12 months driving an effort in this direction, drawing upon my own knowledge of technical AI safety, AI evaluations, nuclear command and control, verification, and international politics, and my connections with experts in relevant fields, in order to identify the pathways and relevant cruxes, evaluate the likelihood of these pathways and identify the possible outcomes, and map out possible interventions to defend against these pathways.
I intend to spend approximately 50% of my time, frontloaded into the initial 5-6 months, on creating and refining detailed threat models with concrete pathways, mechanisms, and quantified assessments of likelihood and impacts, across three pathways: delegation-driven escalation, bargaining breakdown, and commitment failure. About 30% of my time, mostly in months 2-9, will go into generating the measurements those models need, which in this domain cannot simply be looked up. About 20% of my time will go into engaging other people in the field and in adjacent expert communities, in order to get feedback, ensure minimal duplication of effort, and build a shared picture.
The measurements come from two accepted SPAR research teams starting on September 14, five mentees across a technical project and a governance project, with $3,000 of project funding already secured. They build an open testbed for two-agent bargaining and crisis interaction and run experiments on it, and I set the direction, review their work, and integrate the results as parameters. Mentees who lead a set of experiments are first authors on its write-up.
At a minimum, I will publish four pieces of public-facing writing analysing specific facets of different pathways to crisis instability, plus a comprehensive document detailing the pathways and potential defenses, circulated with relevant people in the community and shared publicly if doing so is not too risky. I will lead the project at half of my time, with the two SPAR teams joining in September; for relevant expertise, I can draw on my existing collaborations with the Alva Myrdal Centre and the Oxford Martin AI Governance Initiative.
AI-enabled crisis instability seems to be one of the most critical threats we might face over the coming decade, and one of the ones we are least prepared for. A crisis between major powers escalating past what either side intended would be catastrophic and largely irreversible. In addition, the pressure to delegate faster than a rival would also likely lead to reckless deployment into military and diplomatic functions, increasing the risks of loss-of-control and potentially great power conflict.
My work is valuable as detailed threat models are the necessary inputs for understanding the nature of the threat, and identifying and prioritizing the potential interventions. Hence, the outputs of this project will be useful to funders and people working on military AI, crisis stability, and arms control, and would contribute to building an actionable agenda in order to rapidly scale future work in an effective manner.
This is a twelve-month program which will cost a total of $70,000. We are able to draw on $3,000 of project funding from SPAR. The remainder of the funding is not currently available, and we are asking for assistance to cover it. We are pricing my own time as half of the total, $4,000 per month, as this program runs in parallel with the power-concentration program; together they enable us to fund one full-time person.
Lead stipend - $48,000 ($4,000 per month, half of my salary for this period)
Compute and model API beyond the $3,000 we have already secured - $12,000
Contracting out engineering and a portion of the empirical work to be continued beyond the end of the SPAR period in December - $10,000
The minimum we would be seeking is therefore $14,000 (three months of my time at half time plus a small margin). This would be sufficient to specify the set of scenarios we intend to explore, complete first-pass pathway models, and keep the two teams of individuals on a consistent direction through the autumn.
I lead the project, working with two SPAR research teams from September 14: three mentees on the empirical work and two on the analytical work, selected through SPAR's Fall 2026 process.
Contributor to ControlArena (UK AISI control evaluations) and the HCAST benchmark (arXiv:2503.17354) at METR.
Evaluating and Understanding Scheming Propensity in LLM Agents (arXiv:2603.01608), co-authored at LASR Labs, supervised by David Lindner (Google DeepMind).
Applying LLMs to Political Challenges in Nuclear Verification (in prep., with Giacomo Cassano, Uppsala). Presented at the Politics of Verification in Frontier AI workshop (Stockholm, April 2026) and the AMC Multidisciplinary Conference 2026; accepted at SGRI 2026.
Opportunities and Risks from Advanced AI in Nuclear Verification, forthcoming chapter with Sophia Hatz in the Routledge Handbook on Nuclear Verification.
Member of the Working Group on International AI Governance, Alva Myrdal Centre for Nuclear Disarmament, Uppsala University.
Ongoing collaboration with the Oxford Martin AI Governance Initiative on strategic parameters for AI agreements.
Public writing on AI, nuclear strategy, and strategic stability at thenextfrontier.blog.
The main downside risks are that some of the work might be dual-use in nature and could potentially provide a plausible guide to exploiting weaker counterparts, and there is a risk that published evaluations end up in training data and stop measuring what they claim to. In order to avoid this, I will extensively red-team outputs, seek private feedback from trusted experts, and redact anything too sensitive to disclose, restricting distribution.
$3,000 in project funding is secured for the empirical half through SPAR Fall 2026. An application covering the wider delegation programme is pending with the Effective Institutions Project and Cooperative AI Foundation joint call, submitted in July, and a proposal covering the testbed component was submitted to the Schmidt Sciences and Cooperative AI Foundation call on multi-agent AI safety in August. My threat-modelling work on AI-enabled extreme power concentration is a distinct programme with its own Manifund page and its own pending applications. I will update this page as soon as any commitment lands.
There are no bids on this project.