Preprint
AI Safety & Alignment

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

Elias Fernández Domingos, The Anh Han
July 28, 2026

0

Citations

0

Influential Citations

Venue

2026

Year

Abstract

Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.

Analysis

Why This Paper Matters

This paper addresses a critical tension in AI development: the race to deploy advanced AI systems often pits speed against safety. While much discussion focuses on individual risk preferences or corporate culture, this work provides experimental evidence that the competitive structure itself—specifically the fear of falling behind—can drive unsafe choices. The findings challenge the common assumption that unsafe AI development is primarily a matter of risk tolerance or poor judgment, and instead highlight systemic incentives that can push even cautious actors toward risky behavior.

The study is particularly timely given the current landscape of AI development, where major companies and nations are racing to deploy increasingly powerful models. The results suggest that without interventions to reduce competitive pressure, unsafe development may be an emergent property of the race itself, not just a consequence of reckless actors. This has direct implications for AI governance and policy design.

Technical Contributions

  • Behavioral experiment design: A framed, repeated-choice game where paired participants choose between Safe (slower, lower payoff, no risk) and Unsafe (faster, higher payoff, accumulating private risk up to a cap) development under an uncertain time horizon.
  • Three treatment conditions: Maximum risk levels of 10%, 60%, and 90% to test whether risk tolerance affects choices.
  • Pre-registered and exploratory analyses: The pre-registered hypotheses about risk level and risk preferences were not supported; exploratory analyses revealed the importance of opponent behavior and relative position.
  • Evolutionary model: A reduced model with four strategies (Always Safe, Always Unsafe, Conditionally Safe, Conditionally Antisocial Safe) that reproduces the treatment effect and shows how conditional unsafe behavior can be favored by competitive dynamics.

Results

  • No significant effect of maximum risk level (10%, 60%, 90%) on unsafe choices.
  • Risk preferences (elicited via a separate task) did not predict unsafe behavior.
  • Participants were significantly more likely to choose Unsafe after their opponent did so.
  • Being ahead in the race reduced unsafe play, while falling behind increased it.
  • First-round choices strongly predicted later behavior, indicating early momentum.
  • The evolutionary model showed that conditional strategies (e.g., being unsafe when behind) can be evolutionarily stable under competitive pressure.

Significance

This work shifts the conversation on AI safety from individual risk attitudes to systemic competitive dynamics. It suggests that even well-intentioned actors may be pushed toward unsafe development if they perceive themselves as falling behind. Policy implications include the need for mechanisms to reduce competitive pressure—such as international agreements, safety standards, or cooperative frameworks—rather than focusing solely on risk education or regulation of individual actors. The evolutionary model provides a theoretical foundation for understanding how unsafe behavior can spread in a competitive environment, offering a tool for designing interventions.