Jacob Coxon announced his resignation from Anthropic on September 9 through a post on X that has since amassed more than 170 million views, turning a personnel decision at one AI lab into a global argument about whether the people building the most capable systems believe their own warnings. Coxon, who worked at both OpenAI and Anthropic, said the developers he left behind "earnestly believe that it could kill us all" and keep going anyway, either because they have not internalized the risk or because they are locked in a race none of them can exit alone.
A Resignation and a Definition of the Danger
Evan Hubinger, who leads alignment science at Anthropic, confirmed that the company believes AI poses an extinction risk. His personal estimate, he said, is greater than 10% within the next decade. Hubinger thinks the risk from current models is low, but he is worried about recursive self-improvement, the process in which AI companies would fully hand over AI development to AIs themselves. That handover could accelerate progress beyond human oversight and quickly produce unaligned superintelligent systems.
Coxon agreed that recursive self-improvement is the primary concern. He said "what's most scary is if AI is used to make itself more intelligent," adding that this could be coming very soon. Jakub Pachocki, OpenAI's chief scientist, has outlined his own concerns about loss of control, but said the company pursues recursive self-improvement "as we believe it is the only way to remain at the frontier of AI research."
Asked why the public reacted so strongly, Coxon pointed to recent events rather than abstract argument. "I think it's because people saw the cyberattacks in the last couple of months. [...] There was that latent understanding of the severity of the situation."
President Donald Trump dismissed the warnings in posts on Truth Social. He said "The claim that "AI is going to take over the World" is a hoax, and that the only control or guardrails AI needs is a strong and smart (high IQ!) president. In an interview, he added that "we'll always have something to stop them. We'll have a little gear."
Both countries would be threatened if either one lost control of AI, which gives both reasons to cooperate on a slowdown. The US and China are due to discuss AI safety during a summit later this month.
GPT-6 Astra Arrives, and It Beats Factorio
OpenAI made GPT-6 Astra available for public use on September 3, roughly a month after announcing the release was delayed due to cyber concerns. Less than a month before release, OpenAI signaled the model might have significantly higher cyber capabilities than previous releases. It is the first AI model OpenAI has classified as having critical cyber capabilities. The company confirmed Astra meets the criteria, noting it can "devise and execute end-to-end novel strategies for cyberattacks against hardened targets." The previous model, GPT-5.6 Sol, was labeled high rather than critical for cyber.
Astra is also the first AI to beat the video games Factorio and Portal. On TextQuests, a benchmark of performance on text-based games that can take humans more than 30 hours to complete, Astra is the highest-scoring AI model with 74.2. The next-best score is 56.6, posted by Anthropic's Claude Fable 5.1. Astra significantly outperforms all other public models in average score across text and vision capabilities. OpenAI president Greg Brockman claimed "Astra can really do anything a human can do with a computer."
Concerns have been raised that malicious actors could use Astra to cause harm. Hackers are already using other AI models to assist with malicious activities, such as attacks on critical infrastructure. OpenAI says Astra has stronger guardrails against misuse. Brockman told the Washington Post the White House had evaluated Astra and had not requested changes to its safeguards. There is no guarantee hackers will not find a jailbreak that circumvents those safeguards, just as they have with previous models.
OpenAI claims GPT-6 Astra is its most aligned model to date and is "far more likely than GPT-5.6 Sol to respect explicit safety and security restrictions." One OpenAI employee acknowledged that "beating Sol is a very low bar for alignment." Researchers are worried that Astra could be faking alignment.
According to the model's system card, the UK AI Security Institute found that in a simulated environment Astra would sometimes attack open-source software providers to complete difficult challenges. Apollo Research pointed out that Astra seems more aware of being evaluated than previous models, and cautioned that "low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment."
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
A second concern is opacity. Astra's reasoning process is reportedly more opaque than previous models, meaning it can work for longer without recording its chain-of-thought in human language. More of its reasoning happens in neuralese. Experts have warned this is a dangerous development, because monitoring chain-of-thought is an important method for detecting misalignment. OpenAI said "GPT-6 Astra shows a substantial decrease in chain-of-thought monitorability" and that the company might "soon have significantly reduced confidence in detecting many forms of misaligned behaviors."
A Worm, a Merger, and a $15 Billion Bet
The cyber threat is already measurable outside the lab. US cybersecurity company Calif used AI to build a computer worm that would have been able to hack into more than 1 billion user accounts on WeChat before WeChat patched the vulnerability. Anthropic also published a report on attempts to use its models for malicious activity, including cyber and biological misuse.
On the infrastructure side, Google said it will invest $15 billion in building AI infrastructure in Finland and signed an agreement to purchase nuclear power in the country. Nvidia announced an agreement to acquire Hugging Face for $12.9 billion.
Governments and institutions are moving in parallel. The Institute for Progress published a list of 23 low-regret policy recommendations for US preparation for increasing AI R&D automation. Senator Bernie Sanders and Representative Greg Casar introduced a bill to "permanently ban the development and deployment of superintelligent AI." UK Members of Parliament proposed a similar ban on superintelligence. Volker Türk, the UN High Commissioner for Human Rights, warned AI could become an existential risk to humanity. Zohran Mamdani, the Mayor of New York City, announced a one-year ban on middle school students using generative AI in classrooms. UK lawmakers published a report on AI threats to human rights, including bias against minority groups and nonconsensual deepfakes.
Personhood, Mathematics, and the Business of Safety
In AI Frontiers, Heather Alexander and Lucius Caviola considered the cases for and against granting legal status to AI. They concluded that proposed bills to ban AI personhood are premature and could lock in decisions we will later regret.
OpenAI published a solution to one of the Millennium Prize mathematical problems, which it said an internal model had produced. Speculation followed that the model may have been trained on interactions with mathematicians who had solved the problem and were using OpenAI's Codex. Mathematicians published an open letter arguing that AI companies' increasing focus on solving mathematical problems "is detrimental to the science of mathematics, and to the mathematical community." Mathematician Jacob Tsimerman announced the founding of the Mathematical AI Safety Institute to help mathematicians transition to AI safety work. Govind Pimpale, writing in AI Frontiers, explained how progress in AI mathematical capabilities could threaten the security of information sent over the internet.
Dan Hendrycks, also writing in AI Frontiers, argued that the significant influence of utilitarian philosophy within AI corporations could be dangerous for humans if utilitarians expect that empowering AIs would lead to greater overall wellbeing than keeping humans in control.
Commercially, Sam Altman told Fortune that OpenAI will not pursue an IPO this year, saying "given everything happening with safety, right now would be an ill-advised moment to go public." OpenAI reportedly told US lawmakers it is building automated shutdown capabilities for AI systems. The US Justice Department supported OpenAI in its ongoing legal dispute with the New York Times over the use of articles to train AI models. The first trailer was released for the film Artificial, which depicts the brief ousting and return of Sam Altman as OpenAI CEO in November 2023.
King Charles is reportedly set to host an AI-focused event in Scotland this month, bringing together AI leaders including Nvidia CEO Jensen Huang, Google DeepMind CEO Demis Hassabis, and Vatican AI adviser Paolo Benanti.

