Developer

OpenAI Rolls Out Lockdown Mode to Prevent Data Theft from Prompt Injection

OpenAI has launched Lockdown Mode, a security feature that blocks data exfiltration during prompt injection attacks. The mode limits outbound network requests and is now rolling out to Free, Go, Plus, Pro, and self-serve ChatGPT Business accounts. It does not prevent injections from appearing in content but cuts off the final stage of data theft.

Neura News

Neura News

Neura Market Editorial

June 6, 20263 min read

Originally reported by simonwillison.net

OpenAI Rolls Out Lockdown Mode to Prevent Data Theft from Prompt Injection

OpenAI has activated a new security feature called Lockdown Mode, designed to prevent the final stage of data theft from prompt injection attacks. The feature first teased by the company in February is now rolling out to eligible personal accounts, including Free, Go, Plus, and Pro tiers, as well as self-serve ChatGPT Business accounts.

Lockdown Mode works by limiting outbound network requests that could transfer sensitive data to an attacker. It directly targets the exfiltration vector, the channel through which stolen data leaves a system. According to a blog post by Simon Willison, Lockdown Mode does not stop prompt injections from appearing in content that ChatGPT processes. For example, an injection could still appear in cached web content or an uploaded file and might affect the behavior or accuracy of a response.

What Lockdown Mode Does

Lockdown Mode is a deterministic mechanism. It does not rely on AI systems to evaluate threats, which means it cannot be subverted by sufficiently devious attacks that might trick an AI-based defense. This is a crucial design choice because any security measure that itself uses an LLM could be manipulated. By cutting off the exfiltration path directly, OpenAI provides a layer of protection that is harder to bypass.

The feature is now live and rolling out across the supported account types. OpenAI describes Lockdown Mode as a way to combat the "Lethal Trifecta," a term for the combination of three elements that enable data theft in LLM systems: access to private data, exposure to untrusted content, and a way to steal and transmit data back to an attacker.

The Lethal Trifecta and How Lockdown Mode Breaks It

The Lethal Trifecta occurs when an LLM system has access to all three of these components. To stop an attack, a system must remove at least one leg of the triad. According to the analysis shared by Willison, the easiest leg to restrict without making LLM systems far less useful is the exfiltration vector. Lockdown Mode directly addresses that leg.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

By limiting outbound network requests, the mode prevents the final step of a prompt injection attack. Even if an attacker successfully places a prompt injection in content the model processes, and even if the model is tricked into outputting sensitive data, that data cannot leave the system. The attacker cannot receive the stolen information.

Implications for ChatGPT Security

The existence of Lockdown Mode implies that ChatGPT, in its default settings, does not provide robust protection against sufficiently determined data exfiltration attacks. OpenAI is essentially acknowledging that without this additional measure, the system could be vulnerable to cases where a user or process feeds untrusted content to the model while it has access to private data.

This move is significant for developers and businesses that rely on ChatGPT for handling sensitive information. Users who handle confidential data should consider enabling Lockdown Mode to reduce the risk of data theft through prompt injection. The feature is available without extra cost to eligible accounts.

Lockdown Mode does not affect the model's ability to process content or generate responses. It only restricts outbound network requests. This makes it a relatively low-impact security enhancement that could be applied broadly without degrading functionality.

The rollout is ongoing. OpenAI has not specified a timeline for full availability, but the feature is now accessible to the account types listed.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read