Developer

OpenAI Expands Codex with Screen Control Coding Agent

OpenAI has updated its Codex developer tool to include a background computer use feature that lets the AI view screens, click, and type directly. The agent can now handle tasks autonomously over days or weeks and includes new plugins and image generation. This positions Codex against competitors like Anthropic's Claude Code, with rollout starting on macOS.

Neura News

Neura News

Neura Market Editorial

April 16, 20263 min read

Originally reported by the-decoder.com

OpenAI Expands Codex with Screen Control Coding Agent

OpenAI Expands Codex with Screen Control Coding Agent

OpenAI launched a major update to Codex, turning the AI coding assistant into one that controls computers directly. The new background computer use feature allows Codex to observe screens, click elements, and input text using its own cursor. This change moves Codex past its earlier limits as a simple terminal and editor tool.

Multiple Codex agents can operate at the same time on a Mac. They run without disrupting the user's other applications. Developers find this helpful for tasks like refining front-end code, testing software, or using programs without APIs. For now, the feature works only on macOS.

Browser Integration and Workflow Tools

The Codex desktop app now has a built-in browser. Users can add comments straight on web pages to guide the agent. This setup suits front-end development and game creation most. OpenAI intends to grow this so Codex fully manages browsers, not just local web apps.

Codex supports more parts of the development process. It edits comments on GitHub reviews, handles several terminal tabs together, and links to remote development environments through SSH in early testing. Conversation history from past chats carries over for better context.

Autonomy and Team Automations

Codex schedules its own future work and restarts automatically to finish long projects. OpenAI notes this can span days or weeks. Teams apply these for managing pull requests, following tasks, or checking messages in Slack, Gmail, and Notion.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

OpenAI, founded in 2015 as a research lab focused on safe artificial general intelligence, developed Codex as part of its push into developer productivity tools. Earlier versions integrated with platforms like GitHub Copilot, helping generate code from natural language. This update builds on that by adding persistent, screen-aware operations.

Anthropic's Claude Code serves as a direct rival, offering similar code assistance. OpenAI's changes target that space with stronger autonomy and device control.

Image Tools and Plugin Expansion

Image creation comes via gpt-image-1.5 in Codex. Paired with screen captures and code, it helps teams build product ideas, front-end layouts, mockups, and game visuals in one flow.

Over 90 new plugins arrived, covering skills, app connections, and servers. Examples include Atlassian Rovo for JIRA, CircleCI, CodeRabbit, GitLab Issues, Microsoft Suite, Neon by Databricks, Remotion, Render, and Superpowers. These let Codex access data from various tools and act on it.

Rollout Schedule

Updates roll out right away for Codex desktop users signed in with ChatGPT accounts. Features like personalized suggestions and memory come later for Enterprise, Education, EU, and UK accounts. Screen control stays on macOS initially, with EU and UK access delayed.

This positions Codex as a complete partner for software work, handling everything from code edits to long-running automations.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read