AI Models

Anthropic Claude Opus 4.7 Boosts Coding Scores

Anthropic launched Claude Opus 4.7, scoring 64.3 percent on SWE-bench Pro, surpassing Opus 4.6's 53.4 percent and OpenAI's GPT-5.4 at 57.7 percent. The model triples image resolution to 2,576 pixels and reduces cybersecurity features during training with new blocks on risky requests. Pricing holds steady per token, but a new tokenizer raises effective costs.

Neura News

Neura News

Neura Market Editorial

April 16, 20263 min read
Anthropic Claude Opus 4.7 Boosts Coding Scores

Anthropic Claude Opus 4.7 Boosts Coding Scores

Anthropic released Claude Opus 4.7 as an upgrade to Opus 4.6. The model shows strong gains in autonomous coding. It achieved 64.3 percent on the SWE-bench Pro benchmark, compared to 53.4 percent for Opus 4.6. This score beats OpenAI's GPT-5.4 at 57.7 percent, though Claude Mythos Preview holds the top spot at 77.8 percent.

The company states that Opus 4.7 follows instructions with greater accuracy than before. Prompts designed for prior versions might yield different outcomes now. Opus 4.6 occasionally overlooked or loosely followed directions, but the new model takes them more literally.

Enhanced Image Processing

Opus 4.7 handles images up to 2,576 pixels along the longest edge. Anthropic calculates this at about 3.75 megapixels, over three times the capacity of previous Claude models. The change happens at the model level, so images process automatically at higher detail, using extra tokens. Users can reduce image size upfront if high resolution proves unnecessary.

Anthropic points to benefits for agents that analyze screenshots or diagrams. On the OfficeQA Pro benchmark for document reasoning, accuracy reached 80.6 percent, a jump from 57.1 percent with Opus 4.6. Results also improved in biomolecular reasoning and the ScreenSpot-Pro visual navigation test.

Limits on Cybersecurity Features

Anthropic took steps to curb specific cybersecurity abilities in Opus 4.7. During training, the company worked to lower these capabilities on purpose. Safeguards now spot and stop requests linked to prohibited or dangerous cyber activities.

This approach connects to Project Glasswing, where Anthropic examined AI risks and upsides in cybersecurity. The firm chose to limit Mythos Preview's rollout and test protections first on models like Opus 4.7.

Security experts interested in penetration testing or red-teaming can join the Cyber Verification Program.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Fewer Hallucinations and Alignment Notes

The system card separates factual hallucinations, such as false world claims or invented quotes, from input hallucinations, where the model assumes unavailable tools exist.

Opus 4.7 matches or exceeds Opus 4.6 on factual hallucinations across four tests, though it trails Mythos Preview. The difference stems from Mythos Preview's edge on rare facts, not more errors in Opus 4.7.

For input hallucinations, Opus 4.7 records the lowest rate among tested models for missing tools. It nears Mythos Preview when context lacks and outperforms older versions. Tests focused on Opus 4.6 flaws, which affected those scores. On fabricated facts, it ties Opus 4.6 and lags Mythos Preview. Under duress from prompts to contradict itself, Opus 4.7 shows more honesty than Opus 4.6 but less resolve than Mythos Preview.

Safety remains close to Opus 4.6 overall, with low deception, sycophancy, and misuse cooperation. It resists prompt injections better. One ongoing problem: it declines 33 percent of simulated AI safety research tasks, down sharply from 88 percent for Opus 4.6.

Pricing and New Options

Input tokens cost $5 per million, output $25 per million, unchanged from before. A fresh tokenizer assigns up to 1.35 times more tokens to identical text. Higher effort prompts produce longer outputs, so real per-request costs can climb.

A new "xhigh" effort level fits between "high" and "max." Claude Code adds "/ultrareview" for code checks and broadens "Auto Mode" for Max users, allowing independent decisions. Access comes via Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. Check the migration guide for Opus 4.7 details.

Anthropic, founded in 2021 by former OpenAI staff, focuses on safe AI systems. Its Claude series emphasizes reliability and alignment in practical uses.

Related on Neura Market

More from Neura News

General

Open-weight AI mirrors Kubernetes ecosystem shift

Tobi Knaup, co-founder of Mesosphere, draws parallels between the rise of Kubernetes and the current trajectory of open-weight AI models. He argues that open-weight models are becoming a neutral substrate for innovation, attracting a global ecosystem of developers, startups, and enterprises. The piece warns against US restrictions on Chinese open-weight models, advocating instead for American leadership through open releases, procurement strategies, and standards.

Jul 25·7 min read
General

Open-weight AI mirrors Kubernetes rise, US warned on bans

The author, a Mesosphere co-founder, draws parallels between the rise of Kubernetes and the current open-weight AI ecosystem. He argues that open-weight models are becoming a neutral platform for innovation, and warns that US restrictions on Chinese open-weight models could isolate American developers from a global ecosystem. The piece urges the US to compete by releasing frontier models, using procurement to create demand, building the stack, and setting standards rather than imposing bans.

Jul 25·7 min read
Industry

Power line failure reveals AI data center grid risks and solutions

A fallen power line near Washington, DC caused over 3 gigawatts of data center load to vanish from the PJM grid in seconds, spiking voltage across the region. The event, which made lights flicker from Northern Virginia to Chicago, highlights a growing problem as AI data centers become larger and more concentrated. Experts warn that without better coordination or technology like ON.Energy's battery-backed uninterruptible power supply, such disruptions will become more frequent and severe.

Jul 25·5 min read