At 10:30 last night,Claude Opus4.7Claude Opus 4.7 arrived without warning.
The launch was not entirely unexpected, though. Claude had been unusually unstable over the previous few days, suggesting that a major release might be imminent.
Users saw slowdowns, outages, and inconsistent performance. The new model turned out to be what all that disruption had been leading up to.
Within hours, Opus 4.7 was everywhere across the AI community, including our internal channels.

GeekOnUp's AI trend-monitoring channel
Opus 4.7 brings meaningful improvements across advanced software engineering, visual understanding, professional workflows, memory, long-context reasoning, and agentic tasks.

Here is a closer look at what changed.
Advanced software engineering
The most visible gains are in advanced software engineering. This is not simply a benchmark improvement with little practical difference; the change is noticeable in real work.
Anthropic's launch materials include extensive feedback from early testers, and the results point in the same direction.
- CursorBench rose from 58% with Opus 4.6 to 70% with Opus 4.7.
- Performance across 93 GitHub coding benchmarks improved by 13%.
- Notion reported a 14% improvement in complex, multi-step workflows over Opus 4.6, while tool-use errors fell to one-third of their previous level.
- On Rakuten-SWE-Bench, Opus 4.7 completed three times as many production tasks as Opus 4.6.
The model also designs a validation approach before reporting that a task is complete. That detail matters.
In real projects, the most frustrating failure is often not that a model cannot write code. It is that the model confidently declares success without running tests, reviewing errors, or checking edge cases.
A model that can define and execute its own validation strategy is much closer to an agent that belongs in a production engineering workflow.
Several supporting capabilities reinforce that direction.
- Opus 4.7 introduces an xhigh effort level between high and max.
- Claude Code now defaults every plan to the xhigh effort level.
- The API also offers task budgets in public beta, giving developers control over Claude's token budget for long-running tasks.
Claude Code has also added the /ultrareview command, which opens a dedicated code-review session to identify bugs and design issues. Pro and Max users receive three free trials.
Individually, these may look like minor updates. Together, they reveal a consistent product direction.
Anthropic wants Claude to perform more reliably on long-running work.
That is why Opus 4.7 feels increasingly like a serious execution engine.

One detail in the launch materials is particularly revealing.
Anthropic directly compared Opus 4.6 with GPT-5.4.
Opus 4.6 trailed GPT-5.4 on most programming benchmarks.
In effect, this is the first public acknowledgment that the previous generation was behind GPT-5.4 in coding.
Opus 4.7 closes that gap and moves ahead on several measures.
Visual capabilities
Vision may be the most dramatic part of this release.
Opus 4.7 supports higher-resolution images, with a maximum edge length of 2,576 pixels and roughly 3.75 megapixels, more than three times the resolution supported by earlier Claude models.

This is more than a sharper view of an image. It changes what the model can do in several practical settings.
Examples include information-dense screenshots in computer-use workflows, small terminal text, Figma screens, complex charts, architecture diagrams, financial tables, chemical structures, and patent drawings.
Previously, the model could often grasp the general picture.Now it can begin to read the details that determine the correct next step.
The result on XBOW's visual acuity evaluation is especially striking.Opus 4.7 scored 98.5%,compared with 54.5% for Opus 4.6.

That result is significant for a specific reason.
XBOW builds autonomous penetration-testing systems.
The model must interpret browsers, admin consoles, developer tools, network requests, dialog boxes, and error messages. These interfaces contain a high density of information.Misreading a single line of small text can derail every inference that follows.
Better vision is therefore not just about helping everyday users understand images.
For an agent, vision is the gateway to understanding real-world work interfaces.
If it cannot see accurately, it cannot act accurately.
That matters to teams like ours that design products, systems, and interfaces.
When Claude receives an admin screenshot, a complex dashboard, or a Figma design, its ability to identify the page structure, component states, and meaning of the data directly affects the quality of every subsequent analysis.
With Opus 4.7, vision is not merely a gap being closed. In a meaningful sense,Anthropic is giving its agents better eyes.
Professional workflows
Opus 4.7 also improves on real-world work, particularly financial analysis, professional presentations, office productivity, and synthesis across multiple tasks.
It is also among the leading models on GDPval-AA.

Claude has long been well suited to knowledge work.
Greater reliability in document reasoning, financial analysis, legal materials, and professional reporting would deliver clear practical value.
The outcome still depends on the task, however.
For reading complex files, organizing data, and supporting professional judgment, the new model may be substantially stronger.
For writing with a distinctive human voice or shaping brand language, the result may be less convincing.
This has become one of the most debated points in early hands-on reviews.
Many users say Opus 4.7 is more capable, but also feels more like GPT.
It is clearly stronger at coding, vision, and agentic execution.
Yet some users find it less effective than earlier Claude models for content creation.
Much of Claude's appeal came from its judgment in language and its natural voice. With Opus 4.7, that quality feels less consistent.
It would be a real loss if stronger engineering execution gradually came at the expense of that strength.
Memory
Opus 4.7 improves file-system memory, allowing it to retain important instructions across long periods and multiple sessions. It can carry that knowledge into new tasks, reducing the amount of context users must provide at the outset.
This is another essential step toward capable agents.
In the past, every new session with a model felt like onboarding a new colleague from scratch.
You had to explain the project background, file structure, coding conventions, business rules, previous decisions, and mistakes that should not be repeated.
If file-system memory becomes dependable, Claude starts to resemble a collaborator that can contribute continuously.
But stronger memory makes governance more important as well.
Teams will need to decide what should be retained, what should expire, what constitutes a project rule, what is merely a temporary preference, and what information could distort future work.
That is why we believe production agents should not simply remember more.
Their memory must be manageable and governed.
Security and access
Some context is useful here. Last week, Anthropic introduced Project Glasswing,
demonstrating the cybersecurity capabilitiesMythos Previewof a specialized model.
That model substantially outperformed Opus 4.7 across the evaluations, but Anthropic does not intend to release it publicly.

Opus 4.7's cybersecurity capabilities have been deliberately constrained.
During training, Anthropicexperimentally reduced capabilities associated with cybersecurity tasks。
and launched automated safeguards to detect and block high-risk requests. Legitimate security researchers who need Opus 4.7 for vulnerability research, penetration testing, or red-team exercises can apply to the Cyber Verification Program.
This design has substantial long-term value as advanced models move into enterprise use.
As model capabilities grow, providers face an unavoidable reality.
Unrestricted access is not viable, but blanket restrictions are not viable either.
Security research, enterprise red teaming, vulnerability reproduction, and compliance testing are all legitimate requirements.
If a model cannot distinguish authorized research from a malicious attack, its only option is to reject requests indiscriminately.
Identity verification, tiered permissions, and dedicated access paths will therefore become increasingly important.
Today, the issue is cybersecurity.
Tomorrow, it may involve healthcare, finance, biology, legal services, supply-chain simulation, or enterprise risk management.
Advanced models cannot enter regulated industries under a single, universal access policy.
They need more granular permission systems.
What to consider before migrating
There are two potential pitfalls to understand before switching.
The first is the new tokenizer.
The same input may use more tokens, typically 1.0 to 1.35 times as many depending on the content.
The second is that higher effort levels involve more reasoning,particularly in later turns of agentic workflows, and can generate more output tokens.

The published price has not changed.
Input remains $5 per million tokens and output $25 per million tokens.
Actual cost, however, depends on how the model is used.
- For difficult tasks,fewer retries, fewer tool errors, and a higher first-pass success rate could make Opus 4.7 more economical overall.
- For lightweight work,or routine tasks where the quality improvement is limited, the additional token consumption will be more noticeable.
It would be a mistake to assume that unchanged pricing means unchanged cost, or that higher token use automatically amounts to a hidden price increase.
Both conclusions are too simplistic.
A better approach is to benchmark the model against your own production workloads.
Run the same requirement through Opus 4.6 and Opus 4.7, then compare output quality, rework, total token use, and time to completion before deciding whether to migrate.
For development teams,start with high or xhigh in Claude Code. There is little reason to begin at max or use the heaviest setting for every task.
For content teams,use Opus 4.7 first for source analysis, long-document review, and structural planning, then assess whether its writing matches the voice you need.
There is no need to move every creative workflow immediately.

Opus 4.7 is powerful, but it is not the default answer for every task.
Models are becoming increasingly specialized professional tools.
The more capable they become, the more carefully they need to be applied.
Final thoughts
Claude is becoming better suited to the agent era.
We only hope that as it becomes better at execution, it does not lose the natural voice that made earlier versions distinctive.
Every generation involves trade-offs. Our job is to use the model where it is strongest and compensate thoughtfully where it has regressed.
That may be the mindset every technology professional needs to maintain in this period of rapid change.
Tools will keep changing. Judgment must remain in human hands.


