Search results:

Back

GPT-5.5 Puts OpenAI Back on Top

Strong early feedback and broad capability gains put OpenAI firmly back at the front of the AI race.

AI Summary

Strong early feedback and broad capability gains put OpenAI firmly back at the front of the AI race. Two days ago, OpenAI releasedGPT Image 2, and it broke well beyond the usual AI audience. Text rendering, layout, and real-world knowledge have long been stubborn problems for AI image generation. This model solves nearly all of them. The impact has been immediate. Images made with it are already everywhere, and one thing is clear: OpenAI has once again pulled the wider internet into its orbit.

Two days ago, OpenAI releasedGPT Image 2, and it broke well beyond the usual AI audience.


Text rendering, layout, and real-world knowledge have long been stubborn problems for AI image generation. This model solves nearly all of them.


The impact has been immediate. Images made with it are already everywhere, and one thing is clear: OpenAI has once again pulled the wider internet into its orbit.


Before that momentum had even faded,

GPT-5.5 arrived.



OpenAI reignited excitement at the product layer, then followed with a major model upgrade. Coding, agents, knowledge work, document processing, and scientific research all moved forward at once.

Early community feedback has been exceptionally strong.


OpenAI is once again the clear leader in AI.


It leads the Artificial Analysis Intelligence Index by three points,breaking the previous three-way tie with Anthropic and Google.



OpenAI describes GPT-5.5 as itssmartest and most practicalmodel yet,

built for a new way of working with computers.


Agentic Coding


The biggest story starts with coding.



GPT-5.5's most significant gain is in agentic coding.


Terminal-Bench 2.0 evaluatesend-to-end agent engineering capability.



GPT-5.5 scores 82.7, up from GPT-5.4's 75.1, while Claude Opus 4.7 reaches 69.4.

On OpenAI's internal benchmark for longer-running engineering work, GPT-5.5 scores 73.1 versus 68.5 for GPT-5.4.


That is a meaningful step forward.


A few examples show what this looks like in practice:


1

Artemis Mission Visualization

Prompt:

Build a new application with WebGL and Vite using real data from the Artemis II mission. Test it thoroughly until every feature works and the visual result matches the reference image. Pay particular attention to rendering the planets and flight path. The 3D scene must be interactive, and the orbital mechanics should be credible.



2

UFO Shooter

Prompt:

Create a 3D game with Three.js in which the player controls a tank and shoots down UFOs flying overhead.

Think step by step and take a moment before responding. Restate the problem first.

Imagine you are writing an implementation guide for a junior developer who is about to build the game. Provide clear, specific instructions on which files to inspect, what to change, and what needs to be fixed.

Then write all of the code. Use an attractive low-poly visual style.

Remember that you are an agent. Continue until the user's request is fully resolved before ending your turn. Break the request into every necessary subtask, verify that each one is complete, and do not stop after delivering only a partial solution. End only when you are confident the problem has been solved.

Before making further tool calls, plan against the workflow and carefully evaluate each result so that the user's request and every related subtask are fully addressed.



Beyond benchmark scores, GPT-5.5 also demonstrates astronger understanding of system architecture: why something fails, where a fix belongs, and which other parts of the codebase may be affected.


Itsreasoning and autonomyare notably stronger than GPT-5.4 and Claude Opus 4.7. It can surface issues early and anticipate testing and review needs without being explicitly prompted.

For example, when asked to redesign the commenting system in a collaborative Markdown editor, GPT-5.5 returned a nearly complete twelve-part patch stack.


Knowledge Work


The next major area is knowledge work.


GDPval measures how well AI performs structured knowledge work across 44 occupations.

GPT-5.5 scores 84.9%, ahead of Opus 4.7 at 80.3% and Gemini 3.1 Pro at 67.3%.


On Tau2-bench Telecom, which tests complex customer-service workflows, GPT-5.5 reaches 98.0% without prompt tuning.



This improvement is easy to underestimate.

Model upgrades are usually discussed through coding demos, benchmark scores, and generated games.


But the capabilities businesses will consistently pay for are often not the most eye-catching demos. They are theessential, repetitive tasks that consume hours every day.


Finding information, reading files, building spreadsheets, cleaning data, organizing source material, producing a client-ready document, or turning fragmented business input into a plan a team can actually execute.


OpenAI's own teams are already putting these gains to work.85%Employees across software engineering, finance, communications, marketing, data science, and product management use Codex every week.


Some analyze six months of speaking-request data. Others review 24,771 tax forms spanning 71,637 pages. Another team automated its weekly business reporting,saving five to ten hours every week.


GPT-5.5 is not only getting better for developers.It is becoming a serious presence across a much broader range of business operations.


Scientific Research


Scientific research is another important leap.


These workflows demand more than a correct answer to a difficult question. Researchers need to explore ideas, collect evidence, test hypotheses, interpret results, and decide what to investigate next.GPT-5.5 is better at completing that full cycle than competing models.


Researchers are already using it in several ways:


1.  It helped identify a new proof involving a Ramsey number, which was subsequently verified in Lean.

2. Bartosz Naskręcki, Assistant Professor of Mathematics at Adam Mickiewicz University in Poznań, used GPT-5.5 in Codex to build an algebraic geometry application in just11 minutesfrom a single prompt. The application visualizes intersections of quadratic surfaces and converts the resulting curves into Weierstrass models.



3.  Derya Unutmaz, Professor of Immunology and researcher at The Jackson Laboratory for Genomic Medicine, used it to analyze a gene-expression dataset containing 62 samples and nearly 28,000 genes, producing a detailed research report.

The reportdid more than summarize the findings. It surfaced important questions and new insights.

Unutmaz said the same work would otherwise have taken his team months.


OpenAI captures GPT-5.5's role in research with a simple phrase:a research partner.


Inference Efficiency


The conventional assumption has been straightforward:more capable models are usually slower.


How did GPT-5.5 become both more capable and faster?


The answer goes beyond model-level optimization. OpenAI redesigned the entire inference system.

GPT-5.5 was co-designed, trained, and deployed with NVIDIA GB200 and GB300 NVL72 systems.


There is another notable layer: after analyzing weeks of real production traffic, Codex wrote new load-balancing and partitioning heuristics,increasing token generation speed by more than 20%.


The model is no longer just generating outputs. It is becoming part of the operating system around the work itself.


Safety


This area is less dramatic than coding or research, but no less important.


GPT-5.5 advances beyond GPT-5.4 in cybersecurity capability, and OpenAI has strengthened its safeguards accordingly.


That includes stricter classifiers, tighter controls around high-risk activity, additional protection against repeated abuse, and Trusted Access for Cyber, which gives verified defenders access to advanced cybersecurity capabilities with fewer restrictions.

Applications are now open.


OpenAI's approach increasingly resembles an enterprise-grade access model: stricter defaults, with a dedicated path for legitimate and compliant defensive use.


This framework is unlikely to remain limited to cybersecurity. Other high-risk fields with legitimate requirements will probably move in the same direction.


For powerful models to enter real industries,raw capability is not enough. They also need precise, accountable access controls.


Pricing


Finally, there is the question everyone asks: what does it cost?

GPT-5.5 API pricing is $5 per million input tokens and $30 per million output tokens.

GPT-5.4 costs $2.50 and $15 respectively.


The price has doubled.


GPT-5.5 Pro is substantially more expensive at $30 for input and $180 for output.


OpenAI argues that GPT-5.5 is both more capable and more token-efficient, with most Codex users consuming fewer tokens to complete the same task.


The reasoning makes sense, but every team still needs to run its own numbers.


• For complex, engineering-intensive workloads, fewer retries and less rework may keep the total cost commercially reasonable.

• For lightweight daily use, the higher unit price will be difficult to ignore.


Overall, GPT-5.5 marks a decisive return to form for OpenAI.


Competition is shifting to a more consequential layer: the companies that integrate AI into real workflows first will have the best chance of defining how the next generation of computing is used.


Benchmark scores prove capability. Agent-powered work is the much larger market.


A Final Thought


Looking back at the past week, new models arrived one after another at remarkable speed.


That familiar pressure—the kind that makes an entire industry sit up and move faster—is back.


Everyone is pushing forward.Leadership is never permanent, and neither is falling behind.


In the AI era, we should stay curious, keep creating, and remain open to what we do not yet know.

May these new capabilities help all of us turn ideas into outcomes people can see, use, and value.


Talk About Digital Transformation?

Talk About Digital Transformation?

Discuss digital transformation with GeekOnUp

一同向上生长

Growing Upward Together

Let's Talk About Your Ideas

  • Custom APP Development
  • Custom Mini Program Development
  • Custom Web Development
  • AI Agent Development
  • Enterprise Digital Transformation
  • Other