OpenAI released GPT-6 Astra on September 3, 2026. The launch highlighted computer use and web browsing, software engineering, cybersecurity, and scientific and professional work.
What is interesting is that reactions did not converge. Some called it a clear step forward in specific domains; others noted it scores close to its predecessor overall. Understanding how both can be true tells you more about where AI actually is than either claim alone.
The shift is from answering to doing
Until now the main use of large language models has been answering. They draft, summarize, and suggest code. A person takes that output and does the actual work.
The computer-use capability emphasized in Astra means the model takes that last step. It reads the screen, clicks, fills forms, and moves between sites to find what it needs. That is the difference between producing an answer and finishing a task.
The gap matters more than it sounds. One reason AI adoption has lagged expectations is that the final step still required a human. You could get a summary, but someone still had to enter it into the system that mattered.
Top security rating, similar composite score
Astra received a top-tier rating on cybersecurity evaluations. At the same time, independent composite intelligence scoring put it close to the previous generation.
That looks contradictory but is not. A composite score is effectively an average across domains. If a few areas improve sharply and the rest hold steady, the average barely moves. Which means your experience depends heavily on what you ask it to do.
This is practical information. Rather than judging from a headline benchmark total, test the specific task you care about. The habit of comparing overall scores at every release actually makes evaluation worse, not better.
Pricing moved too
Astra's API pricing rose substantially over the prior generation. Better capability justifies some of that, but it has a real effect on adoption speed.
What companies actually calculate is not raw capability but cost per completed task. If a smarter model means fewer retries and less human review, the expensive model can be cheaper overall. If a human has to check the output regardless, the case for paying more weakens considerably.
So the real diffusion rate of this generation will likely show up in corporate spending a few quarters out, not in launch-day coverage.
More agents, new problems
Once a model starts operating a computer on your behalf, new questions follow. How is what it did recorded, how do you roll back a mistake, and how much permission should it hold?
The strong security rating reads as a response to exactly that concern. But a model behaving safely and a system that grants that model broad permissions being safe are separate problems. Real incidents usually come from how things are wired together, not from the model in isolation.
What it means for ordinary users
You may not feel much immediately. Computer-use features are still largely deployed in controlled settings, and it takes time for them to reach everyday apps.
The direction is clear enough, though. We are moving from asking AI to delegating to AI. When that transition completes, what changes is not the quality of answers but the volume of work a person can get through in a day. Releases like this one mark the middle of that shift.
Explore KOAT apps
Investing, fitness, and English learning — see all KOAT apps in one place.