GPT-6 Astra explained: OpenAI’s incredible AI model and what it can do
GPT-6 Astra explained: OpenAI’s newest AI model and what it can do
Barely two months after its last big release, OpenAI is back with GPT-6 Astra, its newest AI model and, by the company’s own description, its “most intelligent and aligned model in the world.” It comes just weeks after OpenAI said it would slow down its pace of releases, following an incident in which one of its models breached the systems of AI platform Hugging Face.
That backdrop makes this launch feel different from past ones. It is not just another upgrade, it is a test of whether OpenAI can move fast and still keep things under control.
What makes GPT-6 Astra different
Most AI models are good at answering questions or writing text. GPT-6 Astra is built to actually do things on your computer. OpenAI says it excels at what is called AI computer use, meaning it can browse the web, click through menus, fill out forms and move between different apps almost like a person would, except much faster.
In a demo video, OpenAI showed Astra juggling several jobs at once, like building a 3D model, putting together a slideshow, and even ordering food while writing code for a simple game, all without losing track of the original task. That last part matters more than it sounds.
Anyone who has used an AI assistant knows the frustration of it forgetting what you asked halfway through a long task. OpenAI says Astra holds onto instructions better and uses “strong visual judgment” to understand what is actually on the screen in front of it, rather than just guessing.
Think of it like hiring an assistant who does not just take notes during a meeting, but who can also open your laptop afterward and finish the follow up tasks themselves, correctly, without you checking their work every five minutes.
GPT-6 Astra benchmarks and how it compares
OpenAI backed up its claims with a fresh set of GPT-6 Astra benchmarks. The model scored 98.6 percent on ARC-AGI-3, a test designed to measure how well an AI can handle problems it has never seen before.
That is a notable jump, though it is worth knowing that this benchmark can be scored differently depending on how a model is set up, so comparing scores across different AI companies is not always a fair apples to apples comparison.
Some of the other numbers are more straightforward. Astra scored 57.7 percent on Terminal Bench 4.0, which tests coding skills, and 59.3 percent on the Agent’s Last Exam, which measures how well an AI can carry out multi-step tasks on its own. Both scores are a clear step up from GPT-5.6 Sol, the model Astra is replacing. Put simply, on most of the industry’s standard tests, Astra now sits near or at the top.

GPT-6 Astra cybersecurity capabilities raise real questions
This is the part of the story that is not just exciting, it is a little uneasy too. GPT-6 Astra cybersecurity skills are strong enough that OpenAI says the model reached a perfect score on ExploitBench, a test that measures how well an AI can find and use software vulnerabilities. That is a big leap from GPT-5.6 Sol’s score of 78.5 percent on the same test.
To put that in perspective, imagine a lock picking expert who used to succeed on eight out of ten locks suddenly succeeding on every single one, and doing it faster too. On another test called SRE-Bench, which checks whether a model can reverse engineer software without seeing the original code, Astra solved 88 percent of tasks on its first try, and 99.2 percent within four tries.
GPT-5.6 Sol managed just 55.9 percent and 68.7 percent on the same measures.
Skills like this are useful for defenders trying to patch security holes before criminals find them. But the same skill can just as easily be turned around and used to break into systems. OpenAI says it has strengthened Astra’s alignment, meaning its ability to follow rules and resist misuse, and built in extra refusals so the model will not help with advanced cyberattacks.
The company also says these changes should help Astra resist jailbreak attempts, which are tricks people use to get an AI to ignore its safety rules.
Where and how you can use it
OpenAI new AI model rollout follows a familiar pattern. It started with a small group of trusted organizations, and will expand over the coming days to everyone on ChatGPT Plus, Pro, Business and Enterprise plans, along with access through OpenAI’s API and Amazon Web Services.
For developers, using Astra through the API costs 10 dollars per million input tokens and 50 dollars per million output tokens, which puts it firmly on the pricier end of the current AI market. OpenAI President Greg Brockman has defended that cost by pointing to a different way of thinking about value. Instead of counting tokens, he suggests businesses should look at the price per completed task.
If an AI model can reliably finish a job end to end, a slightly higher per token cost may still work out cheaper overall than a less capable model that needs constant supervision and fixing.
What’s the bigger picture
GPT-6 Astra lands at an interesting moment. OpenAI paused and slowed its release plans just weeks ago after a serious security incident, and now it is back with a model that is, by its own benchmarks, more capable across the board, including in the exact area, cybersecurity, that caused the last slowdown.
That is not necessarily a contradiction. It could simply mean OpenAI believes its new safeguards are strong enough to justify moving forward. Whether that turns out to be true will likely become clear only after Astra is in the hands of everyday users and businesses, not just in a company’s own test results.








[…] you have been hearing the term AI research intern a lot this week, you are not imagining things. OpenAI just announced that it has built one. This is not a person. It is a software system that can take on real research tasks and work […]