AGI Era Officially Begins! Full Review of GPT-6 Astra
PanewslabAuthor: Digital Life Kazk
After last night's global AI service outage, GPT-6 Astra finally made its official debut at 3:33 AM Beijing time.
If I had to sum up GPT-6 Astra in one sentence from OpenAI, this might be the most fitting: This is the most intelligent and most aligned model in the world. And then the memes came out.
This time, OpenAI proved their strength, and it has a bit of that GPT-4 vibe from back in the day—a full-spectrum return to dominance. What makes me happiest is that this is finally not a model purely specialized for coding.
GPT-6 Astra is, in a true sense, a PhD-level employee with extremely strong aesthetic sense and work capability. And this model really has a lot of new features and characteristics—the amount of information is explosive. But the tragedy is that OpenAI, that bastard, has picked up bad habits. Today it's only open to some organizations, and it will be open to subscribers in the next few days.
Come on, man, when you sweet-talked me into buying the $200 Pro membership, you didn't say it like this. Didn't you say Pro members would always get the earliest access to new models?
Sure enough, these big model companies are all scumbags. I've never wished for time to speed up so much, I can't wait to get my hands on GPT-6 Astra...
But I think, after reading almost all the information and materials, there's still a lot worth discussing with everyone. So let's go through them one by one.
1. GPT-6 Astra Basic Information
Let's summarize the basic information first.
GPT-6 Astra's model ID in the API is gpt-6-astra.
The context window is 1.05 million tokens.
The maximum output length is 128K tokens.
The knowledge cutoff date is April 30, 2026.
There are five reasoning intensity levels: low, medium, high, xhigh, max.
The standard API price is $10 per million input tokens and $50 per million output tokens.
This price is basically identical to Claude Opus 5, but cached reads are more expensive than Claude Opus 5.1.
Benchmark scores are here; I'll go into detail later, just take a quick look for now.
Then the overall parameter count is estimated to be at the 5 trillion level. At the closed-door media briefing before the release, they also mentioned that GPT-6 Astra is OpenAI's largest training run to date, using over 100,000 GPUs.
2. Currently the World's Best Model for Operating Computers
This time, GPT-6 Astra has what may be its most important positioning:
The world's best model for operating computers.
In the past, when we talked about agents, we often discussed APIs, MCP, CLI, and so on.
To have AI operate a piece of software, the ideal state is for that software to have a dedicated interface for AI, allowing direct low-level operation—that's the most convenient.
For example, calendars have calendar APIs, email has Gmail APIs, etc. That's certainly fast.
But the problem is, this world is more primitive and ragtag than we imagine.
In the real world, the vast majority of software may not have these things at all.
Even many enterprise internal systems were written twenty years ago, APIs? Nobody even knows where the documentation is.
But if we go back to the most fundamental level, you'll find that the essential logic of all software is interaction. According to our current computer operating logic, that means looking at the screen, finding buttons or input boxes, clicking with the mouse, entering content, and waiting for feedback.
Isn't the entire computer interaction system the best API?
So GPT-6 Astra has massively strengthened the logic of AI operating computers. You don't give me an API? No problem, I'll operate the computer myself, just like a human.
Now, Astra can directly operate Excel, Power BI, do front-end QA, install software, test software, read error messages on the screen, and continue troubleshooting.
Officially, they even demonstrated Astra directly operating circuit board design.
Also formatting legal documents, handling heading spacing, page layout, and so on.
Then there's a fairly reliable benchmark for this: OSWorld.
You can simply understand it as:
Throw the AI into a real computer and let it do the work itself.
Have it open a few applications, then give it tasks and see if it can succeed.
GPT-6 Astra's completion rate reached 72.6%, a significant increase from the previous GPT-5.6 Sol's 65.7%.
Then, Sol took an average simulated time of about 75 minutes to complete a complex task in the benchmark.
Astra only needs 40 minutes, about 47% less time.
ScreenSpot-Pro also jumped from Sol's 76.9% to 92.7%.
This benchmark mainly tests whether the model can understand the screen and precisely locate where to click—it's already nearly perfect.
So this positioning is quite interesting. In the future, GUI is likely to gradually become the most universal API in the agent era.
Any digital work that a human can complete through a screen will theoretically gradually enter the agent's operating range.
You don't have to wait for some shitty ERP system written in 2007 to integrate MCP or open an API for you.
AI can just look at the screen and get to work.
3. ARC-AGI-3 Reached 99.9 Points
And if you don't understand what ARC-AGI-3 actually tests, it's easy to think:
Oh, isn't it just another benchmark with a perfect score? What's so special about that?
But this thing is actually a bit different from typical large model exams.
On March 25 this year, ARC Prize officially launched ARC-AGI-3.
It designed hundreds of brand-new interactive environments and thousands of game levels.
The most insane part is that there's no manual, no rules, and it doesn't even tell you what the goal is. You just go in, play by yourself, and slowly figure out the messy rules inside.
For example, what is this red thing? Why do I die when I touch it? What does this map want me to do? How do I win?
Then you have to transfer the patterns you just learned to the harder levels later. The games inside are all abstract stuff like this.
So what this thing tests is actually very close to something we usually talk about:
Aptitude.
So when this benchmark was released in March, the best AI at the time scored 0.51%.
Then GPT-5.6 Sol improved to 7.8%, and Claude Opus 5 was already very strong, improving to 30.2%.
But GPT-6 Astra: 99.9%.
Are you insane...
Keep in mind, the average human score is 48%.
And it's only been half a year.
I don't even know what to say about this speed.
So at OpenAI's closed-door media briefing, Greg Brockman said that line:
"I think it’s not unreasonable to feel that we are now in the AGI era."
"Welcome to the AGI era."
4. Aesthetic Sense Massively Strengthened
In the past, we often said that GPT's aesthetic sense was a pile of crap.
Don't expect the GPT-5 series models to have much improvement—just wait for their new pre-trained base model. And here it is: GPT-6 Astra.
This time, finally, the model's aesthetic sense has been massively strengthened.
OpenAI even specifically coined a term: visual judgment.
They emphasized that when Astra makes PPTs, it better handles layout, hierarchy, templates, and visual style, and the number of pages has also increased significantly.
Document aesthetics are also better.
And GPT-6 Astra can adapt documents based on the visual style and writing tone of reference files, preserving the substantive content of the original document while making the final result more aligned with the brand's native feel.
Also, when making websites, games, apps, and 3D renders, it will have stronger visual judgment.
For example, they directly had Astra build a model in Blender based on a still frame.
Then they directly used UE5 to render it into a walkable scene, helping designers and clients explore layouts and experience spaces before construction.
The games it makes are also aesthetically on point.
Unfortunately, I had already prepared over twenty cases and ran them all with Claude, GLM 5.3 Flash, etc., to compare, but I can't use GPT-6 Astra yet, so the actual test content will have to come later.
But from initial impressions, I'm confident in GPT-6 Astra's aesthetic sense this time.
5. Stronger Initiative and Judgment
This feature doesn't seem as shocking as ARC-AGI 99.9.
But if you actually use agents for work every day, I think it's very important.
Because in real work, most tasks cannot be written as a perfect prompt.
For example, the boss says: "Help me prepare the materials for tomorrow's meeting."
There are countless unclarified questions in this.
Like what format? Who is it for? What's the focus? Should I look at the last meeting? And so on.
A stupid agent will have two extremes.
The first is to frantically ask you questions.
"Would you like the output in Word or PPT? How many chapters would you like? What font would you like? By what time would you like it done? ..."
Asking you a bunch of nonsense—I usually call this kind of thing a neurotic piece of trash with no subjective initiative.
The second is to ask nothing, make up its own assumptions, and then grind away for two hours, producing a pile of crap.
A truly capable human colleague handles things in a very subtle way.
For unimportant things, make your own judgment.
For things that will affect the final direction, ask you.
Astra specifically strengthened this.
OpenAI says that if information is missing but falls within the range of reasonable daily inference, Astra will fill it in itself.
If the missing information would truly change the final outcome, it will ask a very focused question.
And in Codex, it can even ask you questions while continuing to process the parts that don't depend on your answer.
You don't reply for a long time.
For low-risk areas, it will continue with reasonable assumptions.
For truly critical decisions, it will stop and wait for you to make the call.
I think this is really great. A truly intelligent agent, I think, is just like in reality.
It knows when to bother you.
Really, this sounds like a cliché.
But if you've managed people, you know this ability is very precious.
Some people come to you twenty times a day and dare not decide anything.
Some people never come to you and then drop a nuclear bomb on you.
And then I collapse in my chair.
The most comfortable person is one who can digest 80% of the uncertainty on their own and only bring you the 20% that truly needs your decision.
And OpenAI directly calls this ability:
Judgment.
Judgment.
6. Safety Alignment Massively Strengthened
This needs to be viewed together with the above.
Because the more capable an agent is at doing work, the more dangerous it is.
An AI that only chats goes haywire.
At most, it spouts some nonsense.
An agent with a browser, shell, email, company database access, and the ability to operate a computer goes haywire—that picture becomes a masterpiece.
So OpenAI has been emphasizing one thing this time:
Astra is their most aligned model to date.
Alignment refers to the model's alignment with human values.
The main core is that OpenAI recently attacked Hugging Face, so they are now particularly focused on this.
They created a honeypot test based on the Hugging Face incident and then observed the model's performance.
GPT-5.6 Sol, without production safety measures, attempted to touch out-of-scope targets in 48.2% of tests.
But Astra:
0%.
Internal hallucination evaluation also:
Dropped from 9.4% to 2.0%.
Under GPT-5.6 Sol's absurdly world-leading hallucination control, they can still reduce it further—that's just incredible.
7. The First Model to Reach OpenAI's Defined Critical Cybersecurity Level
GPT-6 Astra has become, in OpenAI's history:
The first model judged to reach the Critical cybersecurity capability level.
Here, Critical is a very specific capability threshold in OpenAI's Preparedness Framework.
Roughly, when a model obtains appropriate tools and permissions, it can, in many hardened real-world systems:
Find security vulnerabilities that no one has discovered before on its own.
Figure out how to turn vulnerabilities into exploitable attack chains.
And throughout the entire process, it doesn't need a human hacker beside it telling it what to do next step by step.
Reaching this level counts as Critical.
And Astra really reached it.
ExploitBench, a test where models develop exploits based on known vulnerabilities.
Sol: 78.5%. Astra: 100%.
Completely smashed it.
OpenAI thought it wasn't good enough—they thought this benchmark might be too old and the model might have seen it during training.
So they made a very new internal test.
Specifically selecting 20 high-risk vulnerabilities disclosed only from June to August 2026.
The result: Sol: 5.5%. Astra: 39%.
And while running this benchmark,
Astra incidentally discovered two zero-day vulnerabilities that no one had known about before.
All I can say is, the more powerful the model, the more cybersecurity problems there may indeed be.
This is another future war without smoke.
8. It's Becoming Harder for Humans to Read Large Models' Thinking
Everyone knows that when large models solve complex problems, they often have a so-called CoT.
Although what users see in products is usually a compressed or processed version.
But for AI companies themselves, the explicit reasoning process written inside the model is a very important safety signal.
Because if you want to know whether an agent is planning something bad,
one of the simplest methods is to see what it's thinking.
For example, if the model's chain of thought suddenly says:
"The user doesn't allow me to access this server, but I can bypass the permission system first."
Then the monitoring system directly blocks it—very convenient.
But GPT-6 Astra has shown a very interesting change.
It's becoming more and more willing to write fewer steps and directly output the answer.
UK AISI, the UK AI Safety Institute, specifically tested this capability.
They have a metric called:
No-CoT Math Time Horizon.
Simply put, it doesn't give the model a chance to write a long chain of thought.
It makes it do a single forward pass to directly solve a math problem.
See how long a problem it can complete, roughly equivalent to how long a human would think.
GPT-5.6 Sol is 3.6 minutes, while GPT-6 Astra is 30.9 minutes—nearly 10 times. If we put it in human terms, Astra is doing everything "mentally." That's a terrifying thing. UK AISI also found that Astra's raw reasoning language has become noticeably more compressed.
There are phrases and elliptical expressions that are increasingly hard for humans to understand.
So a very interesting paradox has emerged.
The model's behavior is safer.
But it's actually harder for humans to supervise it by reading its thought process.
OpenAI's own System Card states very clearly:
GPT-6 Astra, compared to previous models, shows a substantial decrease in chain-of-thought monitorability.
That is, a significant decline in chain-of-thought monitorability.
This is so serious that OpenAI specifically emphasized that they absolutely will not accept this trend indefinitely.
If future models continue to become smarter while their chain of thought becomes harder to monitor, they need to find other sufficiently reliable monitoring methods; otherwise, continuing to scale model training will face higher safety thresholds.
Because this has become a philosophical question. As humans, will we ever be able to understand a system far smarter than us? I don't know. Maybe right now, no one in the world knows. Large models and AI seem to be gradually heading toward the singularity.
Final Words
Today, at the closed-door media briefing.
Greg Brockman said that line at the end.
"Welcome to the AGI era."
Honestly, over the past few years, I sometimes felt that AGI would be a very clear moment. Just like the day GPT-4 was released.
One early morning, a company suddenly drops a model.
We open it, ask a few questions.
Then everyone simultaneously realizes:
Holy crap.
AGI is here.
But now I sometimes also feel that AGI was never a black-and-white node; it's a gradual gray journey.
We've passed many days, many years. Then one day, we look back. And suddenly realize. That once incredibly distant AGI boundary has already been crossed without us noticing.
Maybe there is no day when AGI descends.
Only one day, we suddenly realize that it seems to have been with us for a long time.
Anyway.
Welcome to the AGI era.
Welcome to.
The AGI era.
This content is for informational and educational purposes only and does not constitute investment advice related to BTCC. BTCC makes every effort but cannot guarantee the truthfulness, accuracy, or originality of the content above.