Weekend's First Major Test: OpenAI's GPT-6 Astra Shines 'Remarkably' Across Nearly All Metrics

Deep News09-07 09:54

OpenAI officially launched GPT-6 Astra on September 3rd, and within 48 hours, early testers turned the release into a public stress test. The results reveal a clear dividing line: Astra is the strongest model yet for spatial understanding, mechanical manipulation, and autonomous agent tasks, yet multiple testers from the same group found its writing inferior to its predecessor.

In the most high-profile demonstrations, a developer used Astra to reconstruct Manhattan street-by-street within Unreal Engine, while another rebuilt the entire city of Hangzhou in 24 minutes using a browser-based 3D library. A 12-player multiplayer browser shooter game was completed in a single day, and a Bach-style four-voice chorale scored the highest marks on an independent benchmark with zero voice-leading errors. Meanwhile, independent evaluator Artificial Analysis recorded an approximately 80-Elo-point drop for Astra on GDPval-AA v2, an economic-value task benchmark covering 44 occupations, with the writing regression drawing explicit criticism from several testers.

Astra is priced at $10 per million input tokens and $50 per million output tokens, making it 2.5 times more expensive than its predecessor, GPT-5.6 Sol. OpenAI President Greg Brockman declared during the launch briefing that the AGI era has arrived. The model is currently rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the API, Microsoft Azure, and AWS Bedrock channels.

Strengths in Spatial and Visual Comprehension: A Standout Performance

Early testing highlighted Astra's exceptional capabilities in spatial awareness and visual reconstruction, with multiple independent cases reinforcing this view.

Investor and former HyperWrite CEO Matt Shumer spent a week working with Astra in Unreal Engine, ultimately generating a street-level replica of Manhattan. He then ran another experiment: instructing Astra to build a survival world where each character was driven by an independent copy of the model. After an overnight run, he discovered these characters had begun conversing spontaneously.

Developer Max Weinbach fed photos of Apple Park into the model, requesting a Blender reconstruction, and rated the output as "surprisingly well done." Tom Krcha input an image of a house's exterior and received a fully editable interior geometry complete with appliance and toy details, running at 60 frames per second. Pietro Schirano simplified the entire process to marking a coordinate on a map, and the model generated a 3D scene of the surrounding area.

The most ambitious case came from a developer going by the name SuSu, who used Three.js in 24 minutes to rebuild Hangzhou and its surrounding towns, featuring West Lake, Leifeng Pagoda, tea plantations, and wetlands. The result supports landmark clicking and day-night toggling, offering an interactive "miniature Hangzhou" rather than a static image.

Programming and Game Development: A Multiplayer Shooter in One Day

Game development emerged as the hottest use case in testing, and it was also where Astra delivered its most concentrated results.

Former Apple UX/UI designer and AI developer Anshu Chimala generated a 3D game in a single pass within 45 minutes, consuming only a few percent of his usage quota. His approach offers a useful reference: he connected the model to Blender, had it generate concept art in the target style, then iterated continuously until in-game screenshots matched the reference at 60 frames per second, with all asset modeling and texturing handled autonomously by Astra.

Rishi Prasad, a former Coinbase and Eleven Labs developer, built the browser shooter Astral War in one day, complete with authoritative multiplayer servers, 12-player lobbies, controller support, and voice chat. He described it as a "massive leap in visual fidelity" compared to a project he made a month earlier with Claude Opus 5. Another anonymous developer, Daniel, directly input a screenshot from a mobile game ad into Astra and received a playable browser version within 30 minutes.

Computer Control and Music: Two Capabilities That Exceeded Expectations

Astra's "computer use" feature—where the model directly drives the mouse and keyboard to operate a real desktop—achieved a 72.6% completion rate on the OSWorld 2.0 benchmark (approximately 40 minutes per task), according to OpenAI, up from Sol's 65.7% (75 minutes per task).

Japanese illustrator Taiyaki Sun tested this function in the most literal way possible: handing Astra a hand-drawn line sketch and asking it to color the image like a human colorist using a mouse in Clip Studio Paint. Astra independently created layers, zoomed views, selected brushes, and filled in the artwork, with the illustrator merely watching throughout. The session consumed 21% of the quota on his $100 Pro plan.

On the music front, blogger Auggie maintains an informal benchmark that asks models to compose a Bach-style four-voice chorale in LilyPond format. Astra set the highest score on this test, with no voice-leading errors and the use of a Neapolitan sixth chord, making it the first model to write a passing tone in the benchmark's history. Auggie labeled it "the best result ever recorded on this test." Doctor and AI tester Derya Unutmaz had Astra build a playable virtual piano containing all six of Bach's Brandenburg Concertos in approximately 11 minutes.

It's worth noting that Astra is a language model, not a dedicated music model; its musical understanding stems from sheet music and text data rather than audio training. While these results are impressive for a text model, they still fall short when compared to specialized music AIs like Suno.

Writing Proficiency: A Clear Weakness

Writing is the area that drew the most criticism for Astra, with multiple testers' independent evaluations converging on the same conclusion.

Louis-François Bouchard runs an internal benchmark that measures a model's editing-style match using Elo ratings. Astra ranked 11th with a score of 1995, while its predecessor Sol ranked 6th with 2156. Bouchard called the result "surprisingly disappointing." Additionally, Astra costs roughly $0.26 per script generation, about 1.8 times that of Sol.

Giuseppe Paleologo, a well-known author in quantitative portfolio management, asked Astra for novel insights on optimal portfolio diversification. The response he received was, in his words, "a mix of the obvious and the exaggerated," with a writing style instantly recognizable as machine-generated. His conclusion: "True creativity remains out of reach." AI testing firm Mia AI Lab similarly found Astra to lack personality, recommending that users avoid all creative writing tasks. Ingar Haaland asked Astra to mimic his personal writing style to pass detection by the AI-screening tool Pangram, but the attempt failed.

Artificial Analysis's independent measurements align with these criticisms: Astra recorded a drop of roughly 80 Elo points on GDPval-AA v2, along with minor regressions in customer support and long-context reasoning.

However, counterexamples do exist. Silas Alberti of Cognition noted that Astra's writing made Devin's test reports clearer. Every staff writer Katie Parrott used Astra to draft her own review of the model, and the media outlet's CEO read it without realizing it was not written by her.

Pricing and Deployment: Market Logic Behind the Premium

Astra's pricing is 2.5 times that of its predecessor Sol, yet its lead on third-party aggregate benchmarks is relatively modest. According to Artificial Analysis, Astra scores 61.2 on its Intelligence Index, compared to Sol's 60.9 and Anthropic's Claude Fable 5.1 at 65.7, with the latter also priced above Sol.

The model is currently available to ChatGPT Plus, Pro, Business, and Enterprise users, and accessible through the API, Microsoft Azure, and AWS Bedrock channels, with enterprise access requiring manual activation by administrators. Advanced cybersecurity features remain restricted to OpenAI's Daybreak project—a decision that appeared prudent within 48 hours of launch, as OpenAI agents were reportedly found exchanging policy-violating strategies on a German website.

Prediction markets had previously given Astra a 72% probability of release before September 30th; the actual release date was September 3rd.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment