Claude Opus 5.5 vs GPT-6 Sol: Creating a Ferrari F1 Car in Blender
Claude Opus 5.5 arrived on September 22, 2026, promising Fable 5.1-level performance at 40% lower cost than Opus 5. GPT-6 Sol had already landed as part of September’s frontier-model wave. Both are positioned as strong coding and creative models, but benchmark tables only tell part of the story.
A harder test is to give both models the same complex task and see what they can actually produce.
For this comparison, both models were asked to recreate the Ferrari SF-26 Formula 1 car in Blender, complete with animations. The task combines coding, 3D modeling, visual judgment, iteration, and tool use, which makes it a much more demanding test than a standard coding benchmark.
To keep the comparison fair, both models ran at xhigh effort and received the same prompts in the same order. The results were then compared across three practical measures: output quality, total cost, and wall-clock time.
That gives a clearer view of where Claude Opus 5.5 and GPT-6 Sol actually differ when the task is long, visual, and technically demanding.
The setup
Opus 5.5 ran through Claude Code v2.1.288. GPT-6 Sol ran through the Codex terminal interface.
The effort setting matters here. Both models were set to xhigh, which instructs the model to spend more compute and reasoning on the task. For a complex multi-step project like 3D modeling and animation, lower effort tends to produce rushed or incomplete output. Both models ran three sequential prompts:
- Build a high-detail 3D model of the Ferrari 2026 F1 car in Blender using web references
- Create a build animation showing the car being assembled from parts
- Create a pit stop animation
The "F2" correction
The initial prompt intentionally included a typo: "ferrari 2026 f2 car." Ferrari doesn't build a Formula 2 car. Both models caught this through their web research phase and handled it differently.
Opus 5.5 corrected directly: "Since Ferrari doesn't build an F2 car, I'll go with the Ferrari SF-26, their 2026 F1 car."
GPT-6 Sol took a more literal path: "The reference check shows that F2 uses a spec Dallara chassis; Ferrari's SF-26 is an F1 car. I'm treating your request as a Ferrari-styled 2026 F2 concept, using the current F2 proportions and Ferrari's 2026 red, white, and carbon look."
The difference is subtle but meaningful. Opus assumed the user wanted the real Ferrari F1 car and corrected accordingly. Sol tried to honor the literal "F2" request while adapting it to something plausible. Both are reasonable interpretations; they just show different default assumptions about user intent.
Both models then ran web searches for "Ferrari SF-26 2026 Formula 1 car launch design details," fetched content from Wikipedia and formula1.com, and downloaded official launch photos covering multiple angles before generating any code.
Opus 5.5: build and pit stop animations
Opus 5.5's build animation used a blueprint-style aesthetic. Parts appeared piece by piece in the correct assembly order, with each component rendered in a technical drawing style before the full car materialized. The effect was visually striking and showed genuine understanding of the car's mechanical structure.
The pit stop animation introduced motion blur as the car entered the frame, giving it a sense of speed appropriate for a motorsport setting. The car stopped in the pit box, the wheel changes animated correctly, and the camera work was cinematic. The one oddity was the old tires flying upward unrealistically on removal, but this fulfilled the prompt literally while missing the physics.
Overall Opus 5.5 produced a polished, professional result across all three prompts with consistent visual quality.
GPT-6 Sol: same prompts, different results
GPT-6 Sol completed all three tasks but with significantly less polish.
The assembly animation was disjointed: parts appeared in less organized fashion with gaps in the geometry where suspension arms didn't connect cleanly to the chassis. The overall model felt generic, closer to a Formula 3 car than an F1 machine in terms of detail and proportion. Logo placement had errors, including the IBM logo on the rear wing appearing misoriented.
The pit stop animation had a specific glitch: the replacement tires floated alongside the car as it drove into the pit box rather than being staged correctly. A tire also randomly appeared at the bottom of the frame mid-sequence. The animation lacked the motion blur and cinematic composition of Opus 5.5's version.
Sol understood the prompts and executed them at a functional level. The gap to Opus 5.5 was quality, not capability.
Expanding to four models
GPT-6 Astra and Fable 5.1 ran the same three prompts.
GPT-6 Astra produced arguably the most realistic static model of the four. The car's proportions and front wing geometry were more accurate than Opus 5.5's, where some areas veered toward a stylized rather than literal recreation. Astra's pit stop animation was the strongest in the comparison: it correctly showed the car being raised on jacks before the tire change, a mechanical detail no other model included. Astra cost $32.47 and took 1 hour 49 minutes, making it the most expensive and slowest of the four.
Fable 5.1 produced a competent car with some visual artifacts and a less impressive livery. The build animation was less distinctive than Opus's blueprint approach. The pit stop animation included a nice realism touch: the car's suspension compressed correctly as it accelerated out of the pit box. Fable 5.1 cost $42.44 and took 1 hour 13 minutes, making it the most expensive model in the test despite not producing the top result.
Cost and performance
| Model | Cost | Time | Output tokens |
|---|---|---|---|
| GPT-6 Sol | $4.50 | 48 min 44 sec | 57K |
| Claude Opus 5.5 | $27.86 | 1 hr 7 min | 329K |
| GPT-6 Astra | $32.47 | 1 hr 49 min | — |
| Claude Fable 5.1 | $42.44 | 1 hr 13 min | — |
Sol's cost advantage is real: $4.50 versus $27.86 for Opus 5.5. The token count explains the gap. Opus 5.5 generated 329K output tokens versus Sol's 57K. It was writing substantially more code, spending more computation on geometry, and investing more in the aesthetic quality of the output. The lower output quality from Sol reflects a genuinely shorter thinking process, not just a pricing difference.
What this test shows
For this specific class of task, creative 3D generation with animations requiring spatial reasoning and mechanical understanding, the ranking came out:
Top tier: Opus 5.5 and GPT-6 Astra, very close. Opus 5.5 produced more polished visual output and a more striking build animation. Astra produced more geometrically accurate proportions and a mechanically superior pit stop animation. Which one wins depends on what you're optimizing for.
Mid tier: Fable 5.1. Competent but not exceptional, and notably the most expensive despite not producing the best result. Anthropic's own positioning of Opus 5.5 as performing at Fable 5.1 level at 40% lower cost is consistent with what this test found.
Efficiency tier: GPT-6 Sol. The weakest visual quality but by far the fastest and cheapest. A legitimate choice for rapid iteration and prototyping where you need to explore directions quickly before committing to a high-quality final render.
Opus 5.5 is available on the Claude API as claude-opus-5-5, with a 1M token context window, 128K maximum output tokens, and zero data retention. It's also on AWS Bedrock, Google Cloud, and Azure. Sonnet 5.5 followed on September 28 as the next model in the 5.5 family.