GPT-6 Astra
A frontier option for demanding reasoning and agentic work.
57.9% on Terminal-Bench 4.0 at high effort, as reproduced in Anthropic’s comparison on this page. See OpenAI’s model documentation for capabilities.
We keep Globster current with leading frontier models from top AI labs, adding new releases as they become available and ready for your agents. Choose the model that fits your work, with your agent’s tools and context around it.
One view of how models handle multi-step terminal tasks. Use it to inform a shortlist, then try the work that matters to your team.
Published September 22, 2026Read the original comparison| Model | Score |
|---|---|
| Claude Opus 5.5Coming Soon | 66.4% |
| GPT-6 AstraComing Soon | 57.9% |
| Claude Fable 5.1Coming Soon | 55.8% |
| Claude Opus 5 | 52.3% |
| GPT-5.6 Sol | 37.3% |
Provider-reported results, not Globster tests. From one published comparison; model settings differ. Opus 5.5 uses xhigh effort and Astra high effort. OpenAI scores are reproduced by Anthropic. Standard error is ±2.6 points for Opus 5.5 and ±1.6–2 for other Claude models. Production safeguards can route some tasks to other Claude models. Read the source for full methodology.
Access depends on your plan and workspace. The app shows your current choices. Coming Soon models are planned additions.
17 models
A frontier option for demanding reasoning and agentic work.
57.9% on Terminal-Bench 4.0 at high effort, as reproduced in Anthropic’s comparison on this page. See OpenAI’s model documentation for capabilities.
The flagship GPT-5.6 option for complex tasks and production workflows.
37.3% on Terminal-Bench 4.0 in the published comparison on this page. OpenAI documents its role in the GPT-5.6 family.
A balance of intelligence and cost for everyday agent work.
OpenAI’s guidance positions Terra between Sol and Luna. No comparable numeric benchmark is reproduced here.
A compact choice for frequent, well-defined tasks.
OpenAI recommends Luna for cost-sensitive, high-volume workloads. Evaluate quality on representative tasks before trading capability for efficiency.
For complex coding, research, and work that takes multiple steps.
66.4% on Terminal-Bench 4.0 at xhigh effort. The release report also covers knowledge work and computer use.
A capable option for challenging reasoning and agent tasks.
55.8% on Terminal-Bench 4.0. The report includes coding, reasoning, and safety evaluations with its deployment safeguards.
For in-depth analysis and complex agent workflows.
52.3% on Terminal-Bench 4.0 in the comparison on this page. Its release report also examines coding, automation, and computer use.
A balanced option for coding, research, and practical agent work.
Anthropic reports browsing and computer-use evaluations, including BrowseComp and OSWorld-Verified. Effort settings affect quality and cost.
NVIDIA’s open-weight option for reasoning and agentic tasks.
NVIDIA’s technical report evaluates reasoning, coding, long context, and agent tasks. Compare results with the report’s reasoning budgets and serving setup.
A lighter Nemotron option for responsive everyday work.
76.89% on GPQA Diamond without tools in NVIDIA’s BF16 model card. That tested variant may differ from a hosted configuration.
For work that combines broad context with reasoning and tools.
93.5% on GPQA Diamond at the reported max setting. Moonshot’s card also evaluates coding and long-context tasks.
An open-weight option for coding and tool-assisted tasks.
90.5% on GPQA Diamond in Moonshot’s model card. The report includes separate reasoning, vision, and coding evaluations.
For demanding reasoning and coding tasks.
90.1% GPQA Diamond pass@1 at max effort. DeepSeek’s report also covers coding and tool use under documented evaluation settings.
A smaller V4 option for frequent agent tasks.
88.1% GPQA Diamond pass@1 at max effort. Consult the card for the separate quality and efficiency measurements.
For fast responses without a separate reasoning phase.
xAI’s family report covers tool use. We do not assign its reasoning-mode benchmark scores to this non-reasoning variant.
An earlier Opus option for reasoning and agent work.
Anthropic’s report covers coding, agent skills, and knowledge work, with links to its system card and external tester feedback.
A compact Claude option for responsive everyday tasks.
Anthropic evaluates coding and computer use and publishes a safety assessment. Its report emphasizes speed and cost efficiency.
These are provider reports, not an independent Globster ranking. Different benchmarks, tools, effort settings, and model variants measure different things.
Try the same brief, research question, or workflow with a few models. Check the result against what success means to you.
Compare accuracy, useful tool actions, time, and credit use. A higher benchmark score may not improve a simple task.
Read the original methodology. Provider capabilities do not automatically mean every feature is enabled in Globster.
Catalog and references reviewed .
A clear starting point for choosing and evaluating.
Still have a question? Talk to our team
Give your model the tools, memory, and continuity of a Globster agent.
Meet the Agent Harness