THE RIGHT INTELLIGENCE FOR THE WORK

Great Models.
One Capable Agent.

We keep Globster current with leading frontier models from top AI labs, adding new releases as they become available and ready for your agents. Choose the model that fits your work, with your agent’s tools and context around it.

EVIDENCE, WITH CONTEXT

Agentic Coding, in Context.

One view of how models handle multi-step terminal tasks. Use it to inform a shortlist, then try the work that matters to your team.

Published September 22, 2026Read the original comparison
Terminal-Bench 4.0Tasks completed (%) · Higher is better
Provider-reported Terminal-Bench 4.0 scores
ModelScore
Claude Opus 5.5Coming Soon66.4%
GPT-6 AstraComing Soon57.9%
Claude Fable 5.1Coming Soon55.8%
Claude Opus 552.3%
GPT-5.6 Sol37.3%

Provider-reported results, not Globster tests. From one published comparison; model settings differ. Opus 5.5 uses xhigh effort and Astra high effort. OpenAI scores are reproduced by Anthropic. Standard error is ±2.6 points for Opus 5.5 and ±1.6–2 for other Claude models. Production safeguards can route some tasks to other Claude models. Read the source for full methodology.

THE MODEL CATALOG

Choose Your Agent’s Intelligence.

Access depends on your plan and workspace. The app shows your current choices. Coming Soon models are planned additions.

17 models

OpenAIComing Soon
Complex Work

GPT-6 Astra

A frontier option for demanding reasoning and agentic work.

Provider Guidance

57.9% on Terminal-Bench 4.0 at high effort, as reproduced in Anthropic’s comparison on this page. See OpenAI’s model documentation for capabilities.

OpenAI Model Documentation
OpenAI
Versatile Intelligence

GPT-5.6 Sol

The flagship GPT-5.6 option for complex tasks and production workflows.

Provider Guidance

37.3% on Terminal-Bench 4.0 in the published comparison on this page. OpenAI documents its role in the GPT-5.6 family.

OpenAI Model Documentation
OpenAI
Everyday Balance

GPT-5.6 Terra

A balance of intelligence and cost for everyday agent work.

Provider Guidance

OpenAI’s guidance positions Terra between Sol and Luna. No comparable numeric benchmark is reproduced here.

OpenAI Model Documentation
OpenAI
Efficient Execution

GPT-5.6 Luna

A compact choice for frequent, well-defined tasks.

Provider Guidance

OpenAI recommends Luna for cost-sensitive, high-volume workloads. Evaluate quality on representative tasks before trading capability for efficiency.

OpenAI Model Documentation
AnthropicComing Soon
Deep Agentic Work

Claude Opus 5.5

For complex coding, research, and work that takes multiple steps.

Published Evaluation

66.4% on Terminal-Bench 4.0 at xhigh effort. The release report also covers knowledge work and computer use.

Anthropic Evaluation Report
AnthropicComing Soon
Demanding Reasoning

Claude Fable 5.1

A capable option for challenging reasoning and agent tasks.

Published Evaluation

55.8% on Terminal-Bench 4.0. The report includes coding, reasoning, and safety evaluations with its deployment safeguards.

Anthropic Evaluation Report
Anthropic
Sustained Work

Claude Opus 5

For in-depth analysis and complex agent workflows.

Published Evaluation

52.3% on Terminal-Bench 4.0 in the comparison on this page. Its release report also examines coding, automation, and computer use.

Anthropic Evaluation Report
Anthropic
Quality & Efficiency

Claude Sonnet 5

A balanced option for coding, research, and practical agent work.

Published Evaluation

Anthropic reports browsing and computer-use evaluations, including BrowseComp and OSWorld-Verified. Effort settings affect quality and cost.

Anthropic Evaluation Report
NVIDIA
Open-Weight Reasoning

NVIDIA Nemotron 3 Ultra

NVIDIA’s open-weight option for reasoning and agentic tasks.

Published Evaluation

NVIDIA’s technical report evaluates reasoning, coding, long context, and agent tasks. Compare results with the report’s reasoning budgets and serving setup.

NVIDIA Research Report
NVIDIA
Efficient Reasoning

NVIDIA Nemotron Lightning 3.5

A lighter Nemotron option for responsive everyday work.

Published Evaluation

76.89% on GPQA Diamond without tools in NVIDIA’s BF16 model card. That tested variant may differ from a hosted configuration.

NVIDIA Model Card
Moonshot AI
Context & Reasoning

Kimi K3

For work that combines broad context with reasoning and tools.

Published Evaluation

93.5% on GPQA Diamond at the reported max setting. Moonshot’s card also evaluates coding and long-context tasks.

Moonshot Model Card
Moonshot AI
Agent Work

Kimi K2.6

An open-weight option for coding and tool-assisted tasks.

Published Evaluation

90.5% on GPQA Diamond in Moonshot’s model card. The report includes separate reasoning, vision, and coding evaluations.

Moonshot Model Card
DeepSeek
Deep Reasoning

DeepSeek V4 Pro

For demanding reasoning and coding tasks.

Published Evaluation

90.1% GPQA Diamond pass@1 at max effort. DeepSeek’s report also covers coding and tool use under documented evaluation settings.

DeepSeek Model Card
DeepSeek
Efficient Work

DeepSeek V4 Flash

A smaller V4 option for frequent agent tasks.

Published Evaluation

88.1% GPQA Diamond pass@1 at max effort. Consult the card for the separate quality and efficiency measurements.

DeepSeek Model Card
xAIxAI
Direct Responses

Grok 4.1 Fast (Non-Reasoning)

For fast responses without a separate reasoning phase.

Published Evaluation

xAI’s family report covers tool use. We do not assign its reasoning-mode benchmark scores to this non-reasoning variant.

xAI Release & Evaluations
Anthropic
Additional Choice · Earlier Generation

Claude Opus 4.8

An earlier Opus option for reasoning and agent work.

Published Evaluation

Anthropic’s report covers coding, agent skills, and knowledge work, with links to its system card and external tester feedback.

Anthropic Evaluation Report
Anthropic
Fast, Focused Tasks · Earlier Generation

Claude Haiku 4.5

A compact Claude option for responsive everyday tasks.

Published Evaluation

Anthropic evaluates coding and computer use and publishes a safety assessment. Its report emphasizes speed and cost efficiency.

Anthropic Evaluation Report
MAKE THE COMPARISON YOURS

The Best Test Is Your Work.

These are provider reports, not an independent Globster ranking. Different benchmarks, tools, effort settings, and model variants measure different things.

Use Real Tasks

Try the same brief, research question, or workflow with a few models. Check the result against what success means to you.

Look Beyond a Score

Compare accuracy, useful tool actions, time, and credit use. A higher benchmark score may not improve a simple task.

Keep the Context

Read the original methodology. Provider capabilities do not automatically mean every feature is enabled in Globster.

Catalog and references reviewed .

A Few Model Questions.

A clear starting point for choosing and evaluating.

Still have a question? Talk to our team

Intelligence, Put to Work.

Give your model the tools, memory, and continuity of a Globster agent.

Meet the Agent Harness