# Globster Models and Published Evaluations

Globster is by monday.com.

Stay current with leading frontier models from top AI labs. Explore Globster’s model lineup, compare published evaluations, and choose the intelligence for your agents.

We keep Globster current with leading frontier models from top AI labs, adding new releases as they become available and ready for your agents.

Reviewed 2026-09-28. Access depends on plan and workspace configuration. Coming Soon models are planned additions, not currently available in Globster.

## Terminal-Bench 4.0

- Claude Opus 5.5 (Coming Soon): 66.4%
- GPT-6 Astra (Coming Soon): 57.9%
- Claude Fable 5.1 (Coming Soon): 55.8%
- Claude Opus 5: 52.3%
- GPT-5.6 Sol: 37.3%

Provider-reported results, not Globster tests. From one published comparison; model settings differ. Opus 5.5 uses xhigh effort and Astra high effort. OpenAI scores are reproduced by Anthropic. Standard error is ±2.6 points for Opus 5.5 and ±1.6–2 for other Claude models. Production safeguards can route some tasks to other Claude models. Read the source for full methodology.

[Original comparison, September 22, 2026](https://www.anthropic.com/claude-opus-5-5)

## GPT-6 Astra (Coming Soon)

OpenAI. Complex Work.

A frontier option for demanding reasoning and agentic work.

57.9% on Terminal-Bench 4.0 at high effort, as reproduced in Anthropic’s comparison on this page. See OpenAI’s model documentation for capabilities.

[OpenAI Model Documentation](https://developers.openai.com/api/docs/models/gpt-6-astra)

## GPT-5.6 Sol

OpenAI. Versatile Intelligence.

The flagship GPT-5.6 option for complex tasks and production workflows.

37.3% on Terminal-Bench 4.0 in the published comparison on this page. OpenAI documents its role in the GPT-5.6 family.

[OpenAI Model Documentation](https://developers.openai.com/api/docs/models/gpt-5.6-sol)

## GPT-5.6 Terra

OpenAI. Everyday Balance.

A balance of intelligence and cost for everyday agent work.

OpenAI’s guidance positions Terra between Sol and Luna. No comparable numeric benchmark is reproduced here.

[OpenAI Model Documentation](https://developers.openai.com/api/docs/models/gpt-5.6-terra)

## GPT-5.6 Luna

OpenAI. Efficient Execution.

A compact choice for frequent, well-defined tasks.

OpenAI recommends Luna for cost-sensitive, high-volume workloads. Evaluate quality on representative tasks before trading capability for efficiency.

[OpenAI Model Documentation](https://developers.openai.com/api/docs/models/gpt-5.6-luna)

## Claude Opus 5.5 (Coming Soon)

Anthropic. Deep Agentic Work.

For complex coding, research, and work that takes multiple steps.

66.4% on Terminal-Bench 4.0 at xhigh effort. The release report also covers knowledge work and computer use.

[Anthropic Evaluation Report](https://www.anthropic.com/claude-opus-5-5)

## Claude Fable 5.1 (Coming Soon)

Anthropic. Demanding Reasoning.

A capable option for challenging reasoning and agent tasks.

55.8% on Terminal-Bench 4.0. The report includes coding, reasoning, and safety evaluations with its deployment safeguards.

[Anthropic Evaluation Report](https://www.anthropic.com/claude-fable-and-mythos-5-1)

## Claude Opus 5

Anthropic. Sustained Work.

For in-depth analysis and complex agent workflows.

52.3% on Terminal-Bench 4.0 in the comparison on this page. Its release report also examines coding, automation, and computer use.

[Anthropic Evaluation Report](https://www.anthropic.com/news/claude-opus-5)

## Claude Sonnet 5

Anthropic. Quality & Efficiency.

A balanced option for coding, research, and practical agent work.

Anthropic reports browsing and computer-use evaluations, including BrowseComp and OSWorld-Verified. Effort settings affect quality and cost.

[Anthropic Evaluation Report](https://www.anthropic.com/news/claude-sonnet-5)

## NVIDIA Nemotron 3 Ultra

NVIDIA. Open-Weight Reasoning.

NVIDIA’s open-weight option for reasoning and agentic tasks.

NVIDIA’s technical report evaluates reasoning, coding, long context, and agent tasks. Compare results with the report’s reasoning budgets and serving setup.

[NVIDIA Research Report](https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/)

## NVIDIA Nemotron Lightning 3.5

NVIDIA. Efficient Reasoning.

A lighter Nemotron option for responsive everyday work.

76.89% on GPQA Diamond without tools in NVIDIA’s BF16 model card. That tested variant may differ from a hosted configuration.

[NVIDIA Model Card](https://catalog.ngc.nvidia.com/orgs/nim/nvidia/models/nemotron-3.5-lightning/hf-3db7814)

## Kimi K3

Moonshot AI. Context & Reasoning.

For work that combines broad context with reasoning and tools.

93.5% on GPQA Diamond at the reported max setting. Moonshot’s card also evaluates coding and long-context tasks.

[Moonshot Model Card](https://huggingface.co/moonshotai/Kimi-K3)

## Kimi K2.6

Moonshot AI. Agent Work.

An open-weight option for coding and tool-assisted tasks.

90.5% on GPQA Diamond in Moonshot’s model card. The report includes separate reasoning, vision, and coding evaluations.

[Moonshot Model Card](https://huggingface.co/moonshotai/Kimi-K2.6)

## DeepSeek V4 Pro

DeepSeek. Deep Reasoning.

For demanding reasoning and coding tasks.

90.1% GPQA Diamond pass@1 at max effort. DeepSeek’s report also covers coding and tool use under documented evaluation settings.

[DeepSeek Model Card](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro)

## DeepSeek V4 Flash

DeepSeek. Efficient Work.

A smaller V4 option for frequent agent tasks.

88.1% GPQA Diamond pass@1 at max effort. Consult the card for the separate quality and efficiency measurements.

[DeepSeek Model Card](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash)

## Grok 4.1 Fast (Non-Reasoning)

xAI. Direct Responses.

For fast responses without a separate reasoning phase.

xAI’s family report covers tool use. We do not assign its reasoning-mode benchmark scores to this non-reasoning variant.

[xAI Release & Evaluations](https://x.ai/news/grok-4-1-fast)

## Claude Opus 4.8

Anthropic. Additional Choice.

An earlier Opus option for reasoning and agent work.

Anthropic’s report covers coding, agent skills, and knowledge work, with links to its system card and external tester feedback.

[Anthropic Evaluation Report](https://www.anthropic.com/news/claude-opus-4-8)

## Claude Haiku 4.5

Anthropic. Fast, Focused Tasks.

A compact Claude option for responsive everyday tasks.

Anthropic evaluates coding and computer use and publishes a safety assessment. Its report emphasizes speed and cost efficiency.

[Anthropic Evaluation Report](https://www.anthropic.com/news/claude-haiku-4-5)

## Will Globster Keep Adding New Models?

We keep Globster current with leading frontier models from top AI labs, adding new releases as they become available and ready for your agents. We continue to expand support across model providers. Timing depends on provider availability and compatibility checks; the app’s model selector shows what you can use today.

## How Should I Choose a Model?

Start with a few real tasks. Compare answer quality, tool use, time to completion, and credit use. A benchmark can narrow your shortlist, but your team’s work is the most useful test.

## Are These Globster Benchmark Results?

No. The scores are published by model providers. They use specific tools, effort settings, and evaluation environments. They are reference points, not a prediction of an agent’s success rate in Globster.

## Can Every Workspace Use Every Model?

Model access depends on your plan and workspace configuration. The model selector in the app shows your current choices. Models labeled Coming Soon are planned additions and are not yet available in Globster.

## What Does the Harness Add to the Model?

The model supplies intelligence. The Globster Agent Harness brings context, tools, memory, goals, and review into the work around it. Choose a capable model, then give your agent the context and boundaries its task needs.

[Explore the model catalog](https://globster.ai/models/). [Meet the Harness](https://globster.ai/agent-harness/).
