A capable model is only useful when its strengths fit your workflow. Compare specific models on the work they suit, the trade-offs they introduce and the control you retain.
8 model assessments·Primary sources reviewed 20 September 2026
How to read this: capabilities are grounded in the linked provider documentation. Pros, cons and fit are our editorial assessment, not results from our own benchmark. Hosted prices, availability and model aliases change; verify them before committing.
OpenAI · General-purpose reasoning model
GPT-6 Astra
Hosted API; model ID gpt-6-astra. No published weights for local deployment.
Best fit: difficult research, coding and document workflows with many dependent steps.
+ The advantages
1.05M-token context and 128K output accommodate large working sets.
Responses API supports search, shell, computer use, MCP and structured outputs.
Selectable reasoning effort lets difficult steps receive more computation.
− The limitations
Premium hosted rates make long, repeated reasoning runs costly.
No fine-tuning support; adaptation relies on instructions, retrieval and tools.
No native audio or video input; use separate media processing.
OUR ASSESSMENT
A capable central planner with broad tool support. The application still supplies permissions, execution and success checks.
Decision rule: Escalate to Astra when a cheaper model repeatedly fails your real end-to-end tasks and the improvement justifies its cost.
OpenAI · General-purpose reasoning model
GPT-5.6 Terra
Hosted API; model ID gpt-5.6-terra. No published weights for local deployment.
Best fit: everyday tool-using agents and repeated production tasks.
+ The advantages
Supports the same broad Responses tool categories, including MCP and computer use.
Reasoning can be disabled or increased, allowing task-specific effort.
1.05M-token context and lower standard token prices than Astra.
− The limitations
A cost-balanced tier: test difficult planning before assigning unattended work.
No fine-tuning support and no local weight deployment.
Very large prompts incur higher rates; tool calls add their own costs.
OUR ASSESSMENT
Useful as a default worker in a routed system, with harder or repeatedly failing cases sent to a stronger model.
Decision rule: Start with Terra for a repeatable workload, measure completion cost including retries, then compare Astra on the failing cases.
Anthropic · General-purpose reasoning model
Claude Opus 5
Active hosted model; claude-opus-5 on the Claude API and supported cloud platforms.
Best fit: complex coding, enterprise analysis and iterative project work.
+ The advantages
1M context and 128K output support substantial code and document work.
Adaptive thinking and adjustable effort let the agent spend effort per task.
Mid-conversation tool changes can preserve the prompt cache (beta).
− The limitations
Higher token prices and comparative latency than Sonnet 5.
Thinking is enabled by default; disabling it requires effort high or below.
Text/image input and text output; separate services are needed for native speech.
OUR ASSESSMENT
Anthropic's recommended general starting point for complex agent work. A Messages API integration still needs its own execution loop.
Decision rule: Use Opus as the Claude baseline; increase effort before paying for Fable, and compare both on task completion.
Anthropic · General-purpose reasoning model
Claude Fable 5.1
Active hosted model, released 1 September 2026; Claude API ID claude-fable-5-1.
Best fit: demanding multistep research and long-running coding or office tasks.
+ The advantages
Designed for long-horizon work across code, documents, spreadsheets and slides.
1M-token context with 128K output.
Per-message effort and readable between-tool progress updates are available in beta.
− The limitations
Higher token prices and slower comparative latency than Opus 5.
Adaptive thinking is always on, limiting lightweight execution choices.
Forced tool use errors; older models cannot read its thinking blocks, complicating migration.
OUR ASSESSMENT
An escalation model for difficult jobs. Progress updates help humans follow work but do not prove that the result is correct.
Decision rule: Choose Fable when Opus at higher effort still misses important cases in your own evaluation set.
Google · Multimodal general-purpose model
Gemini 3.8 Flash
Stable Gemini API model; ID gemini-3.8-flash. September 2026 update.
Best fit: agents combining documents, images, audio or video with grounded tools.
+ The advantages
Accepts text, images, video, audio and PDFs in one model.
Supports search grounding, file search, code execution and function calls.
About 1M input tokens plus caching support large multimodal working sets.
− The limitations
Produces text; image/audio generation and Live API are separate models.
Computer use remains a preview capability despite the stable base model.
Supports low/medium/high thinking; requesting minimal produces an error.
OUR ASSESSMENT
A practical candidate for agents that must inspect varied media before choosing tools. Media support is a capability, not a measured accuracy ranking.
Decision rule: Shortlist Flash when multimodal input is central, then test grounding accuracy and full-workflow latency on your own files.
DeepSeek · Open-weight multimodal reasoning model
DeepSeek V4.1 Flash
Hosted API ID deepseek-flash; MIT-licensed weights at deepseek-ai/DeepSeek-V4.1-Flash. Released 10 September 2026.
Best fit: input-heavy coding agents and teams evaluating self-hosted infrastructure.
+ The advantages
Native image/text input, tool calls and a 1M context window.
Supports Responses and Anthropic-compatible APIs, easing harness integration.
Published weights and inference/evaluation resources enable deployment control.
− The limitations
Hundreds of billions of total parameters make local hosting a substantial infrastructure task.
deepseek-flash is a moving service name; legacy V4 names now route to V4.1.
API token prices do not represent the hardware, serving and operations cost of self-hosting.
OUR ASSESSMENT
Works within an agent harness; downloadable weights do not supply tool execution or a complete production system.
Decision rule: Compare hosted total task cost first; consider self-hosting when control and workload volume justify the infrastructure.
Qwen / Alibaba · Open-weight multimodal model
Qwen3.8-Flash-Next
Released weights at Qwen/Qwen3.8-Flash-Next under Qwen Community 1.0; experimental architecture preview.
Best fit: teams building custom multimodal and coding agents with control of serving.
+ The advantages
Published weights support self-hosting and inspection of the model configuration.
Vision encoder plus tool-oriented examples support multimodal agent work.