Reasoning effort decision dashboard

AI Model Atlas

Explore OpenAI, Anthropic and Google models, compare their reasoning controls, and choose a starting effort for your task.

Model docs checked 15 Sep 2026
Choose a model. Calibrate your effort.
More thinking is a tradeoff.

Use the catalog to check supported controls. Try the advisor, then validate the result on your own tasks. Effort names are not equivalent across providers.

10 models3 providersWorks offline • No API keys
About this dashboard

A comparison and planning tool. It does not run AI models. Provider settings are documented; advisor suggestions are editorial heuristics.

The original Astra/Sol benchmark snapshot is preserved below, labeled as supplied data. No scores have been invented for added models.

Model catalog

Documented API controls • Availability in chat apps may differ

Astra / Sol benchmark snapshot

Supplied with the original dashboard, dated 14 September 2026; not independently reverified.

GPT-6 Astra GPT-5.6 Sol

Artificial Analysis Intelligence Index v4.3

Higher is better. Same effort label does not imply the same amount of compute across models.
Key calibration: Astra Low (46) is already above Sol High (42), while Astra Medium (50) exceeds Sol Max (47) on this aggregate index.

Effort ladder

Original supplied Artificial Analysis v4.3 figures; not independently reverified

EffortAstra indexSol indexAstra cost/taskSol cost/taskAstra advantagePractical reading
Low4634$0.82$0.26+12Strong default for routine professional work
Medium5039$1.54$0.50+11Best general-purpose balance
High5142$1.72$0.81+9Use when extra checking or iteration matters
XHigh5344$2.31$1.18+9Hard debugging, architecture, difficult research
Max5347$3.26$1.99+6Escalation level; aggregate returns flatten

Interactive effort advisor

Model-aware decision aid — editorial guidance, not a benchmark

Suggestions use only supported levels. They are starting points for testing, not provider defaults or measured equivalences. Review important outputs independently.

ASTRA MEDIUM

Recommended starting point: Medium

Enough headroom for analysis and troubleshooting without paying the steep compute premium of the top settings.

Escalate only if the task misses requirements, needs deeper verification, or lower effort has already failed.

A simple mental model

Use higher settings because the task needs them, not because they exist.

Low

“Do it.”

Clear tasks, explanations, XML/PowerShell edits, normal admin work.

Medium

“Think about it.”

Troubleshooting, code review, implementation planning, most serious daily work.

High

“Be careful.”

Subtle failures, several interacting systems, verification matters.

XHigh

“Investigate.”

Difficult debugging, large refactors, complex architecture or research.

Max

“Exhaust it.”

Escalation setting for unusually hard tasks or after lower levels fail.

Sources and methodology

Primary sources first; independent benchmark second.

OpenAI Help Center — Managing usage with GPT-6 Astra

Official guidance on reasoning effort, allowance use, and the explicit Astra Low vs Sol High calibration.

Open source ↗
OpenAI — GPT-6 Astra launch

Official model announcement and evidence that higher effort can buy more iterations, verification and code execution in agentic coding.

Open source ↗
OpenAI API — Models

Official effort levels, context window and current model/API specifications.

Open source ↗
Artificial Analysis — Intelligence Index v4.3

Independent benchmark. Original figures were supplied with the uploaded dashboard and have not been independently reverified for this edition.

Open source ↗

Important: Artificial Analysis “cost per task” is a weighted API benchmark workload, not a ChatGPT/Codex credit charge. The Intelligence Index is an aggregate benchmark and should not be interpreted as a guaranteed quality score for every individual task.