Explore OpenAI, Anthropic and Google models, compare their reasoning controls, and choose a starting effort for your task.
Use the catalog to check supported controls. Try the advisor, then validate the result on your own tasks. Effort names are not equivalent across providers.
A comparison and planning tool. It does not run AI models. Provider settings are documented; advisor suggestions are editorial heuristics.
The original Astra/Sol benchmark snapshot is preserved below, labeled as supplied data. No scores have been invented for added models.
Documented API controls • Availability in chat apps may differ
Supplied with the original dashboard, dated 14 September 2026; not independently reverified.
Original supplied Artificial Analysis v4.3 figures; not independently reverified
| Effort | Astra index | Sol index | Astra cost/task | Sol cost/task | Astra advantage | Practical reading |
|---|---|---|---|---|---|---|
| Low | 46 | 34 | $0.82 | $0.26 | +12 | Strong default for routine professional work |
| Medium | 50 | 39 | $1.54 | $0.50 | +11 | Best general-purpose balance |
| High | 51 | 42 | $1.72 | $0.81 | +9 | Use when extra checking or iteration matters |
| XHigh | 53 | 44 | $2.31 | $1.18 | +9 | Hard debugging, architecture, difficult research |
| Max | 53 | 47 | $3.26 | $1.99 | +6 | Escalation level; aggregate returns flatten |
Model-aware decision aid — editorial guidance, not a benchmark
Suggestions use only supported levels. They are starting points for testing, not provider defaults or measured equivalences. Review important outputs independently.
Enough headroom for analysis and troubleshooting without paying the steep compute premium of the top settings.
Use higher settings because the task needs them, not because they exist.
Clear tasks, explanations, XML/PowerShell edits, normal admin work.
Troubleshooting, code review, implementation planning, most serious daily work.
Subtle failures, several interacting systems, verification matters.
Difficult debugging, large refactors, complex architecture or research.
Escalation setting for unusually hard tasks or after lower levels fail.
Primary sources first; independent benchmark second.
Official guidance on reasoning effort, allowance use, and the explicit Astra Low vs Sol High calibration.
Open source ↗Official model announcement and evidence that higher effort can buy more iterations, verification and code execution in agentic coding.
Open source ↗Official effort levels, context window and current model/API specifications.
Open source ↗Independent benchmark. Original figures were supplied with the uploaded dashboard and have not been independently reverified for this edition.
Open source ↗Important: Artificial Analysis “cost per task” is a weighted API benchmark workload, not a ChatGPT/Codex credit charge. The Intelligence Index is an aggregate benchmark and should not be interpreted as a guaranteed quality score for every individual task.