AI Model Index
Pricing sourced from Artificial Analysis. Intelligence Index, hallucination rates, and Coding Agent Index from Artificial Analysis.
Last updated: September 16, 2026
53 models
Models with a published Artificial Analysis Coding Agent Index — fully benchmarked for coding agent use.
Stable
Colored cells grade each metric in quartiles relative to other models in this table — top 25% green, bottom 25% red; shade intensity reflects position within the quartile. Sorted by coding index (highest first).
| Model | $/1M | Cost/Task | Intel. | Omni Acc. | Recall | Coding | Tok/s | Time/task |
|---|---|---|---|---|---|---|---|---|
| $10.00 / $50.00 | $7.63 | 53 | 67.2% | 85.3% | 81.6 | 65 tok/s | 20.0 min | |
| $10.00 / $50.00 | $5.98 | 53 | 66.2% | 83.0% | 80.7 | 58 tok/s | 17.3 min | |
| $10.00 / $50.00 | $3.91 | 51 | 64.9% | 83.7% | 79.1 | 51 tok/s | 12.5 min | |
| $4.00 / $20.00 | $1.18 | 44 | 58.8% | 82.3% | 78.3 | 64 tok/s | 5.1 min | |
| $5.00 / $25.00 | $5.86 | 51 | 60.9% | 79.3% | 78 | 49 tok/s | 24.5 min | |
| $4.00 / $20.00 | $1.99 | 47 | 59.4% | 84.0% | 77.4 | 65 tok/s | 7.6 min | |
| $4.00 / $20.00 | $0.81 | 43 | 58.4% | 81.7% | 77.2 | 57 tok/s | 3.8 min | |
| $10.00 / $50.00 | $1.72 | 51 | 61.1% | 80.0% | 77.1 | 51 tok/s | 3.8 min | |
| $10.00 / $50.00 | $2.98 | 49 | 63.1% | 84.7% | 77.1 | 51 tok/s | 9.1 min | |
| $5.00 / $25.00 | $4.88 | 50 | 59.5% | 80.3% | 77 | 49 tok/s | 20.4 min | |
| $10.00 / $50.00 | $3.26 | 53 | 62.6% | 80.7% | 76.9 | 53 tok/s | 8.6 min | |
| $2.00 / $6.00 | $1.86 | 44 | 48.2% | 80.3% | 76.8 | 58 tok/s | 10.4 min | |
| $10.00 / $50.00 | $1.54 | 50 | 60.6% | 79.7% | 76.7 | 49 tok/s | 3.2 min | |
| $2.00 / $12.00 | $1.40 | 42 | 46.8% | 83.0% | 76.7 | 99 tok/s | 6.6 min | |
| $10.00 / $50.00 | $8.75 | 50 | 65.4% | 82.3% | 76.5 | 65 tok/s | 17.0 min | |
| $5.00 / $25.00 | $3.61 | 48 | 58.9% | 79.0% | 76.5 | 50 tok/s | 15.5 min | |
| $1.25 / $4.25 | $1.37 | 45 | 41.5% | 83.0% | 76.5 | 220 tok/s | 4.1 min | |
| $0.75 / $3.75 | $1.24 | 41 | 54.6% | 81.3% | 76.3 | 332 tok/s | 3.6 min | |
| $4.00 / $20.00 | $0.51 | 40 | 57.8% | 80.3% | 76.3 | 60 tok/s | 2.2 min | |
| $2.00 / $6.00 | $5.41 | 45 | 31.7% | 80.3% | 76.2 | 40 tok/s | 45.2 min | |
| $3.00 / $15.00 | $2.00 | 44 | 47.6% | 88.7% | 76.2 | 35 tok/s | 23.2 min | |
| $0.75 / $3.75 | $0.93 | 39 | 55.3% | 81.7% | 76.1 | 289 tok/s | 3.4 min | |
| $10.00 / $50.00 | $2.31 | 53 | 61.9% | 80.0% | 75.9 | 53 tok/s | 5.4 min | |
| $2.00 / $6.00 | $2.32 | 44 | 43.0% | 81.0% | 75.9 | 56 tok/s | 11.2 min | |
| $1.25 / $4.25 | $1.60 | 48 | 43.6% | 83.0% | 75.8 | 226 tok/s | 4.4 min | |
| $10.00 / $50.00 | $0.82 | 46 | 59.5% | 80.0% | 75.7 | 50 tok/s | 1.5 min | |
| $10.00 / $50.00 | $2.37 | 47 | 60.2% | 82.3% | 75.2 | 50 tok/s | 7.1 min | |
| $5.00 / $30.00 | $2.63 | 39 | 58.0% | 84.3% | 74.9 | 94 tok/s | 4.2 min | |
| $1.40 / $4.40 | $2.01 | 45 | 33.9% | 79.7% | 74.8 | 67 tok/s | 17.6 min | |
| $2.00 / $6.00 | $1.50 | 43 | 41.9% | 81.0% | 74.4 | 54 tok/s | 8.7 min | |
| $5.00 / $25.00 | $2.19 | 45 | 57.1% | 82.0% | 74.3 | 48 tok/s | 10.1 min | |
| $5.00 / $25.00 | $4.08 | 42 | 48.8% | 77.7% | 74.3 | 56 tok/s | 20.9 min | |
| $0.75 / $3.75 | $0.93 | 40 | 53.0% | 84.0% | 74.1 | — | 21.2 min | |
| $5.00 / $25.00 | — | 41 | 48.9% | 78.7% | 73.6 | 45 tok/s | 20.4 min | |
| $0.75 / $3.75 | — | 34 | 52.2% | 80.7% | 73.5 | — | 22.0 min | |
| $0.15 / $0.47 | $0.37 | 40 | 24.5% | 79.7% | 73.1 | 52 tok/s | 34.8 min | |
| $2.00 / $6.00 | $1.04 | 39 | 51.6% | 79.3% | 72.4 | 60 tok/s | 7.5 min | |
| $1.25 / $4.25 | $0.97 | 40 | 45.4% | 79.0% | 72.2 | 172 tok/s | 4.6 min | |
| $3.00 / $15.00 | $1.15 | 31 | 45.8% | 79.3% | 72 | 33 tok/s | 6.1 min | |
| $2.00 / $6.00 | $2.16 | 40 | 31.3% | 80.3% | 71.9 | 40 tok/s | 28.4 min | |
| $0.15 / $0.50 | $0.25 | 42 | 27.5% | 80.0% | 71.5 | 114 tok/s | 10.0 min | |
| $0.20 / $1.20 | $0.18 | 38 | 42.7% | 83.7% | 71.4 | 115 tok/s | 6.0 min | |
| $0.44 / $1.32 | $0.22 | 35 | 40.4% | 79.7% | 69.1 | 215 tok/s | 4.8 min | |
| $1.32 / $3.96 | $0.67 | 36 | 49.1% | 80.3% | 68.8 | 94 tok/s | 9.7 min | |
| $0.20 / $1.20 | $0.09 | 35 | 42.5% | 81.7% | 68.6 | 118 tok/s | 3.3 min | |
| $0.50 / $3.00 | $0.82 | 34 | 15.6% | 82.0% | 68.1 | 44 tok/s | 25.2 min | |
| $2.50 / $7.50 | $1.15 | 30 | 31.1% | 79.0% | 66 | 195 tok/s | 2.8 min | |
| $0.20 / $1.20 | $0.04 | 32 | 41.8% | 80.3% | 63.3 | 109 tok/s | 2.1 min | |
| $3.00 / $15.00 | $2.49 | 31 | 40.9% | 80.0% | 63 | 45 tok/s | 27.5 min | |
| $0.60 / $3.60 | $0.62 | 22 | 19.6% | 77.3% | 53.7 | 57 tok/s | 9.5 min | |
| $0.20 / $1.20 | $0.02 | 26 | 40.7% | 75.0% | 50.7 | 112 tok/s | 40s | |
| $1.00 / $5.00 | $0.21 | 18 | 18.0% | 74.3% | 43.9 | 85 tok/s | 3.6 min | |
| $0.38 / $2.25 | $0.48 | 19 | 18.8% | 71.7% | 41.9 | 127 tok/s | 4.5 min |
Notes
- Intelligence Index = Artificial Analysis Intelligence Index v4.0 (composite of 9 evals including GDPval-AA, Terminal-Bench, SciCode, AA-Omniscience). Higher is better. Scores reflect max-effort / adaptive reasoning mode where applicable.
- AA-Omni Accuracy = AA-Omniscience Accuracy (%). Higher is better. Proportion of correctly answered questions out of all questions. See artificialanalysis.ai/evaluations/omniscience.
- Recall = Artificial Analysis Long Context Reasoning (AA-LCR) recall accuracy (%). Higher is better. “—” means AA has not measured the model. See artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning.
- Cost per Task = Weighted average pay-per-token API cost (USD) per Artificial Analysis Intelligence Index task. Lower is better. Reflects input, cache, reasoning, and output token pricing across index evaluations.
- Coding Agent Index = Artificial Analysis Coding Index — composite pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Higher is better. Scores shown are per-model (best published harness per model from AA's language models dataset); this differs from the separate agents leaderboard (model+harness pairs). "—" means AA has not yet evaluated that model. See artificialanalysis.ai/agents/coding-agents.
- Model lists are synced from Artificial Analysis via `pnpm sync:model-index` (stable = top coding agent scores; local = curated Qwen open-weight list; experiment = top intelligence scores without coding benchmark, limited to trusted providers).
- Accuracy rates are sourced from AA-Omniscience via SSR scrape snapshot (`aa-omniscience-snapshot.json`, refreshed with `pnpm fetch:aa-omniscience`). Not included in the free AA sync API — see memory/development/ai-model-index-omniscience-scrape.md.
- Recall is sourced from the public AA-LCR SSR snapshot (`aa-lcr-snapshot.json`, refreshed with `pnpm fetch:aa-lcr`) because the free AA sync API omits per-benchmark fields. Null means AA has not measured the model.
- Pricing reflects standard (non-cached) token rates from Artificial Analysis.
- OpenAI cut GPT-5.6 Luna API pricing 80% and Terra 20% on July 30, 2026 (Luna $0.20/$1.20, Terra $2/$12 per 1M; Sol unchanged).
- Gemini 3.1 Pro Preview pricing doubles to $4/$18 per 1M above 200K tokens per prompt.
- Claude Opus 4.7 and Claude Sonnet 4.6 both support a 1M token context window with 128K max output tokens.
- Claude Fable 5.1 leads the AA Intelligence Index v4.1.1 at 66 (Sept 2026), ahead of Opus 5 at 63 and Fable 5 at 62.
- Claude Fable 5.1 and Claude Mythos 5.1 released September 1, 2026; Fable 5.1 is on the Experiment tab until AA publishes a Coding Agent Index.
- Claude Fable 5 and Claude Mythos 5 (Fable 5's unrestricted-cyber/bio counterpart) launched June 9, 2026. Access was suspended June 12–July 1, 2026 for U.S. export-control compliance; current pricing/scores reflect the restored, generally available model.
- GPT-5.6 ships as three durable capability tiers: Sol (flagship, AA-Omni index 21.7 / accuracy 58.5%), Terra (balanced, -0.2 / 45.9%), and Luna (fastest/cheapest, -11.2 / 41.5%).
- Grok 4.5 (high) is xAI's first release since its SpaceXAI rebrand and IPO; it's notably cost-efficient per completed task with an AA-Omniscience accuracy of 52.1% (index 26.4).
- Grok 4.6 (high) scores 61 on the AA Intelligence Index (tying GPT-5.6 Sol max) and 76.8 on the Coding Agent Index — a +4.4-point coding gain over Grok 4.5 at the same $2/$6 per 1M pricing. AA-Omniscience accuracy: 48.2% (index 30.5).
- Open-weight models (self-hostable): Kimi K3 (weights releasing July 27, 2026), Kimi K2.7 Code, Kimi K2.6, MiMo-V2.5-Pro, GLM 5.3, GLM 5.3 Flash, DeepSeek V4 Pro, DeepSeek V4 Flash, Hy3-preview, Nemotron 3 Super.
- DeepSeek V4 Pro and V4 Flash OpenRouter pricing shown; first-party DeepSeek API pricing differs ($1.74/$3.48 for Pro, $0.14/$0.28 for Flash).
- Nemotron 3 Super is also available free-tier on OpenRouter (rate-limited).
Changelog
| Date | Change |
|---|---|
| 2026-09-16 | Expanded the Stable and Experiment tabs from the top 25 to the top 50 models each via the model sync limits; refreshed model lists and benchmark snapshots. |
| 2026-09-16 | Refreshed Stable, Experiment, and Local model lists and benchmark snapshots from Artificial Analysis; added DeepSeek V4.1 Flash and Claude Sonnet 5 effort variants (xhigh, high, medium, low) to the Experiment tab. |
| 2026-09-04 | Added AA-LCR Recall as a separate benchmark. Configured Grok fast variants use the base Grok 4 recall while retaining their own source pricing and identity. |
| 2026-09-03 | Added Grok 4 Fast (Reasoning and Non-reasoning) as separate Experiment rows with duplicated Grok 4 benchmark scores and Artificial Analysis speed-specific token pricing. |
| 2026-09-03 | Refreshed Stable, Experiment, and Local model data from Artificial Analysis; added Claude Fable 5.1 effort variants to Stable and Claude 4.1 Opus (Reasoning) to Experiment. |
| 2026-09-02 | Added Claude Fable 5.1 (Anthropic) to Experiment tab — AA Intelligence Index 66 (#1). Pinned Claude Fable 5 on Stable tab; refreshed Fable 5 benchmarks. |
| 2026-09-01 | Refreshed pricing and cost-per-task metrics from Artificial Analysis for all Stable, Experiment, and Local tab models. |
| 2026-08-27 | Added GLM 5.3 Flash on Stable tab — AA Intelligence Index 58, Coding Agent Index 71.5. |
| 2026-08-20 | Added GLM 5.3 (max) benchmarks on Stable tab — AA Intelligence Index 60, Coding Agent Index 74.8. |
| 2026-08-20 | Refreshed DeepSeek V4 Pro 0813 benchmarks on Stable tab — latest GA checkpoint (deepseek-v4-pro). |
| 2026-08-18 | Added Local tab with curated Qwen 3.6/3.8 open-weight models. |
| 2026-08-15 | Refreshed AA time-per-task snapshot — all tracked models now have Time/task metrics including Luna effort variants; fixed prior timeout failures. |
| 2026-08-15 | Added GPT-5.6 Luna thinking-effort variants (xhigh, high, medium) to Stable tab alongside existing Luna (max). |
| 2026-08-14 | Restored DeepSeek V4 Pro 0813 on Stable tab — AA slug renamed to deepseek-v4-pro. |
| 2026-08-15 | Models with announced pricing changes show a warning icon beside $/1M — tap or click for details. |
| 2026-08-14 | Replaced sort dropdown with Coding / Cost / Speed buttons — single-click primary dimension for model list ordering. |
| 2026-08-14 | Added Tok/s and Time/task columns to desktop table; card and table now share a single field registry so metrics stay in sync. |
| 2026-08-14 | Added AA time-per-task scores via SSR snapshot — median output speed and intelligence-index output tokens per task scraped from Artificial Analysis model pages. |
| 2026-08-14 | Refreshed AA-Omniscience accuracy snapshot for all 50 tracked models — 49/50 populated. Claude 4.1 Opus (Reasoning) remains unbenchmarked on AA (omniscience null; deprecated). |
| 2026-08-14 | Pinned GLM 5.3 (Z.ai) for Stable tab — awaiting Artificial Analysis benchmark publication. |
| 2026-08-13 | Added DeepSeek V4 Pro 0813 benchmarks (intelligence 53, coding index 68.8) on Stable tab. |
| 2026-08-13 | Added Grok 4.6 (high) to Stable benchmarks — AA Intelligence Index 61, Coding Agent Index 76.8. |
| 2026-08-03 | Added Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) to Stable benchmarks — AA Coding Agent Index 63. |
| 2026-08-01 | Added output-price sort to model table. Pinned Claude Haiku 4.5 (Reasoning) and GPT-5.6 Luna on Stable tab. |
| 2026-08-01 | Synced GPT-5.6 Luna and Terra price reductions (Luna $0.20/$1.20, Terra $2/$12 per 1M; Sol unchanged at $5/$30) following OpenAI July 30, 2026 API price cut. |
| 2026-07-31 | Updated DeepSeek V4 Flash 0731 benchmarks (intelligence 50, coding index 69.1) on Stable tab. |
| 2026-07-25 | Wired AA-Omniscience hallucination rates for all 50 tracked models (stable + experiment) from SSR scrape snapshot. |
| 2026-07-25 | Fixed benchmark column to show AA-Omniscience Accuracy instead of conditional hallucination rate — values now match AA's prominently displayed metric (e.g. GPT-5.6 Sol ~58.5%, Claude Opus 5 ~54.2%). |
| 2026-07-24 | Synced Claude Opus 5 (max, xhigh, high, medium) from Artificial Analysis — Opus 5 (max) leads Intelligence Index at 61 and Coding Index at 78. |
| 2026-07-20 | Added Qwen3.7 Max (Alibaba) to Stable benchmarks. Qwen 3.8 is not yet published on Artificial Analysis — pinned model IDs will pick it up automatically on the next sync once AA adds it. |
| 2026-07-19 | Added Cost/Task column with relative color scale on Stable and Experiment model cards and tables. |
| 2026-07-18 | Switched to AA-synced stable/experiment model lists (sync:model-index) instead of manual models.json. |
| 2026-07-17 | Added Kimi K2.7 Code (Moonshot AI). |
| 2026-07-17 | Replaced Agentic Index and AA-Omni index with Intelligence Index, hallucination rate, and Coding Agent Index columns. |
| 2026-07-17 | Added Kimi K3 (Moonshot AI). |
| 2026-07-10 | Added Grok 4.5 (xAI). |
| 2026-07-10 | Added Claude Fable 5 (Anthropic). |
| 2026-07-10 | Added GPT-5.6 Sol, Terra, and Luna (OpenAI). |
| 2026-05-17 | Added MiniMax M2.7 (MiniMax). |
| 2026-05-17 | Added Hy3-preview (Tencent). |
| 2026-05-17 | Filled in AA-Omni for Gemini 3 Flash (12), GPT-5.5 xhigh (18, corrected from 20), Grok 4.3 high (18). |
| 2026-05-17 | Added GPT-5.5 xhigh (OpenAI). |
| 2026-05-17 | Added Grok 4.3 high (xAI). |
| 2026-05-17 | Added Gemini 3 Flash Preview (Google). |
| 2026-05-17 | Filled in AA-Omni scores for MiMo-V2.5-Pro (4), Qwen3.6 Plus (3), Nemotron 3 Super (-42). |
| 2026-05-17 | Added MiMo-V2.5-Pro (Xiaomi), Qwen3.6 Plus (Alibaba), Nemotron 3 Super (NVIDIA). |
| 2026-05-17 | Added Gemini 3.1 Pro Preview. Filled in AA-Omni scores for Claude Sonnet 4.6 (12) and GLM 5.1 (2). |
| 2026-05-17 | Added AA-Omni column (AA-Omniscience Index from Artificial Analysis). |
| 2026-05-17 | Added DeepSeek V4 Pro and DeepSeek V4 Flash. |
| 2026-05-17 | Removed Context and Tier columns. |
| 2026-05-17 | Added Kimi K2.6 (Moonshot AI) and GLM 5.1 (Z.ai). Added sort-order comment to table. |
| 2026-05-17 | Initial document created. Added Claude Opus 4.7 and Claude Sonnet 4.6. |