Skip to main content

AI Model Index

Pricing sourced from Artificial Analysis. Intelligence Index, hallucination rates, and Coding Agent Index from Artificial Analysis.
Last updated: September 16, 2026

53 models

Models with a published Artificial Analysis Coding Agent Index — fully benchmarked for coding agent use.

Stable

Colored cells grade each metric in quartiles relative to other models in this table — top 25% green, bottom 25% red; shade intensity reflects position within the quartile. Sorted by coding index (highest first).

Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)

Anthropic

Intelligence & accuracy

Intel
53
Omni Acc
67.2%
Recall
85.3%
Coding
81.6

Speed

Tok/s
65 tok/s
Time/task
20.0 min

Price

$/1M
$10.00 / $50.00
Cost
$7.63
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)

Anthropic

Intelligence & accuracy

Intel
53
Omni Acc
66.2%
Recall
83.0%
Coding
80.7

Speed

Tok/s
58 tok/s
Time/task
17.3 min

Price

$/1M
$10.00 / $50.00
Cost
$5.98
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)

Anthropic

Intelligence & accuracy

Intel
51
Omni Acc
64.9%
Recall
83.7%
Coding
79.1

Speed

Tok/s
51 tok/s
Time/task
12.5 min

Price

$/1M
$10.00 / $50.00
Cost
$3.91
GPT-5.6 Sol (xhigh)

OpenAI

Intelligence & accuracy

Intel
44
Omni Acc
58.8%
Recall
82.3%
Coding
78.3

Speed

Tok/s
64 tok/s
Time/task
5.1 min

Price

$/1M
$4.00 / $20.00
Cost
$1.18
Claude Opus 5 (Adaptive Reasoning, Max Effort)

Anthropic

Intelligence & accuracy

Intel
51
Omni Acc
60.9%
Recall
79.3%
Coding
78

Speed

Tok/s
49 tok/s
Time/task
24.5 min

Price

$/1M
$5.00 / $25.00
Cost
$5.86
GPT-5.6 Sol (max)

OpenAI

Intelligence & accuracy

Intel
47
Omni Acc
59.4%
Recall
84.0%
Coding
77.4

Speed

Tok/s
65 tok/s
Time/task
7.6 min

Price

$/1M
$4.00 / $20.00
Cost
$1.99
GPT-5.6 Sol (high)

OpenAI

Intelligence & accuracy

Intel
43
Omni Acc
58.4%
Recall
81.7%
Coding
77.2

Speed

Tok/s
57 tok/s
Time/task
3.8 min

Price

$/1M
$4.00 / $20.00
Cost
$0.81
GPT-6 Astra (high)

OpenAI

Intelligence & accuracy

Intel
51
Omni Acc
61.1%
Recall
80.0%
Coding
77.1

Speed

Tok/s
51 tok/s
Time/task
3.8 min

Price

$/1M
$10.00 / $50.00
Cost
$1.72
Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)

Anthropic

Intelligence & accuracy

Intel
49
Omni Acc
63.1%
Recall
84.7%
Coding
77.1

Speed

Tok/s
51 tok/s
Time/task
9.1 min

Price

$/1M
$10.00 / $50.00
Cost
$2.98
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)

Anthropic

Intelligence & accuracy

Intel
50
Omni Acc
59.5%
Recall
80.3%
Coding
77

Speed

Tok/s
49 tok/s
Time/task
20.4 min

Price

$/1M
$5.00 / $25.00
Cost
$4.88
GPT-6 Astra (max)

OpenAI

Intelligence & accuracy

Intel
53
Omni Acc
62.6%
Recall
80.7%
Coding
76.9

Speed

Tok/s
53 tok/s
Time/task
8.6 min

Price

$/1M
$10.00 / $50.00
Cost
$3.26
Grok 4.6 (high)

SpaceXAI

Intelligence & accuracy

Intel
44
Omni Acc
48.2%
Recall
80.3%
Coding
76.8

Speed

Tok/s
58 tok/s
Time/task
10.4 min

Price

$/1M
$2.00 / $6.00
Cost
$1.86
GPT-6 Astra (medium)

OpenAI

Intelligence & accuracy

Intel
50
Omni Acc
60.6%
Recall
79.7%
Coding
76.7

Speed

Tok/s
49 tok/s
Time/task
3.2 min

Price

$/1M
$10.00 / $50.00
Cost
$1.54
GPT-5.6 Terra (max)

OpenAI

Intelligence & accuracy

Intel
42
Omni Acc
46.8%
Recall
83.0%
Coding
76.7

Speed

Tok/s
99 tok/s
Time/task
6.6 min

Price

$/1M
$2.00 / $12.00
Cost
$1.40
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

Anthropic

Intelligence & accuracy

Intel
50
Omni Acc
65.4%
Recall
82.3%
Coding
76.5

Speed

Tok/s
65 tok/s
Time/task
17.0 min

Price

$/1M
$10.00 / $50.00
Cost
$8.75
Claude Opus 5 (Adaptive Reasoning, High Effort)

Anthropic

Intelligence & accuracy

Intel
48
Omni Acc
58.9%
Recall
79.0%
Coding
76.5

Speed

Tok/s
50 tok/s
Time/task
15.5 min

Price

$/1M
$5.00 / $25.00
Cost
$3.61
Muse Spark 1.3 (xhigh)

Meta

Intelligence & accuracy

Intel
45
Omni Acc
41.5%
Recall
83.0%
Coding
76.5

Speed

Tok/s
220 tok/s
Time/task
4.1 min

Price

$/1M
$1.25 / $4.25
Cost
$1.37
Gemini 3.8 Flash (high)

Google

Intelligence & accuracy

Intel
41
Omni Acc
54.6%
Recall
81.3%
Coding
76.3

Speed

Tok/s
332 tok/s
Time/task
3.6 min

Price

$/1M
$0.75 / $3.75
Cost
$1.24
GPT-5.6 Sol (medium)

OpenAI

Intelligence & accuracy

Intel
40
Omni Acc
57.8%
Recall
80.3%
Coding
76.3

Speed

Tok/s
60 tok/s
Time/task
2.2 min

Price

$/1M
$4.00 / $20.00
Cost
$0.51
Qwen3.8 Max (0902)

Alibaba

Intelligence & accuracy

Intel
45
Omni Acc
31.7%
Recall
80.3%
Coding
76.2

Speed

Tok/s
40 tok/s
Time/task
45.2 min

Price

$/1M
$2.00 / $6.00
Cost
$5.41
Kimi K3 (max)

Kimi

Intelligence & accuracy

Intel
44
Omni Acc
47.6%
Recall
88.7%
Coding
76.2

Speed

Tok/s
35 tok/s
Time/task
23.2 min

Price

$/1M
$3.00 / $15.00
Cost
$2.00
Gemini 3.7 Flash (high)

Google

Intelligence & accuracy

Intel
39
Omni Acc
55.3%
Recall
81.7%
Coding
76.1

Speed

Tok/s
289 tok/s
Time/task
3.4 min

Price

$/1M
$0.75 / $3.75
Cost
$0.93
GPT-6 Astra (xhigh)

OpenAI

Intelligence & accuracy

Intel
53
Omni Acc
61.9%
Recall
80.0%
Coding
75.9

Speed

Tok/s
53 tok/s
Time/task
5.4 min

Price

$/1M
$10.00 / $50.00
Cost
$2.31
Grok 4.6 (xhigh)

SpaceXAI

Intelligence & accuracy

Intel
44
Omni Acc
43.0%
Recall
81.0%
Coding
75.9

Speed

Tok/s
56 tok/s
Time/task
11.2 min

Price

$/1M
$2.00 / $6.00
Cost
$2.32
Muse Spark 1.3 (max)

Meta

Intelligence & accuracy

Intel
48
Omni Acc
43.6%
Recall
83.0%
Coding
75.8

Speed

Tok/s
226 tok/s
Time/task
4.4 min

Price

$/1M
$1.25 / $4.25
Cost
$1.60
GPT-6 Astra (low)

OpenAI

Intelligence & accuracy

Intel
46
Omni Acc
59.5%
Recall
80.0%
Coding
75.7

Speed

Tok/s
50 tok/s
Time/task
1.5 min

Price

$/1M
$10.00 / $50.00
Cost
$0.82
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)

Anthropic

Intelligence & accuracy

Intel
47
Omni Acc
60.2%
Recall
82.3%
Coding
75.2

Speed

Tok/s
50 tok/s
Time/task
7.1 min

Price

$/1M
$10.00 / $50.00
Cost
$2.37
GPT-5.5 (xhigh)

OpenAI

Intelligence & accuracy

Intel
39
Omni Acc
58.0%
Recall
84.3%
Coding
74.9

Speed

Tok/s
94 tok/s
Time/task
4.2 min

Price

$/1M
$5.00 / $30.00
Cost
$2.63
GLM-5.3 (max)

Z AI

Intelligence & accuracy

Intel
45
Omni Acc
33.9%
Recall
79.7%
Coding
74.8

Speed

Tok/s
67 tok/s
Time/task
17.6 min

Price

$/1M
$1.40 / $4.40
Cost
$2.01
Grok 4.6 (medium)

SpaceXAI

Intelligence & accuracy

Intel
43
Omni Acc
41.9%
Recall
81.0%
Coding
74.4

Speed

Tok/s
54 tok/s
Time/task
8.7 min

Price

$/1M
$2.00 / $6.00
Cost
$1.50
Claude Opus 5 (Adaptive Reasoning, Medium Effort)

Anthropic

Intelligence & accuracy

Intel
45
Omni Acc
57.1%
Recall
82.0%
Coding
74.3

Speed

Tok/s
48 tok/s
Time/task
10.1 min

Price

$/1M
$5.00 / $25.00
Cost
$2.19
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)

Anthropic

Intelligence & accuracy

Intel
42
Omni Acc
48.8%
Recall
77.7%
Coding
74.3

Speed

Tok/s
56 tok/s
Time/task
20.9 min

Price

$/1M
$5.00 / $25.00
Cost
$4.08
Gemini 3.8 Flash (medium)

Google

Intelligence & accuracy

Intel
40
Omni Acc
53.0%
Recall
84.0%
Coding
74.1

Speed

Tok/s
Time/task
21.2 min

Price

$/1M
$0.75 / $3.75
Cost
$0.93
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)

Anthropic

Intelligence & accuracy

Intel
41
Omni Acc
48.9%
Recall
78.7%
Coding
73.6

Speed

Tok/s
45 tok/s
Time/task
20.4 min

Price

$/1M
$5.00 / $25.00
Cost
Gemini 3.8 Flash (low)

Google

Intelligence & accuracy

Intel
34
Omni Acc
52.2%
Recall
80.7%
Coding
73.5

Speed

Tok/s
Time/task
22.0 min

Price

$/1M
$0.75 / $3.75
Cost
Qwen3.8-Flash-Next

Alibaba

Intelligence & accuracy

Intel
40
Omni Acc
24.5%
Recall
79.7%
Coding
73.1

Speed

Tok/s
52 tok/s
Time/task
34.8 min

Price

$/1M
$0.15 / $0.47
Cost
$0.37
Grok 4.5 (high)

SpaceXAI

Intelligence & accuracy

Intel
39
Omni Acc
51.6%
Recall
79.3%
Coding
72.4

Speed

Tok/s
60 tok/s
Time/task
7.5 min

Price

$/1M
$2.00 / $6.00
Cost
$1.04
Muse Spark 1.2 (xhigh)

Meta

Intelligence & accuracy

Intel
40
Omni Acc
45.4%
Recall
79.0%
Coding
72.2

Speed

Tok/s
172 tok/s
Time/task
4.6 min

Price

$/1M
$1.25 / $4.25
Cost
$0.97
Kimi K3 (low)

Kimi

Intelligence & accuracy

Intel
31
Omni Acc
45.8%
Recall
79.3%
Coding
72

Speed

Tok/s
33 tok/s
Time/task
6.1 min

Price

$/1M
$3.00 / $15.00
Cost
$1.15
Qwen3.8 2.4T A95B

Alibaba

Intelligence & accuracy

Intel
40
Omni Acc
31.3%
Recall
80.3%
Coding
71.9

Speed

Tok/s
40 tok/s
Time/task
28.4 min

Price

$/1M
$2.00 / $6.00
Cost
$2.16
GLM-5.3-Flash

Z AI

Intelligence & accuracy

Intel
42
Omni Acc
27.5%
Recall
80.0%
Coding
71.5

Speed

Tok/s
114 tok/s
Time/task
10.0 min

Price

$/1M
$0.15 / $0.50
Cost
$0.25
GPT-5.6 Luna (max)

OpenAI

Intelligence & accuracy

Intel
38
Omni Acc
42.7%
Recall
83.7%
Coding
71.4

Speed

Tok/s
115 tok/s
Time/task
6.0 min

Price

$/1M
$0.20 / $1.20
Cost
$0.18
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

DeepSeek

Intelligence & accuracy

Intel
35
Omni Acc
40.4%
Recall
79.7%
Coding
69.1

Speed

Tok/s
215 tok/s
Time/task
4.8 min

Price

$/1M
$0.44 / $1.32
Cost
$0.22
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)

DeepSeek

Intelligence & accuracy

Intel
36
Omni Acc
49.1%
Recall
80.3%
Coding
68.8

Speed

Tok/s
94 tok/s
Time/task
9.7 min

Price

$/1M
$1.32 / $3.96
Cost
$0.67
GPT-5.6 Luna (xhigh)

OpenAI

Intelligence & accuracy

Intel
35
Omni Acc
42.5%
Recall
81.7%
Coding
68.6

Speed

Tok/s
118 tok/s
Time/task
3.3 min

Price

$/1M
$0.20 / $1.20
Cost
$0.09
Qwen3.8 27B (xhigh)

Alibaba

Intelligence & accuracy

Intel
34
Omni Acc
15.6%
Recall
82.0%
Coding
68.1

Speed

Tok/s
44 tok/s
Time/task
25.2 min

Price

$/1M
$0.50 / $3.00
Cost
$0.82
Qwen3.7 Max

Alibaba

Intelligence & accuracy

Intel
30
Omni Acc
31.1%
Recall
79.0%
Coding
66

Speed

Tok/s
195 tok/s
Time/task
2.8 min

Price

$/1M
$2.50 / $7.50
Cost
$1.15
GPT-5.6 Luna (high)

OpenAI

Intelligence & accuracy

Intel
32
Omni Acc
41.8%
Recall
80.3%
Coding
63.3

Speed

Tok/s
109 tok/s
Time/task
2.1 min

Price

$/1M
$0.20 / $1.20
Cost
$0.04
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)

Anthropic

Intelligence & accuracy

Intel
31
Omni Acc
40.9%
Recall
80.0%
Coding
63

Speed

Tok/s
45 tok/s
Time/task
27.5 min

Price

$/1M
$3.00 / $15.00
Cost
$2.49
Qwen3.6 27B (Reasoning)

Alibaba

Intelligence & accuracy

Intel
22
Omni Acc
19.6%
Recall
77.3%
Coding
53.7

Speed

Tok/s
57 tok/s
Time/task
9.5 min

Price

$/1M
$0.60 / $3.60
Cost
$0.62
GPT-5.6 Luna (medium)

OpenAI

Intelligence & accuracy

Intel
26
Omni Acc
40.7%
Recall
75.0%
Coding
50.7

Speed

Tok/s
112 tok/s
Time/task
40s

Price

$/1M
$0.20 / $1.20
Cost
$0.02
Claude 4.5 Haiku (Reasoning)

Anthropic

Intelligence & accuracy

Intel
18
Omni Acc
18.0%
Recall
74.3%
Coding
43.9

Speed

Tok/s
85 tok/s
Time/task
3.6 min

Price

$/1M
$1.00 / $5.00
Cost
$0.21
Qwen3.6 35B A3B (Reasoning)

Alibaba

Intelligence & accuracy

Intel
19
Omni Acc
18.8%
Recall
71.7%
Coding
41.9

Speed

Tok/s
127 tok/s
Time/task
4.5 min

Price

$/1M
$0.38 / $2.25
Cost
$0.48

Notes

  • Intelligence Index = Artificial Analysis Intelligence Index v4.0 (composite of 9 evals including GDPval-AA, Terminal-Bench, SciCode, AA-Omniscience). Higher is better. Scores reflect max-effort / adaptive reasoning mode where applicable.
  • AA-Omni Accuracy = AA-Omniscience Accuracy (%). Higher is better. Proportion of correctly answered questions out of all questions. See artificialanalysis.ai/evaluations/omniscience.
  • Recall = Artificial Analysis Long Context Reasoning (AA-LCR) recall accuracy (%). Higher is better. “—” means AA has not measured the model. See artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning.
  • Cost per Task = Weighted average pay-per-token API cost (USD) per Artificial Analysis Intelligence Index task. Lower is better. Reflects input, cache, reasoning, and output token pricing across index evaluations.
  • Coding Agent Index = Artificial Analysis Coding Index — composite pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Higher is better. Scores shown are per-model (best published harness per model from AA's language models dataset); this differs from the separate agents leaderboard (model+harness pairs). "—" means AA has not yet evaluated that model. See artificialanalysis.ai/agents/coding-agents.
  • Model lists are synced from Artificial Analysis via `pnpm sync:model-index` (stable = top coding agent scores; local = curated Qwen open-weight list; experiment = top intelligence scores without coding benchmark, limited to trusted providers).
  • Accuracy rates are sourced from AA-Omniscience via SSR scrape snapshot (`aa-omniscience-snapshot.json`, refreshed with `pnpm fetch:aa-omniscience`). Not included in the free AA sync API — see memory/development/ai-model-index-omniscience-scrape.md.
  • Recall is sourced from the public AA-LCR SSR snapshot (`aa-lcr-snapshot.json`, refreshed with `pnpm fetch:aa-lcr`) because the free AA sync API omits per-benchmark fields. Null means AA has not measured the model.
  • Pricing reflects standard (non-cached) token rates from Artificial Analysis.
  • OpenAI cut GPT-5.6 Luna API pricing 80% and Terra 20% on July 30, 2026 (Luna $0.20/$1.20, Terra $2/$12 per 1M; Sol unchanged).
  • Gemini 3.1 Pro Preview pricing doubles to $4/$18 per 1M above 200K tokens per prompt.
  • Claude Opus 4.7 and Claude Sonnet 4.6 both support a 1M token context window with 128K max output tokens.
  • Claude Fable 5.1 leads the AA Intelligence Index v4.1.1 at 66 (Sept 2026), ahead of Opus 5 at 63 and Fable 5 at 62.
  • Claude Fable 5.1 and Claude Mythos 5.1 released September 1, 2026; Fable 5.1 is on the Experiment tab until AA publishes a Coding Agent Index.
  • Claude Fable 5 and Claude Mythos 5 (Fable 5's unrestricted-cyber/bio counterpart) launched June 9, 2026. Access was suspended June 12–July 1, 2026 for U.S. export-control compliance; current pricing/scores reflect the restored, generally available model.
  • GPT-5.6 ships as three durable capability tiers: Sol (flagship, AA-Omni index 21.7 / accuracy 58.5%), Terra (balanced, -0.2 / 45.9%), and Luna (fastest/cheapest, -11.2 / 41.5%).
  • Grok 4.5 (high) is xAI's first release since its SpaceXAI rebrand and IPO; it's notably cost-efficient per completed task with an AA-Omniscience accuracy of 52.1% (index 26.4).
  • Grok 4.6 (high) scores 61 on the AA Intelligence Index (tying GPT-5.6 Sol max) and 76.8 on the Coding Agent Index — a +4.4-point coding gain over Grok 4.5 at the same $2/$6 per 1M pricing. AA-Omniscience accuracy: 48.2% (index 30.5).
  • Open-weight models (self-hostable): Kimi K3 (weights releasing July 27, 2026), Kimi K2.7 Code, Kimi K2.6, MiMo-V2.5-Pro, GLM 5.3, GLM 5.3 Flash, DeepSeek V4 Pro, DeepSeek V4 Flash, Hy3-preview, Nemotron 3 Super.
  • DeepSeek V4 Pro and V4 Flash OpenRouter pricing shown; first-party DeepSeek API pricing differs ($1.74/$3.48 for Pro, $0.14/$0.28 for Flash).
  • Nemotron 3 Super is also available free-tier on OpenRouter (rate-limited).

Changelog

DateChange
2026-09-16Expanded the Stable and Experiment tabs from the top 25 to the top 50 models each via the model sync limits; refreshed model lists and benchmark snapshots.
2026-09-16Refreshed Stable, Experiment, and Local model lists and benchmark snapshots from Artificial Analysis; added DeepSeek V4.1 Flash and Claude Sonnet 5 effort variants (xhigh, high, medium, low) to the Experiment tab.
2026-09-04Added AA-LCR Recall as a separate benchmark. Configured Grok fast variants use the base Grok 4 recall while retaining their own source pricing and identity.
2026-09-03Added Grok 4 Fast (Reasoning and Non-reasoning) as separate Experiment rows with duplicated Grok 4 benchmark scores and Artificial Analysis speed-specific token pricing.
2026-09-03Refreshed Stable, Experiment, and Local model data from Artificial Analysis; added Claude Fable 5.1 effort variants to Stable and Claude 4.1 Opus (Reasoning) to Experiment.
2026-09-02Added Claude Fable 5.1 (Anthropic) to Experiment tab — AA Intelligence Index 66 (#1). Pinned Claude Fable 5 on Stable tab; refreshed Fable 5 benchmarks.
2026-09-01Refreshed pricing and cost-per-task metrics from Artificial Analysis for all Stable, Experiment, and Local tab models.
2026-08-27Added GLM 5.3 Flash on Stable tab — AA Intelligence Index 58, Coding Agent Index 71.5.
2026-08-20Added GLM 5.3 (max) benchmarks on Stable tab — AA Intelligence Index 60, Coding Agent Index 74.8.
2026-08-20Refreshed DeepSeek V4 Pro 0813 benchmarks on Stable tab — latest GA checkpoint (deepseek-v4-pro).
2026-08-18Added Local tab with curated Qwen 3.6/3.8 open-weight models.
2026-08-15Refreshed AA time-per-task snapshot — all tracked models now have Time/task metrics including Luna effort variants; fixed prior timeout failures.
2026-08-15Added GPT-5.6 Luna thinking-effort variants (xhigh, high, medium) to Stable tab alongside existing Luna (max).
2026-08-14Restored DeepSeek V4 Pro 0813 on Stable tab — AA slug renamed to deepseek-v4-pro.
2026-08-15Models with announced pricing changes show a warning icon beside $/1M — tap or click for details.
2026-08-14Replaced sort dropdown with Coding / Cost / Speed buttons — single-click primary dimension for model list ordering.
2026-08-14Added Tok/s and Time/task columns to desktop table; card and table now share a single field registry so metrics stay in sync.
2026-08-14Added AA time-per-task scores via SSR snapshot — median output speed and intelligence-index output tokens per task scraped from Artificial Analysis model pages.
2026-08-14Refreshed AA-Omniscience accuracy snapshot for all 50 tracked models — 49/50 populated. Claude 4.1 Opus (Reasoning) remains unbenchmarked on AA (omniscience null; deprecated).
2026-08-14Pinned GLM 5.3 (Z.ai) for Stable tab — awaiting Artificial Analysis benchmark publication.
2026-08-13Added DeepSeek V4 Pro 0813 benchmarks (intelligence 53, coding index 68.8) on Stable tab.
2026-08-13Added Grok 4.6 (high) to Stable benchmarks — AA Intelligence Index 61, Coding Agent Index 76.8.
2026-08-03Added Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) to Stable benchmarks — AA Coding Agent Index 63.
2026-08-01Added output-price sort to model table. Pinned Claude Haiku 4.5 (Reasoning) and GPT-5.6 Luna on Stable tab.
2026-08-01Synced GPT-5.6 Luna and Terra price reductions (Luna $0.20/$1.20, Terra $2/$12 per 1M; Sol unchanged at $5/$30) following OpenAI July 30, 2026 API price cut.
2026-07-31Updated DeepSeek V4 Flash 0731 benchmarks (intelligence 50, coding index 69.1) on Stable tab.
2026-07-25Wired AA-Omniscience hallucination rates for all 50 tracked models (stable + experiment) from SSR scrape snapshot.
2026-07-25Fixed benchmark column to show AA-Omniscience Accuracy instead of conditional hallucination rate — values now match AA's prominently displayed metric (e.g. GPT-5.6 Sol ~58.5%, Claude Opus 5 ~54.2%).
2026-07-24Synced Claude Opus 5 (max, xhigh, high, medium) from Artificial Analysis — Opus 5 (max) leads Intelligence Index at 61 and Coding Index at 78.
2026-07-20Added Qwen3.7 Max (Alibaba) to Stable benchmarks. Qwen 3.8 is not yet published on Artificial Analysis — pinned model IDs will pick it up automatically on the next sync once AA adds it.
2026-07-19Added Cost/Task column with relative color scale on Stable and Experiment model cards and tables.
2026-07-18Switched to AA-synced stable/experiment model lists (sync:model-index) instead of manual models.json.
2026-07-17Added Kimi K2.7 Code (Moonshot AI).
2026-07-17Replaced Agentic Index and AA-Omni index with Intelligence Index, hallucination rate, and Coding Agent Index columns.
2026-07-17Added Kimi K3 (Moonshot AI).
2026-07-10Added Grok 4.5 (xAI).
2026-07-10Added Claude Fable 5 (Anthropic).
2026-07-10Added GPT-5.6 Sol, Terra, and Luna (OpenAI).
2026-05-17Added MiniMax M2.7 (MiniMax).
2026-05-17Added Hy3-preview (Tencent).
2026-05-17Filled in AA-Omni for Gemini 3 Flash (12), GPT-5.5 xhigh (18, corrected from 20), Grok 4.3 high (18).
2026-05-17Added GPT-5.5 xhigh (OpenAI).
2026-05-17Added Grok 4.3 high (xAI).
2026-05-17Added Gemini 3 Flash Preview (Google).
2026-05-17Filled in AA-Omni scores for MiMo-V2.5-Pro (4), Qwen3.6 Plus (3), Nemotron 3 Super (-42).
2026-05-17Added MiMo-V2.5-Pro (Xiaomi), Qwen3.6 Plus (Alibaba), Nemotron 3 Super (NVIDIA).
2026-05-17Added Gemini 3.1 Pro Preview. Filled in AA-Omni scores for Claude Sonnet 4.6 (12) and GLM 5.1 (2).
2026-05-17Added AA-Omni column (AA-Omniscience Index from Artificial Analysis).
2026-05-17Added DeepSeek V4 Pro and DeepSeek V4 Flash.
2026-05-17Removed Context and Tier columns.
2026-05-17Added Kimi K2.6 (Moonshot AI) and GLM 5.1 (Z.ai). Added sort-order comment to table.
2026-05-17Initial document created. Added Claude Opus 4.7 and Claude Sonnet 4.6.