Skip to main content
Open research · Model benchmarks

Know which models work best for coding

Independent benchmarks across intelligence, hallucination rates, and coding agent performance for coding model selection.

103

Models tracked

September 16, 2026

Last updated

3

Benchmark dimensions

OpenRouter · AA

Data sources

Benchmarks

Leaderboard snapshot

Top models by Artificial Analysis Intelligence Index. Last updated September 16, 2026.

Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)

Anthropic

#1

Intelligence Index

53

AA-Omni Accuracy
67.2%low
Coding Agent Index
81.6
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)

Anthropic

#2

Intelligence Index

53

AA-Omni Accuracy
66.2%low
Coding Agent Index
80.7
GPT-6 Astra (max)

OpenAI

#3

Intelligence Index

53

AA-Omni Accuracy
62.6%low
Coding Agent Index
76.9

What we build

Open, data-driven tooling for developers working with generative AI — starting with model selection benchmarks.

Model benchmarks

Living index of pricing, intelligence scores, hallucination rates, and coding agent performance.

Explore benchmarks
Composable packages
Coming soon

Early-stage @ck-ai/* monorepo packages — filesystem helpers, HTTP gateways, and SDK primitives for AI workflows.

Research & evaluation
Coming soon

Composable generative-AI tasks, experiments, and model evaluation pipelines — promote stable work into shared packages.

Packages

Early-stage libraries in the CK-AI monorepo — work-in-progress primitives for AI-powered developer tools.

In development
@ck-ai/sdk

Composable filesystem and agent-tool helpers for local workspaces.

listFilesTreereadTextFiles
In development
@ck-ai/gateway

HTTP gateway server — entry point for running services against the monorepo.

Bun runtime
In development
@ck-ai/research

R&D: composable generative-AI tasks, benchmarks, and model evaluation pipelines.

BenchmarksExperiments