Know which models work best for coding
Independent benchmarks across intelligence, hallucination rates, and coding agent performance for coding model selection.
103
Models tracked
September 16, 2026
Last updated
3
Benchmark dimensions
OpenRouter · AA
Data sources
Leaderboard snapshot
Top models by Artificial Analysis Intelligence Index. Last updated September 16, 2026.
Anthropic
Intelligence Index
53
- AA-Omni Accuracy
- 67.2%low
- Coding Agent Index
- 81.6
Anthropic
Intelligence Index
53
- AA-Omni Accuracy
- 66.2%low
- Coding Agent Index
- 80.7
OpenAI
Intelligence Index
53
- AA-Omni Accuracy
- 62.6%low
- Coding Agent Index
- 76.9
What we build
Open, data-driven tooling for developers working with generative AI — starting with model selection benchmarks.
Living index of pricing, intelligence scores, hallucination rates, and coding agent performance.
Explore benchmarksEarly-stage @ck-ai/* monorepo packages — filesystem helpers, HTTP gateways, and SDK primitives for AI workflows.
Composable generative-AI tasks, experiments, and model evaluation pipelines — promote stable work into shared packages.
Packages
Early-stage libraries in the CK-AI monorepo — work-in-progress primitives for AI-powered developer tools.
Composable filesystem and agent-tool helpers for local workspaces.
listFilesTreereadTextFilesHTTP gateway server — entry point for running services against the monorepo.
Bun runtimeR&D: composable generative-AI tasks, benchmarks, and model evaluation pipelines.
BenchmarksExperiments