Independent AI analysis
The most complete AI ranking of 2026, with 727 active LLMs compared by AA Intelligence Index, AA Coding Index and LMArena Elo — covering reasoning, coding, speed and cost — plus latency and per-token pricing metrics. Use this ranking to find the best AI models of 2026 by category.
Luis Fernando Roquette · SWEN · methodology described at the bottom of this page · last updated: Sep 01, 2026
Use now
Top 10 · AA Intelligence Index
Top 10 · Output tokens/second
Top 10 · USD / 1M tokens input
Y: Intelligence Index · X: $/1M tokens · Tamanho: Tok/s
Canto superior esquerdo é o ponto ideal: inteligência alta + preço baixo. Tamanho do círculo proporcional à velocidade de output.
Eixo X em escala logarítmica. Apenas modelos com Intelligence Index ≥ 30 e preço público.
Ranking by Artificial Analysis composite score (0–100). Top 30 benchmark models.
Daily progression of the Intelligence Index for the top 8 models.
According to the AA Intelligence Index — a composite index aggregating GPQA Diamond, MMLU-Pro, AIME, HLE and LiveCodeBench — Claude Opus 5 (Anthropic) leads the ranking in 2026 with a score of 63.1/100, followed by Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (62.1) and GPT-5.6 Sol (max) (60.9). The Intelligence Index is calculated by Artificial Analysis based on independent evaluations and reflects real technical capability in reasoning, math, science and coding. It differs from LMArena ELO, which measures human preference in open conversations. For tasks requiring deep reasoning, code or scientific analysis, models at the top of the Intelligence Index typically perform best. For everyday conversations and creativity, ELO is a more representative guide. See the updated ranking for real-time positions.
ELO comes from LMArena (Chatbot Arena), where real users compare responses from two anonymized models and pick the best one. It is a measure of subjective human preference — reflecting naturalness, usefulness and perceived quality in everyday conversations. A model with a high ELO may not be the most accurate on technical tasks, but it is what people prefer to use. The AA Intelligence Index, calculated by Artificial Analysis, is objective: it aggregates results from standardized benchmarks such as GPQA Diamond (PhD-level questions), MMLU-Pro (broad academic knowledge), AIME (olympiad math), HLE (frontier scientific knowledge) and LiveCodeBench (programming). The higher the score, the more technical capability the model demonstrated in controlled evaluations. Use ELO to choose a general conversational assistant; use the Intelligence Index to select models for technical or scientific pipelines.
For coding, the most relevant benchmarks are LiveCodeBench — code challenges evaluated with real execution — and the AA Coding Index. In 2026, GPT-5.6 Sol (batch) leads the coding ranking (78.3/100), with Claude Opus 5 in second and Grok 4.6 in third. The ideal choice depends on context: for code generation via API, cost per token and context window matter as much as accuracy. For interactive IDE development (Cursor, VS Code), latency is critical. For multi-file projects, context windows above 100K tokens are required. See the full table to compare coding models by score, price and speed.
The SWEN ranking is updated automatically and continuously from three main sources. Artificial Analysis benchmark data (Intelligence Index, Coding Index, Math Index, inference speed) is synced every 6 hours via automated integration. API pricing — input and output per 1M tokens — is updated daily via OpenRouter, reflecting provider changes in near real time. LMArena ELO (Chatbot Arena) is synced weekly. The page revalidates its cache every 5 minutes via ISR (Incremental Static Regeneration): when a new model enters or a score changes, the ranking updates within 5 minutes without a manual rebuild. The last sync occurred on Sep 01, 2026.
Google's Gemini 3 family does not follow sequential linear numbering. Google released Gemini 3 Flash, Gemini 3.1 Pro/Flash Lite and Gemini 3.5 Flash — without publishing an official “Gemini 3.2”. Each number denotes a distinct technical generation: 3.1 brought reasoning improvements; 3.5 expanded capability at an intermediate cost. Gemini 3.1 Pro costs $2.00/1M tokens with a 1-million-token context window, positioning itself as an alternative to GPT-4o and Claude 3.7. See the full Gemini 3 family comparison →
“Gemini Spark” is a name circulating online that Google has never officially launched as a product. The term appeared in APK teardowns linked to a possible ultra-lightweight version of Gemini for edge devices. Google's confirmed lightweight models are: Gemini Nano (on-device, Pixel 8 Pro/Pixel 9) and Gemini Flash(via API, $0.075/1M tokens). Any prediction about “Gemini Spark” is speculation until official confirmation. Read what is known about Gemini Spark →
“Better” depends entirely on the task. For general conversation and versatility, ChatGPT (GPT-5.6) remains the most popular choice with the largest user base. For long-form reasoning and technical writing, Claude Opus 5 scores higher on the AA Intelligence Index (60.7 vs 58.9). For coding, both are nearly tied on LiveCodeBench. For cost-efficiency, DeepSeek V4 and Gemini Flash deliver strong quality at significantly lower prices. The “best” AI depends on what you need it for. Compare ChatGPT vs Claude vs Gemini side by side →
According to the Andreessen Horowitz (a16z) Top 100 AI Apps report and Forbes data, the three most used AI platforms globally are ChatGPT (OpenAI), Gemini (Google), and DeepSeek, followed by Canva and Perplexity. However, popularity doesn’t equal technical capability — Claude Opus 5 leads the AA Intelligence Index despite not being in the top 3 by user count. For professional and technical use cases, benchmark scores (Intelligence Index, LiveCodeBench, ELO) are more relevant than raw usage statistics. See the SWEN technical ranking →
The most comprehensive and rigorous AI ranking in the world is the AA Intelligence Index by Artificial Analysis, which aggregates 10 standardized evaluations (GPQA Diamond, MMLU-Pro, AIME, HLE, LiveCodeBench, and others) into a single 0–100 score. As of 2026, the global top 5 is: Claude Opus 5 (Anthropic, 60.7), GPT-5.6 (OpenAI, 58.9), Gemini 3.5 Pro (Google, 57.1), Claude Fable 5 (Anthropic, 59.9 batch), and GPT-5.6 Sol Pro (OpenAI, 57.7). See the full ranking updated every 3h →
Artificial Analysis — provides Intelligence Index and Coding Index. Synced every 6h via automated cron (our account is on the Free tier — individual sub-benchmarks like GPQA, MMLU-Pro and AIME are Pro-only and not synced).
LMArena — Human preference ELO in blind side-by-side comparisons. Updated weekly.
OpenRouter — provider pricing in USD per 1M tokens. Updated daily.
Historical snapshots — daily score capture at 06:30 UTC to feed temporal evolution charts. Started on Sep 01, 2026.
Benchmarks are indicative — always test on your specific use case before deciding. Performance varies by inference provider (same model, different latency).