Independent analysis of +750 AI models from top companies. Chatbot Arena ELO, Intelligence Index, pricing and specs. Updated daily.
SWEN.AI is the largest independent AI benchmark portal in Brazil, tracking over +750 large language models (LLMs) with data updated every 6 hours. The SWEN benchmark ranks AI models using the Artificial Analysis Intelligence Index — a composite score from 0 to 100 that aggregates standardized evaluations including GPQA Diamond (PhD-level science), MMLU-Pro (broad academic knowledge), AIME (olympiad mathematics), HLE (frontier scientific reasoning) and LiveCodeBench (programming). Unlike popularity rankings based on user votes or website traffic, SWEN's methodology prioritizes objective technical capability measured through controlled, reproducible benchmarks. Claude Opus 5 (Anthropic) currently leads with an AA Score of 60.7 out of 100, followed by GPT-5.6 (OpenAI) and Gemini 3.5 Pro (Google). The ranking also includes pricing in Brazilian Real (BRL), inference speed in tokens per second, and context window size — allowing Brazilian professionals and companies to compare AI models based on technical merit, cost and performance rather than marketing claims. All data is sourced from Artificial Analysis, LMArena and OpenRouter, with full methodology publicly documented.
By Luis Fernando Roquette • Last updated: September 01, 2026
771 models • 630 with benchmarks • 573 with AA Score • Synced: September 01, 2026
What is the best LLM today?
Most Intelligent
AA Score — Artificial Analysis · updated every 6h
Fastest?
Fastest
Output tokens/second · higher is better
Best value for money?
Most Value per Dollar
AA Score per US$ (blended 3:1)
Score AA — Artificial Analysis · top 20
Score AA = AA Intelligence Index from Artificial Analysis. Updated every 6h. Click any model for detailed benchmarks.
771
Models
108
Companies
573
With AA Score
103
Reasoning
130
Open Source
219
Multimodal
Classification based on the AA Intelligence Index by Artificial Analysis — composite score updated every 6 hours.
An AI benchmark is an independent, standardized test that measures and compares the real-world capability of artificial intelligence models across specific tasks. Unlike marketing claims from AI companies, benchmarks provide objective, verifiable metrics that help developers and businesses choose the right model for their use case.
The 3 criteria that matter most when comparing AI models:
ELO: Human preference ranking from LMArena (Chatbot Arena). Real users compare blind model responses and vote. Higher = people prefer it for everyday conversations.
Intel. (AA Intelligence Index): Composite score 0-100 by Artificial Analysis aggregating GPQA Diamond, MMLU-Pro, AIME, HLE, LiveCodeBench and more. Measures real technical capability — reasoning, math, coding, science.
Input $/1M: Price in USD per 1 million input tokens (≈750,000 words). What you pay to send text to the model. Output tokens cost separately.
Context: Maximum tokens the model can process at once — determines how much text (code, documents, conversation history) fits in a single request.
Tokens/s: Generation speed — how many tokens the model outputs per second. Higher = faster responses, lower latency, better for real-time apps.
| # | Model | ELO | Input $/1M |
|---|---|---|---|
| 🥇 | Claude Opus 5 | 1492 | $5.00 |
| 🥈 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | 1507 | $10.00 |
| · | Claude Fable 5 (batch) | — | $5.00 |
| 4 | Grok 4.6New | 1461 | $2.00 |
| · | GPT-5.6 Sol (max) | — | $4.00 |
| · | Kimi K3 | — | $3.00 |
| 7 | GLM-5.3New | 1484 | $1.40 |
| 8 | GPT-5.6 Sol (xhigh) | 1482 | $4.00 |
| · | GPT-5.6 Sol (batch) | — | $1.00 |
| · | Claude Opus 5 (batch) | — | $2.50 |
| · | Claude Opus 5 (Fast) | — | $10.00 |
| · | Qwen3.8 MaxNew | — | $2.00 |
| · | Qwen3.8 2.4T A95BNew | — | $2.00 |
| 14 | GLM-5.3-FlashNew Z AI | 1469 | $0.15 |
| 15 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | 1481 | $5.00 |
| · | Claude Opus 4.8 (batch) | — | $2.50 |
| · | GPT-5.6 Sol (high) | — | $4.00 |
| 18 | Muse Spark 1.2 (xhigh)New | 1498 | $1.25 |
| · | GPT-5.6 Terra (max) | — | $2.00 |
| 20 | GPT-5.5 | 1482 | $5.00 |
Prices in USD per 1M input tokens. See methodology for details.
OS = Open Source • MM = Multimodal • R = Reasoning •AA Score: Artificial Analysis • Intel.: Artificial Analysis •Pricing: OpenRouter •View full methodology
Daily progression of the AA Intelligence Index for top models. Higher = more capable.
Tokens per second — top 15
Speed measured in tokens/second via API. TTFT = Time to First Token (latency until the first response).
Claude Opus 5 is the most intelligent AI model in 2026 with an AA Score of 63.1, according to the AA Intelligence Index by Artificial Analysis — a composite score updated every 6 hours. The language model (LLM) market in 2026 is dominated by a race between OpenAI, Anthropic, Google DeepMind, Meta AI and labs like DeepSeek, Alibaba (Qwen) and xAI (Grok). With over 771 models available via API, choosing the right model for each use case has become a complex decision involving quality (measured by benchmarks like AA Intelligence Index, MMLU and SWE-bench), token pricing, inference speed, context and multimodal capabilities.
Intelligence Index methodology — composite score composition. Sources: Artificial Analysis, LMArena, OpenRouter. Updated every 6h. CC BY 4.0.
According to the AA Intelligence Index by Artificial Analysis, Claude Opus 5 leads with a Score of 63.1. However, "best" depends on use case: for value, OpenAI: GPT-5 Nano (batch) offers excellent quality at low cost.
The AA Score is the Intelligence Index by Artificial Analysis — a composite score (0-100) combining multiple reasoning, code, math and science benchmarks. It is automatically updated every 6 hours from the public Artificial Analysis API.
Among models with good quality (AA Score > 40), Agnes 2.5 Pro Beta is the most affordable at $0.10/1M input tokens.
Benchmarks are indicative, not definitive. The AA Intelligence Index is considered robust because it combines multiple standardized evaluations. Individual benchmarks (MMLU, GPQA) can suffer from contamination. We recommend testing on your specific use case.
ChatGPT (OpenAI) is the most popular and versatile. Claude (Anthropic) excels at long reasoning and quality text. Gemini (Google) integrates well with the Google ecosystem. The ideal choice depends on your use case — use our comparator at /benchmark/comparar.