The July 2026 AI model rankings from the AI Rank leaderboard - DataCore's real-time aggregator pulling from Artificial Analysis, OpenRouter, METR, and Aider Polyglot - shows a decisive shift in the competitive landscape. Anthropic's Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) now sits at #1 with a composite score of 65.8. More striking: Anthropic holds five of the top ten spots outright - positions 1, 2, 4, and 5 - while OpenAI's GPT-5.5 (xhigh) is the only non-Claude model in the top 6.

Top 10 this month
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) - Anthropic - 65.8
- Claude Opus 4.8 (Adaptive Reasoning, Max Effort) - Anthropic - 62.3
- GPT-5.5 (xhigh) - OpenAI - 61.9
- Claude Opus 4.7 (Adaptive Reasoning, Max Effort) - Anthropic - 60.6
- Claude Sonnet 5 (Adaptive Reasoning, Max Effort) - Anthropic - 59.8
- GPT-5.5 - OpenAI - 59.6
- GPT-5.4 (xhigh) - OpenAI - 58.4
- GLM-5.2 (max) - Zhipu AI (open weights) - 57.3
- Gemini 3.1 Pro Preview - Google - 54.4
- GPT-5.2 (xhigh) - OpenAI - 54.4
How these AI model rankings are calculated
The AI model rankings combine two primary axes: an intelligence index covering reasoning, knowledge, and language tasks, and a coding index measuring how well each model writes, edits, and refactors real code. Claude Fable 5 posts 59.9 on intelligence and 76.5 on coding, which explains its composite lead - its coding advantage over the rest of the field is currently the widest gap on the board.
Beyond the composite score, the leaderboard tracks specialist categories: generation speed in tokens per second, Aider Polyglot completion rates, METR time horizon, math accuracy, context length, and price per million tokens. A model that wins the composite is not automatically the right pick for a latency-sensitive chatbot or a document-heavy retrieval pipeline, so the category views often matter more than the headline number.

The open-weights picture
GLM-5.2 (max) by Zhipu AI is the only open-weights model in the top 10, scoring 57.3 at just $2.15 per million output tokens - a compelling price-to-performance ratio. The AI model rankings now track 222 open-weights models out of 534 total (~41.6%), so the ecosystem is deep even if the very top remains dominated by proprietary models.
Specialized benchmarks
GPT-5 Chat leads the Aider Polyglot code-editing leaderboard at 88.0%, edging Claude on that specific metric. Claude Opus 4.6 holds the longest METR time horizon at 14 hours - meaning it can autonomously complete tasks that would take a human 14 hours, more reliably than any other model. Mercury 2 remains the speed king at 1,019 tokens per second. GPT-5.2 (xhigh) scores a near-perfect 99 on math. Llama 4 Scout (Meta, open weights) tops the board for context length at 10 million tokens. Best value this month goes to HyperNova 60B 2605 - the highest intelligence-per-dollar on the board.
The big picture
Anthropic's new generation is pulling ahead on the composite score. OpenAI is competitive but fragmented across many GPT-5.x variants. Google is holding steady in the top 10 with Gemini 3.1 Pro Preview. Open-weights models continue to close the gap on specialized tasks and price - and pricing overall keeps compressing, with several frontier-class models now available well under $5 per million output tokens.
AI model rankings top ten, model by model
Claude Fable 5 (65.8) takes the top spot on the strength of its coding index, and its adaptive reasoning mode with an Opus 4.8 fallback gives it unusual consistency across task types. Claude Opus 4.8 (62.3) remains the safest pick for teams that want proven behavior over peak scores.
GPT-5.5 xhigh (61.9) is OpenAI's strongest entry and the only non-Claude model in the top six, though the xhigh effort setting carries a meaningful latency and cost premium. Claude Opus 4.7 (60.6) and Claude Sonnet 5 (59.8) round out the Anthropic block, with Sonnet 5 offering the best cost-to-composite ratio of the five.
GPT-5.5 (59.6) and GPT-5.4 xhigh (58.4) show how tightly OpenAI's variants cluster - fragmentation that makes model selection inside the GPT line its own project. GLM-5.2 max (57.3) is the story of the month for value buyers at $2.15 per million output tokens.
Gemini 3.1 Pro Preview (54.4) keeps Google inside the top ten while its next generation is in preview, and GPT-5.2 xhigh (54.4) ties it - notable mainly for its near-perfect 99 math score, which still makes it a specialist pick for quantitative workloads.
Where the numbers come from
The AI model rankings aggregate benchmark data from Artificial Analysis, OpenRouter, METR, and Aider Polyglot, normalized into the composite and category scores shown above. Aggregation smooths out single-benchmark quirks, but no leaderboard replaces evals on your own workload - treat the AI model rankings as a map, not the territory.
What the July AI model rankings mean for your stack
If you are building on large language models today, three signals in this month's AI model rankings deserve attention. First, the spread between the top five composite scores is about six points - small enough that task-specific benchmarks should drive model selection rather than overall rank. Second, pricing keeps compressing: with several frontier models under $5 per million output tokens and GLM-5.2 at $2.15, the build-versus-buy math for self-hosting changes almost quarterly. We covered the hardware side of that equation in our analysis of China's LineShine CPU supercomputer topping the TOP500.
Third, reliability metrics are becoming differentiators. METR's 14-hour time horizon for Claude Opus 4.6 matters more than raw benchmark points for teams deploying long-running agents, and vendor concentration risk deserves a line in any procurement review - a topic closely related to the data security questions we examined in the FireAnt MetaKit incident analysis.

A quick checklist for reading the AI model rankings
- Set a hard budget ceiling per million output tokens and filter the board before comparing scores.
- Match the index to the job: coding index for developer tools, intelligence index for analysis and advisory work.
- Check tokens-per-second if your product needs real-time responses.
- Compare context length against your median document size, not your largest outlier.
- Shortlist open-weights options when data residency rules apply to your workloads.
- Re-run your own evals on every shortlisted model before signing anything long-term.

Reading the AI model rankings from Vietnam
For teams in Vietnam and the wider Southeast Asian market, two practical notes apply this month. Data residency requirements increasingly shape which entries on the AI model rankings are actually deployable: open-weights models that can run inside domestic infrastructure carry a compliance advantage that no composite score captures. That makes the depth of the open-weights ecosystem - now more than two fifths of all tracked models - as important a headline as who sits at number one.
Vietnamese-language quality also remains unevenly distributed across the board. Models that sit within a point or two of each other on the composite can differ visibly on Vietnamese summarization, extraction, and tone control. Budget a short internal benchmark - even fifty representative prompts - before committing, and re-run it whenever the AI model rankings show a new entrant in your price band.
Finally, watch the cadence of releases rather than any single month. The July snapshot shows Anthropic ahead, but the six-point spread across the top five means the order can change with one release cycle. Teams that architect for model portability - an abstraction layer over the provider API, evals in version control, and a monthly review of the AI model rankings - capture price drops and capability jumps without re-platforming.
FAQ: AI model rankings
How often do the AI model rankings change? The leaderboard updates in real time as new benchmark runs land, and we publish a snapshot commentary each month. Positions inside the AI model rankings top ten routinely shift after major model releases.
Do the rankings reflect performance in languages other than English? Only partially. Most standardized benchmarks are English-first, so teams working in Vietnamese or other languages should treat the AI model rankings as a shortlisting tool and validate finalists on their own multilingual data.
What is the difference between the composite score and category leaderboards? The composite blends intelligence and coding into a single number for quick comparison, while category leaderboards isolate one dimension such as speed, math, or context length. Procurement decisions usually hinge on two or three categories that map directly to the workload, so start with those and use the composite as a tiebreaker.
Are open-weights models catching up in the AI model rankings? Steadily. With 222 of 534 tracked models now open-weights and GLM-5.2 holding a top-ten composite spot at a fraction of proprietary pricing, the gap is narrowing on specialized tasks even though the very top of the board remains proprietary this month.
Can I compare models across different effort settings? Yes, but read carefully: entries like GPT-5.5 and GPT-5.5 xhigh are the same base model at different reasoning-effort levels, with materially different latency and cost. The AI model rankings list them separately precisely because their real-world behavior diverges - budget-sensitive deployments often prefer the standard setting even when the xhigh variant scores a point or two higher.
Track the full AI model rankings in real time - intelligence, coding, math, speed, and cost - at https://airank.datacore.vn






Để lại một bình luận
You must be logged in to post a comment.