Alibaba
Down 0.8 since 23 Sep, rising over 30 days (Direction ▲ 2.0). Latest driver: Alibaba publishes Qwen-Image-2.1 on Hugging Face.
Score over time
32 daily closes, the earlier ones reconstructed. Download CSV
Reconstructed Rebuilt with the live code from evidence dated on or before 24 September 2026. How
The eight pillars
Each pillar 0 to 100 on fixed goalposts, with its rank among the 13 ranked labs. The tick marks the median lab; points are what the pillar adds to Strength.
Why Alibaba is #4
Most of its Strength comes from Capability and Reliability. Points are contributions to the Score.
- Behind Google DeepMind by 13.4 points. Biggest differences: Usage and adoption −5.2, Release execution −2.8, Trust, safety and privacy −2.7; Direction −0.5.
- Ahead of Z.ai by 1.9 points. Biggest differences: Trust, safety and privacy +3.3, Ecosystem −2.1, Usage and adoption +1.9; Direction +0.2.
With every pillar weight moved up to 50% either way, it ranks 4 to 5 in 90% of recalculations; with equal weights, #4.
Since 23 Sep
Score down 0.8 (60.9 to 60.1). Usage and adoption −0.5, Value −0.3.
What was measuredAll 25 readings behind this close, with source, date and goalposts
Filled: no fresh reading, so the ranked labs' median was used. Carried: a source that keeps no history, read later.
| Indicator | Reading | Scaled | Goalposts | As of |
|---|---|---|---|---|
| Arena text rating Arena (arena.ai) text leaderboard: rating of the lab's best model with 1,000 votes or more | 1480.6 qwen3.8-max | 77 | 1250 to 1550 rating | 24 Sep |
| Benchmarks Epoch AI Benchmarking Hub: the lab's best score on each of GPQA Diamond, FrontierMath tiers 1 to 3, SWE-bench Verified and OTIS Mock AIME, each placed between fixed goalposts, averaged; runs from the last 365 days, at least 3 of the 4 | 81.9 GPQA diamond 92.7% (Qwen3.8 Max (xhigh)); FrontierMath-Tiers-1-3-v2-Private 74.7% (Qwen3.8 Max (xhigh)); SWE-Bench verified 77.3% (Qwen3.7 Max); OTIS Mock AIME 2024-2025 100.0% (Qwen3.8 Max (0902) (xhigh)) | 82 | already 0 to 100 | 24 Sep |
| Paid adoption Published spend and paid-use evidence, latest edition of each: Ramp AI Index (share of US businesses paying the vendor, monthly), Menlo Ventures (enterprise API spend share), a16z CIO survey (share of LLM wallet), ICONIQ (share of AI builders using the provider), each placed on its own scale and averaged by weight; facts and sources at /data/adoption.json | 1.8 Menlo Ventures, State of Generative AI in the Enterprise (2025-12): not named, below its threshold; a16z enterprise CIO survey (2026-01): not named, below its threshold; ICONIQ, State of AI 2026: The Builder's Economy (2026-07): 6% | 2 | already 0 to 100 | 24 Sep |
| Assistant users Users of the lab's assistants (chat apps and coding agents) as the lab or its parent last disclosed them, or a panel measurement (QuestMobile) where a lab discloses none; weekly figures count as monthly (a lower bound); dated by disclosure, and only disclosures from the last 12 months count; not applicable to a lab with no figure; facts and sources at /data/assistants.json | 300000000 Qwen, all platforms: consumer-facing Qwen has surpassed 300 million monthly active users across all platforms (February, the Chinese New Year peak) (Alibaba, December quarter results, 2026-03-19) | 75 | 1 million to 2 billion monthly users | 24 Sep |
| Web traffic Cloudflare Radar: the lab's best rank among the top 20 generative-AI services by traffic, 21 minus the rank | 0 not in the top 20 generative-AI services | 0 | 0 to 20 rank points | 24 Sep |
| API spend share OpenRouter rankings: the lab's share of all dollars spent on OpenRouter in the last 30 days, across every model (not token volume, which cheap bulk models dominate) | 2.4973 2.5% of $95.1M spent on OpenRouter in 30 days (76 models) | 27 | 0% to 35% of spend | 24 Sep |
| App charts App Store top free charts in 7 countries | 14 CN #3 | 14 | 0 to 100 chart points | 24 Sep |
| Coding agent downloads npm and PyPI downloads of the lab's official coding agents (such as Claude Code, Codex, Gemini CLI, Qwen Code, Mistral Vibe), last 30 days; not applicable to a lab with none; standalone IDEs without public counts (such as Antigravity) are not seen | 235968 npm:@qwen-code/qwen-code 235968 | 34 | 10 thousand to 100 million a month | 24 Sep |
| Coding extension installs VS Code Marketplace: installs of the lab's official coding extensions; not applicable to a lab with none | 2712940 Carried Carried back from 2026-09-24: Alibaba-Cloud.tongyi-lingma 2712940 | 53 | 100 thousand to 50 million installs | 24 Sep |
| SDK downloads PyPI and npm downloads of the official SDKs, last 30 days | 2497548 pypi:dashscope 2497548 | 35 | 100 thousand to 1 billion a month | 24 Sep |
| Open-model downloads Hugging Face downloads of the lab's models, last 30 days | 310014458 Qwen | 90 | 10 thousand to 1 billion a month | 24 Sep |
| Shipping cadence Vikshy Wire: confirmed launches and updates of materiality 2 or more, last 90 days | 4 4 confirmed launches and updates in 90 days | 37 | 0 to 30 in 90 days | 24 Sep |
| Biggest launch Vikshy Wire: the biggest launch in the last 90 days, fading with age | 2.83 | 57 | materiality 0 to 5 | 24 Sep |
| Rating per dollar Arena rating of the lab's best-value model, less 30 points for each doubling of its blended list price (3 parts input to 1 part output), with prices below $0.30 per million counted as $0.30 | 1480.7 qwen3.7-plus: 1456 at $0.56/M blended | 58 | 1250 to 1650 points | 24 Sep |
| Cloud availability Amazon Bedrock, Google Cloud and Microsoft Foundry: whether each offers the lab's current models as a managed, pay-per-use API (the flagship or the model one step behind it counts 1, an older generation only 0.5, open-weight side lines 0), dated by availability; facts and sources at /data/distribution.json | 1 Amazon Bedrock: Qwen3 235B, Coder, Next; no Qwen 3.5 or later (older generation, half); Google Cloud: Qwen3 (deprecated, retiring 21 Oct 2026) (older generation, half) | 33 | 0 to 3 platforms | 24 Sep |
| Open-weight hosts OpenRouter: distinct providers serving the lab's open-weight models; not applicable to a lab that publishes none (a closed model's reach is its cloud availability) | 20 Alibaba, Novita, SiliconFlow, Venice, Modal, DeepInfra, Together, Wafer, Reka, Darkbloom, DekaLLM, Phala | 73 | 1 to 60 providers | 24 Sep |
| Developer interest Hugging Face likes on the lab's models | 136747 Qwen | 85 | 1 thousand to 316 thousand likes | 24 Sep |
| Hacker News Hacker News points on stories about the lab, last 30 days | 2597 7 stories; top: Qwen Image 2.1 (737) | 76 | 0 to 30,000 points | 24 Sep |
| Press and analysts Vikshy Wire: press and analyst pieces on the lab's developments, last 30 days | 4 | 26 | 0 to 60 pieces | 24 Sep |
| Official API uptime The lab's own status page: 30-day uptime of the components that make up its model API, counting major outages in full and partial outages at 30%, the pages' own formula (degraded performance is not counted); labs label outages differently, so a measured cross-check sits beside it | No fresh reading: filled with the median | 86 | 99% to 100% over 30 days | |
| Measured uptime OpenRouter: the median uptime of the lab's own API endpoints, an independent cross-check of the status page, the mean of daily readings over 30 days | 99.9919 mean of 30 daily readings, 2026-08-26 to 2026-09-24 | 100 | 90% to 100% | 24 Sep |
| Independent assessments Independent assessments published by then, each put on 0 to 100: the Future of Life Institute AI Safety Index (40%), the Stanford Foundation Model Transparency Index (30%) and the SaferAI risk-management ratings (30%), averaged over those that cover the lab; a lab none covers is not independently assessed, and filled with the median like any gap | 22.7 2 of 3 assessments: FLI AI Safety Index (2026-07) 0.87; Foundation Model Transparency Index (2025-12) 26 | 23 | 0 to 100 | 24 Sep |
| Customer data commitments From the lab's own documentation: API and business data not used for training by default; a zero-data-retention option; a choice of data region | 2 no training on API data by default, choice of data region | 67 | 0 to 3 commitments | 24 Sep |
| Security certifications From the lab's trust center: SOC 2 Type II, ISO/IEC 27001, ISO/IEC 42001 (AI management) covering its model API | 2 SOC 2 Type II, ISO/IEC 27001 | 67 | 0 to 3 certifications | 24 Sep |
| Security and privacy incidents Security breaches, data leaks and regulatory privacy actions (bans, fines, orders) affecting the lab's AI services or caused by its AI systems in the last 12 months, from reputable reporting and regulators, counted once per public disclosure, and half when the lab itself disclosed harm from its AI with no legal duty to; every fact, and what was not counted and why, is at /data/trust.json | 0 no security or privacy incident or regulatory action found in 12 months | 100 | 0 to 3 in 12 months, fewer is better | 24 Sep |
What moved Direction
The developments that count most in the last 30 days. Direction ▲ 2.0.
- +1.2Alibaba publishes Qwen-Image-2.1 on Hugging Face14 Sep · New open-weight model
- +0.5Qwen releases Qwen-Drive-1.0-4B on Hugging Face27 Aug · New open-weight model
- +0.3Alibaba releases Qwen 3.8 Omni Flash17 Sep · New open-weight model
All 3 developments, line by line
Points for materiality, times recency, times diminishing returns (or the setback cap).
| Development | Materiality | Points | Recency | Diminish or cap | Counts |
|---|---|---|---|---|---|
| Qwen releases Qwen-Drive-1.0-4B on Hugging Face 27 Aug · New open-weight model | 3 | 1.5 | 0.53 | 0.63 | +0.5 |
| Alibaba publishes Qwen-Image-2.1 on Hugging Face 14 Sep · New open-weight model | 3 | 1.5 | 0.83 | 1 | +1.24 |
| Alibaba releases Qwen 3.8 Omni Flash 17 Sep · New open-weight model | 2 | 0.75 | 0.9 | 0.45 | +0.31 |
Latest from Alibaba
What Alibaba shipped, broke and changed lately. The full Alibaba timeline, with every model it has released.
- 12 AugQwen3.8-27B-FP8 quantized model published on Hugging FaceModels · New-generation model · 1M context · $0.42 in / $2.55 out per 1M
- 3 AugQwen3.8-27B released on Hugging FaceModels · New model generation from a lab · 1M context · $2.00 in / $6.00 out per 1M