Method v2.3 · 2026-09-24

How the Vikshy Score works.

The Vikshy Score is a daily, open ranking of the leading AI labs, built only from public evidence: how strong each lab is today, which way it is heading, and exactly why.

Strength is where a lab stands, from eight measured pillars. Direction is where it is heading, from the last 30 days of confirmed developments. Score = Strength + 0.18 × Direction. It closes daily at 06:00 UTC; every close is kept with the method that made it, and everything below is also published as data, so anyone can recompute a close.

What Strength measures

Eight pillars, each measured from public data. Weights add to 100.

Capability20

How strong the models are on independent evaluations.

  • Arena (arena.ai) text leaderboard: rating of the lab's best model with 1,000 votes or more (goalposts 1250 to 1550 rating)
  • Epoch AI Benchmarking Hub: the lab's best score on each of GPQA Diamond, FrontierMath tiers 1 to 3, SWE-bench Verified and OTIS Mock AIME, each placed between fixed goalposts, averaged; runs from the last 365 days, at least 3 of the 4 (goalposts already 0 to 100)

Usage and adoption16

How many people and companies use them, from API share, app charts and downloads.

  • Published spend and paid-use evidence, latest edition of each: Ramp AI Index (share of US businesses paying the vendor, monthly), Menlo Ventures (enterprise API spend share), a16z CIO survey (share of LLM wallet), ICONIQ (share of AI builders using the provider), each placed on its own scale and averaged by weight; facts and sources at /data/adoption.json (goalposts already 0 to 100)
  • Users of the lab's assistants (chat apps and coding agents) as the lab or its parent last disclosed them, or a panel measurement (QuestMobile) where a lab discloses none; weekly figures count as monthly (a lower bound); dated by disclosure, and only disclosures from the last 12 months count; not applicable to a lab with no figure; facts and sources at /data/assistants.json (goalposts 1 million to 2 billion monthly users)
  • Cloudflare Radar: the lab's best rank among the top 20 generative-AI services by traffic, 21 minus the rank (goalposts 0 to 20 rank points)
  • OpenRouter rankings: the lab's share of all dollars spent on OpenRouter in the last 30 days, across every model (not token volume, which cheap bulk models dominate) (goalposts 0% to 35% of spend)
  • App Store top free charts in 7 countries (goalposts 0 to 100 chart points)
  • npm and PyPI downloads of the lab's official coding agents (such as Claude Code, Codex, Gemini CLI, Qwen Code, Mistral Vibe), last 30 days; not applicable to a lab with none; standalone IDEs without public counts (such as Antigravity) are not seen (goalposts 10 thousand to 100 million a month)
  • VS Code Marketplace: installs of the lab's official coding extensions; not applicable to a lab with none (goalposts 100 thousand to 50 million installs)
  • PyPI and npm downloads of the official SDKs, last 30 days (goalposts 100 thousand to 1 billion a month)
  • Hugging Face downloads of the lab's models, last 30 days (goalposts 10 thousand to 1 billion a month)

Release execution6

How often a lab ships, and how big the releases are.

  • Vikshy Wire: confirmed launches and updates of materiality 2 or more, last 90 days (goalposts 0 to 30 in 90 days)
  • Vikshy Wire: the biggest launch in the last 90 days, fading with age (goalposts materiality 0 to 5)

Value15

Capability per dollar at list prices.

  • Arena rating of the lab's best-value model, less 30 points for each doubling of its blended list price (3 parts input to 1 part output), with prices below $0.30 per million counted as $0.30 (goalposts 1250 to 1650 points)

Ecosystem11

How widely the models are hosted, and developer interest in them.

  • Amazon Bedrock, Google Cloud and Microsoft Foundry: whether each offers the lab's current models as a managed, pay-per-use API (the flagship or the model one step behind it counts 1, an older generation only 0.5, open-weight side lines 0), dated by availability; facts and sources at /data/distribution.json (goalposts 0 to 3 platforms)
  • OpenRouter: distinct providers serving the lab's open-weight models; not applicable to a lab that publishes none (a closed model's reach is its cloud availability) (goalposts 1 to 60 providers)
  • Hugging Face likes on the lab's models (goalposts 1 thousand to 316 thousand likes)

Reception4

How much attention the community and the press give the lab.

  • Hacker News points on stories about the lab, last 30 days (goalposts 0 to 30,000 points)
  • Vikshy Wire: press and analyst pieces on the lab's developments, last 30 days (goalposts 0 to 60 pieces)

Reliability11

Measured uptime of the lab's own API, over 30 days.

  • The lab's own status page: 30-day uptime of the components that make up its model API, counting major outages in full and partial outages at 30%, the pages' own formula (degraded performance is not counted); labs label outages differently, so a measured cross-check sits beside it (goalposts 99% to 100% over 30 days)
  • OpenRouter: the median uptime of the lab's own API endpoints, an independent cross-check of the status page, the mean of daily readings over 30 days (goalposts 90% to 100%)

Trust, safety and privacy17

Independent safety and transparency assessments, security certifications, commitments on customer data, and incidents.

  • Independent assessments published by then, each put on 0 to 100: the Future of Life Institute AI Safety Index (40%), the Stanford Foundation Model Transparency Index (30%) and the SaferAI risk-management ratings (30%), averaged over those that cover the lab; a lab none covers is not independently assessed, and filled with the median like any gap (goalposts 0 to 100)
  • From the lab's own documentation: API and business data not used for training by default; a zero-data-retention option; a choice of data region (goalposts 0 to 3 commitments)
  • From the lab's trust center: SOC 2 Type II, ISO/IEC 27001, ISO/IEC 42001 (AI management) covering its model API (goalposts 0 to 3 certifications)
  • Security breaches, data leaks and regulatory privacy actions (bans, fines, orders) affecting the lab's AI services or caused by its AI systems in the last 12 months, from reputable reporting and regulators, counted once per public disclosure, and half when the lab itself disclosed harm from its AI with no legal duty to; every fact, and what was not counted and why, is at /data/trust.json (goalposts 0 to 3 in 12 months, fewer is better)

Why these weights

Weights in a composite score are value judgements [1]; what makes them defensible is where they come from, and whether each pillar's actual influence on the ranking matches what it is meant to have [2]. Ours come from three steps, and every number below is recomputed by published code from the sources listed.

1. What buyers and developers say

Six published surveys of how organizations and developers choose AI model providers [7] [8] [9] [11] [12]: every criterion mapped to the pillar that measures it (accuracy to Capability, cost to Value, APIs and fine-tuning to Ecosystem, latency and uptime to Reliability, security, privacy and transparency to Trust), combined with more weight for larger and more recent surveys. Leaving any one survey out gives the ranges shown.

2. What adoption science adds

Buyers are not asked about installed base or attention. Meta-analyses of technology adoption put habit and social influence at about a quarter of what drives use [14] [15], so Usage and Reception share 25 points, 80 to 20. Release gets a floor of 6: users move to new models within weeks, and labs that stop shipping lose share [17] [16].

3. Calibrated to real influence

A pillar's influence depends on how much labs differ on it, not only on its weight. The weights used are chosen inside each pillar's evidence range so each pillar's actual influence comes as close as it can to its target [3], the approach of the EU's statistical audits of composite indices [6].

Always beside it: equal weights

Every close also ranks the labs with all eight pillars weighted equally [5]. Each lab page shows both, so any rank that depends on the weighting is visible, and rank ranges test weights moved up to 50% [4].

PillarTargetEvidence rangeWeightInfluence
Capability1713 to 20.12017.8
Usage and adoption1912.8 to 34.21625
Release execution66 (floor)68.6
Value1310.8 to 15.6155.9
Ecosystem148 to 17.4117.8
Reception70.8 to 13.8421.7
Reliability106.6 to 11.7110.9
Trust, safety and privacy1410.3 to 17.61712.4

Target: the share of the ranking each pillar should drive, from the evidence. Weight: the weight the Score uses, in percent. Influence: each pillar's actual share of what separates the labs, measured on the closes so far (the squared correlation between the pillar and Strength across labs, normalised, averaged over days) [2]. Where a weight reaches the edge of its evidence range before its influence reaches the target (Value and Reliability today), the gap is shown rather than forced. A re-run on the rebuilt closes found a slightly closer fit that did not hold when each half of the days, or each lab in turn, was left out, so the weights stayed; the influence shown is measured on the rebuilt closes.

Direction's weight: 0.18

Score = Strength + 0.18 × Direction reads as a forecast of Strength about 30 days ahead, in the textbook form of a level plus a damped trend [19] [18]. The weight is the Strength a point of Direction is expected to become: ρ × SD(30-day change in Strength) ÷ SD(Direction) = 0.2 × 1.723 ÷ 1.897, with Strength here leaving out Release and Reception, which already count news. Scaling by volatility follows the Conference Board's leading indicators [20]. ρ, how well Direction predicts later Strength, is set at a deliberately sceptical 0.2 [21] until enough live closes exist to measure it (about eight months); it is then estimated, and the weight updated as a new method version.

Sources

  1. OECD, European Union and EC-JRC (2008). Handbook on Constructing Composite Indicators: Methodology and User Guide. OECD Publishing. Link
  2. Paruolo, P., Saisana, M. and Saltelli, A. (2013). Ratings and rankings: voodoo or science? Journal of the Royal Statistical Society A 176(3), 609-634. Link
  3. Becker, W., Saisana, M., Paruolo, P. and Vandecasteele, I. (2017). Weights and importance in composite indicators: closing the gap. Ecological Indicators 80, 12-22. Link
  4. Saisana, M., Saltelli, A. and Tarantola, S. (2005). Uncertainty and sensitivity analysis techniques as tools for the quality assessment of composite indicators. JRSS A 168(2), 307-323. Link
  5. Greco, S., Ishizaka, A., Tasiou, M. and Torrisi, G. (2019). On the methodological framework of composite indices: a review of the issues of weighting, aggregation, and robustness. Social Indicators Research 141(1), 61-94. Link
  6. European Commission JRC (2025). Statistical audit of the 2025 Global Innovation Index. In WIPO, Global Innovation Index 2025, Appendix II. Link
  7. ICONIQ (2025). The Builder's Playbook: 2025 State of AI Report (n = 265). Link
  8. ICONIQ (2026). 2026 State of AI (n = 305). Link
  9. Menlo Ventures (2024). 2024: The State of Generative AI in the Enterprise (n = 600). Link
  10. Menlo Ventures (2025). 2025: The State of Generative AI in the Enterprise. Link
  11. Artificial Analysis (2025). AI Adoption Survey H1 2025 (n = 1,006). Link
  12. Stack Overflow (2025). 2025 Developer Survey (n = 35,897). Link
  13. Andreessen Horowitz (2025). How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025. Link
  14. Tamilmani, K., Rana, N. P. and Dwivedi, Y. K. (2021). Consumer acceptance and use of information technology: a meta-analytic evaluation of UTAUT2. Information Systems Frontiers 23(4), 987-1005. Link
  15. Blut, M., Chong, A. Y. L., Tsiga, Z. and Venkatesh, V. (2022). Meta-analysis of the Unified Theory of Acceptance and Use of Technology (UTAUT). Journal of the Association for Information Systems 23(1), 13-95. Link
  16. Aubakirova, M., Atallah, A., Clark, C., Summerville, J. and Midha, A. (2026). State of AI: an empirical 100 trillion token study with OpenRouter. arXiv:2601.10088. Link
  17. Fradkin, A. (2025). Demand for LLMs: descriptive evidence on substitution, market expansion, and multihoming. arXiv:2504.15440. Link
  18. Gardner, E. S. Jr. and McKenzie, E. (1985). Forecasting trends in time series. Management Science 31(10), 1237-1246. Link
  19. Holt, C. C. (2004, first issued 1957). Forecasting seasonals and trends by exponentially weighted moving averages. International Journal of Forecasting 20(1), 5-10. Link
  20. The Conference Board. Calculating the composite indexes (component standardization factors). Link
  21. Armstrong, J. S., Green, K. C. and Graefe, A. (2015). Golden rule of forecasting: be conservative. Journal of Business Research 68(8), 1717-1731. Link
  22. Wan, A., Klyman, K., Kapoor, S. et al. (2025). The 2025 Foundation Model Transparency Index. Stanford CRFM. arXiv:2512.10169. Link
  23. Future of Life Institute (2026). AI Safety Index, Summer 2026. Link

How readings become a score

Fixed goalposts

Heavy-tailed readings (downloads, likes, hosts, points) go on a log scale, shares and counts on a square-root scale. Each is then placed between two goalposts fixed for the method version: 0 at the low one, 100 at the high one. Because the goalposts are fixed, a lab moves only when its own readings move, never because another lab joined, left or had a good day.

Not applicable is not missing

A lab with no open models has no Hugging Face downloads: that reading does not apply, and drops out of its weights. A reading that should exist but is not fresh is filled with the median of the ranked labs that day, flagged, and counts against coverage, so a gap neither helps nor hurts a lab.

Confidence

Coverage is the share of a lab's applicable weight that is measured and fresh; a reading carried back from a later day counts half. High at 85% or more, medium from 70%, provisional below. A provisional lab stays on the board, flagged.

Same rules for every lab

Every indicator is read the same way for every lab. Incidents move Direction only when an independent outlet reports them, so a lab with a detailed public status page is not penalised for its openness, and one without is not spared.

Direction

What counts

Confirmed developments from the Wire in the last 30 days. Launches, updates, research and usage limits raised add. Setbacks subtract, at most 4 points in all: material incidents reported by an independent outlet, usage limits cut, and price rises of 30% or more. Money, policy and model retirements count zero; minor items (materiality 1) count zero. Points by materiality: materiality 1 = 0, materiality 2 = 0.75, materiality 3 = 1.5, materiality 4 = 2.5, materiality 5 = 4.

Recency and diminishing returns

A development counts in full on the day, falling to half at 30 days. A lab's positive developments are then taken biggest first, and the k-th counts 1 / (1 + 0.6k), so a stream of small posts is worth less than one major launch, and the order they arrived in does not matter.

Materiality as it stood

A development's materiality (1 to 5) comes from the rulebook applied to the facts in each report, not from the model's opinion. A close uses it as it stood that day: the lab's own reports decide when there are any; otherwise the most common reading among the reports, so more coverage does not mean a higher reading.

Going quiet, and live moves

A lab with no confirmed development of materiality 2 or more in 30 days gets minus 1. Direction runs from minus 10 to plus 10. Big confirmed events (materiality 3 or more) move Direction on the page within minutes, marked live, until the next close folds them in. Shared numbers are always closes.

Every lab page lists each development that counted, with its materiality, recency and diminishing-returns factors and its points, and the ones that counted zero and why.

Rank ranges

Each close is recomputed 500 times, from a fixed seeded sequence so any close reproduces, with every pillar weight moved up to 50% each way and the weight on Direction drawn anywhere from 0 to 1. A lab's range runs from the 5th to the 95th percentile of its ranks across those draws. A range of 1 to 2 means that in at least 90% of the draws the lab ranked first or second.

History and reconstructed closes

Closes before 24 September 2026 are reconstructed: the same code run for each past day, using evidence dated on or before it where the source keeps history (benchmark runs by run date, Arena by archived snapshots, SDK downloads by day, stories by publication time). Some sources keep no history (App Store charts, Hugging Face downloads and likes, OpenRouter hosts and uptime, and part of OpenRouter's usage share); for those, reconstructed closes use our first reading, taken on 23 September, marked as carried back. Carried readings count half toward coverage and such closes are at most medium confidence. Reconstructed closes are labelled everywhere they appear.

How the Wire's stories are classified

News items are classified by a language model (currently DeepSeek's V4.1 Flash, from one of the ranked labs) with a fixed prompt whose fingerprint is published with the method. The model only reports facts (is this a launch, which lab, is it a new model, the price); the rulebook in our code turns facts into materiality. A check of the classifier against hand-labelled items will be published here.

Versions and corrections

The method version is stored with every close. A change to a weight, an indicator or a formula is a new version with a changelog entry; old closes keep the version that made them. A wrong number is corrected beside the original, never by silently rewriting it, and every correction is listed on the corrections log.

Changelog

v2.3 · 2026-09-24

  • Reliability comes from the labs' own status pages: the 30-day uptime of each model API, counting major outages in full and partial at 30%, as the pages do; the uptime measured on OpenRouter stays beside it as an independent check, now the median of a lab's endpoints (retired listings at 0% had read as outages).
  • Usage counts paid adoption: the share of US businesses paying each vendor (Ramp AI Index), enterprise spend and wallet shares (Menlo Ventures, a16z) and builder adoption (ICONIQ), plus each lab's share of the dollars spent on OpenRouter, which replaces the token share of its top 20 models (cheap bulk models dominate token volume).
  • Ecosystem counts where a lab's current models are offered as a managed API on Amazon Bedrock, Google Cloud and Microsoft Foundry; hosting breadth counts only open-weight models, and does not apply to a lab with none.
  • Trust: harm from a lab's AI that the lab disclosed itself, with no legal duty to, counts half, so openness is not penalised.
  • Pillar weights re-checked on the rebuilt closes and kept: the re-run's optimum was not stable when each half of the days or each lab in turn was left out. Direction's weight re-derived on the corrected history: 0.18, down from 0.24. The first live close moves to 25 Sep 2026.

v2.2 · 2026-09-23

  • An eighth pillar, Trust, safety and privacy: independent safety and transparency assessments, security certifications, customer data commitments, and security and privacy incidents.
  • Pillar weights from published evidence: buyer and developer surveys weighted by sample size and recency, and adoption science for usage and attention, each with an uncertainty range (research/weights/derive.mjs).
  • Nominal weights calibrated inside those ranges so each pillar's actual influence on the ranking comes as close as it can to its evidence-based target (Paruolo, Saisana and Saltelli 2013; Becker et al. 2017).
  • Weights move from v2.1: Release execution 15 to 6 (buyers rarely name it, and launches also move Direction) and Reception 10 to 4 (it spreads the labs widely, so it needs little weight to reach its share); Capability 25 to 20 and Usage 20 to 16, to make room for Trust at 17; Value 12 to 15, Reliability 8 to 11, Ecosystem 10 to 11.
  • Direction's weight is derived, not chosen: the Strength points a point of Direction is expected to become over 30 days, with a stated sceptical predictive correlation until live closes can estimate it. It is 0.24, down from 0.5.
  • An equal-weights ranking is published beside every close as a standing comparison.
  • Published before the first public close.

v2.1 · 2026-09-23

  • Fixed goalposts for every indicator instead of scaling across the day's labs, so a lab moves only when its own readings move.
  • A missing reading is filled with the ranked labs' median and flagged, instead of silently taking the lab's own average.
  • Benchmarks: each benchmark on fixed goalposts, runs from the last 365 days, at least 3 of the 4.
  • Reliability is measured API uptime over 30 days. Incident minutes are dropped: only some labs publish comparable status pages.
  • Open-weights share is dropped from Ecosystem: it came from our own catalogue, not a measured source.
  • Value counts list prices below $0.30 per million as $0.30.
  • Direction: minor items count zero; diminishing returns apply biggest first, in any order; money, policy and model retirements count zero; usage limits raised add; usage limits cut and price rises of 30% or more subtract; incidents need an independent report; each development's materiality is as it stood at the close.
  • Rank ranges vary pillar weights by up to 50% and the Direction weight from 0 to 1, and report the 5th to 95th percentile.
  • Reconstructed closes that use carried readings are at most medium confidence.
  • Published before the first public close; no close under v2.0 was published.

v2.0 · 2026-09-23

  • First version of the measured Score (never published).