The latest OpenRouter ranking is a useful antidote to the usual AI-model scoreboard. It does not ask which system had the highest benchmark score last week. It shows what developers routed through OpenRouter’s API, measured in tokens, and the result is a market with several strong suppliers rather than one obvious winner.
The snapshot supplied to us was taken on 2 September 2026 and covers usage through 1 September. In the current-week model table, DeepSeek V4 Flash 0731 led with 12.1 trillion tokens. GLM 5.3 Flash followed on 10 trillion, and OpenAI’s GPT-5.6 Luna was third on 9.52 trillion. Xiaomi’s MiMo-V2.5 and two Tencent models were next.
That is already different from the conversation most business owners hear. The headlines tend to turn AI into a contest between a small number of household names. The usage record looks more like a routing market. Different teams are sending different jobs to different models, and the provider matters less than whether a model is a good fit for the work in front of it.
The market is divided by work, not by one universal winner
OpenRouter’s task view divides current spend into four broad groups: general work at 32.4%, code at 30.3%, agents at 28.5%, and data work at 8.8%. General use, coding and agent work are close enough that none can be treated as a side case.
That matters because the three jobs make very different demands. A short classification task, a code review and an agent that has to use tools across several steps are not the same test. A model that feels excellent in a chat window may be too slow or too expensive for a high-volume process. A cheaper model can be the better operational choice if it gets the routine work right often enough. The report cannot tell us which choice is correct for a particular business, but it does show why a single league table is not enough.
The classification list makes the point from another direction. Claude Opus 5 held 9.0% of category spend, GPT-5.6 Sol 6.9%, and Gemini 3.1 Pro Preview 6.5%. There is no contradiction between those models appearing here and the Flash-named models leading the all-model token table. They are different views of a market where the task changes the answer.
Fast models are carrying a lot of the everyday load
The first two models in the weekly table, along with the seventh-place DeepSeek model, are explicitly labelled Flash. Several of the leaders came from providers that barely register in mainstream business coverage. That tells us something practical: a lot of real AI usage is not waiting for the most celebrated model to write a perfect answer. It is getting work through a system quickly and at a workable cost.
OpenRouter’s cost-per-session panel reinforces the same point. Its least expensive agent-session entries begin at $0.0001, with the exact cost changing by tool, session length and the model used. The sensible lesson is not to choose the cheapest line in a chart. It is to cost the whole workflow. A model that requires repeated correction, retries or human clean-up is not cheap merely because its token price is low.
For a marketing team, that usually means separating the work before choosing a model. Research, extraction, first-pass organisation and routine transformations can be measured for speed and accuracy. Brand claims, factual calls, publication decisions and work that changes a live website still need a person responsible for the final answer. We have written before about where that human gate belongs in an AI-assisted SEO workflow. The adoption data here gives the operational reason for it: the tools are becoming normal infrastructure, not an autopilot.
No provider owns the OpenRouter market
At provider level, DeepSeek led the report’s market-share view with 23.0%, followed by Google at 20.7% and OpenAI at 19.6%. Those three account for 63.3% together. Z-AI held 10.2%, followed by Qwen on 5.0%, Tencent on 3.5%, and a long tail of other providers.
Even in one API marketplace, then, this is not a winner-takes-all picture. The leader has less than a quarter of the displayed share. For anybody building a process around AI, that is a good reason to avoid tying the whole workflow to a single provider if the work can be tested elsewhere. Keep prompts, source material and evaluation criteria portable. Compare a replacement before a price change or model retirement turns into an emergency.
It is also a reminder that market-share headlines can hide a great deal. OpenRouter ranks model variants separately. A free variant and a paid variant of what people would casually call the same model can land in different places. The top-ten chart also groups every model outside the top ten into an Others series. Those are sensible reporting choices, but they are not a census of brands or users.
What these figures do not prove
OpenRouter is unusually clear about this. Its rankings measure tokens routed through OpenRouter. They do not measure the whole AI market, traffic sent directly to a provider’s API, requests kept private, accuracy, reasoning ability or benchmark performance. Token volume is not a count of customers, requests or dollars spent either. Models vary in how much text they produce and how they tokenize it.
That caveat is not small print. A model with more tokens on this chart has not automatically won more users, earned more revenue or produced better work. It has handled more token volume in this particular marketplace during the selected period. The rankings are useful precisely because they are a live picture of actual routing. They become misleading when somebody asks them to answer a different question.
The trend column needs the same care. OpenRouter compares the trailing seven days with the seven days before it and only includes models with at least one million current-period tokens, which prevents tiny bases from creating dramatic percentage jumps. GPT-5.6 Luna, for example, was up 129% in the supplied weekly view. That is a meaningful signal to watch. It is not a forecast.
How to use the data without chasing every leaderboard
Start with a piece of work, not a model name. Define what success means for one repeatable job: accurate extraction from a document, useful topic clustering, a code task that passes its tests, or a support reply that needs minimal correction. Measure the output and the clean-up time.
Keep a comparison set. One dependable default and one or two alternatives are enough for most teams. Re-run the same small set of real examples when a model changes or pricing moves. This is more useful than reacting to a weekly ranking, because it tells you whether the change affects your work.
Keep people at the points where a mistake becomes expensive. AI can make a fast first pass over a large set of material. It cannot be left to invent client claims, decide what deserves publishing, or alter a site without review. The route from a plausible answer to a correct one still has an editor in it.
Read adoption and quality as separate questions. Usage data can tell you where developers are putting volume today. A task-specific test can tell you whether a model meets your standard. Both are useful. Neither replaces the other.
Frequently asked questions
Is DeepSeek the best AI model because it leads OpenRouter?
No. It leads this snapshot of token volume routed through OpenRouter, which is an adoption measure. OpenRouter states that its rankings do not establish model accuracy, reasoning ability or benchmark performance.
Does OpenRouter show the whole AI market?
No. It reports traffic through OpenRouter’s API and excludes requests users or applications keep private. Direct use of providers’ own APIs is outside this dataset.
Why are Flash models so prominent in the weekly ranking?
The ranking shows that models with Flash in their names occupy many current leading positions. The data itself does not say why any team selected them, so it should not be used to claim a particular model is better for every task.
Should a business switch models whenever the ranking changes?
No. Rankings are worth watching for possible changes in supply, cost and adoption. A switch should follow a test against the business’s own work, including the human time needed to check and correct the output.
The source and the useful conclusion
This article uses OpenRouter’s AI Model Rankings, captured on 2 September 2026 with usage data through 1 September 2026. OpenRouter says the rankings data is licensed under CC BY 4.0. Its page explains the measurement method and the limits described above.
AI is becoming a collection of specialised infrastructure choices. The data does not crown a universal winner. It shows a busy market in which general work, coding and agent activity all matter, providers compete closely, and the best choice depends on what the work actually is. That is the more useful question to take into a planning meeting.