LLM Statistics in 2026

- Key LLM statistics for 2026
- Read LLM statistics by measurement family
- Adoption is broad, but it is not one market metric
- Spending estimates and market forecasts answer different questions
- The open-model ecosystem is large and highly concentrated
- Routed tokens reveal a different kind of demand
- What the statistics mean for testing LLM-backed features
- Why LLM numbers disagree
- Methodology and source notes
- Conclusion
A product leader preparing a budget, a writer building a deck, and an engineer choosing a model can all search for LLM statistics—and get numbers that appear to disagree. They may be looking at organizations that report using AI, people using generative-AI products, a provider’s weekly active users, tokens routed through one platform, or a commercial market forecast. Those are different measurements.The short answer: LLM use is expanding across business, consumer, developer, and open-model contexts, but there is no single credible number for “the LLM market.” This report separates the measures so you can choose a statistic that matches your decision and cite it without turning a survey response into a user count or a forecast into current revenue.
Key LLM statistics for 2026
78% of surveyed organizations reported using AI in 2024, up from 55% in 2023. In the same Stanford HAI synthesis, 71% reported using generative AI in at least one business function in 2024, versus 33% in 2023. These are organizational survey measures of AI and generative AI, not LLM inference volumes. Stanford HAI’s 2025 AI Index Economy chapter is the primary research synthesis.
Corporate AI investment reached $252.3 billion in 2024, including $33.9 billion in private generative-AI investment, according to Stanford HAI; it reported the private generative-AI figure was up 18.7% from 2023. These are investment measures, not LLM-provider revenue. Stanford HAI reports the 2024 figures.
16.3% of the world’s population used generative-AI products in H2 2025, up from 15.1% in H1 2025, according to the Microsoft AI Economy Institute’s telemetry-based adoption report. This is a population-level diffusion estimate, not a count of unique LLM accounts. Microsoft’s Global AI Adoption in 2025 report provides the methodology and estimate.
Among the working-age population in H2 2025, the Microsoft AI Economy Institute reported generative-AI adoption of 24.7% in the Global North and 14.1% in the Global South. The population is working-age people, not all residents or organizations. Microsoft’s report is the first-party source.
Companies spent an estimated $37 billion on generative AI in 2025, up from $11.5 billion in 2024, according to Menlo Ventures’ U.S.-enterprise survey and bottom-up market model. It estimated $19 billion of 2025 spend went to user-facing products and software. This is a modeled enterprise-spend estimate, not audited global revenue. Menlo Ventures’ 2025 enterprise report explains its scope.
Menlo Ventures estimated 2025 U.S.-enterprise LLM spend at 40% Anthropic, 27% OpenAI, and 21% Google, and estimated those providers represented 88% of enterprise LLM API usage. Its model draws on production API-usage shares, weighted survey responses, and public-financial triangulation; it is not a global provider-revenue ranking. Menlo Ventures surveyed 495 U.S. enterprise AI decision-makers from November 7–25, 2025.
ChatGPT reached 900 million weekly active users in February 2026, according to OpenAI’s announcement as reported by TechCrunch. Weekly active users are not directly comparable with a monthly-user metric. TechCrunch’s report on the announcement identifies the metric and date.
Google reported that the Gemini app surpassed 1 billion monthly active users on August 11, 2026. This is a company-reported monthly active-user figure for one product, not a market-wide count. Google’s Gemini announcement is the first-party source.
OpenRouter analyzed more than 100 trillion tokens of real-world LLM interactions on its multi-model platform in its State of AI 2025 study. The study primarily covers a rolling 13-month period ending in November 2025, making this platform telemetry rather than a census of global inference. OpenRouter’s study documents the dataset.
In OpenRouter’s open-source-model traffic, roleplay accounted for about 52% of token usage, with programming the second-largest category. Because consistent task tags begin in May 2025, this is a mid-2025-onward category analysis, not a full rolling-year result. OpenRouter states that limitation.
Hugging Face reported that public Hub model repositories increased from 2.43 million in January 2026 to 2.96 million in August 2026; datasets rose from 711,000 to 1 million, and Spaces from 1.00 million to 1.44 million. These are public Hub repository and Space counts, not unique production deployments. Hugging Face’s Summer 2026 observations are first-party platform analysis.
On the same Hub, roughly 85.6% of models had fewer than 200 lifetime downloads, while 1.5% of repositories accounted for 99.2% of downloads as of the summer 2026 analysis. Hugging Face also counted 151,448 Qwen-based derivatives. These describe concentration and downstream activity on the Hub, not model quality or enterprise deployment. Hugging Face provides the underlying ecosystem analysis.
In Stack Overflow’s 2025 developer survey, more than 84% of respondents said they were using or planning to use AI tools, while 29% said they trusted AI tools. “Using or planning” is broader than current use, and the population is survey respondents rather than all developers. Stack Overflow’s survey analysis reports both measures.
JetBrains reported that 90% of more than 15,000 professional developers in its 2026 globally representative survey used AI coding agents at work at least weekly, while 68% used them daily. The category includes local and remote cloud agents; it is not a measure of all LLM users. JetBrains Research describes the survey and definitions.
Precedence Research estimates the global large-language-model market at $7.77 billion in 2025 and projects $10.57 billion in 2026 and $149.89 billion by 2035, a projected 34.44% CAGR for 2026–2035. This is a commercial market forecast with a 2025 base year, not a disclosed current-revenue total. Precedence Research’s market page is the publisher’s forecast.

Get the Mobile Testing Playbook Used by 800+ QA Teams
Discover 50+ battle-tested strategies to catch critical bugs before production and ship 5-star apps faster.
Read LLM statistics by measurement family
The useful question is not which figure is largest. It is what the figure measures. The statistics above fall into six measurement families:
Organizational adoption surveys show whether respondents say AI has entered business functions.
Population and product reach show use of generative-AI products, either across a population or within a named application.
Investment and spend models show money committed to an ecosystem under a stated methodology.
Developer surveys show reported behavior and sentiment within a defined respondent group.
Platform telemetry shows activity on a particular service.
Repository and download data shows publication and reuse within an open-model platform.
This map is the difference between a usable citation and a misleading one. For example, ChatGPT’s weekly active-user figure and Gemini’s monthly active-user figure both indicate enormous product reach, but neither establishes which product has more users on a common time basis. Similarly, a Hub repository count shows publishing activity; it does not establish production adoption.
Adoption is broad, but it is not one market metric
Stanford HAI’s organizational results and the Microsoft AI Economy Institute’s population-level findings point at different stages of diffusion. One asks organizations about their use of AI. The other estimates adoption of generative-AI products across people. Both are valuable, but neither is a direct count of LLM API calls, paid seats, or deployed applications.The geographic split is a practical reason to retain the denominator. “Global adoption” can conceal substantial differences in access, infrastructure, language support, and product fit. If you are using an adoption figure in a market plan, use the source’s population definition rather than shortening it to a universal claim.Developer data adds another layer. Reported use can become widespread before trust does. That gap matters when you interpret claims about AI-assisted development: adoption tells you that a tool has entered a workflow, while a trust measure tells you something different about reliance and review.For a related look at how AI changes quality workflows, see Quash’s AI testing statistics and adoption data. It addresses software-testing evidence rather than treating developer AI-tool use as a proxy for test quality.
Spending estimates and market forecasts answer different questions
Menlo Ventures’ estimate is most useful as a view of enterprise generative-AI spending under a disclosed model. Its scope includes foundation models, model-training infrastructure, AI infrastructure, and AI applications, while excluding chips, inference and model serving, and AI features embedded in existing software. Those boundaries are part of the statistic, not fine print.The provider-share estimate should be read with the same restraint. It is an estimate of U.S.-enterprise LLM spend and API-usage concentration, not a worldwide league table. A provider can have a different position in consumer use, open-source deployment, or a region absent from the survey.Precedence Research’s figure serves another job: it offers a commercial forecast. Forecasts are useful for scenario planning, but they should never be presented as money already earned. No consensus global LLM-market total is reconstructed here because commercial estimates can use materially different scopes and methods.
The open-model ecosystem is large and highly concentrated
Hugging Face’s repository growth shows that publishing and remixing activity remain intense in 2026. Yet its download distribution provides the needed counterweight: an expansive catalog does not imply that attention is evenly distributed across it.The Qwen derivative measure is best read as ecosystem momentum. Developers creating derivatives and publishing them to the Hub are making a visible downstream commitment to a model family. That activity does not prove the models are superior, nor does it measure corporate production use. It does show why repository counts, downloads, and commercial spend must remain separate columns in any serious LLM statistics report.
Routed tokens reveal a different kind of demand
OpenRouter’s token study captures observed interaction volume rather than survey intent. It is useful precisely because it reports what passed through a multi-model routing platform. It is also bounded precisely by that fact: OpenRouter users and workloads are not a representative sample of all LLM use.The roleplay finding is a reminder that LLM demand is not reducible to workplace productivity. But it is a task-mix result from an open-source-model subset with a defined tagging window, not a conclusion about all consumers or all models. Keep that context attached when you cite it.
What the statistics mean for testing LLM-backed features
No first-party Quash data covers LLM usage, model failure rates, or the testing burden of LLM-backed features. Public sources also do not provide a single statistic that can tell you which failure type your product will encounter.The practical implication is to collect product-specific evidence rather than borrowing confidence from market growth. If you ship an LLM feature, define expected outputs and workflow outcomes, then test degraded inputs, latency, fallback behavior, permissions, and changes to the surrounding product. Broad adoption does not validate any one implementation.This is especially relevant when AI-assisted code or agent behavior reaches production. Quash’s report on vibe coding adoption, tools, and risks explores that verification problem from the software-delivery side, while its 2026 QA automation trends report covers the broader QA context. Neither should be used as evidence for the market figures in this article; they answer adjacent testing questions.
Why LLM numbers disagree
The numbers disagree only when you ask them to answer a question they were not designed to answer.
Measurement | What it can tell you | What it cannot establish |
Organization survey | Reported AI use in a surveyed business population | Consumer reach or inference volume |
Product active users | Reach of a named product over its stated time period | A comparable count for another product with a different active-user period |
Enterprise-spend model | Estimated spending within the model’s scope | Audited global provider revenue |
Routed tokens | Observed activity on a defined platform | Global LLM inference |
Hub repositories and downloads | Open-model publication and distribution on that Hub | Production deployment or model quality |
Commercial forecast | A publisher’s projected market value | Current disclosed market revenue |
Use the source whose population and metric fit the decision in front of you. A product-reach claim needs a product metric. A budget discussion needs a clearly scoped spend figure. A technology-roadmap discussion may need developer or repository activity instead. The label is part of the evidence.
Methodology and source notes
This report prioritizes first-party product announcements, original research publishers, and source-owned platform analyses. TechCrunch is used for the ChatGPT metric because it reports the specific OpenAI announcement; secondary statistics roundups are not used as evidence.Each included statistic retains its source, period, population, and evidence type. Current observations, survey responses, modeled estimates, platform telemetry, company-reported activity, and commercial forecasts are intentionally separated. That approach leaves out many tempting numbers, but it avoids creating a false consensus from incompatible measures.
Conclusion
The most defensible 2026 LLM statistic depends on the decision you need to make. Use a survey to describe reported adoption, a company metric to describe product reach, telemetry to describe activity on a platform, and a forecast only to describe a forecast. Once you preserve those boundaries, the numbers become more useful—and much harder to misuse.



