Skip to content

Private AI

Open AI models now handle most of corporate America's AI traffic. Here is what business owners should know before assuming a paid subscription is the only option.

Open-weight AI models, the kind a business can download, run on its own infrastructure, and modify freely, now account for 58% of US enterprise AI model usage, up from roughly 10% a year ago, according to usage data from OpenRouter reported by The New York Times. AT&T is the clearest example: its chief data and AI officer says open models grew from 20% to 40% of the company's AI use since May and may reach 60% within months, cutting AI costs by up to 80% along the way. The direct answer for a business owner: this is real, verified evidence that a large share of everyday AI work no longer requires a premium closed-model subscription to get comparable results. It is not evidence that any business should rip out its current AI vendor this quarter. What changed, and what it takes to actually capture the savings, is worth understanding before reacting to either the headline or the hype.

By Fabio Rabelo · Founder, ATLACIS ·

What happened

The New York Times reported on September 6, 2026 that open-weight AI models, ones with publicly available weights that a business can run on its own servers instead of calling a vendor's API, accounted for 58% of US enterprise AI model usage in August 2026, according to data from OpenRouter, a platform that routes requests across multiple AI providers. A year earlier, that figure was roughly 10%. AT&T is the article's central example. Chief data and AI officer Andy Markus told the Times that open models made up about 20% of AT&T's AI use in May 2026, have since risen to about 40%, and could reach 60% within months. Markus said the shift is saving AT&T up to 80% on AI costs compared with earlier in the year. AT&T's own engineering blog, published in July, describes the mechanism behind that number: a proprietary AI Gateway that routes each task to the most cost-effective model across roughly 45 billion tokens processed per day, up from about 8 billion a year earlier, cutting AI costs by as much as 90% in production. AT&T also trained its own open-source telecom-specific model, OTel 2.0, and reports cutting costs on coding and other advanced tasks by as much as 56% using open-source model routing tools, with only a 2% quality decline, according to a separate report on the same effort. The Times reports Airbnb and Deloitte are also increasing their use of open models. AT&T researches Chinese open-weight models but does not use them, citing regulatory and data-privacy concerns, relying instead on Google's Gemma and Meta's Llama.

Why it matters for business owners

For most of the last few years, the practical choice for a business adopting AI was a paid subscription or API from a closed vendor: OpenAI, Anthropic, Google, or Microsoft. Open-weight alternatives existed but were seen as a hobbyist or research option, a step behind on capability and not something a normal business would build a real workflow around. That gap has narrowed enough that a major enterprise is now routing a majority of its AI traffic to open models in production, on real customer-facing work like call transcription and customer service tooling, and reporting real cost reductions doing it. This is not a single vendor's marketing claim. It is a market-wide usage shift, visible in independent platform data, plus one company's detailed, named account of how it got there. For a business owner paying a growing monthly AI bill, that combination is worth ten minutes of attention: a meaningful share of routine AI work may no longer need to run on the most expensive model available.

What owners should not misunderstand

AT&T did not flip a switch. It built a custom routing system, employs a dedicated data science and AI infrastructure team, and processes 45 billion tokens a day, a scale no small or medium business will approach. Its 80% and 90% savings figures are Markus's own account and AT&T's own blog post, not numbers audited by an outside party, and they measure AT&T's specific mix of workloads, not a universal result. AT&T also is not moving everything to open models indiscriminately. It is deliberately migrating AI use cases older than a year, the ones where an open model has had time to catch up to whatever closed model originally handled the task, and it keeps newer, harder work on premium models. And despite Chinese open-weight models being cheaper still, AT&T does not use them, choosing Google's and Meta's US-developed open models instead over regulatory and data-privacy concerns. Cheapest is not the same as right for a given business's risk tolerance, and that judgment call matters as much as the cost math. 'Open' also does not mean free. Running an open model still requires servers or GPU capacity, someone who can deploy and maintain the routing and infrastructure, and ongoing attention as workloads and models change. AT&T's savings come from replacing a metered API bill with infrastructure and engineering it already had the team to build. A business without that team is comparing a different set of costs, not a free upgrade.

The operational lesson

The useful takeaway is not "switch to open models." It is that the market has now validated, at real production scale, that well-defined, repetitive, or older AI workloads can run on open models without a meaningful quality loss, which changes what is worth questioning in a current AI bill. A task a business has been running on a premium closed model for a year or more, especially something repetitive like transcription, categorization, or first-draft customer responses, is exactly the kind of workload AT&T is actively migrating off its most expensive models. The judgment that still has to happen locally is routing: which tasks are safe to shift to a cheaper or self-hosted model, which need to stay on a frontier model for quality or reasoning reasons, and which involve data sensitive enough that the deployment question (cloud, private cloud, or on-premise) matters more than the price per token. That is a workload-by-workload evaluation, not a single decision to make once.

What a serious business should do next

Do not switch AI vendors or models this week because of a headline statistic. Start by identifying which of your current AI-powered tasks have been running the same way for a year or more, since those are the workloads most likely to have a viable, cheaper open-model alternative today. Separate your AI spend by task, not just by total monthly bill. A single blended number hides which specific workloads are expensive because they need a frontier model and which are expensive simply because nothing has revisited the choice since it was first set up. Before evaluating any specific open model, weigh its country of origin and licensing terms against your own data-sensitivity and compliance requirements, the same way AT&T rules out cheaper Chinese models it has evaluated. Price is one input, not the only one. Price the real cost of self-hosting or routing to an open model honestly: server or GPU capacity, the engineering time to deploy and maintain it, and ongoing monitoring, against the full cost of your current API bill for that specific workload, not against a headline savings percentage from a company operating at a different scale.

The Atlacis view

A market-wide shift in enterprise AI usage is a real signal, not something to copy from a headline or ignore because it happened at a telecom company running billions of tokens a day. The workload-level question underneath it, which specific tasks in your business are paying for capability they no longer need, and which genuinely require a frontier model, is exactly the kind of evaluation that determines whether a private or open-model deployment actually saves money or just adds complexity. Atlacis helps business owners map their real AI workloads and spend, test whether an open, private, or on-premise model can safely and honestly replace part of a closed vendor bill, and make that call based on the business's own data sensitivity and budget, not on what a much larger company with a dedicated AI team is doing.

The short version

  • Open-weight AI models now account for 58% of US enterprise AI usage on OpenRouter as of August 2026, up from roughly 10% a year earlier, according to The New York Times.
  • AT&T's chief data and AI officer reports open models grew from 20% to 40% of the company's AI use since May 2026 and may reach 60% within months, cutting AI costs by up to 80%, using a proprietary, cache-aware routing system built by a dedicated internal team.
  • AT&T deliberately migrates AI use cases older than a year to open models rather than switching everything at once, and continues to run newer or harder tasks on premium closed models.
  • AT&T avoids cheaper Chinese open-weight models over regulatory and data-privacy concerns, using US-developed open models (Google's Gemma, Meta's Llama) instead, showing that lowest cost is not the only factor in model selection.
  • AT&T's savings figures are self-reported and reflect its own infrastructure and scale, not an independently audited, universal result a smaller business should expect to match without comparable investment.
  • The useful move for most businesses is auditing AI spend by task to find older, repetitive workloads worth testing against an open or private model, not switching vendors wholesale because of a market-wide statistic.
Tags:private AIopen source AIAI cost optimizationvendor dependencymodel selectionon-premise AIbusiness AIAI decision support
FAQ

Common questions

Does the rise in open-model AI usage mean my business should stop paying for ChatGPT, Claude, or another closed AI subscription?
No. The data shows a market-wide shift in where AI workloads run, not that closed models no longer have value. AT&T, the leading example in this story, still runs its newer and harder tasks on premium closed models and only migrates well-defined, older use cases to open alternatives. Review your own workloads before changing anything.
Can a small or medium business realistically get AT&T's 80% cost savings by switching to open models?
Not without comparable investment. AT&T's savings come from a custom-built, cache-aware routing system, a dedicated data science and AI infrastructure team, and a scale of 45 billion tokens processed per day. A smaller business considering open or private models should price its own infrastructure and engineering costs honestly against its current AI bill, workload by workload, rather than expecting the same percentage.
If open models are cheaper, why does AT&T avoid the cheapest Chinese open-weight models?
AT&T researches Chinese open-weight models but does not use them, citing regulatory and data-privacy concerns, and relies instead on US-developed open models from Google and Meta. It is a reminder that a model's cost is one factor in a buying decision, alongside where the model comes from, its licensing terms, and how that fits a business's own compliance and data-sensitivity requirements.

Make better AI decisions, starting with one call.

Book a free AI Fit Call. We will tell you what to use, what to avoid, and where to start. No jargon, no pressure.