Apple began delivering the new generation of Mac Studio on September 22 with an unusual proposition for a desktop: convincing companies that part of artificial intelligence can leave the cloud and return to in-house infrastructure. With the M5 Ultra, multiple machines can be clustered via Thunderbolt 5 and RDMA to run large open-weight models, turning a variable inference expense into a hardware investment.

The question is not whether a Mac can run a large model. The question is when buying in-house capacity starts to make more economic sense than continuously paying for AI usage in the cloud.
Apple is trying to turn tokens into a fixed cost
The M5 Ultra supports up to 512 GB of unified memory and 1.2 TB/s of bandwidth. Apple also added support for Mac Studio clusters connected via Thunderbolt 5 with RDMA, technology that enables low-latency communication directly between the memory of the machines. According to the company's own tests, four systems can achieve up to three times the performance of a single Mac Studio in distributed inference.
Reuters attended a demonstration in which four Mac Studios ran a model with 1 trillion parameters to locate and fix a problem in graphics code. The setup ran plugged into a single wall outlet.
That is the point at which the strategy stops being merely technical.
Hosted AI services typically turn usage into operating expense: the more tokens, calls, or computing capacity an application consumes, the larger the bill tends to be. In in-house infrastructure, the cost does not disappear, but it changes its nature. The company pays upfront for the capacity and can reuse it without an additional charge for each inference.
It is this difference that Apple is trying to exploit.
"No cost per token" does not mean free AI
The wording used by Apple is appealing, but it needs an important caveat.
The Mac Studio with M5 Ultra starts at US$ 5,499 in the United States, while full configurations can approach US$ 20,000, according to Reuters. A cluster with several units, therefore, can require tens of thousands of dollars before running its first AI task.
To this are added energy, storage, administration, model updates, observability, redundancy, and the risk of the equipment sitting idle.
The economic argument becomes stronger when utilization is high, continuous, and relatively predictable. A company that runs internal agents, code generation, or document processing for much of the day can spread the investment across a growing volume of inferences.
The scenario is different for applications with irregular demand. The cloud allows capacity to be increased or reduced without buying machines to meet only peaks.
Therefore, the relevant comparison is not simply "paid token versus free token." It is total cost per useful work produced, considering investment, operation, and utilization rate. Researchers have been proposing metrics that combine CAPEX and OPEX precisely because price per token or per GPU hour, in isolation, does not capture the entire economics of an AI deployment.
Open-weight makes the math possible
There is also a structural difference between running a model locally and contracting an API.
Companies cannot simply download proprietary models from providers such as OpenAI or Anthropic and transfer them to a Mac Studio. Apple's argument depends mainly on the growth of open-weight models, whose weights can be run on infrastructure controlled by the customer itself.
Apple states that the clustering of Mac Studios creates enough capacity to load some of the largest open-weight models available. MLX, the company's machine learning framework, already supports distributed computing and a backend called JACCL for RDMA communication via Thunderbolt, including for techniques such as tensor parallelism.

This opens up an intermediate alternative between two structures that until recently were well separated: fully managed APIs on one side and specialized servers with data center GPUs on the other.
The Mac Studio tries to occupy the space of AI infrastructure that can literally sit inside the office.
For companies that work with proprietary code, internal documents, or sensitive data, there is also a second incentive. Keeping inference local can reduce the amount of information sent to external infrastructure and give greater control over storage and processing.
The biggest obstacle may not be the chip
Having enough memory to load a model is only part of the problem.
Traditional AI clusters were developed around a mature ecosystem of GPUs, high-speed networks, orchestration, monitoring, and data center tools. Apple's alternative still needs to demonstrate that it can offer comparable operational simplicity.
MLX's own documentation shows that RDMA over Thunderbolt still requires specific configuration. To enable it, for example, it is necessary to enter macOS recovery mode, enable the feature, and subsequently configure communication between the nodes.

This kind of friction matters more in corporate environments than in technical demonstrations.
Apple also starts from a small position in the enterprise computer market. According to IDC data cited by Reuters, Macs account for about 4.6% of the corporate desktop and laptop market, versus 91.3% for Windows machines.
To turn the Mac Studio into AI infrastructure, therefore, the company must win over not only developers but also IT teams accustomed to completely different ecosystems.
The Mac Studio does not need to replace the cloud to change the math
The most plausible scenario is not a complete migration of enterprise inference to desktops.
It is a hybrid architecture.
Applications with variable demand, the need for global scale, or dependence on the most advanced proprietary models tend to continue favoring external services. Predictable workloads based on open-weight models, meanwhile, may begin to justify in-house infrastructure when volume makes recurring charges sufficiently relevant.
In this scenario, the company could maintain a permanent local capacity for the base load and turn to the cloud for peaks or specialized models.

It is a smaller change than abandoning the data center, but economically important. Part of the AI budget stops growing directly with the number of tokens processed and starts depending on the utilization of assets already acquired.
The next signals will come from outside Apple's demonstrations. It will be necessary to observe independent benchmarks of throughput, latency, and power consumption, in addition to the adoption of Mac clusters by companies in real workloads and the evolution of tools for managing multiple machines.
There will also be an immediate hardware test. Although the new Mac Studios have begun reaching customers on September 22, the M5 Ultra configuration with 512 GB of unified memory is only expected at the end of October.
If organizations begin keeping agents, development tools, and other intensive workloads running continuously on these systems, Apple's economic proposition will gain concrete evidence.
Until then, the Mac Studio does not eliminate the cloud. It makes a question that companies will have to ask with increasing frequency more plausible: which AI workloads still need to be rented per token, and which already justify in-house capacity?



