Apple isn’t just introducing two new chips: the M6 and the M5 Ultra represent two different approaches to bringing artificial intelligence as close to the device as possible.
The first relies on the density of a 2-nanometer process and a dual Neural Engine; the second expands unified memory up to 512 gigabytes.
The focus is therefore no longer just on speed, but on the ability to run resource-intensive models without relying solely on the cloud.
On August 25, 2026, Apple unveiled two architectures that significantly shift the center of gravity in personal and professional computing. The M6 chip marks Apple’s first use of a 2-nanometer process and combines a 12-core CPU, a 12-core GPU with Neural Accelerators, and a dual 16-core Neural Engine. Its memory bandwidth can reach 170 GB/s, while the unified memory capacity goes up to 32 GB.
What matters here is not so much the sheer number of figures as their intended use. Apple is explicitly positioning this generation for the local execution of large language models, agent-based tasks, and creative processing. The M6’s GPU incorporates a neural accelerator in each core and claims peak AI performance nearly 30 percent higher than that of the M5. The dual Neural Engine is said to be up to twice as fast as that of the previous generation. This combination gives the machine a different purpose: to perform more computations on-device, in an environment where data does not necessarily have to leave the computer.
The M5 Ultra takes this concept much further. It is based on an architecture featuring four chips connected by a new generation of UltraFusion. The interconnect bandwidth exceeds 4.4 TB/s, with a connection density reported to be more than six times that of the previous generation. The CPU can have up to thirty-six cores, the GPU up to eighty cores, and the unified memory bandwidth reaches 1.2 TB/s.
But the most defining feature lies elsewhere: 512 GB of unified memory. This capacity allows Apple to position the Mac as a workstation capable of loading AI models with hundreds of billions of parameters locally. In a market where the cost of inference and dependence on cloud infrastructure are becoming strategic issues, this architectural choice is no small matter. It brings a form of workstation sovereignty back to the table: keeping certain models, certain datasets, and certain sensitive operations on the machine itself.
The M5 Ultra features a 32-core Neural Engine and a GPU with Neural Accelerators in each core. Apple reports AI performance up to 4.5 times faster than the M3 Ultra and more than six times faster than the M1 Ultra. However, these figures are based on Apple’s internal testing protocols and should be viewed as comparative benchmarks, not independent measurements.
The same caution applies to energy efficiency, which the California-based company heavily emphasizes but without providing complete details on actual power consumption based on workloads. The most compelling benefit, therefore, is architectural: more unified memory, more bandwidth, greater matrix acceleration, and a clear commitment to running models locally.
This shift goes beyond the Mac alone. By strengthening Core AI, Core ML, Metal, MLX, and Xcode around a hardware architecture explicitly designed for AI, Apple aims to keep developers and creatives in an environment where hardware, frameworks, and data remain tightly integrated. This approach echoes what has historically been the company’s strength: not separating the device from its system.
The question in the coming years will be less about whether a computer can run an AI model locally and more about determining to what extent this local execution can replace the cloud. With the M6 and M5 Ultra, Apple isn’t putting an end to the debate; it’s taking it to a whole new hardware level.

Cette publication est également disponible en :
