AI Newsway

Apple Puts Its First 2nm Chip in a Mac mini Built for Local AI

M6 and the quad-die M5 Ultra push desktop Macs toward on-device inference, at noticeably higher prices

|4 min read0
AI Summary
Apple launched a refreshed Mac mini with the M6, its first 2-nanometer chip, and a Mac Studio powered by the quad-die M5 Ultra, shipping September 22. The M6 mixes super, performance and efficiency cores with a dual 16-core Neural Engine and claims roughly four times the AI performance of the M4 mini, while the M5 Ultra reaches 512GB of unified memory at 1.2TB/s. Apple pitches both explicitly for running LLMs locally, with Thunderbolt 5 clustering.
The compact aluminum Mac mini enclosure, carried over into the new M6 model that Apple positions as an always-on machine for agentic AI workloads.
The compact aluminum Mac mini enclosure, carried over into the new M6 model that Apple positions as an always-on machine for agentic AI workloads.

Apple has put its first 2-nanometer processor inside its cheapest desktop. The M6 debuts in a refreshed Mac mini. A new M5 Ultra powers the top Mac Studio. Pre-orders opened August 25, and most configurations ship September 22.

The pitch this time is not video editing. It is inference. Apple is selling both machines as places to run large language models locally, and for once it says so outright.

What Is Actually New in M6

M6 is the first Apple chip to mix all three of its CPU core types in one design. The 12-core complex pairs two super cores with four performance cores and six efficiency cores. Super cores arrived earlier this year in M5 Pro and M5 Max. This is the first time they share a die with both of the other types.

The GPU also grows to 12 cores. Each one carries a Neural Accelerator. Apple pairs that with a Dual 16-core Neural Engine, a first for the line, and system frameworks can drive both engines simultaneously.

Memory bandwidth reaches 170GB/s against a 32GB ceiling. Apple claims 40 percent faster CPU work than the M4 mini and roughly four times its AI performance. In LM Studio, the company measured prompt processing at up to 4.8x the M4 and 13.5x the M1.

M5 Ultra Goes Quad-Die

The Mac Studio chip is the stranger engineering story. M5 Ultra fuses two dual-die M5 Max parts into a single quad-die package using UltraFusion. Inter-die bandwidth clears 4.4TB/s. Connection density rises more than sixfold. Software sees the four dies as one processor.

That yields up to 36 CPU cores, up to 80 GPU cores, and a 32-core Neural Engine. Unified memory tops out at 512GB moving at 1.2TB/s. Apple puts peak AI GPU compute at 4.5 times the M3 Ultra.

A catch is buried in the naming. M5 Ultra is still a 3nm part. Ultra chips are assembled from Max dies, so they trail the newest process node by a generation. Apple is also skipping M6 Pro and M6 Max entirely, jumping straight to M7 in mid-2027.

Clustering Is the Real Feature

Mac Studio now officially supports clustering over Thunderbolt 5 using RDMA. Memory pools across machines. Apple says four linked Studios can reach triple the inference throughput of a single unit.

This mostly formalizes something users were already doing. macOS 26.2 added low-latency communication between Thunderbolt 5 hosts last December for distributed MLX inference. Developers have been daisy-chaining minis ever since to hold models no single box can fit.

The economic logic is easy to follow. Coding-agent bills keep climbing. Open-weight models such as Qwen and DeepSeek cover a good share of that work. At some point a capital purchase beats a metered one.

Every Number Here Is Apple's

Readers should hold the benchmarks loosely. All figures come from Apple's internal testing. All use "up to" phrasing. All lean on scattered baselines — M1, M4, M2 Pro, M3 Ultra — rather than a single consistent reference point. Independent reviews do not exist yet.

Pricing is the harder sell. The M6 mini starts at $899 and the M5 Pro mini at $1,699, each $100 above its predecessor. Mac Studio runs $2,499 with M5 Max and $5,499 with M5 Ultra. Apple's entry desktop now costs roughly 50 percent more than it did two months ago.

Memory supply is the reason. DRAM prices jumped about 90 percent quarter over quarter in early 2026, then climbed more than 50 percent again. IDC frames it as a structural reallocation of wafer capacity toward HBM for AI. The 512GB configuration will not ship until late October.

There is a footnote of corporate timing as well. Tim Cook hands the chief executive role to John Ternus on August 31. These are the last Macs launched on his watch.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Apple Pushes On-Device AI With 512GB Mac Studio and M5 Ultra
Tech & Business

Apple Pushes On-Device AI With 512GB Mac Studio and M5 Ultra

Apple introduced a new Mac Studio on August 25 built around the M5 Max and an all-new M5 Ultra, and the configuration Apple is pushing hardest is aimed squarely...

Seung Jung7 days ago
OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months
Tech & Business

OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months

AI-generated kernels beat OpenAI expert-written versions by up to 1.8x, as Jalapeño posts its first InferenceX benchmark results.

Seung Jung22 days ago
Google Signs Marvell for Custom AI Silicon, Ending Broadcom's Clean Run
Tech & Business

Google Signs Marvell for Custom AI Silicon, Ending Broadcom's Clean Run

An SEC filing shows Google tapped Marvell for TPU-adjacent custom silicon with a $12.2 billion warrant, ending Broadcom's sole-supplier position.

Seung Jung28 days ago
Amazon Triples Its Nvidia Order to 2 Million GPUs
Tech & Business

Amazon Triples Its Nvidia Order to 2 Million GPUs

Amazon is adding 2 million more Nvidia GPUs to AWS just five months after committing to 1 million, even as it scales its own Trainium and Graviton silicon.

Seung Jung21 days ago
Nvidia Nears $12.9 Billion Deal to Buy Hugging Face
Tech & Business

Nvidia Nears $12.9 Billion Deal to Buy Hugging Face

Nvidia has reportedly agreed to acquire Hugging Face for $12.9 billion, a multiple of roughly 80x revenue that buys the developer graph, not the income statement.

Seung Jung21 days ago
Perplexity's New Agent Runs on Your Own GPU and Bills You Nothing to Do It
Tech & Business

Perplexity's New Agent Runs on Your Own GPU and Bills You Nothing to Do It

Perplexity and Nvidia shipped Portable Computer, an agent running Qwen or PPLX 27B on local RTX hardware with no metered tokens until the user allows a cloud step.

Seung Jung22 days ago