Skip to main content

Local AI hardware

AI PCs — NPU laptops vs GPU desktops vs Mac

Honest trade-offs between Copilot+ NPUs, discrete-GPU towers, and Apple silicon for local open-weight models — with links to GPUs, Macs, and cloud inference.

GPU catalogCloud GPU inferenceMac buying guideOpen-weight models

The planning rule

Pick the largest model you need to run locally, then buy the memory architecture that can hold it — NPU marketing comes after that.

No invented benchmark league tables. Capacity guidance follows the same VRAM/unified-memory arithmetic as the GPU and Mac catalogs. Amazon links are product-class searches, not fixed listings.

As an Amazon Associate I earn from qualifying purchases. Amazon UK links may earn us a commission; this does not change the products we include. Affiliate disclosure.

Form factors

Four AI PC paths — and what each is actually for

NPU TOPS and GPU TFLOPS are not interchangeable scores. Memory capacity is the first filter.

AI PC form factor comparison
Form factorComputeMemoryLocal AIGamingPortabilitySearch
Copilot+ / NPU laptop
Copilot+ / NPU laptopNPU TOPS are not comparable to GPU TFLOPS or Mac tokens/sec. 16GB RAM caps serious local LLM work — verify memory before buying. UK Copilot+ Snapdragon entry examples sit around ~£800 (e.g. Zenbook A14 class at John Lewis); memory-crunch pricing can push cheaper non-Copilot configs — check live list.
CPU + integrated graphics + dedicated NPU (vendor-quoted TOPS)Shared system RAM — often 16–32GB fixed at purchaseStrong for background AI features, transcription, and small on-device models; weak for large open-weight LLMsLight esports and older titles — not a discrete-GPU replacementHigh — battery-powered daily driverSearch Amazon UK(paid link) (opens in new tab)
Discrete-GPU desktop
Discrete-GPU desktopVRAM capacity decides what runs, not the marketing name on the box. Power, thermals, and noise are real ownership costs. Flagship mobile AI-PC anchors (e.g. MSI Stealth A16 AI+ with HX 370 + RTX 5090 laptop GPU) sit around ~£4.1–4.3k UK — verify the exact GPU suffix; laptop 5090 ≠ desktop 5090.
Dedicated GPU VRAM + CPU — CUDA/ROCm/oneAPI depending on cardSeparate system RAM and GPU VRAM — VRAM is the local-AI ceilingBest path for open-weight models when VRAM fits the quantisation — scales with card tierFull gaming PC — the GPU does double dutyNone — tower or SFF on a deskSearch Amazon UK(paid link) (opens in new tab)
Apple Silicon Mac
Apple Silicon MacMemory is soldered — buy the RAM you need on day one. Large models are a capacity decision before a chip decision. MacBook Neo (Apple UK £699–£799) is an entry everyday Mac with an 8GB floor — not a local-LLM path.
Unified memory shared by CPU, GPU, and Neural EngineFixed unified pool — 24GB minimum for serious local LLM useExcellent memory-per-pound for medium models; bandwidth-limited vs top discrete GPUs at equal spendModerate — macOS library and port quality vary by titleHigh on MacBook; Mac mini / Studio are desk-boundSearch Amazon UK(paid link) (opens in new tab)
Compact GPU workstation / mini PC
Compact GPU workstation / mini PCCompact thermals throttle sustained loads. Upgrade paths are often worse than a standard ATX build.
Mobile or compact discrete GPU in a small chassisOften soldered RAM with limited GPU VRAM — read the spec sheet carefullyUseful middle ground when desk space is tight but NPU laptops are too limitedVaries widely — many compact boxes use laptop-class GPUsLow — small desk footprint, not a laptopSearch Amazon UK(paid link) (opens in new tab)

Comparison notes last reviewed August 10, 2026. Figures describe product classes, not benchmark winners.

NPU vs GPU — different jobs

Copilot+ NPUs accelerate vendor-defined on-device tasks — live captions, background blur, small bundled models. They do not replace a 12GB+ VRAM GPU for loading a 32B open-weight checkpoint. Comparing NPU TOPS to GPU TFLOPS is marketing arithmetic, not a planning metric.

When cloud inference is the honest answer

If the target model needs more memory than any machine you are willing to buy, renting a cloud GPU or using an API is cheaper than overspending on hardware you will still outgrow. The GPU inference catalog compares providers without pretending local always wins.

Software stack matters as much as silicon

The same weights run differently on llama.cpp, Ollama, vLLM, MLX, and vendor tools. A Mac with MLX can feel faster than a higher-TFLOPS card with a mismatched backend. Start from the model catalog, then pick hardware that fits the runtime you will actually use.

Best for

Who each path suits

Match the hardware class to the workload — NPU laptops, discrete GPUs, and Macs solve different jobs.

Copilot+ / NPU laptop

  • · Windows laptops with Copilot+ features
  • · Meetings, notes, and lightweight on-device AI
  • · Buyers who will rent cloud inference for big models

Discrete-GPU desktop

  • · Running 7B–70B open models locally (VRAM permitting)
  • · Gamers who also want local inference
  • · Upgradeable buyers who can swap GPUs later

Apple Silicon Mac

  • · Developers already on macOS
  • · Unified-memory workloads up to practical ~48–64B quantised ceilings
  • · Quiet desk inference without building a tower

Compact GPU workstation / mini PC

  • · Homelab and inference beside the main desk
  • · Buyers who cannot host a full tower
  • · Secondary inference node while a laptop is the daily driver

Common questions

AI PC buying questions

Answered from the verified figures on this page rather than general guidance.

Is a Copilot+ PC enough to run local LLMs?

For small bundled models and Windows AI features, often yes. For the open-weight models in our catalog at useful quantisations, usually no — system RAM and the lack of large dedicated VRAM cap practical model size well below what a discrete-GPU desktop or high-memory Mac can hold.

Mac or PC for local AI in 2026?

Mac wins on unified memory per pound and quiet desk operation when the model fits in RAM. PC wins on raw GPU choice, upgradeability, and gaming. Neither wins every workload — decide on model size first, then ecosystem second.

How much RAM or VRAM do I need?

Use the GPU catalog rule of thumb: parameters × bits ÷ 8 plus headroom. A 32B model at Q4 needs on the order of 20GB+ of usable memory. That is why 16GB laptops — NPU or not — are planning dead-ends for serious open-weight work.

Should I build a PC or buy prebuilt for AI?

Building lets you pick the GPU tier deliberately. Prebuilts often pair a mid GPU with flashy RAM marketing — verify the exact card model and PSU, then compare to the DIY parts lists. The prebuilt guide covers what to check without inventing SKUs.

Keep exploring

Go deeper on the stack

VRAM fit, unified memory, and rental inference each answer a different constraint.