Local AI hardware
Best GPU for Local AI
Compare GPUs for local AI with a guided tool that starts from VRAM arithmetic — model size at Q4 — before raster charts or brand loyalty.
Best starting point
GPU Finder
Built for builders running open models locally. Use this guide for context, then run the tool to turn those priorities into a clearer shortlist.
Explained methodology
Each tool and guide makes the decision criteria and fit logic visible.
Clear disclosure
Commercial relationships are disclosed so readers can judge with context.
Ongoing updates
Important guides and tools are reviewed as products and categories change.
Overview
The best GPU for local AI is usually the one with enough usable memory for the models you will actually run, not the card that wins a synthetic raster leaderboard. This guide separates capacity-first builds from gaming-first cards before you overspend.
Start with whether the model fits
Local inference fails as memory planning first. At Q4, a rough planning rule is parameters × bits ÷ 8, plus headroom for context and the OS. A 70B model needs about 42GB under that heuristic — which is why a 24GB card is a mid-size workhorse, not a 70B free pass.
Quantization and context length change the answer more than brand loyalty. A 7B–13B chat model at Q4 fits comfortably on 8–12GB with room for a long prompt. A 32B–34B coding model wants 20GB+ before you start fighting swap. Multi-modal or long-context workloads eat VRAM that parameter-count tables never show. Write down the largest model you actually intend to keep resident, then size the card to that — not to a leaderboard you will never run.
Capacity tiers that match real workloads
Treat these as planning buckets, not product rankings:
- 12–16GB — small and mid chat models, image tools that stay under the limit, light coding assistants. Fine for experimenting; painful if every useful model spills to CPU RAM.
- 24GB — the practical sweet spot for many builders: mid-size open weights, batch jobs, and a second model occasionally in memory. Still the wrong expectation for comfortable single-GPU 70B.
- 48GB+ discrete or 64GB+ unified memory — where 70B-class and long-context work stop being a science project. A 32GB flagship (for example RTX 5090) is still short of the ~42GB Q4 weight floor once OS and context land. Apple Silicon with 64GB+ unified memory is often cheaper capacity than chasing multi-GPU desktops.
A used 24GB Ada card can beat a new gaming SKU with less memory for local AI specifically. Do not buy a 16GB “faster” raster champion if your workload is “does the model fit.”
Then pick the ecosystem you can live with
CUDA remains the path most guides, Docker images, and tutorials assume. NVIDIA wins on “will this random repo run on day one,” which matters if you are not already fluent in ROCm debugging.
AMD often wins on street price per gigabyte — especially used 24GB cards — with thinner day-one tooling. That trade is acceptable when you already know the stack you will run and have verified ROCm or Vulkan paths for it. It is a bad surprise when your first three GitHub READMEs only list CUDA.
Apple unified memory trades discrete gaming leadership for large-model capacity inside a Mac. Buy that path when the Mac form factor and memory size are the product; do not buy it as a substitute for a Windows gaming PC that also “does some AI.”
Power, case, and cooling are not footnotes
Local AI cards that sit at high utilization for hours stress the PSU and thermals harder than a gaming session that spikes and recovers. Check connector count, case clearance, and whether the card will throttle under sustained load in your chassis. A card that fits the VRAM math but browns out the power supply is not a bargain.
Laptop GPUs with the same marketing name usually have less usable memory bandwidth and worse sustained clocks than the desktop twin. If the machine must be portable, size for laptop reality — not desktop charts.
What to judge before spending
- What is the largest model (parameters + quantization + context) you will keep loaded day to day?
- Is day-one CUDA compatibility worth more than raw GB per pound?
- Will this machine also game, or is it an inference box?
- New vs used: is software support still active for the generation you are buying?
- Desktop discrete, workstation, or Mac unified memory — which form factor you will actually keep on the desk?
Run the finder, then verify on the tables
Use the GPU finder for a shortlist keyed to your constraints, then confirm fit on the /gpus size tables and /macs capacity notes before you buy. Street prices move weekly; treat launch MSRP as an anchor only. If two cards both clear the VRAM bar, prefer the one whose software path you already use — capacity that sits idle behind broken drivers is still a failed purchase.
Top recommendations
GeForce RTX 5090
Top pick32GB consumer VRAM with very high bandwidth — the usual single-card ceiling for mid-size Q4 models, not a 70B host.
View offer (paid link)Affiliate disclosure: this link may earn AI Choice Engine a commission at no extra cost to you.
GeForce RTX 4090
24GB CUDA workhorse24GB Ada card — comfortable for mid-size Q4 models. Does not hold 70B Q4.
View offer (paid link)Affiliate disclosure: this link may earn AI Choice Engine a commission at no extra cost to you.
Apple M4 Max (64GB)
Tight Mac 70B path64GB unified memory (~48GB usable) — a tight Mac path for 70B Q4, not comfortable headroom after OS and runtime overhead.
View offer (paid link)Affiliate disclosure: this link may earn AI Choice Engine a commission at no extra cost to you.
Best-fit GPU profile
Answer 3 short prompts to get a logic-based recommendation plus strong alternatives.
- 4 questions, under 2 minutes
- Workload-first: gaming, creative, or local models
- Covers NVIDIA, AMD, and Mac unified-memory paths
Current status
Question 1 of 3
State is saved locally, so refreshing keeps your progress intact.
PC building · Graphics
What is the main job for this GPU?
VRAM and software ecosystems diverge hard between games, creative apps, and local models. Pick the job that would make you regret the wrong card.
Restoring your saved answers...
Frequently asked questions
How much VRAM do I need for a 70B model?+−
About 42GB at Q4 under a simple parameters×bits planning rule, before generous context. Single 24GB cards are the wrong expectation for comfortable 70B local use.
Is the RTX 4090 still worth it for local AI?+−
Often yes on the used market: 24GB still covers many mid-size workloads, and Ada software support is mature. It is not a substitute for 48GB+ discrete or 64GB+ unified memory when 70B is the real target.
Should I buy a Mac instead of a desktop GPU?+−
Buy a Mac when you want that form factor and enough unified memory for your models. Do not buy one as a Windows gaming substitute — check the Macs capacity table for usable-GB after OS headroom.
Suggested tools
Tech & Devices
Local AI Model Fit Finder
Estimate minimum VRAM from model size and quantisation, then shortlist GPUs and Mac configs that can hold the weights — with honest tight-fit caveats.
Open →
Tech & Devices
CPU Finder
Match workload, budget, platform preference, and upgrade appetite to the processor that fits your actual build — gaming-first, balanced, or workstation.
Open →
Tech & Devices
Best Laptop Finder
Match budget, workload, portability needs, and performance expectations to the laptop shortlist that best fits your real-world use case.
Open →
Related guides
Related guide
Best CPU for Gaming and Streaming
Compare processors for a PC that games and broadcasts at once, with a guided decision flow that weighs refresh-rate targets, budget splits, and socket longevity.
Open →
Related guide
Best Mac for Local LLMs
Choose a Mac for local LLMs by unified memory capacity and sustained load — not by chip marketing alone — then shortlist with the local AI model fit finder.
Open →
Related guide
Buy a GPU vs Rent Cloud GPUs for AI
Decide when to buy a local GPU for AI versus renting cloud GPUs — using utilization, VRAM needs, and ops overhead rather than sticker price alone.
Open →
AI & tech clusters
AI models
Compare AI models
Frontier, balanced, and budget tiers with pricing and benchmark receipts.
Open →
GPUs
GPU and local model fit
Which local models fit each card's VRAM and where the memory ceiling bites.
Open →
Open weights
Open-weight models
Licence, deployment, and hardware notes for open checkpoints.
Open →