Local AI Capacity First
Your answers favor memory headroom for open models: buy VRAM (or usable unified memory) before chasing peak raster scores.
Best for: Builders running mid-to-large open models locally, People who already know CUDA or MLX tooling, Buyers who will regret undersized VRAM more than lower frame rates
Watch out: A 70B Q4 model needs ~35GB of weights (~42GB with runtime overhead) — 24GB and 32GB discrete cards do not host it. Prefer 48GB+ discrete or 64GB+ unified memory; verify street prices and power/cooling before you buy.
Shortlist examples: RTX PRO 6000 Blackwell, GeForce RTX 5090, GeForce RTX 4090