Skip to main content

Cloud inference

Where to rent GPUs — or skip them for a managed API

Three product shapes solve “I need inference”: rent a GPU and run your stack, use serverless GPU functions, or call a managed model API. Pick the shape first; then compare providers inside it.

GPU marketplace

You choose the card and image. Best when control and $/hour matter more than an API SLA.

Serverless GPU

You ship code; the platform scales hardware. Best for bursty jobs without babysitting VMs.

Managed inference API

You pay for tokens or predictions. Best when time-to-product beats owning the serving stack.

Start here

Pick the path before you pick the logo

Most bad shortlists mix a marketplace pod, a serverless function, and a token API as if they were the same product.

Rent a GPU (marketplace / cluster)

When: You need a specific SKU, custom serving stack, or sustained utilization that beats token APIs.

Next: Start with RunPod or Vast.ai for flexible rental; size VRAM on /gpus first. Remember stop ≠ terminate — tear down storage when the experiment ends.

Serverless GPU functions

When: Work arrives in bursts, your code is container-shaped, and you refuse to babysit VMs overnight.

Next: Compare Modal and Fal.ai against cold-start tolerance for your workload; treat DGX Cloud Lepton as NVIDIA marketplace capacity, not a Modal clone.

Managed inference API

When: You want tokens or predictions tomorrow, not CUDA drivers — and the model is already on a catalog.

Next: Shortlist Together, Fireworks, Groq, Replicate, Hugging Face Endpoints, or Baseten by model fit and latency.

Buy local hardware instead

When: Privacy, offline use, or always-on personal inference makes cloud rental look like a subscription you cannot pause.

Next: Use the /gpus capacity tables and GPU finder before you rent indefinitely. Match VRAM to model size, then pick rental vs buy.

Before you rent

Preflight checklist

Five questions that prevent the usual bill shocks and cold-start surprises.

  1. 1Name the product shape first: marketplace GPU, serverless function, managed API, or local card — do not mix them in one shortlist.
  2. 2Write down cold-start tolerance in user-visible seconds. If the answer is “near zero,” scale-to-zero is a product risk.
  3. 3Forecast spend from utilization: idle pods, token volume, or per-prediction size — not from a homepage hero number.
  4. 4Confirm the exact model ID or container image before architecture lock-in.
  5. 5Plan a failover provider for anything user-facing.
  6. 6Size VRAM on /gpus before you rent — a 70B Q4 fit check beats discovering OOM after the pod is warm.
  7. 7On marketplaces: stop ≠ terminate. Tear down volumes/disks when the experiment ends or storage keeps billing.
  8. 8Prefer interruptible/community only for jobs that can checkpoint; use secure/on-demand for anything user-facing.

Compare chart

Fit by decision axis

Bars show how each provider scores on the axis that usually drives the choice — not a single overall winner.

At a glance

Provider matrix

Billing shape and GPU control differ more than marketing copy. Scan the matrix, then open a head-to-head.

GPU inference providers compared by category and billing
ProviderCategoryBillingGPU choiceBest for
RunPodGPU marketplacePer-second / per-hour GPU rentalBroad consumer and datacenter SKUsTeams that want to pick a specific GPU SKU and control the container stack
Vast.aiGPU marketplaceMarketplace hourly bidsWide, host-dependent inventoryCost-sensitive batch jobs and experiments that can tolerate host variability
ModalServerless GPUServerless CPU/GPU timePlatform-managed GPU classesDevelopers who want GPU code as functions with autoscaling
Together AIManaged inference APIToken / request API pricingHidden behind the inference APIServing open-weight models without operating your own GPU fleet
Fireworks AIManaged inference APIToken / request API pricingManaged serving stackProduction chat and agent backends that need snappy open-model inference
ReplicateManaged inference APIPer-prediction / hardware timeTied to the model’s declared hardwareProduct teams shipping model-backed features without owning GPU ops
LambdaCloud GPU clusterOn-demand / reserved instancesDatacenter NVIDIA focusTraining and sustained inference on reserved or on-demand GPU instances
Hugging Face EndpointsManaged inference APIDedicated endpoint + usageSelect hardware tiers per endpointTeams already living in the Hub who want a managed endpoint per model
GroqManaged inference APIToken / request API pricingFixed LPU-backed serving — no SKU selectionLatency-sensitive chat, agents, and batch inference on Groq’s supported model set
Fal.aiServerless GPUServerless per-run / GPU-secondPlatform-managed GPU classes per modelProduct teams shipping image, video, or audio features without operating GPU fleets
BasetenManaged inference APIDedicated endpoint + autoscale usageSelect hardware per deploymentML teams shipping bespoke models who want managed serving without building the ops stack
CoreWeaveCloud GPU clusterReserved / on-demand cluster capacityDatacenter NVIDIA at cluster scaleTeams that need large, predictable GPU clusters without building their own datacenter
NVIDIA DGX Cloud LeptonGPU marketplacePartner GPU marketplace / DGX Cloud usageNVIDIA and partner GPU classes exposed through DGX Cloud LeptonTeams already standardised on NVIDIA DGX Cloud / partner GPU inventory who need marketplace access rather than a standalone serverless brand

Some outbound links use a first-party redirect hop for click counting. Commission is only claimed when a partner programme is active for that specific link — most vendor hops here are not paid placements. Affiliate disclosure.

Providers

13 places teams actually rent or buy inference

Grouped by product shape — marketplace rental, serverless GPU, managed API, or cluster. Each card now shows GPU choice, cold-start reality, and the main watch-out.

GPU marketplace

Per-second / per-hour GPU rental

RunPod

7.4/10 avg

Rent individual GPUs or pods by the hour, with community and secure cloud options for training and inference workloads.

GPU choice
Broad consumer and datacenter SKUs
Cold start
Depends on pod spin-up and image cache

Best for: Teams that want to pick a specific GPU SKU and control the container stack

Watch out: You still own ops: images, scaling, and idle spend. Stopping a pod is not the same as terminating it — stopped storage can keep billing until you tear the volume down.

Marketplace hourly bids

Vast.ai

6.4/10 avg

A marketplace of host-offered GPUs with competitive spot-style pricing for interruptible or carefully filtered capacity.

GPU choice
Wide, host-dependent inventory
Cold start
Instance provisioning + your stack

Best for: Cost-sensitive batch jobs and experiments that can tolerate host variability

Watch out: Host quality and availability vary — filter carefully for production inference. Interruptible/community instances are cheaper but can be reclaimed; stop ≠ terminate (disk can still bill).

Partner GPU marketplace / DGX Cloud usage

NVIDIA DGX Cloud Lepton

6.0/10 avg

NVIDIA-operated GPU marketplace / AI platform (formerly Lepton AI) connecting developers to NVIDIA’s partner compute ecosystem — not an independent Modal-style serverless vendor.

GPU choice
NVIDIA and partner GPU classes exposed through DGX Cloud Lepton
Cold start
Depends on the selected partner capacity and deployment shape

Best for: Teams already standardised on NVIDIA DGX Cloud / partner GPU inventory who need marketplace access rather than a standalone serverless brand

Watch out: Product identity is DGX Cloud Lepton after NVIDIA’s acquisition — treat legacy lepton.ai marketing as historical; verify current regions, SKUs, and billing on NVIDIA’s pages

Serverless GPU

Managed inference API

Token / request API pricing

Together AI

7.2/10 avg

Managed inference and fine-tuning oriented around open models, with API-shaped endpoints rather than raw GPU rental.

GPU choice
Hidden behind the inference API
Cold start
Low for warm endpoints

Best for: Serving open-weight models without operating your own GPU fleet

Watch out: You trade GPU-level control for API pricing and catalog coverage

Token / request API pricing

Fireworks AI

7.0/10 avg

Low-latency inference API focused on fast serving of popular open and partner models.

GPU choice
Managed serving stack
Cold start
Typically warm API

Best for: Production chat and agent backends that need snappy open-model inference

Watch out: Compare latency and price per token against peers for your exact model

Per-prediction / hardware time

Replicate

6.6/10 avg

Run public and private models via simple API predictions — popular for image, video, and community model demos.

GPU choice
Tied to the model’s declared hardware
Cold start
Cold models can add latency

Best for: Product teams shipping model-backed features without owning GPU ops

Watch out: Per-prediction economics and cold models can surprise at scale

Dedicated endpoint + usage

Hugging Face Endpoints

7.0/10 avg

Deploy Hub models to dedicated inference endpoints with Hugging Face’s hosting and scaling controls.

GPU choice
Select hardware tiers per endpoint
Cold start
Scaled-to-zero endpoints can cold-start

Best for: Teams already living in the Hub who want a managed endpoint per model

Watch out: Endpoint sizing and autoscaling choices drive cost more than list price alone

Token / request API pricing

Groq

6.4/10 avg

Ultra-low-latency inference API built on Groq LPU hardware — strong when raw tokens-per-second on supported models matters more than picking a GPU SKU.

GPU choice
Fixed LPU-backed serving — no SKU selection
Cold start
Typically warm API endpoints

Best for: Latency-sensitive chat, agents, and batch inference on Groq’s supported model set

Watch out: Model catalog and deployment options are narrower than general GPU rental; verify your model is supported before committing architecture

Dedicated endpoint + autoscale usage

Baseten

6.8/10 avg

Deploy custom models to production-grade inference endpoints with autoscaling, tracing, and hardware tier selection — between DIY GPU rental and a closed model API.

GPU choice
Select hardware per deployment
Cold start
Scaled-to-zero endpoints can cold-start

Best for: ML teams shipping bespoke models who want managed serving without building the ops stack

Watch out: Pricing spans dedicated and shared tiers — forecast from your traffic shape, not list rates alone

Cloud GPU cluster

Head-to-head

Comparisons that clarify the trade-off

Marketplace vs serverless, API vs API, and cluster vs rental — each pair answers a real shortlist question.

FAQ

Questions teams ask before the first invoice

Qualitative answers only — verify live pricing on each provider.

Is renting a GPU cheaper than a managed API?

Only when utilization is high enough that hardware hours beat token or per-prediction pricing — and when you can operate the stack. Spiky or low-traffic products usually lose money on warm pods and win on APIs or serverless.

What is the difference between RunPod and Together AI?

RunPod rents GPUs so you run your own serving stack. Together AI sells managed inference for open models. One is hardware control; the other is time-to-endpoint.

When do cold starts matter?

Whenever a user is waiting. Scale-to-zero serverless and cold Replicate/HF endpoints can add seconds. Batch jobs and async media pipelines tolerate cold starts; interactive chat often does not.

Should I rent cloud GPUs or buy a local card?

Rent for burst capacity, team sharing, and SKUs you will not keep busy for years. Buy local when privacy, offline use, or always-on personal inference makes cloud hours look like rent. Many teams do both.

Does stopping a rented GPU stop the bill?

Not always. Stopping a pod can leave disks or network volumes attached. Terminate the instance and confirm storage teardown — especially on RunPod and Vast.ai — or you keep paying after the GPU is idle.

How do I rent a GPU for a coding agent or local model without getting burned?

Size the model on /gpus first, pick marketplace vs managed API, start with a short prepaid or capped budget, pull a known image, run one smoke prompt, then shut down fully (stop + delete storage). Keep a failover provider for anything user-facing.

Are the fit scores prices?

No. They are editorial planning ratings for product shape. Always verify current pricing, regions, and model availability on the provider site before you commit spend.

Decision rule

If you need a specific GPU SKU, rent. If you need a model endpoint tomorrow, buy an API.

Mixing those goals in one shortlist is how teams overpay. Use the fit chart for shape, the matrix for ops reality, and a head-to-head before you commit spend.